{"thread":{"id":"57568","subject":"[PATCH 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","startedAt":"2022-03-15T21:31:07Z","lastAt":"2022-05-24T12:31:31Z","messageCount":175,"participants":["Neeraj K. Singh via GitGitGadget","Neeraj Singh via GitGitGadget","Junio C Hamano","Patrick Steinhardt","Neeraj Singh","Bagas Sanjaya","Ævar Arnfjörð Bjarmason","nksingh85@gmail.com","Johannes Schindelin"],"isPatch":true,"patchVersion":1,"patchTotal":7},"messages":[{"id":"451440","messageId":"pull.1134.git.1647379859.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":null,"subject":"[PATCH 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:52Z","receivedAt":"2022-03-15T21:31:07Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"When core.fsync includes loose-object, we issue an fsync after every written\nobject. For a 'git-add' or similar command that adds a lot of files to the\nrepo, the costs of these fsyncs adds up. One major factor in this cost is\nthe time it takes for the physical storage controller to flush its caches to\ndurable media.\n\nThis series takes advantage of the writeout-only mode of git_fsync to issue\nOS cache writebacks for all of the objects being added to the repository\nfollowed by a single fsync to a dummy file, which should trigger a\nfilesystem log flush and storage controller cache flush. This mechanism is\nknown to be safe on common Windows filesystems and expected to be safe on\nmacOS. Some linux filesystems, such as XFS, will probably do the right thing\nas well. See [1] for previous discussion on the predecessor of this patch\nseries.\n\nThis series is important on Windows, where loose-objects are included in the\nfsync set by default in Git-For-Windows. In this series, I'm also setting\nthe default mode for Windows to turn on loose object fsyncing with batch\nmode, so that we can get CI coverage of the actual git-for-windows\nconfiguration upstream. We still don't actually issue fsyncs for the test\nsuite since GIT_TEST_FSYNC is set to 0, but we exercise all of the\nsurrounding batch mode code.\n\nThis work is based on 'seen' at 367f447f0f0cf39e9830c865e8373e42a3c45303.\nIt's dependent on ns/core-fsyncmethod.\n\n[1]\nhttps://lore.kernel.org/git/2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com/\n\nNeeraj Singh (7):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncmethod: batched disk flushes for loose-objects\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  core.fsyncmethod: tests for batch mode\n  core.fsyncmethod: performance tests for add and stash\n\n Documentation/config/core.txt |  5 ++\n builtin/unpack-objects.c      |  3 ++\n builtin/update-index.c        |  6 +++\n bulk-checkin.c                | 89 +++++++++++++++++++++++++++++++----\n bulk-checkin.h                |  2 +\n cache.h                       | 12 ++++-\n compat/mingw.h                |  3 ++\n config.c                      |  4 +-\n git-compat-util.h             |  2 +\n object-file.c                 |  2 +\n t/lib-unique-files.sh         | 36 ++++++++++++++\n t/perf/p3700-add.sh           | 59 +++++++++++++++++++++++\n t/perf/p3900-stash.sh         | 62 ++++++++++++++++++++++++\n t/perf/perf-lib.sh            |  4 +-\n t/t3700-add.sh                | 22 +++++++++\n t/t3903-stash.sh              | 17 +++++++\n t/t5300-pack-object.sh        | 32 ++++++++-----\n 17 files changed, 335 insertions(+), 25 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 367f447f0f0cf39e9830c865e8373e42a3c45303\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1134%2Fneerajsi-msft%2Fns%2Fbatched-fsync-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1134/neerajsi-msft/ns/batched-fsync-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1134\n-- \ngitgitgadget\n"},{"id":"451441","messageId":"a77d02df626ed6dff485e1342ff7affd6999ec44.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:53Z","receivedAt":"2022-03-15T21:31:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nPreparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex e988a388b65..93b1dc5138a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -278,21 +278,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"451442","messageId":"d38f20b4430bada1d0dccc1e600e6f0b098f3767.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:54Z","receivedAt":"2022-03-15T21:31:11Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin plugging and unplugging functionality,\ntakes advantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have seen the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nBatch mode is only enabled if core.fsyncObjectFiles is false or unset.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  5 +++\n bulk-checkin.c                | 67 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 ++\n cache.h                       |  8 ++++-\n config.c                      |  2 ++\n object-file.c                 |  2 ++\n 6 files changed, 85 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 062e5259905..c041ed33801 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -628,6 +628,11 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then a single full fsync to trigger\n+  the disk cache flush at the end of the operation. This mode is expected to\n+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+  and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 93b1dc5138a..5c13fe17802 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -80,6 +86,34 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t}\n+\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +308,24 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tassert(bulk_fsync_objdir);\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -288,6 +340,19 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * A temporary object directory is used to hold the files\n+\t * while they are not fsynced.\n+\t */\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -297,4 +362,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex d347d0757f7..4d07691e791 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH\n };\n \n extern enum fsync_method fsync_method;\n@@ -1766,6 +1767,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex 261ee7436e0..0b28f90de8b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 295cb899e22..ef6621ffe56 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1894,6 +1894,8 @@ static void close_loose_object(int fd)\n \n \tif (fsync_object_files > 0)\n \t\tfsync_or_die(fd, \"loose object file\");\n+\telse if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n \t\t\t\t       \"loose object file\");\n-- \ngitgitgadget\n\n"},{"id":"451443","messageId":"99e3a61b9191e4102529c8b90b36e0f96f2b23bf.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 4/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:56Z","receivedAt":"2022-03-15T21:31:13Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..c55b6616aed 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"451444","messageId":"b0480f0c814ae1c726d51af172bf5833fc74f010.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:55Z","receivedAt":"2022-03-15T21:31:14Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete.  This\nusage is unlikely, since any tool invoking update-index and expecting to\nsee objects would have to synchronize with the update-index process\nafter passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 75d646377cc..38e9d7e88cb 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1190,6 +1194,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"451445","messageId":"4e56c58c8cb812b244feaf814b43b7dc28879f9a.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 5/7] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:57Z","receivedAt":"2022-03-15T21:31:18Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the test suite, since GIT_TEST_FSYNC is\nset to 0. However, we do exercise all of the surrounding batch mode code\nsince GIT_TEST_FSYNC merely makes the maybe_fsync wrapper always appear\nto succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n\nchange fsyncmethod to batch as well\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex 4d07691e791..04193d87246 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 0b28f90de8b..c76443dc556 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 0892e209a2f..fffe42ce7c1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \ngitgitgadget\n\n"},{"id":"451446","messageId":"88e47047d790f76ffacaa61ed8041b733e30f45a.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 6/7] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:58Z","receivedAt":"2022-03-15T21:31:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nWe aren't actually issuing any fsyncs in these tests, since\nGIT_TEST_FSYNC is 0, but we still exercise all of the tmp_objdir logic\nin bulk-checkin.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 22 ++++++++++++++++++++++\n t/t3903-stash.sh       | 17 +++++++++++++++++\n t/t5300-pack-object.sh | 32 +++++++++++++++++++++-----------\n 4 files changed, 96 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..1f349f52ad3 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,26 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit $BATCH_CONFIGURATION add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 4abbc8fccae..877276c1ca3 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1410,6 +1411,22 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..8e2f73cc68f 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,25 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +190,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +204,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"451447","messageId":"876741f1ef9a5b8af28f73948a3e9ddc16d88c6d.1647379859.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH 7/7] core.fsyncmethod: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-15T21:30:59Z","receivedAt":"2022-03-15T21:31:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings. This shows the benefit of batch\nmode relative to an ordinary stash command.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 59 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 62 +++++++++++++++++++++++++++++++++++++++++++\n t/perf/perf-lib.sh    |  4 +--\n 3 files changed, 123 insertions(+), 2 deletions(-)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..2ea78c9449d\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,59 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\ttest_perf \"add $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..3526f06cef4\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,62 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG stash push -u -- files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex 932105cd12c..d270d1d962a 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -98,8 +98,8 @@ test_perf_create_repo_from () {\n \tmkdir -p \"$repo/.git\"\n \t(\n \t\tcd \"$source\" &&\n-\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" 2>/dev/null ||\n-\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\"; } &&\n+\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" ||\n+\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\" 2>/dev/null;} &&\n \n \t\t# common_dir must come first here, since we want source_git to\n \t\t# take precedence and overwrite any overlapping files\n-- \ngitgitgadget\n"},{"id":"451455","messageId":"xmqqbky6bgw0.fsf@gitster.g","threadId":"57568","inReplyTo":"a77d02df626ed6dff485e1342ff7affd6999ec44.1647379859.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-16T05:33:51Z","receivedAt":"2022-03-16T05:34:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Preparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n>\n> * Rename 'state' variable to 'bulk_checkin_state', since we will later\n>   be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n>   find in the debugger, since the name is more unique.\n>\n> * Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n>   static variable. Doing this avoids resetting the variable in\n>   finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n>   seem to unintentionally disable the plugging functionality the first\n>   time a new packfile must be created due to packfile size limits. While\n>   disabling the plugging state only results in suboptimal behavior for\n>   the current code, it would be fatal for the bulk-fsync functionality\n>   later in this patch series.\n\nSorry, but I am confused.  The bulk-checkin infrastructure is there\nso that we can send many little objects into a single packfile\ninstead of creating many little loose object files.  Everything we\nthrow at object-file.c::index_stream() will be concatenated into the\nsingle packfile while we are \"plugged\" until we get \"unplugged\".\n\nMy understanding of what you are doing in this series is to still\ncreate many little loose object files, but avoid the overhead of\nhaving to fsync them individually.  And I am not sure how well the\noriginal idea behind the bulk-checkin infrastructure to avoid\noverhead of having to create many loose objects by creating a single\npackfile (and presumably having to fsync at the end, but that is\njust a single .pack file) with your goal of still creating many\nloose object files but synching them more efficiently.\n\nIs it just the new feature is piggybacking on the existing bulk\ncheckin infrastructure, even though these two have nothing in\ncommon?\n\n"},{"id":"451456","messageId":"YjGSOJlZnDSske3s@ncase","threadId":"57568","inReplyTo":"d38f20b4430bada1d0dccc1e600e6f0b098f3767.1647379859.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-16T07:31:04Z","receivedAt":"2022-03-16T07:31:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Mar 15, 2022 at 09:30:54PM +0000, Neeraj Singh via GitGitGadget wrote:\n> From: Neeraj Singh <neerajsi@microsoft.com>\n> \n> When adding many objects to a repo with `core.fsync=loose-object`,\n> the cost of fsync'ing each object file can become prohibitive.\n> \n> One major source of the cost of fsync is the implied flush of the\n> hardware writeback cache within the disk drive. This commit introduces\n> a new `core.fsyncMethod=batch` option that batches up hardware flushes.\n> It hooks into the bulk-checkin plugging and unplugging functionality,\n> takes advantage of tmp-objdir, and uses the writeout-only support code.\n> \n> When the new mode is enabled, we do the following for each new object:\n> 1. Create the object in a tmp-objdir.\n> 2. Issue a pagecache writeback request and wait for it to complete.\n> \n> At the end of the entire transaction when unplugging bulk checkin:\n> 1. Issue an fsync against a dummy file to flush the hardware writeback\n>    cache, which should by now have seen the tmp-objdir writes.\n> 2. Rename all of the tmp-objdir files to their final names.\n> 3. When updating the index and/or refs, we assume that Git will issue\n>    another fsync internal to that operation. This is not the default\n>    today, but the user now has the option of syncing the index and there\n>    is a separate patch series to implement syncing of refs.\n> \n> On a filesystem with a singular journal that is updated during name\n> operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\n> we would expect the fsync to trigger a journal writeout so that this\n> sequence is enough to ensure that the user's data is durable by the time\n> the git command returns.\n> \n> Batch mode is only enabled if core.fsyncObjectFiles is false or unset.\n> \n> _Performance numbers_:\n> \n> Linux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\n> Mac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\n> Windows - Same host as Linux, a preview version of Windows 11.\n> \n> Adding 500 files to the repo with 'git add' Times reported in seconds.\n> \n> object file syncing | Linux | Mac   | Windows\n> --------------------|-------|-------|--------\n>            disabled | 0.06  |  0.35 | 0.61\n>               fsync | 1.88  | 11.18 | 2.47\n>               batch | 0.15  |  0.41 | 1.53\n> \n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  Documentation/config/core.txt |  5 +++\n>  bulk-checkin.c                | 67 +++++++++++++++++++++++++++++++++++\n>  bulk-checkin.h                |  2 ++\n>  cache.h                       |  8 ++++-\n>  config.c                      |  2 ++\n>  object-file.c                 |  2 ++\n>  6 files changed, 85 insertions(+), 1 deletion(-)\n> \n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 062e5259905..c041ed33801 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -628,6 +628,11 @@ core.fsyncMethod::\n>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>    filesystem and storage hardware, data added to the repository may not be\n>    durable in the event of a system crash. This is the default mode on macOS.\n> +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> +  updates in the disk writeback cache and then a single full fsync to trigger\n> +  the disk cache flush at the end of the operation. This mode is expected to\n> +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> +  and on Windows for repos stored on NTFS or ReFS filesystems.\n\nThis mode will not be supported by all parts of our stack that use our\nnew fsync infra. So I think we should both document that some parts of\nthe stack don't support batching, and say what the fallback behaviour is\nfor those that don't.\n\n>  core.fsyncObjectFiles::\n>  \tThis boolean will enable 'fsync()' when writing object files.\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index 93b1dc5138a..5c13fe17802 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -3,14 +3,20 @@\n>   */\n>  #include \"cache.h\"\n>  #include \"bulk-checkin.h\"\n> +#include \"lockfile.h\"\n>  #include \"repository.h\"\n>  #include \"csum-file.h\"\n>  #include \"pack.h\"\n>  #include \"strbuf.h\"\n> +#include \"string-list.h\"\n> +#include \"tmp-objdir.h\"\n>  #include \"packfile.h\"\n>  #include \"object-store.h\"\n>  \n>  static int bulk_checkin_plugged;\n> +static int needs_batch_fsync;\n> +\n> +static struct tmp_objdir *bulk_fsync_objdir;\n>  \n>  static struct bulk_checkin_state {\n>  \tchar *pack_tmp_name;\n> @@ -80,6 +86,34 @@ clear_exit:\n>  \treprepare_packed_git(the_repository);\n>  }\n>  \n> +/*\n> + * Cleanup after batch-mode fsync_object_files.\n> + */\n> +static void do_batch_fsync(void)\n> +{\n> +\t/*\n> +\t * Issue a full hardware flush against a temporary file to ensure\n> +\t * that all objects are durable before any renames occur.  The code in\n> +\t * fsync_loose_object_bulk_checkin has already issued a writeout\n> +\t * request, but it has not flushed any writeback cache in the storage\n> +\t * hardware.\n> +\t */\n> +\n> +\tif (needs_batch_fsync) {\n> +\t\tstruct strbuf temp_path = STRBUF_INIT;\n> +\t\tstruct tempfile *temp;\n> +\n> +\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> +\t\ttemp = xmks_tempfile(temp_path.buf);\n> +\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> +\t\tdelete_tempfile(&temp);\n> +\t\tstrbuf_release(&temp_path);\n> +\t}\n> +\n> +\tif (bulk_fsync_objdir)\n> +\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n> +}\n> +\n\nWe never unset `bulk_fsync_objdir` anywhere. Shouldn't we be doing that\nwhen we unplug this infrastructure?\n\nPatrick\n\n>  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>  {\n>  \tint i;\n> @@ -274,6 +308,24 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \treturn 0;\n>  }\n>  \n> +void fsync_loose_object_bulk_checkin(int fd)\n> +{\n> +\t/*\n> +\t * If we have a plugged bulk checkin, we issue a call that\n> +\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t * command. Later on we will issue a single hardware flush\n> +\t * before as part of do_batch_fsync.\n> +\t */\n> +\tif (bulk_checkin_plugged &&\n> +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> +\t\tassert(bulk_fsync_objdir);\n> +\t\tif (!needs_batch_fsync)\n> +\t\t\tneeds_batch_fsync = 1;\n> +\t} else {\n> +\t\tfsync_or_die(fd, \"loose object file\");\n> +\t}\n> +}\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags)\n> @@ -288,6 +340,19 @@ int index_bulk_checkin(struct object_id *oid,\n>  void plug_bulk_checkin(void)\n>  {\n>  \tassert(!bulk_checkin_plugged);\n> +\n> +\t/*\n> +\t * A temporary object directory is used to hold the files\n> +\t * while they are not fsynced.\n> +\t */\n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> +\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> +\t\tif (!bulk_fsync_objdir)\n> +\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n> +\n> +\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> +\t}\n> +\n>  \tbulk_checkin_plugged = 1;\n>  }\n>  \n> @@ -297,4 +362,6 @@ void unplug_bulk_checkin(void)\n>  \tbulk_checkin_plugged = 0;\n>  \tif (bulk_checkin_state.f)\n>  \t\tfinish_bulk_checkin(&bulk_checkin_state);\n> +\n> +\tdo_batch_fsync();\n>  }\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index b26f3dc3b74..08f292379b6 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -6,6 +6,8 @@\n>  \n>  #include \"cache.h\"\n>  \n> +void fsync_loose_object_bulk_checkin(int fd);\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags);\n> diff --git a/cache.h b/cache.h\n> index d347d0757f7..4d07691e791 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1040,7 +1040,8 @@ extern int use_fsync;\n>  \n>  enum fsync_method {\n>  \tFSYNC_METHOD_FSYNC,\n> -\tFSYNC_METHOD_WRITEOUT_ONLY\n> +\tFSYNC_METHOD_WRITEOUT_ONLY,\n> +\tFSYNC_METHOD_BATCH\n>  };\n>  \n>  extern enum fsync_method fsync_method;\n> @@ -1766,6 +1767,11 @@ void fsync_or_die(int fd, const char *);\n>  int fsync_component(enum fsync_component component, int fd);\n>  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n>  \n> +static inline int batch_fsync_enabled(enum fsync_component component)\n> +{\n> +\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n> +}\n> +\n>  ssize_t read_in_full(int fd, void *buf, size_t count);\n>  ssize_t write_in_full(int fd, const void *buf, size_t count);\n>  ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\n> diff --git a/config.c b/config.c\n> index 261ee7436e0..0b28f90de8b 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n>  \t\telse if (!strcmp(value, \"writeout-only\"))\n>  \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n> +\t\telse if (!strcmp(value, \"batch\"))\n> +\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n>  \t\telse\n>  \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n>  \n> diff --git a/object-file.c b/object-file.c\n> index 295cb899e22..ef6621ffe56 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1894,6 +1894,8 @@ static void close_loose_object(int fd)\n>  \n>  \tif (fsync_object_files > 0)\n>  \t\tfsync_or_die(fd, \"loose object file\");\n> +\telse if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> +\t\tfsync_loose_object_bulk_checkin(fd);\n>  \telse\n>  \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n>  \t\t\t\t       \"loose object file\");\n> -- \n> gitgitgadget\n> \n"},{"id":"451457","messageId":"CANQDOdcak1nV1Pr9cmyk9dgEjHOH8Au92pUMskJipUodzskzqQ@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqbky6bgw0.fsf@gitster.g","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-16T07:33:15Z","receivedAt":"2022-03-16T07:33:32Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 15, 2022 at 10:33 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Preparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n> >\n> > * Rename 'state' variable to 'bulk_checkin_state', since we will later\n> >   be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n> >   find in the debugger, since the name is more unique.\n> >\n> > * Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n> >   static variable. Doing this avoids resetting the variable in\n> >   finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n> >   seem to unintentionally disable the plugging functionality the first\n> >   time a new packfile must be created due to packfile size limits. While\n> >   disabling the plugging state only results in suboptimal behavior for\n> >   the current code, it would be fatal for the bulk-fsync functionality\n> >   later in this patch series.\n>\n> Sorry, but I am confused.  The bulk-checkin infrastructure is there\n> so that we can send many little objects into a single packfile\n> instead of creating many little loose object files.  Everything we\n> throw at object-file.c::index_stream() will be concatenated into the\n> single packfile while we are \"plugged\" until we get \"unplugged\".\n>\n\nI noticed that you invented bulk-checkin back in 2011, but I don't think your\ndescription matches what the code actually does.  index_bulk_checkin\nis only called from index_stream, which is only called from index_fd. index_fd\ngoes down the index_bulk_checkin path for large files (512MB by default). It\nlooks like the effect of the 'plug/unplug' code is to allow multiple\nlarge blobs to\ngo into a single packfile rather than each getting one getting its own separate\npackfile.\n\n> My understanding of what you are doing in this series is to still\n> create many little loose object files, but avoid the overhead of\n> having to fsync them individually.  And I am not sure how well the\n> original idea behind the bulk-checkin infrastructure to avoid\n> overhead of having to create many loose objects by creating a single\n> packfile (and presumably having to fsync at the end, but that is\n> just a single .pack file) with your goal of still creating many\n> loose object files but synching them more efficiently.\n>\n> Is it just the new feature is piggybacking on the existing bulk\n> checkin infrastructure, even though these two have nothing in\n> common?\n>\n\nI think my new usage is congruent with the existing API, which seems\nto be about combining multiple add operations into a large transaction,\nwhere we can do some cleanup operations once we're finished. In the\npreexisting code, the transaction is about adding a bunch of large objects\nto a single pack file (while leaving small objects loose), and then completing\nthe packfile when the adds are finished.\n\n---\nOn a side note, I've also been thinking about how we could use a packfile\napproach as an alternative means to achieve faster addition of many small\nobjects. It's essentially what you stated above, where we'd send our\nlittle objects\ninto a pack file. But to avoid frequent repacking overhead, we might\nwant to reuse\nthe 'latest' packfile across multiple Git invocations by appending\nobjects to it, with\nan fsync on the file at the end.\n\nWe'd need sufficient padding between objects created by different Git\ninvocations to\nensure that previously synced data doesn't get disturbed by later\noperations.  We'd\nneed to rewrite the pack indexes each time, but that's at least\nderived metadata, so it\ndoesn't need to be fsynced. To make the pack indexes more\nincrementally-updatable,\nwe might want to have the fanout table be checksummed, with\nchecksummed pointers to\nleaf blocks. If we detect corruption during an index lookup, we could\nrecreate the index\nfrom the packfile.\n\nEssentially the above proposal is to move away from storing loose\nobjects in the filesystem\nand instead to index the data within Git itself.\n\nThanks,\nNeeraj\n"},{"id":"451468","messageId":"65998787-15e4-fac4-1343-65df60e971d0@gmail.com","threadId":"57568","inReplyTo":"d38f20b4430bada1d0dccc1e600e6f0b098f3767.1647379859.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2022-03-16T11:50:06Z","receivedAt":"2022-03-16T11:50:17Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 16/03/22 04.30, Neeraj Singh via GitGitGadget wrote:\n> On a filesystem with a singular journal that is updated during name\n> operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\n> we would expect the fsync to trigger a journal writeout so that this\n> sequence is enough to ensure that the user's data is durable by the time\n> the git command returns.\n> \n\nBut what about ext4? Will fsync-ing trigger writing journal?\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"451480","messageId":"xmqq35jhoows.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdcak1nV1Pr9cmyk9dgEjHOH8Au92pUMskJipUodzskzqQ@mail.gmail.com","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-16T16:14:27Z","receivedAt":"2022-03-16T16:14:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> I think my new usage is congruent with the existing API, which seems\n> to be about combining multiple add operations into a large transaction,\n> where we can do some cleanup operations once we're finished. In the\n> preexisting code, the transaction is about adding a bunch of large objects\n> to a single pack file (while leaving small objects loose), and then completing\n> the packfile when the adds are finished.\n\nOK, so it was part me, and part a suboptimal presentation, I guess\n;-)\n\nLet me rephrase the idea to see if I got it right this time.\n\nThe bulk-checkin API has two interesting entry points, \"plug\" that\nsignals that we are about to repeat possibly many operations to add\nnew objects to the object store, and \"unplug\" that signals that we\nare done such adding.  They are meant to serve as a hint for the\nobject layer to optimize its operation.\n\nSo far the only way the hint was used was that the logic that sends\nan overly large object into a packfile (instead of storing it loose,\nwhich leaves it subject to expensive repacking later) can shove more\nthan one such objects in the same packfile.\n\nThis series invents another use of the \"plug\"-\"unplug\" hint.  By\nknowing that many loose object files are created and when the series\nof object creation ended, we can avoid having to fsync each and\nevery one of them on certain filesystems and achieve the same\nrobustness.  The new \"batch\" option to core.fsyncmethod triggers\nthis mechanism.\n\nDid I get it right, more-or-less?\n\nThanks.\n"},{"id":"451492","messageId":"CANQDOde+3QaTWnNWPQzz85iAGH=M-ZhG-HDR9SP4FCJf+=k43A@mail.gmail.com","threadId":"57568","inReplyTo":"xmqq35jhoows.fsf@gitster.g","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-16T17:59:45Z","receivedAt":"2022-03-16T18:00:00Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 16, 2022 at 9:14 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > I think my new usage is congruent with the existing API, which seems\n> > to be about combining multiple add operations into a large transaction,\n> > where we can do some cleanup operations once we're finished. In the\n> > preexisting code, the transaction is about adding a bunch of large objects\n> > to a single pack file (while leaving small objects loose), and then completing\n> > the packfile when the adds are finished.\n>\n> OK, so it was part me, and part a suboptimal presentation, I guess\n> ;-)\n>\n> Let me rephrase the idea to see if I got it right this time.\n>\n> The bulk-checkin API has two interesting entry points, \"plug\" that\n> signals that we are about to repeat possibly many operations to add\n> new objects to the object store, and \"unplug\" that signals that we\n> are done such adding.  They are meant to serve as a hint for the\n> object layer to optimize its operation.\n>\n> So far the only way the hint was used was that the logic that sends\n> an overly large object into a packfile (instead of storing it loose,\n> which leaves it subject to expensive repacking later) can shove more\n> than one such objects in the same packfile.\n>\n> This series invents another use of the \"plug\"-\"unplug\" hint.  By\n> knowing that many loose object files are created and when the series\n> of object creation ended, we can avoid having to fsync each and\n> every one of them on certain filesystems and achieve the same\n> robustness.  The new \"batch\" option to core.fsyncmethod triggers\n> this mechanism.\n>\n> Did I get it right, more-or-less?\n\nYes, that's my understanding as well.\n\nThanks,\nNeeraj\n"},{"id":"451495","messageId":"xmqqpmmllqf5.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOde+3QaTWnNWPQzz85iAGH=M-ZhG-HDR9SP4FCJf+=k43A@mail.gmail.com","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-16T18:10:06Z","receivedAt":"2022-03-16T18:10:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> Did I get it right, more-or-less?\n>\n> Yes, that's my understanding as well.\n\nI guess what I wrote would make a useful material for early part of\nthe log message to help future developers.\n\nThanks.\n"},{"id":"451496","messageId":"CANQDOddkJ_iyUjYQAHs93nqCda1W6ss5JQzbz3uuh-XnoATg5g@mail.gmail.com","threadId":"57568","inReplyTo":"YjGSOJlZnDSske3s@ncase","subject":"Re: [PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-16T18:21:56Z","receivedAt":"2022-03-16T18:22:12Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 16, 2022 at 12:31 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Tue, Mar 15, 2022 at 09:30:54PM +0000, Neeraj Singh via GitGitGadget wrote:\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index 062e5259905..c041ed33801 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -628,6 +628,11 @@ core.fsyncMethod::\n> >  * `writeout-only` issues pagecache writeback requests, but depending on the\n> >    filesystem and storage hardware, data added to the repository may not be\n> >    durable in the event of a system crash. This is the default mode on macOS.\n> > +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> > +  updates in the disk writeback cache and then a single full fsync to trigger\n> > +  the disk cache flush at the end of the operation. This mode is expected to\n> > +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> > +  and on Windows for repos stored on NTFS or ReFS filesystems.\n>\n> This mode will not be supported by all parts of our stack that use our\n> new fsync infra. So I think we should both document that some parts of\n> the stack don't support batching, and say what the fallback behaviour is\n> for those that don't.\n>\n\nCan do. I'm hoping that you'll revive your batch-mode refs change too so that\nwe get batching across the ODB and Refs, which are the two data stores that\nmay receive many updates in a single Git command.  This documentation\ncomment will read:\n```\n* `batch` enables a mode that uses writeout-only flushes to stage multiple\n  updates in the disk writeback cache and then does a single full fsync of\n  a dummy file to trigger the disk cache flush at the end of the operation.\n  Currently `batch` mode only applies to loose-object files. Other repository\n  data is made durable as if `fsync` was specified. This mode is expected to\n  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n  and on Windows for repos stored on NTFS or ReFS filesystems.\n```\n\n\n> >  core.fsyncObjectFiles::\n> >       This boolean will enable 'fsync()' when writing object files.\n> > diff --git a/bulk-checkin.c b/bulk-checkin.c\n> > index 93b1dc5138a..5c13fe17802 100644\n> > --- a/bulk-checkin.c\n> > +++ b/bulk-checkin.c\n> > @@ -3,14 +3,20 @@\n> >   */\n> >  #include \"cache.h\"\n> >  #include \"bulk-checkin.h\"\n> > +#include \"lockfile.h\"\n> >  #include \"repository.h\"\n> >  #include \"csum-file.h\"\n> >  #include \"pack.h\"\n> >  #include \"strbuf.h\"\n> > +#include \"string-list.h\"\n> > +#include \"tmp-objdir.h\"\n> >  #include \"packfile.h\"\n> >  #include \"object-store.h\"\n> >\n> >  static int bulk_checkin_plugged;\n> > +static int needs_batch_fsync;\n> > +\n> > +static struct tmp_objdir *bulk_fsync_objdir;\n> >\n> >  static struct bulk_checkin_state {\n> >       char *pack_tmp_name;\n> > @@ -80,6 +86,34 @@ clear_exit:\n> >       reprepare_packed_git(the_repository);\n> >  }\n> >\n> > +/*\n> > + * Cleanup after batch-mode fsync_object_files.\n> > + */\n> > +static void do_batch_fsync(void)\n> > +{\n> > +     /*\n> > +      * Issue a full hardware flush against a temporary file to ensure\n> > +      * that all objects are durable before any renames occur.  The code in\n> > +      * fsync_loose_object_bulk_checkin has already issued a writeout\n> > +      * request, but it has not flushed any writeback cache in the storage\n> > +      * hardware.\n> > +      */\n> > +\n> > +     if (needs_batch_fsync) {\n> > +             struct strbuf temp_path = STRBUF_INIT;\n> > +             struct tempfile *temp;\n> > +\n> > +             strbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> > +             temp = xmks_tempfile(temp_path.buf);\n> > +             fsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> > +             delete_tempfile(&temp);\n> > +             strbuf_release(&temp_path);\n> > +     }\n> > +\n> > +     if (bulk_fsync_objdir)\n> > +             tmp_objdir_migrate(bulk_fsync_objdir);\n> > +}\n> > +\n>\n> We never unset `bulk_fsync_objdir` anywhere. Shouldn't we be doing that\n> when we unplug this infrastructure?\n>\n\nWill Fix.\n\nThanks,\nNeeraj\n"},{"id":"451506","messageId":"CANQDOdc=df=iBmyEs4chcPEOA1-1S-isWDSeWTvyi09B4ewAfA@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqpmmllqf5.fsf@gitster.g","subject":"Re: [PATCH 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-16T19:50:43Z","receivedAt":"2022-03-16T19:50:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 16, 2022 at 11:10 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> >> Did I get it right, more-or-less?\n> >\n> > Yes, that's my understanding as well.\n>\n> I guess what I wrote would make a useful material for early part of\n> the log message to help future developers.\n>\n> Thanks.\n\nWill do.  I changed the commit message to explain the current\nfunctionality of bulk-checkin and how it's similar to batched-fsync.\n"},{"id":"451508","messageId":"CANQDOdcEOb5zJDJ2GdzTPm_ULvhriX5d5p5go=PeQtHvB6mRPQ@mail.gmail.com","threadId":"57568","inReplyTo":"65998787-15e4-fac4-1343-65df60e971d0@gmail.com","subject":"Re: [PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-16T19:59:11Z","receivedAt":"2022-03-16T19:59:27Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 16, 2022 at 4:50 AM Bagas Sanjaya <bagasdotme@gmail.com> wrote:\n>\n> On 16/03/22 04.30, Neeraj Singh via GitGitGadget wrote:\n> > On a filesystem with a singular journal that is updated during name\n> > operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\n> > we would expect the fsync to trigger a journal writeout so that this\n> > sequence is enough to ensure that the user's data is durable by the time\n> > the git command returns.\n> >\n>\n> But what about ext4? Will fsync-ing trigger writing journal?\n>\n\nThat's a good question. So I did an experiment on ext4 which gives me\nsome confidence:\n\nHere's my ext4 configuration: /dev/sdc on / type ext4\n(rw,relatime,discard,errors=remount-ro,data=ordered)\n\nI added a new mode called core.fsyncMethod=batch-extra-fsync. This\nissues an extra open,fsync,close during migration from the tmp-objdir\n(which I confirmed is really happening using strace).  The added cost\nof this extra operation is relatively small compared to\ncore.fsyncMethod=fsync.  That leads me to believe that (barring fs\nbugs), ext4 thinks that the data is already sufficiently durable that\nit doesn't need to issue an extra disk cache flush.  See\nhttps://github.com/neerajsi-msft/git/commit/131466dd95165efc5c480d971c69ea1e9182657e\nfor the test code.  I don't particularly want to add this as a\nbuilt-in mode at this point since it will be somewhat hard to document\nwhich mode a user should choose.\n\nThanks,\nNeeraj\n"},{"id":"451536","messageId":"YjLLmRvvfkAOR4/6@ncase","threadId":"57568","inReplyTo":"CANQDOddkJ_iyUjYQAHs93nqCda1W6ss5JQzbz3uuh-XnoATg5g@mail.gmail.com","subject":"Re: [PATCH 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2022-03-17T05:48:09Z","receivedAt":"2022-03-17T06:09:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 16, 2022 at 11:21:56AM -0700, Neeraj Singh wrote:\n> On Wed, Mar 16, 2022 at 12:31 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > On Tue, Mar 15, 2022 at 09:30:54PM +0000, Neeraj Singh via GitGitGadget wrote:\n> > > From: Neeraj Singh <neerajsi@microsoft.com>\n> > > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > > index 062e5259905..c041ed33801 100644\n> > > --- a/Documentation/config/core.txt\n> > > +++ b/Documentation/config/core.txt\n> > > @@ -628,6 +628,11 @@ core.fsyncMethod::\n> > >  * `writeout-only` issues pagecache writeback requests, but depending on the\n> > >    filesystem and storage hardware, data added to the repository may not be\n> > >    durable in the event of a system crash. This is the default mode on macOS.\n> > > +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> > > +  updates in the disk writeback cache and then a single full fsync to trigger\n> > > +  the disk cache flush at the end of the operation. This mode is expected to\n> > > +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> > > +  and on Windows for repos stored on NTFS or ReFS filesystems.\n> >\n> > This mode will not be supported by all parts of our stack that use our\n> > new fsync infra. So I think we should both document that some parts of\n> > the stack don't support batching, and say what the fallback behaviour is\n> > for those that don't.\n> >\n> \n> Can do. I'm hoping that you'll revive your batch-mode refs change too so that\n> we get batching across the ODB and Refs, which are the two data stores that\n> may receive many updates in a single Git command.\n\nHuh, I completely forgot that my previous implementation already had\nsuch a mechanism. I may have a go at it again, but it would take me a\nwhile given that I'll be OOO most of April.\n\n> This documentation\n> comment will read:\n> ```\n> * `batch` enables a mode that uses writeout-only flushes to stage multiple\n>   updates in the disk writeback cache and then does a single full fsync of\n>   a dummy file to trigger the disk cache flush at the end of the operation.\n>   Currently `batch` mode only applies to loose-object files. Other repository\n>   data is made durable as if `fsync` was specified. This mode is expected to\n>   be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n>   and on Windows for repos stored on NTFS or ReFS filesystems.\n> ```\n\nReads good to me, thanks!\n\nPatrick\n"},{"id":"451674","messageId":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.git.1647379859.gitgitgadget@gmail.com","subject":"[PATCH v2 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:53Z","receivedAt":"2022-03-20T07:16:08Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"V2 changes:\n\n * Change doc to indicate that only some repo updates are batched\n * Null and zero out control variables in do_batch_fsync under\n   unplug_bulk_checkin\n * Make batch mode default on Windows.\n * Update the description for the initial patch that cleans up the\n   bulk-checkin infrastructure.\n * Rebase onto 'seen' at 0cac37f38f9.\n\n--Original definition-- When core.fsync includes loose-object, we issue an\nfsync after every written object. For a 'git-add' or similar command that\nadds a lot of files to the repo, the costs of these fsyncs adds up. One\nmajor factor in this cost is the time it takes for the physical storage\ncontroller to flush its caches to durable media.\n\nThis series takes advantage of the writeout-only mode of git_fsync to issue\nOS cache writebacks for all of the objects being added to the repository\nfollowed by a single fsync to a dummy file, which should trigger a\nfilesystem log flush and storage controller cache flush. This mechanism is\nknown to be safe on common Windows filesystems and expected to be safe on\nmacOS. Some linux filesystems, such as XFS, will probably do the right thing\nas well. See [1] for previous discussion on the predecessor of this patch\nseries.\n\nThis series is important on Windows, where loose-objects are included in the\nfsync set by default in Git-For-Windows. In this series, I'm also setting\nthe default mode for Windows to turn on loose object fsyncing with batch\nmode, so that we can get CI coverage of the actual git-for-windows\nconfiguration upstream. We still don't actually issue fsyncs for the test\nsuite since GIT_TEST_FSYNC is set to 0, but we exercise all of the\nsurrounding batch mode code.\n\nThis work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n\n[1]\nhttps://lore.kernel.org/git/2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com/\n\nNeeraj Singh (7):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  core.fsyncmethod: batched disk flushes for loose-objects\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  core.fsyncmethod: tests for batch mode\n  core.fsyncmethod: performance tests for add and stash\n\n Documentation/config/core.txt |  7 +++\n builtin/unpack-objects.c      |  3 ++\n builtin/update-index.c        |  6 +++\n bulk-checkin.c                | 92 +++++++++++++++++++++++++++++++----\n bulk-checkin.h                |  2 +\n cache.h                       | 12 ++++-\n compat/mingw.h                |  3 ++\n config.c                      |  4 +-\n git-compat-util.h             |  2 +\n object-file.c                 |  2 +\n t/lib-unique-files.sh         | 36 ++++++++++++++\n t/perf/p3700-add.sh           | 59 ++++++++++++++++++++++\n t/perf/p3900-stash.sh         | 62 +++++++++++++++++++++++\n t/perf/perf-lib.sh            |  4 +-\n t/t3700-add.sh                | 22 +++++++++\n t/t3903-stash.sh              | 17 +++++++\n t/t5300-pack-object.sh        | 32 +++++++-----\n 17 files changed, 340 insertions(+), 25 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: 0cac37f38f94bb93550eb164b5d574cd96e23785\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1134%2Fneerajsi-msft%2Fns%2Fbatched-fsync-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1134/neerajsi-msft/ns/batched-fsync-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1134\n\nRange-diff vs v1:\n\n 1:  a77d02df626 ! 1:  9c2abd12bbb bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n     @@ Metadata\n       ## Commit message ##\n          bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n      \n     -    Preparation for adding bulk-fsync to the bulk-checkin.c infrastructure.\n     +    This commit prepares for adding batch-fsync to the bulk-checkin\n     +    infrastructure.\n     +\n     +    The bulk-checkin infrastructure is currently used to batch up addition\n     +    of large blobs to a packfile. When a blob is larger than\n     +    big_file_threshold, we unconditionally add it to a pack. If bulk\n     +    checkins are 'plugged', we allow multiple large blobs to be added to a\n     +    single pack until we reach the packfile size limit; otherwise, we simply\n     +    make a new packfile for each large blob. The 'unplug' call tells us when\n     +    the series of blob additions is done so that we can finish the packfiles\n     +    and make their objects available to subsequent operations.\n     +\n     +    Stated another way, bulk-checkin allows callers to define a transaction\n     +    that adds multiple objects to the object database, where the object\n     +    database can optimize its internal operations within the transaction\n     +    boundary.\n     +\n     +    Batched fsync will fit into bulk-checkin by taking advantage of the\n     +    plug/unplug functionality to determine the appropriate time to fsync\n     +    and make newly-added objects available in the primary object database.\n      \n          * Rename 'state' variable to 'bulk_checkin_state', since we will later\n            be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n 2:  d38f20b4430 ! 2:  3ed1dcd9b9b core.fsyncmethod: batched disk flushes for loose-objects\n     @@ Documentation/config/core.txt: core.fsyncMethod::\n         filesystem and storage hardware, data added to the repository may not be\n         durable in the event of a system crash. This is the default mode on macOS.\n      +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n     -+  updates in the disk writeback cache and then a single full fsync to trigger\n     -+  the disk cache flush at the end of the operation. This mode is expected to\n     ++  updates in the disk writeback cache and then does a single full fsync of\n     ++  a dummy file to trigger the disk cache flush at the end of the operation.\n     ++  Currently `batch` mode only applies to loose-object files. Other repository\n     ++  data is made durable as if `fsync` was specified. This mode is expected to\n      +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n      +  and on Windows for repos stored on NTFS or ReFS filesystems.\n       \n     @@ bulk-checkin.c: clear_exit:\n      +\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n      +\t\tdelete_tempfile(&temp);\n      +\t\tstrbuf_release(&temp_path);\n     ++\t\tneeds_batch_fsync = 0;\n      +\t}\n      +\n     -+\tif (bulk_fsync_objdir)\n     ++\tif (bulk_fsync_objdir) {\n      +\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n     ++\t\tbulk_fsync_objdir = NULL;\n     ++\t}\n      +}\n      +\n       static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n 3:  b0480f0c814 = 3:  54797dbc520 update-index: use the bulk-checkin infrastructure\n 4:  99e3a61b919 = 4:  6662e2dae0f unpack-objects: use the bulk-checkin infrastructure\n 5:  4e56c58c8cb ! 5:  03bf591742a core.fsync: use batch mode and sync loose objects by default on Windows\n     @@ Commit message\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     -    change fsyncmethod to batch as well\n     -\n       ## cache.h ##\n      @@ cache.h: enum fsync_component {\n       \t\t\t      FSYNC_COMPONENT_INDEX | \\\n 6:  88e47047d79 = 6:  1937746df47 core.fsyncmethod: tests for batch mode\n 7:  876741f1ef9 = 7:  624244078c7 core.fsyncmethod: performance tests for add and stash\n\n-- \ngitgitgadget\n"},{"id":"451675","messageId":"9c2abd12bbbd27261378ffb6478a3e5db8a5063c.1647760560.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 1/7] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:54Z","receivedAt":"2022-03-20T07:16:11Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit prepares for adding batch-fsync to the bulk-checkin\ninfrastructure.\n\nThe bulk-checkin infrastructure is currently used to batch up addition\nof large blobs to a packfile. When a blob is larger than\nbig_file_threshold, we unconditionally add it to a pack. If bulk\ncheckins are 'plugged', we allow multiple large blobs to be added to a\nsingle pack until we reach the packfile size limit; otherwise, we simply\nmake a new packfile for each large blob. The 'unplug' call tells us when\nthe series of blob additions is done so that we can finish the packfiles\nand make their objects available to subsequent operations.\n\nStated another way, bulk-checkin allows callers to define a transaction\nthat adds multiple objects to the object database, where the object\ndatabase can optimize its internal operations within the transaction\nboundary.\n\nBatched fsync will fit into bulk-checkin by taking advantage of the\nplug/unplug functionality to determine the appropriate time to fsync\nand make newly-added objects available in the primary object database.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex e988a388b65..93b1dc5138a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -278,21 +278,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"451676","messageId":"3ed1dcd9b9ba9b34f26b3012eaba8da0269ee842.1647760560.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:55Z","receivedAt":"2022-03-20T07:16:12Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin plugging and unplugging functionality,\ntakes advantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1. Create the object in a tmp-objdir.\n2. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1. Issue an fsync against a dummy file to flush the hardware writeback\n   cache, which should by now have seen the tmp-objdir writes.\n2. Rename all of the tmp-objdir files to their final names.\n3. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns.\n\nBatch mode is only enabled if core.fsyncObjectFiles is false or unset.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  7 ++++\n bulk-checkin.c                | 70 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  2 +\n cache.h                       |  8 +++-\n config.c                      |  2 +\n object-file.c                 |  2 +\n 6 files changed, 90 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 889522956e4..a3798dfc334 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -628,6 +628,13 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then does a single full fsync of\n+  a dummy file to trigger the disk cache flush at the end of the operation.\n+  Currently `batch` mode only applies to loose-object files. Other repository\n+  data is made durable as if `fsync` was specified. This mode is expected to\n+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+  and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 93b1dc5138a..a702e0ff203 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,14 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n+static int needs_batch_fsync;\n+\n+static struct tmp_objdir *bulk_fsync_objdir;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -80,6 +86,37 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur.  The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware.\n+\t */\n+\n+\tif (needs_batch_fsync) {\n+\t\tstruct strbuf temp_path = STRBUF_INIT;\n+\t\tstruct tempfile *temp;\n+\n+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\t\ttemp = xmks_tempfile(temp_path.buf);\n+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\t\tdelete_tempfile(&temp);\n+\t\tstrbuf_release(&temp_path);\n+\t\tneeds_batch_fsync = 0;\n+\t}\n+\n+\tif (bulk_fsync_objdir) {\n+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n+\t\tbulk_fsync_objdir = NULL;\n+\t}\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +311,24 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void fsync_loose_object_bulk_checkin(int fd)\n+{\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (bulk_checkin_plugged &&\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n+\t\tassert(bulk_fsync_objdir);\n+\t\tif (!needs_batch_fsync)\n+\t\t\tneeds_batch_fsync = 1;\n+\t} else {\n+\t\tfsync_or_die(fd, \"loose object file\");\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -288,6 +343,19 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n+\n+\t/*\n+\t * A temporary object directory is used to hold the files\n+\t * while they are not fsynced.\n+\t */\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\t\tif (!bulk_fsync_objdir)\n+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n+\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+\t}\n+\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -297,4 +365,6 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..08f292379b6 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,8 @@\n \n #include \"cache.h\"\n \n+void fsync_loose_object_bulk_checkin(int fd);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex 3160bc1e489..d1ae51388c9 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH\n };\n \n extern enum fsync_method fsync_method;\n@@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex 261ee7436e0..0b28f90de8b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 5258d9ed827..bdb0a38328f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1895,6 +1895,8 @@ static void close_loose_object(int fd)\n \n \tif (fsync_object_files > 0)\n \t\tfsync_or_die(fd, \"loose object file\");\n+\telse if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n \t\t\t\t       \"loose object file\");\n-- \ngitgitgadget\n\n"},{"id":"451677","messageId":"54797dbc52060b7fa913642cd5266f7e159a5bc9.1647760561.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:56Z","receivedAt":"2022-03-20T07:16:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables bulk-checkin for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete.  This\nusage is unlikely, since any tool invoking update-index and expecting to\nsee objects would have to synchronize with the update-index process\nafter passing it a file path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 75d646377cc..38e9d7e88cb 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \tthe_index.updated_skipworktree = 1;\n \n+\t/* we might be adding many objects to the object database */\n+\tplug_bulk_checkin();\n+\n \t/*\n \t * Custom copy of parse_options() because we want to handle\n \t * filename arguments as they come.\n@@ -1190,6 +1194,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/* by now we must have added all of the new objects */\n+\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"451678","messageId":"6662e2dae0f5d65c158fba785d186885f9671073.1647760561.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 4/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:57Z","receivedAt":"2022-03-20T07:16:18Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling bulk-checkin when unpacking objects, we can take advantage\nof batched fsyncs.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..c55b6616aed 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tplug_bulk_checkin();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"451679","messageId":"624244078c7adc2186941fbfa08cb3afecdece4c.1647760561.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 7/7] core.fsyncmethod: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:16:00Z","receivedAt":"2022-03-20T07:16:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings. This shows the benefit of batch\nmode relative to an ordinary stash command.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 59 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 62 +++++++++++++++++++++++++++++++++++++++++++\n t/perf/perf-lib.sh    |  4 +--\n 3 files changed, 123 insertions(+), 2 deletions(-)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..2ea78c9449d\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,59 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\ttest_perf \"add $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..3526f06cef4\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,62 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG stash push -u -- files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex 932105cd12c..d270d1d962a 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -98,8 +98,8 @@ test_perf_create_repo_from () {\n \tmkdir -p \"$repo/.git\"\n \t(\n \t\tcd \"$source\" &&\n-\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" 2>/dev/null ||\n-\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\"; } &&\n+\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" ||\n+\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\" 2>/dev/null;} &&\n \n \t\t# common_dir must come first here, since we want source_git to\n \t\t# take precedence and overwrite any overlapping files\n-- \ngitgitgadget\n"},{"id":"451680","messageId":"03bf591742a48d750d6b8e6c54b5a8fd954561a5.1647760561.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 5/7] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:58Z","receivedAt":"2022-03-20T07:16:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the test suite, since GIT_TEST_FSYNC is\nset to 0. However, we do exercise all of the surrounding batch mode code\nsince GIT_TEST_FSYNC merely makes the maybe_fsync wrapper always appear\nto succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex d1ae51388c9..4d2131e8f4f 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 0b28f90de8b..c76443dc556 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 0892e209a2f..fffe42ce7c1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \ngitgitgadget\n\n"},{"id":"451681","messageId":"1937746df47eefecfc343e32eb9bf6c0949fb7b9.1647760561.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v2 6/7] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-20T07:15:59Z","receivedAt":"2022-03-20T07:16:26Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added due\nto it already being in the repo.\n\nWe aren't actually issuing any fsyncs in these tests, since\nGIT_TEST_FSYNC is 0, but we still exercise all of the tmp_objdir logic\nin bulk-checkin.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 36 ++++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 22 ++++++++++++++++++++++\n t/t3903-stash.sh       | 17 +++++++++++++++++\n t/t5300-pack-object.sh | 32 +++++++++++++++++++++-----------\n 4 files changed, 96 insertions(+), 11 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a7de4ca8512\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,36 @@\n+# Helper to create files with unique contents\n+\n+\n+# Create multiple files with unique contents. Takes the number of\n+# directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with unique\n+#\t\t\t\t\t contents.\n+\n+test_create_unique_files() {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=$1\n+\tlocal files=$2\n+\tlocal basedir=$3\n+\tlocal counter=0\n+\ttest_tick\n+\tlocal basedata=$test_tick\n+\n+\n+\trm -rf $basedir\n+\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i\n+\n+\t\tmkdir -p \"$dir\"\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1))\n+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..1f349f52ad3 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,26 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit $BATCH_CONFIGURATION add -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files2 &&\n+\tfind fsync-files2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\trm -f fsynced_files2 &&\n+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n+\ttest_line_count = 8 fsynced_files2 &&\n+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 55cd77901a8..5a3996b838f 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1462,6 +1463,22 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 fsync-files &&\n+\tgit $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n+\trm -f fsynced_files &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n+\ttest_line_count = 8 fsynced_files &&\n+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..8e2f73cc68f 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -162,23 +162,25 @@ test_expect_success 'pack-objects with bogus arguments' '\n \n check_unpack () {\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $2 init --bare git2 &&\n+\t(\n+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <obj-list >current &&\n+\tcmp obj-list current\n }\n \n test_expect_success 'unpack without delta' '\n \tcheck_unpack test-1-${packname_1}\n '\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'pack with REF_DELTA' '\n \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n \tcheck_deltas stderr -gt 0\n@@ -188,6 +190,10 @@ test_expect_success 'unpack with REF_DELTA' '\n \tcheck_unpack test-2-${packname_2}\n '\n \n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'pack with OFS_DELTA' '\n \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n \t\t\t<obj-list 2>stderr) &&\n@@ -198,6 +204,10 @@ test_expect_success 'unpack with OFS_DELTA' '\n \tcheck_unpack test-3-${packname_3}\n '\n \n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} \"$BATCH_CONFIGURATION\"\n+'\n+\n test_expect_success 'compare delta flavors' '\n \tperl -e '\\''\n \t\tdefined($_ = -s $_) or die for @ARGV;\n-- \ngitgitgadget\n\n"},{"id":"451728","messageId":"220321.86a6dj9xja.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"3ed1dcd9b9ba9b34f26b3012eaba8da0269ee842.1647760560.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T14:41:22Z","receivedAt":"2022-03-21T14:43:19Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n> [...]\n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> +\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> +\t\tif (!bulk_fsync_objdir)\n> +\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n\nShould camel-case the config var, and we should have a die_errno() here\nwhich tell us why we couldn't create it (possibly needing to ferry it up\nfrom the tempfile API...)\n"},{"id":"451729","messageId":"220321.865yo79wkf.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"54797dbc52060b7fa913642cd5266f7e159a5bc9.1647760561.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T15:01:00Z","receivedAt":"2022-03-21T15:04:06Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The update-index functionality is used internally by 'git stash push' to\n> setup the internal stashed commit.\n>\n> This change enables bulk-checkin for update-index infrastructure to\n> speed up adding new objects to the object database by leveraging the\n> batch fsync functionality.\n>\n> There is some risk with this change, since under batch fsync, the object\n> files will be in a tmp-objdir until update-index is complete.  This\n> usage is unlikely, since any tool invoking update-index and expecting to\n> see objects would have to synchronize with the update-index process\n> after passing it a file path.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/update-index.c | 6 ++++++\n>  1 file changed, 6 insertions(+)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index 75d646377cc..38e9d7e88cb 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -5,6 +5,7 @@\n>   */\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"cache.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"config.h\"\n>  #include \"lockfile.h\"\n>  #include \"quote.h\"\n> @@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \n>  \tthe_index.updated_skipworktree = 1;\n>  \n> +\t/* we might be adding many objects to the object database */\n> +\tplug_bulk_checkin();\n> +\n\nShouldn't this be after parse_options_start()?\n"},{"id":"451735","messageId":"220321.861qyv9rjr.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"3ed1dcd9b9ba9b34f26b3012eaba8da0269ee842.1647760560.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T15:47:11Z","receivedAt":"2022-03-21T16:52:53Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> One major source of the cost of fsync is the implied flush of the\n> hardware writeback cache within the disk drive. This commit introduces\n> a new `core.fsyncMethod=batch` option that batches up hardware flushes.\n> It hooks into the bulk-checkin plugging and unplugging functionality,\n> takes advantage of tmp-objdir, and uses the writeout-only support code.\n>\n> When the new mode is enabled, we do the following for each new object:\n> 1. Create the object in a tmp-objdir.\n> 2. Issue a pagecache writeback request and wait for it to complete.\n>\n> At the end of the entire transaction when unplugging bulk checkin:\n> 1. Issue an fsync against a dummy file to flush the hardware writeback\n>    cache, which should by now have seen the tmp-objdir writes.\n> 2. Rename all of the tmp-objdir files to their final names.\n> 3. When updating the index and/or refs, we assume that Git will issue\n>    another fsync internal to that operation. This is not the default\n>    today, but the user now has the option of syncing the index and there\n>    is a separate patch series to implement syncing of refs.\n\nRe my question in\nhttps://lore.kernel.org/git/220310.86r179ki38.gmgdl@evledraar.gmail.com/\n(which you *partially* replied to per my reading, i.e. not the\nfsync_nth() question) I still don't get why the tmp-objdir part of this\nis needed.\n\nFor \"git stash\" which is one thing sped up by this let's go over what\ncommands/FS ops we do. I changed the test like this:\n\t\n\tdiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\n\tindex 3fc16944e9e..479a495c68c 100755\n\t--- a/t/t3903-stash.sh\n\t+++ b/t/t3903-stash.sh\n\t@@ -1383,7 +1383,7 @@ BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n\t \n\t test_expect_success 'stash with core.fsyncmethod=batch' \"\n\t \ttest_create_unique_files 2 4 fsync-files &&\n\t-\tgit $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n\t+\tstrace -f git $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n\t \trm -f fsynced_files &&\n\t \n\t \t# The files were untracked, so use the third parent,\n\t\nThen we get this output, with my comments, and I snipped some output:\n\t \n\t$ ./t3903-stash.sh --run=1-4,114 -vixd 2>&1|grep --color -e 89772c935031c228ed67890f9 -e .git/stash -e bulk_fsync -e .git/index\n\t[pid 14703] access(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n\t[pid 14703] access(\".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n\t[pid 14703] link(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/tmp_obj_bdUlzu\", \".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n\nHere we're creating the tmp_objdir() files. We then sync_file_range()\nand close() this.\n\n\t[pid 14703] openat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/objects/tmp_objdir-bulk-fsync-rR3AQI/bulk_fsync_HsDRl7\", O_RDWR|O_CREAT|O_EXCL, 0600) = 9\n\t[pid 14703] unlink(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/objects/tmp_objdir-bulk-fsync-rR3AQI/bulk_fsync_HsDRl7\") = 0\n\nThis is the flushing of the \"cookie\" in do_batch_fsync().\n\n\t[pid 14703] newfstatat(AT_FDCWD, \".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, 0) = 0\n\t[pid 14703] link(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n\nHere we're going through the object dir migration with\nunplug_bulk_checkin().\n\n\t[pid 14703] unlink(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n\tnewfstatat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, AT_SYMLINK_NOFOLLOW) = 0\n\t[pid 14705] access(\".git/objects/tmp_objdir-bulk-fsync-0F7DGy/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n\t[pid 14705] access(\".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = 0\n\t[pid 14705] utimensat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", NULL, 0) = 0\n\t[pid 14707] openat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", O_RDONLY|O_CLOEXEC) = 9\n\nWe then update the index itself, first a temporary index.stash :\n\n    openat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", O_RDWR|O_CREAT|O_EXCL|O_CLOEXEC, 0666) = 8\n    openat(AT_FDCWD, \".git/index.stash.19141\", O_RDONLY) = 9\n    newfstatat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, AT_SYMLINK_NOFOLLOW) = 0\n    newfstatat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", {st_mode=S_IFREG|0644, st_size=927, ...}, 0) = 0\n    rename(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141\") = 0\n    unlink(\".git/index.stash.19141\")        = 0\n\nFollowed by the same and a later rename of the actual index:\n\n    [pid 19146] rename(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.lock\", \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index\") = 0\n\nSo, my question is still why the temporary object dir migration part of\nthis is needed.\n\nWe are writing N loose object files, and we write those to temporary\nnames already.\n\nAFAIKT we could do all of this by doing the same\ntmp/rename/sync_file_range dance on the main object store.\n\nThen instead of the \"bulk_fsync\" cookie file don't close() the last file\nobject file we write until we issue the fsync on it.\n\nBut maybe this is all needed, I just can't understand from the commit\nmessage why the \"bulk checkin\" part is being done.\n\nI think since we've been over this a few times without any success it\nwould really help to have some example of the smallest set of syscalls\nto write a file like this safely. I.e. this is doing (pseudocode):\n\n    /* first the bulk path */\n    open(\"bulk/x.tmp\");\n    write(\"bulk/x.tmp\");\n    sync_file_range(\"bulk/x.tmp\");\n    close(\"bulk/x.tmp\");\n    rename(\"bulk/x.tmp\", \"bulk/x\");\n    open(\"bulk/y.tmp\");\n    write(\"bulk/y.tmp\");\n    sync_file_range(\"bulk/y.tmp\");\n    close(\"bulk/y.tmp\");\n    rename(\"bulk/y.tmp\", \"bulk/y\");\n    /* Rename to \"real\" */\n    rename(\"bulk/x\", x\");\n    rename(\"bulk/y\", y\");\n    /* sync a cookie */\n    fsync(\"cookie\");\n\nAnd I'm asking why it's not:\n\n    /* Rename to \"real\" as we go */\n    open(\"x.tmp\");\n    write(\"x.tmp\");\n    sync_file_range(\"x.tmp\");\n    close(\"x.tmp\");\n    rename(\"x.tmp\", \"x\");\n    last_fd = open(\"y.tmp\"); /* don't close() the last one yet */\n    write(\"y.tmp\");\n    sync_file_range(\"y.tmp\");\n    rename(\"y.tmp\", \"y\");\n    /* sync a cookie */\n    fsync(last_fd);\n\nWhich I guess is two questions:\n\n A. do we need the cookie, or can we re-use the fd of the last thing we\n    write?\n B. Is the bulk indirection needed?\n\n> +\t\tfsync_or_die(fd, \"loose object file\");\n\nUnrelated nit: this API is producing sentence lego unfriendly to\ntranslators.\n\nShould be made to take an enum or something, so we can emit the relevant\ntranslated message in fsync_or_die(). Imagine getting:\n\n\tfsync error on '日本語は話せません'\n\nWhich this will do, just the other way around for non-English speakers\nusing the translation.\n\n(The solution is also not to add _() here, since translators will want\nto control the word order.)\n\n> diff --git a/cache.h b/cache.h\n> index 3160bc1e489..d1ae51388c9 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1040,7 +1040,8 @@ extern int use_fsync;\n>  \n>  enum fsync_method {\n>  \tFSYNC_METHOD_FSYNC,\n> -\tFSYNC_METHOD_WRITEOUT_ONLY\n> +\tFSYNC_METHOD_WRITEOUT_ONLY,\n> +\tFSYNC_METHOD_BATCH\n>  };\n>  \n>  extern enum fsync_method fsync_method;\n> @@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n>  int fsync_component(enum fsync_component component, int fd);\n>  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n>  \n> +static inline int batch_fsync_enabled(enum fsync_component component)\n> +{\n> +\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n> +}\n> +\n>  ssize_t read_in_full(int fd, void *buf, size_t count);\n>  ssize_t write_in_full(int fd, const void *buf, size_t count);\n>  ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\n> diff --git a/config.c b/config.c\n> index 261ee7436e0..0b28f90de8b 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>  \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n>  \t\telse if (!strcmp(value, \"writeout-only\"))\n>  \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n> +\t\telse if (!strcmp(value, \"batch\"))\n> +\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n>  \t\telse\n>  \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n>  \n> diff --git a/object-file.c b/object-file.c\n> index 5258d9ed827..bdb0a38328f 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1895,6 +1895,8 @@ static void close_loose_object(int fd)\n>  \n>  \tif (fsync_object_files > 0)\n>  \t\tfsync_or_die(fd, \"loose object file\");\n> +\telse if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> +\t\tfsync_loose_object_bulk_checkin(fd);\n>  \telse\n>  \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n>  \t\t\t\t       \"loose object file\");\n\nThis is related to the above comments about what minimum set of syscalls\nare needed to trigger this \"bulk\" behavior, but it seems to me that this\nwhole API is avoiding just passing some new flags down to object-file.c\nand friends.\n\nFor e.g. update-index that results in e.g. the \"plug bulk\" not being\naware of HASH_WRITE_OBJECT, so with dry-run writes and the like we'll do\nthe whole setup/teardown for nothing.\n\nWhich is another reason I wondered why this couldn't be a flagged passed\ndown to the object writing...\n"},{"id":"451736","messageId":"xmqqpmmf1bm5.fsf@gitster.g","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T17:03:46Z","receivedAt":"2022-03-21T17:03:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> V2 changes:\n>\n>  * Change doc to indicate that only some repo updates are batched\n\nOK.\n\n>  * Null and zero out control variables in do_batch_fsync under\n>    unplug_bulk_checkin\n\nOK.\n\n>  * Make batch mode default on Windows.\n\nI do not care either way ;-)\n\n>  * Update the description for the initial patch that cleans up the\n>    bulk-checkin infrastructure.\n\nOK.\n\n>  * Rebase onto 'seen' at 0cac37f38f9.\n\nThat's unfortunate.  Having to depend on almost everything in 'seen'\nis a guaranteed way to ensure that the topic would never graduate to\n'next'.\n\nFor this topic, ns/core-fsyncmethod is the only thing outside of\n'master' that the previous round needed, so I did an equivalent of\n\n    $ git checkout -b ns/batch-fsync b896f729e2\n    $ git merge ns/core-fsyncmethod \n\nto prepare fd008b1442 and then queued the patches on top, i.e.\n\n    $ git am -s mbox\n\n> This work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n\n\"at .\"?\n\nIn any case, I've applied them on 0cac37f38f9 and then re-applied\nthe result on top of fd008b1442 (i.e. the same base as the previous\nround was queued), which, with the magic of \"am -3\", applied\ncleanly.  Double checking the result was also simple (i.e. the tip of\nsuch an application on top of fd008b1442 can be merged with\n0cac37f38f9 and the result should be identical to the result of\napplying them directly on top of 0cac37f38f9) and seems to have\nproduced the right result.\n\n\\Thanks.\n\n\n"},{"id":"451738","messageId":"xmqqee2v1adc.fsf@gitster.g","threadId":"57568","inReplyTo":"3ed1dcd9b9ba9b34f26b3012eaba8da0269ee842.1647760560.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T17:30:39Z","receivedAt":"2022-03-21T17:30:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> +  updates in the disk writeback cache and then does a single full fsync of\n> +  a dummy file to trigger the disk cache flush at the end of the operation.\n\nIt is unfortunate that we have a rather independent \"unplug\" that is\nnot tied to the \"this is the last operation in the batch\"---if there\nwere we didn't have to invent a dummy but a single full sync on the\nreal file who happened to be the last one in the batch would be\nsufficient.  It would not matter, if the batch is any meaningful\nsize, hopefully.\n\n> +/*\n> + * Cleanup after batch-mode fsync_object_files.\n> + */\n> +static void do_batch_fsync(void)\n> +{\n> +\t/*\n> +\t * Issue a full hardware flush against a temporary file to ensure\n> +\t * that all objects are durable before any renames occur.  The code in\n> +\t * fsync_loose_object_bulk_checkin has already issued a writeout\n> +\t * request, but it has not flushed any writeback cache in the storage\n> +\t * hardware.\n> +\t */\n> +\n> +\tif (needs_batch_fsync) {\n> +\t\tstruct strbuf temp_path = STRBUF_INIT;\n> +\t\tstruct tempfile *temp;\n> +\n> +\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> +\t\ttemp = xmks_tempfile(temp_path.buf);\n> +\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> +\t\tdelete_tempfile(&temp);\n> +\t\tstrbuf_release(&temp_path);\n> +\t\tneeds_batch_fsync = 0;\n> +\t}\n> +\n> +\tif (bulk_fsync_objdir) {\n> +\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n> +\t\tbulk_fsync_objdir = NULL;\n\nThe struct obtained from tmp_objdir_create() is consumed by\ntmp_objdir_migrate() so the only clean-up left for the caller to do\nis to clear it to NULL.  OK.\n\n> +\t}\n\nThis initially made me wonder why we need two independent flags.\nAfter applying this patch but not any later steps, upon plugging, we\ncreate the tentative object directory, and any loose object will be\ncreated there, but because nobody calls the writeout-only variant\nvia fsync_loose_object_bulk_checkin() yet, needs_batch_fsync may not\nbe turned on.  But even in that case, any new loose objects are in\nthe tentative object directory and need to be migrated to the real\nplace.\n\nAnd we may not cover all the existing code paths at the end of the\nseries, or any new code paths right away after they get introduced,\nto be aware of the fsync_loose_object_bulk_checkin() when they\ncreate a loose object file, so it is most likely that these two if\nstatements will be with us forever.\n\nOK.\n\n> @@ -274,6 +311,24 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \treturn 0;\n>  }\n>  \n> +void fsync_loose_object_bulk_checkin(int fd)\n> +{\n> +\t/*\n> +\t * If we have a plugged bulk checkin, we issue a call that\n> +\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t * command. Later on we will issue a single hardware flush\n> +\t * before as part of do_batch_fsync.\n> +\t */\n> +\tif (bulk_checkin_plugged &&\n> +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> +\t\tassert(bulk_fsync_objdir);\n> +\t\tif (!needs_batch_fsync)\n> +\t\t\tneeds_batch_fsync = 1;\n\nExcept for when we unplug, do we ever flip needs_batch_fsync bit\noff, once it is set?  If the answer is no, wouldn't it be clearer to\nunconditionally set it, instead of \"set it only for the first time\"?\n\n> +\t} else {\n> +\t\tfsync_or_die(fd, \"loose object file\");\n> +\t}\n> +}\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags)\n> @@ -288,6 +343,19 @@ int index_bulk_checkin(struct object_id *oid,\n>  void plug_bulk_checkin(void)\n>  {\n>  \tassert(!bulk_checkin_plugged);\n> +\n> +\t/*\n> +\t * A temporary object directory is used to hold the files\n> +\t * while they are not fsynced.\n> +\t */\n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> +\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> +\t\tif (!bulk_fsync_objdir)\n> +\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n> +\n> +\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> +\t}\n> +\n>  \tbulk_checkin_plugged = 1;\n>  }\n>  \n> @@ -297,4 +365,6 @@ void unplug_bulk_checkin(void)\n>  \tbulk_checkin_plugged = 0;\n>  \tif (bulk_checkin_state.f)\n>  \t\tfinish_bulk_checkin(&bulk_checkin_state);\n> +\n> +\tdo_batch_fsync();\n>  }\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index b26f3dc3b74..08f292379b6 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -6,6 +6,8 @@\n>  \n>  #include \"cache.h\"\n>  \n> +void fsync_loose_object_bulk_checkin(int fd);\n> +\n>  int index_bulk_checkin(struct object_id *oid,\n>  \t\t       int fd, size_t size, enum object_type type,\n>  \t\t       const char *path, unsigned flags);\n> diff --git a/cache.h b/cache.h\n> index 3160bc1e489..d1ae51388c9 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1040,7 +1040,8 @@ extern int use_fsync;\n>  \n>  enum fsync_method {\n>  \tFSYNC_METHOD_FSYNC,\n> -\tFSYNC_METHOD_WRITEOUT_ONLY\n> +\tFSYNC_METHOD_WRITEOUT_ONLY,\n> +\tFSYNC_METHOD_BATCH\n>  };\n\nStyle.\n\nThese days we allow trailing comma to enum definitions.  Perhaps\ngive a trailing comma after _BATCH so that the next update patch\nwill become less noisy?\n\nThanks.\n"},{"id":"451740","messageId":"xmqqzgljyz34.fsf@gitster.g","threadId":"57568","inReplyTo":"54797dbc52060b7fa913642cd5266f7e159a5bc9.1647760561.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T17:50:23Z","receivedAt":"2022-03-21T17:50:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index 75d646377cc..38e9d7e88cb 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -5,6 +5,7 @@\n>   */\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"cache.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"config.h\"\n>  #include \"lockfile.h\"\n>  #include \"quote.h\"\n> @@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \n>  \tthe_index.updated_skipworktree = 1;\n>  \n> +\t/* we might be adding many objects to the object database */\n> +\tplug_bulk_checkin();\n> +\n>  \t/*\n>  \t * Custom copy of parse_options() because we want to handle\n>  \t * filename arguments as they come.\n> @@ -1190,6 +1194,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\tstrbuf_release(&buf);\n>  \t}\n>  \n> +\t/* by now we must have added all of the new objects */\n> +\tunplug_bulk_checkin();\n\nI understand read-from-stdin code path would be worth plugging, but\nthe list of paths on the command line?  How many of them would one\nfit?\n\nOf course, the feeder may be expecting for the objects to appear in\nthe object store as it feeds the paths and will be utterly broken by\nthis change, as you mentioned in the proposed log message.  The\nexisting plug/unplug will change the behaviour by making the objects\nsent to the packfile available only after getting unplugged.  This\nseries makes it even worse by making loose objects also unavailable\nuntil unplug is called.\n\nSo, it probably is safer and more sensible approach to introduce a\nnew command line option to allow the bulk checkin, and those who do\nnot care about the intermediate state to opt into the new feature.\n\n"},{"id":"451741","messageId":"xmqqv8w7yyua.fsf@gitster.g","threadId":"57568","inReplyTo":"6662e2dae0f5d65c158fba785d186885f9671073.1647760561.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 4/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T17:55:41Z","receivedAt":"2022-03-21T17:55:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The unpack-objects functionality is used by fetch, push, and fast-import\n> to turn the transfered data into object database entries when there are\n> fewer objects than the 'unpacklimit' setting.\n>\n> By enabling bulk-checkin when unpacking objects, we can take advantage\n> of batched fsyncs.\n\nThis feels confused in that we dispatch to unpack-objects (instead\nof index-objects) only when the number of loose objects should not\nmatter from performance point of view, and bulk-checkin should shine\nfrom performance point of view only when there are enough objects to\nbatch.\n\nAlso if we ever add \"too many small loose objects is wasteful, let's\nsend them into a single 'batch pack'\" optimization, it would create\na funny situation where the caller sends the contents of a small\nincoming packfile to unpack-objects, but the command chooses to\nbunch them all together in a packfile anyway ;-)\n\nSo, I dunno.\n\n\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/unpack-objects.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index dbeb0680a58..c55b6616aed 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -1,5 +1,6 @@\n>  #include \"builtin.h\"\n>  #include \"cache.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"config.h\"\n>  #include \"object-store.h\"\n>  #include \"object.h\"\n> @@ -503,10 +504,12 @@ static void unpack_all(void)\n>  \tif (!quiet)\n>  \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>  \tCALLOC_ARRAY(obj_list, nr_objects);\n> +\tplug_bulk_checkin();\n>  \tfor (i = 0; i < nr_objects; i++) {\n>  \t\tunpack_one(i);\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n> +\tunplug_bulk_checkin();\n>  \tstop_progress(&progress);\n>  \n>  \tif (delta_list)\n"},{"id":"451743","messageId":"CANQDOdfCJy68z0bNrhSmwo_uEVa6=y4V1dY0kZDq7JOTUD+6iQ@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqpmmf1bm5.fsf@gitster.g","subject":"Re: [PATCH v2 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T18:14:53Z","receivedAt":"2022-03-21T18:15:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 10:03 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> >  * Rebase onto 'seen' at 0cac37f38f9.\n>\n> That's unfortunate.  Having to depend on almost everything in 'seen'\n> is a guaranteed way to ensure that the topic would never graduate to\n> 'next'.\n>\n> For this topic, ns/core-fsyncmethod is the only thing outside of\n> 'master' that the previous round needed, so I did an equivalent of\n>\n>     $ git checkout -b ns/batch-fsync b896f729e2\n>     $ git merge ns/core-fsyncmethod\n>\n> to prepare fd008b1442 and then queued the patches on top, i.e.\n>\n>     $ git am -s mbox\n>\n> > This work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n>\n> \"at .\"?\n>\n> In any case, I've applied them on 0cac37f38f9 and then re-applied\n> the result on top of fd008b1442 (i.e. the same base as the previous\n> round was queued), which, with the magic of \"am -3\", applied\n> cleanly.  Double checking the result was also simple (i.e. the tip of\n> such an application on top of fd008b1442 can be merged with\n> 0cac37f38f9 and the result should be identical to the result of\n> applying them directly on top of 0cac37f38f9) and seems to have\n> produced the right result.\n>\n> \\Thanks.\n\nThanks Junio.  I was worried about how to properly represent the dependency\nbetween these two in-flight branches without waiting for ns/core-fsyncmethod to\nget into next.   Now ns/core-fsyncmethod appears to be there, so I'm assuming\nthat branch should have a stable OID until the end of the cycle.\n\nShould I base future versions of this series on the tip of\nns/core-fsyncmethod, or\non the merge point between that branch and 'next'?  I guess it doesn't\nreally matter\nif the merge is clean.\n\nThanks,\nNeeraj\n"},{"id":"451744","messageId":"CANQDOdcPAb31JY6LdMiCn26192_1wvKkHodpNEBkSh3WbD+e4g@mail.gmail.com","threadId":"57568","inReplyTo":"220321.86a6dj9xja.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T18:28:31Z","receivedAt":"2022-03-21T18:28:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 7:43 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> > [...]\n> > +     if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> > +             bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> > +             if (!bulk_fsync_objdir)\n> > +                     die(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n>\n> Should camel-case the config var, and we should have a die_errno() here\n> which tell us why we couldn't create it (possibly needing to ferry it up\n> from the tempfile API...)\n\nThanks for noticing the camelCasing.  The config var name was also\nwrong. Now it will read:\n> > +                     die(_(\"Could not create temporary object directory for core.fsyncMethod=batch\"));\n\nDo you have any recommendations on how to easily ferry the correct\nerrno out of tmp_objdir_create?\nIt looks like the remerge-diff usage has the same die behavior w/o\nerrno, and the builtin/receive-pack.c usage\ndoesn't die, but also loses the errno.  I'm concerned about preserving\nthe errno across the or tmp_objdir_destroy\ncalls.  I could introduce a temp errno var to preserve it across\nthose. Is that what you had in mind?\n\nThanks,\nNeeraj\n"},{"id":"451745","messageId":"xmqqpmmfxigx.fsf@gitster.g","threadId":"57568","inReplyTo":"1937746df47eefecfc343e32eb9bf6c0949fb7b9.1647760561.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 6/7] core.fsyncmethod: tests for batch mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T18:34:38Z","receivedAt":"2022-03-21T18:34:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Add test cases to exercise batch mode for:\n>  * 'git add'\n\nI was wondering why the obviously safe and good candidate 'git add' is\nnot gaining plug/unplug pair in this series.  It is obviously safe,\nunlike 'update-index', that nobody can interact with it, observe its\nintermediate output, and expect anything from it.\n\nI think the stupid reason of the lack of new plug/unplug is because\nwe already had them, which is good ;-).\n\n>  * 'git stash'\n>  * 'git update-index'\n\nAs I said, I suspect that we'd want to do this safely by adding a\nnew option to \"update-index\" and passing it from \"stash\" which knows\nthat it does not care about the intermediate state.\n\n> These tests ensure that the added data winds up in the object database.\n\nIn other words, \"git add $path; git rev-parse :$path\" (and its\ncousins) would be happy?  Like new object files not left hanging in\na tentative object store etc. _after_ the commands finish.\n\nGood.\n\n> In this change we introduce a new test helper lib-unique-files.sh. The\n> goal of this library is to create a tree of files that have different\n> oids from any other files that may have been created in the current test\n> repo. This helps us avoid missing validation of an object being added due\n> to it already being in the repo.\n\nMore on this below.\n\n> We aren't actually issuing any fsyncs in these tests, since\n> GIT_TEST_FSYNC is 0, but we still exercise all of the tmp_objdir logic\n> in bulk-checkin.\n\nShouldn't we manually override that, if it matters?\nNot a suggestion but a question.\n\n> +# Create multiple files with unique contents. Takes the number of\n> +# directories, the number of files in each directory, and the base\n> +# directory.\n\nThis is more honest, compared to the claim made in the proposed log\nmessage, in that the uniqueness guarantee is only among the files\ncreated by this helper.  If we created other test contents without\nusing this helper, that may crash with the ones created here.\n\n> +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> +#\t\t\t\t\t each in my_dir, all with unique\n> +#\t\t\t\t\t contents.\n> +\n> +test_create_unique_files() {\n\nStyle.  SP on both sides of ().  I.e.\n\n\ttest_create_unique_files () {\n\n> +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n> +\n> +\tlocal dirs=$1\n> +\tlocal files=$2\n> +\tlocal basedir=$3\n> +\tlocal counter=0\n> +\ttest_tick\n> +\tlocal basedata=$test_tick\n\nI am not sure if consumption and reliance on tick is a wise thing.\n$basedir must be unique across all the other directories in this\ntest repository (there is no other $basedir)---can't we key\nuniqueness off of it?\n\n> +\trm -rf $basedir\n\nCan $basedir have any $IFS character in it?  We should \"$quote\" it.\n\n> +\tfor i in $(test_seq $dirs)\n> +\tdo\n> +\t\tlocal dir=$basedir/dir$i\n> +\n> +\t\tmkdir -p \"$dir\"\n> +\t\tfor j in $(test_seq $files)\n> +\t\tdo\n> +\t\t\tcounter=$((counter + 1))\n> +\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n\nAn extra SP before \">\"?\n\n> +\t\tdone\n> +\tdone\n> +}\n\nThere is no &&- cascade here, and we expect nothing in this to\nfail.  Is that sensible?\n\n> +test_expect_success 'git add: core.fsyncmethod=batch' \"\n> +\ttest_create_unique_files 2 4 fsync-files &&\n> +\tgit $BATCH_CONFIGURATION add -- ./fsync-files/ &&\n> +\trm -f fsynced_files &&\n> +\tgit ls-files --stage fsync-files/ > fsynced_files &&\n\nStyle.  No SP between redirection operator and its target.  I.e.\n\n\tgit ls-files --stage fsync-files/ >fsynced_files &&\n\nMixture of names-with-dash and name_with_understore looks somewhat\nirritating.\n\n> +\ttest_line_count = 8 fsynced_files &&\n\nThe magic \"8\" matches \"2 4\" we saw earlier for create_unique_files?\n\n> +\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n\nA test helper that takes the name of a file that has \"ls-files -s\" output\nmay prove to be useful.  I dunno.\n\n> diff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\n> index a11d61206ad..8e2f73cc68f 100755\n> --- a/t/t5300-pack-object.sh\n> +++ b/t/t5300-pack-object.sh\n> @@ -162,23 +162,25 @@ test_expect_success 'pack-objects with bogus arguments' '\n>  \n>  check_unpack () {\n>  \ttest_when_finished \"rm -rf git2\" &&\n> -\tgit init --bare git2 &&\n> -\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n> -\tgit -C git2 unpack-objects <\"$1\".pack &&\n> -\t(cd .git && find objects -type f -print) |\n> -\twhile read path\n> -\tdo\n> -\t\tcmp git2/$path .git/$path || {\n> -\t\t\techo $path differs.\n> -\t\t\treturn 1\n> -\t\t}\n> -\tdone\n> +\tgit $2 init --bare git2 &&\n> +\t(\n> +\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n> +\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n> +\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n> +\t) <obj-list >current &&\n> +\tcmp obj-list current\n>  }\n\nI think the change from the old \"the existence and the contents of\nthe object files must all match\" to the new \"cat-file should say\nthat the objects we expect to exist indeed do\" is not a bad thing.\n\nWe used to only depend on the contents of the provided packfile but\nnow we assume that obj-list file gives us the list of objects.  Is\nthat sensible?  I somehow do not think so.  Don't we have the\ncorresponding \"$1.idx\" that we can feed to \"git show-index\", e.g.\n\n\tgit show-index <\"$1.pack\" >expect.full &&\n\tcut -d\" \" -f2 >expect <expect.full &&\n\t... your test in \"$2\", but feeding expect instead of obj-list ...\n\ttest_cmp expect actual\n\nAlso make sure you quote whatever is coming from outside, even if\nyou happen to call the helper with tokens that do not need quoting\nin the current code.  It is a good discipline to help readers.\n\nThanks.\n"},{"id":"451756","messageId":"CANQDOdez2u4oTNeETM0zLQr7Xb6XXbEuoxXPhSqGuurwQWbkHA@mail.gmail.com","threadId":"57568","inReplyTo":"220321.861qyv9rjr.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T20:14:19Z","receivedAt":"2022-03-21T20:14:42Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 9:52 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > One major source of the cost of fsync is the implied flush of the\n> > hardware writeback cache within the disk drive. This commit introduces\n> > a new `core.fsyncMethod=batch` option that batches up hardware flushes.\n> > It hooks into the bulk-checkin plugging and unplugging functionality,\n> > takes advantage of tmp-objdir, and uses the writeout-only support code.\n> >\n> > When the new mode is enabled, we do the following for each new object:\n> > 1. Create the object in a tmp-objdir.\n> > 2. Issue a pagecache writeback request and wait for it to complete.\n> >\n> > At the end of the entire transaction when unplugging bulk checkin:\n> > 1. Issue an fsync against a dummy file to flush the hardware writeback\n> >    cache, which should by now have seen the tmp-objdir writes.\n> > 2. Rename all of the tmp-objdir files to their final names.\n> > 3. When updating the index and/or refs, we assume that Git will issue\n> >    another fsync internal to that operation. This is not the default\n> >    today, but the user now has the option of syncing the index and there\n> >    is a separate patch series to implement syncing of refs.\n>\n> Re my question in\n> https://lore.kernel.org/git/220310.86r179ki38.gmgdl@evledraar.gmail.com/\n> (which you *partially* replied to per my reading, i.e. not the\n> fsync_nth() question) I still don't get why the tmp-objdir part of this\n> is needed.\n>\n\nSorry for not fully answering your question. I think part of the issue might be\nbackground, where it's not clear to me what's different between your\nunderstanding\nand mine, so may not have included something that's questionable to\nyou but not to me.\n\nYour syscall description below makes the issues very concrete, so I\nthink we'll get it this round :).\n\n> For \"git stash\" which is one thing sped up by this let's go over what\n> commands/FS ops we do. I changed the test like this:\n>\n>         diff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\n>         index 3fc16944e9e..479a495c68c 100755\n>         --- a/t/t3903-stash.sh\n>         +++ b/t/t3903-stash.sh\n>         @@ -1383,7 +1383,7 @@ BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n>\n>          test_expect_success 'stash with core.fsyncmethod=batch' \"\n>                 test_create_unique_files 2 4 fsync-files &&\n>         -       git $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n>         +       strace -f git $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n>                 rm -f fsynced_files &&\n>\n>                 # The files were untracked, so use the third parent,\n>\n> Then we get this output, with my comments, and I snipped some output:\n>\n>         $ ./t3903-stash.sh --run=1-4,114 -vixd 2>&1|grep --color -e 89772c935031c228ed67890f9 -e .git/stash -e bulk_fsync -e .git/index\n>         [pid 14703] access(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n>         [pid 14703] access(\".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n>         [pid 14703] link(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/tmp_obj_bdUlzu\", \".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n>\n> Here we're creating the tmp_objdir() files. We then sync_file_range()\n> and close() this.\n>\n>         [pid 14703] openat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/objects/tmp_objdir-bulk-fsync-rR3AQI/bulk_fsync_HsDRl7\", O_RDWR|O_CREAT|O_EXCL, 0600) = 9\n>         [pid 14703] unlink(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/objects/tmp_objdir-bulk-fsync-rR3AQI/bulk_fsync_HsDRl7\") = 0\n>\n> This is the flushing of the \"cookie\" in do_batch_fsync().\n>\n>         [pid 14703] newfstatat(AT_FDCWD, \".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, 0) = 0\n>         [pid 14703] link(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\", \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n>\n> Here we're going through the object dir migration with\n> unplug_bulk_checkin().\n>\n>         [pid 14703] unlink(\".git/objects/tmp_objdir-bulk-fsync-rR3AQI/fb/89772c935031c228ed67890f953c0a2b5c8316\") = 0\n>         newfstatat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, AT_SYMLINK_NOFOLLOW) = 0\n>         [pid 14705] access(\".git/objects/tmp_objdir-bulk-fsync-0F7DGy/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = -1 ENOENT (No such file or directory)\n>         [pid 14705] access(\".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", F_OK) = 0\n>         [pid 14705] utimensat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", NULL, 0) = 0\n>         [pid 14707] openat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", O_RDONLY|O_CLOEXEC) = 9\n>\n> We then update the index itself, first a temporary index.stash :\n>\n>     openat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", O_RDWR|O_CREAT|O_EXCL|O_CLOEXEC, 0666) = 8\n>     openat(AT_FDCWD, \".git/index.stash.19141\", O_RDONLY) = 9\n>     newfstatat(AT_FDCWD, \".git/objects/fb/89772c935031c228ed67890f953c0a2b5c8316\", {st_mode=S_IFREG|0444, st_size=29, ...}, AT_SYMLINK_NOFOLLOW) = 0\n>     newfstatat(AT_FDCWD, \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", {st_mode=S_IFREG|0644, st_size=927, ...}, 0) = 0\n>     rename(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141.lock\", \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.stash.19141\") = 0\n>     unlink(\".git/index.stash.19141\")        = 0\n>\n> Followed by the same and a later rename of the actual index:\n>\n>     [pid 19146] rename(\"/home/avar/g/git/t/trash directory.t3903-stash/.git/index.lock\", \"/home/avar/g/git/t/trash directory.t3903-stash/.git/index\") = 0\n>\n> So, my question is still why the temporary object dir migration part of\n> this is needed.\n>\n> We are writing N loose object files, and we write those to temporary\n> names already.\n>\n> AFAIKT we could do all of this by doing the same\n> tmp/rename/sync_file_range dance on the main object store.\n>\n\nWhy not the main object store? We want to maintain the invariant that any\nname in the main object store refers to a file that durably has the\ncorrect contents.\nIf we do sync_file_range and then rename, and then crash, we now have a file\nin the main object store with some SHA name, whose contents may or may not\nmatch the SHA.  However, if we ensure an fsync happens before the rename,\na crash at any point will leave us either with no file in the main\nobject store or\nwith a file that is durable on the disk.\n\n> Then instead of the \"bulk_fsync\" cookie file don't close() the last file\n> object file we write until we issue the fsync on it.\n>\n> But maybe this is all needed, I just can't understand from the commit\n> message why the \"bulk checkin\" part is being done.\n>\n> I think since we've been over this a few times without any success it\n> would really help to have some example of the smallest set of syscalls\n> to write a file like this safely. I.e. this is doing (pseudocode):\n>\n>     /* first the bulk path */\n>     open(\"bulk/x.tmp\");\n>     write(\"bulk/x.tmp\");\n>     sync_file_range(\"bulk/x.tmp\");\n>     close(\"bulk/x.tmp\");\n>     rename(\"bulk/x.tmp\", \"bulk/x\");\n>     open(\"bulk/y.tmp\");\n>     write(\"bulk/y.tmp\");\n>     sync_file_range(\"bulk/y.tmp\");\n>     close(\"bulk/y.tmp\");\n>     rename(\"bulk/y.tmp\", \"bulk/y\");\n>     /* Rename to \"real\" */\n>     rename(\"bulk/x\", x\");\n>     rename(\"bulk/y\", y\");\n>     /* sync a cookie */\n>     fsync(\"cookie\");\n>\n\nThe '/* Rename to \"real\" */' and '/* sync a cookie */' steps are\nreversed in your above sequence. It should be\n1: (for each file)\n    a) open\n    b) write\n    c) sync_file_range\n    d) close\n    e) rename in tmp_objdir  -- The rename step is not required for\nbulk-fsync. An earlier version of this series didn't do it, but\nJeff King pointed out that it was required for concurrency:\nhttps://lore.kernel.org/all/YVOrikAl%2Fu5%2FVi61@coredump.intra.peff.net/\n\n2: fsync something on the same volume to flush the filesystem log and\ndisk cache. This functions as a \"barrier\".\n3: Rename to final names.  At this point we know that the \"contents\"\nare durable, so if the final name exists, we can read through it to\nget the data.\n\n> And I'm asking why it's not:\n>\n>     /* Rename to \"real\" as we go */\n>     open(\"x.tmp\");\n>     write(\"x.tmp\");\n>     sync_file_range(\"x.tmp\");\n>     close(\"x.tmp\");\n>     rename(\"x.tmp\", \"x\");\n>     last_fd = open(\"y.tmp\"); /* don't close() the last one yet */\n>     write(\"y.tmp\");\n>     sync_file_range(\"y.tmp\");\n>     rename(\"y.tmp\", \"y\");\n>     /* sync a cookie */\n>     fsync(last_fd);\n>\n> Which I guess is two questions:\n>\n>  A. do we need the cookie, or can we re-use the fd of the last thing we\n>     write?\n\nWe can re-use the FD of the last thing we write, but that results in a\ntricker API which\nis more intrusive on callers. I was originally using a lockfile, but\nfound a usage where\nthere was no lockfile in unpack-objects.\n\n>  B. Is the bulk indirection needed?\n>\n\nHopefully the explanation above makes it clear why we need the\nindirection. To state it again,\nwe need a real fsync before creating the final name in the objdir,\notherwise on a crash a name\ncould exist that points at contents which could have been lost, since\nthey weren't durable. I\nupdated the comment in do_batch_fsync to make this a little clearer.\n\n> > +             fsync_or_die(fd, \"loose object file\");\n>\n> Unrelated nit: this API is producing sentence lego unfriendly to\n> translators.\n>\n> Should be made to take an enum or something, so we can emit the relevant\n> translated message in fsync_or_die(). Imagine getting:\n>\n>         fsync error on '日本語は話せません'\n>\n> Which this will do, just the other way around for non-English speakers\n> using the translation.\n>\n> (The solution is also not to add _() here, since translators will want\n> to control the word order.)\n\nThis line is copied from the preexisting version of the same code in\nclose_loose_object.\nIf I'm understanding it correctly, the entire chain of messages is\nuntranslated and would\nremain as english.  fsync_or_die doesn't have a _().  Can we just\nleave it that way, since\nthis is not a situation that should actually happen to many users?\nAlternatively, I think it\nwould be pretty trivial to just pass through the file name, so I'll\njust do that.\n\n> > diff --git a/cache.h b/cache.h\n> > index 3160bc1e489..d1ae51388c9 100644\n> > --- a/cache.h\n> > +++ b/cache.h\n> > @@ -1040,7 +1040,8 @@ extern int use_fsync;\n> >\n> >  enum fsync_method {\n> >       FSYNC_METHOD_FSYNC,\n> > -     FSYNC_METHOD_WRITEOUT_ONLY\n> > +     FSYNC_METHOD_WRITEOUT_ONLY,\n> > +     FSYNC_METHOD_BATCH\n> >  };\n> >\n> >  extern enum fsync_method fsync_method;\n> > @@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n> >  int fsync_component(enum fsync_component component, int fd);\n> >  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n> >\n> > +static inline int batch_fsync_enabled(enum fsync_component component)\n> > +{\n> > +     return (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n> > +}\n> > +\n> >  ssize_t read_in_full(int fd, void *buf, size_t count);\n> >  ssize_t write_in_full(int fd, const void *buf, size_t count);\n> >  ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\n> > diff --git a/config.c b/config.c\n> > index 261ee7436e0..0b28f90de8b 100644\n> > --- a/config.c\n> > +++ b/config.c\n> > @@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n> >                       fsync_method = FSYNC_METHOD_FSYNC;\n> >               else if (!strcmp(value, \"writeout-only\"))\n> >                       fsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n> > +             else if (!strcmp(value, \"batch\"))\n> > +                     fsync_method = FSYNC_METHOD_BATCH;\n> >               else\n> >                       warning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index 5258d9ed827..bdb0a38328f 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1895,6 +1895,8 @@ static void close_loose_object(int fd)\n> >\n> >       if (fsync_object_files > 0)\n> >               fsync_or_die(fd, \"loose object file\");\n> > +     else if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> > +             fsync_loose_object_bulk_checkin(fd);\n> >       else\n> >               fsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n> >                                      \"loose object file\");\n>\n> This is related to the above comments about what minimum set of syscalls\n> are needed to trigger this \"bulk\" behavior, but it seems to me that this\n> whole API is avoiding just passing some new flags down to object-file.c\n> and friends.\n>\n> For e.g. update-index that results in e.g. the \"plug bulk\" not being\n> aware of HASH_WRITE_OBJECT, so with dry-run writes and the like we'll do\n> the whole setup/teardown for nothing.\n>\n> Which is another reason I wondered why this couldn't be a flagged passed\n> down to the object writing...\n\nIn the original implementation [1], I did some custom thing for\nrenaming the files\nrather than tmp_objdir. But you suggested at the time that I use\ntmp_objdir, which\nwas a good decision, since it made access to the objects possible in-process and\nfor descendents in the middle of the transaction.\n\nIt sounds to me like I just shouldn't plug the bulk checkin for cases\nwhere we're not\ngoing to add to the ODB.  Plugging the bulk checkin is always\noptional.  But when\nI wrote the code, I didn't love the result, since it makes arbitrary\ncallers harder. So\nI changed the code to lazily create the tmp objdir the first time an object\nshows up, which has the same effect of avoiding the cost when we\naren't adding any\nobjects.  This also avoids the need to write an error message, since\nfailing to create\nthe tmp objdir will just result in a fsync.  The main downside here is\nthat it's another thing\nthat will have to change if we want to make adding to the ODB multithreaded.\n\nThanks,\nNeeraj\n\n[1] https://lore.kernel.org/all/12cad737635663ed596e52f89f0f4f22f58bfe38.1632176111.git.gitgitgadget@gmail.com/\n"},{"id":"451757","messageId":"CANQDOdckKac7wfrKhskBV7YCVhX6qG0P44Ej552U+LBsE6Mg-g@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqee2v1adc.fsf@gitster.g","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T20:23:27Z","receivedAt":"2022-03-21T20:23:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 10:30 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> > +  updates in the disk writeback cache and then does a single full fsync of\n> > +  a dummy file to trigger the disk cache flush at the end of the operation.\n>\n> It is unfortunate that we have a rather independent \"unplug\" that is\n> not tied to the \"this is the last operation in the batch\"---if there\n> were we didn't have to invent a dummy but a single full sync on the\n> real file who happened to be the last one in the batch would be\n> sufficient.  It would not matter, if the batch is any meaningful\n> size, hopefully.\n>\n\nI'm banking on  a large batch size or the fact that the additional\ncost of creating\nand syncing an empty file to be so small that it wouldn't be\nnoticeable event for\nsmall batches. The current unfortunate scheme at least has a very simple API\nthat's easy to apply to any other operation going forward. For instance\nbuiltin/hash-object.c might be another good operation, but it wasn't clear to me\nif it's used for any mainline scenario.\n\n> > +/*\n> > + * Cleanup after batch-mode fsync_object_files.\n> > + */\n> > +static void do_batch_fsync(void)\n> > +{\n> > +     /*\n> > +      * Issue a full hardware flush against a temporary file to ensure\n> > +      * that all objects are durable before any renames occur.  The code in\n> > +      * fsync_loose_object_bulk_checkin has already issued a writeout\n> > +      * request, but it has not flushed any writeback cache in the storage\n> > +      * hardware.\n> > +      */\n> > +\n> > +     if (needs_batch_fsync) {\n> > +             struct strbuf temp_path = STRBUF_INIT;\n> > +             struct tempfile *temp;\n> > +\n> > +             strbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> > +             temp = xmks_tempfile(temp_path.buf);\n> > +             fsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> > +             delete_tempfile(&temp);\n> > +             strbuf_release(&temp_path);\n> > +             needs_batch_fsync = 0;\n> > +     }\n> > +\n> > +     if (bulk_fsync_objdir) {\n> > +             tmp_objdir_migrate(bulk_fsync_objdir);\n> > +             bulk_fsync_objdir = NULL;\n>\n> The struct obtained from tmp_objdir_create() is consumed by\n> tmp_objdir_migrate() so the only clean-up left for the caller to do\n> is to clear it to NULL.  OK.\n>\n> > +     }\n>\n> This initially made me wonder why we need two independent flags.\n> After applying this patch but not any later steps, upon plugging, we\n> create the tentative object directory, and any loose object will be\n> created there, but because nobody calls the writeout-only variant\n> via fsync_loose_object_bulk_checkin() yet, needs_batch_fsync may not\n> be turned on.  But even in that case, any new loose objects are in\n> the tentative object directory and need to be migrated to the real\n> place.\n>\n> And we may not cover all the existing code paths at the end of the\n> series, or any new code paths right away after they get introduced,\n> to be aware of the fsync_loose_object_bulk_checkin() when they\n> create a loose object file, so it is most likely that these two if\n> statements will be with us forever.\n>\n> OK.\n\nAfter Avarb's last feedback, I've changed this to lazily create the objdir, so\nthe existence of an objdir is a suitable proxy for there being something worth\nsyncing. The potential downside is that the lazy-creation would need to be\nsynchronized if the ODB becomes multithreaded.\n\n>\n> > @@ -274,6 +311,24 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n> >       return 0;\n> >  }\n> >\n> > +void fsync_loose_object_bulk_checkin(int fd)\n> > +{\n> > +     /*\n> > +      * If we have a plugged bulk checkin, we issue a call that\n> > +      * cleans the filesystem page cache but avoids a hardware flush\n> > +      * command. Later on we will issue a single hardware flush\n> > +      * before as part of do_batch_fsync.\n> > +      */\n> > +     if (bulk_checkin_plugged &&\n> > +         git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n> > +             assert(bulk_fsync_objdir);\n> > +             if (!needs_batch_fsync)\n> > +                     needs_batch_fsync = 1;\n>\n> Except for when we unplug, do we ever flip needs_batch_fsync bit\n> off, once it is set?  If the answer is no, wouldn't it be clearer to\n> unconditionally set it, instead of \"set it only for the first time\"?\n>\n\nThis code is now gone. I was stupidly optimizing for a future\nmultithreaded world\nwhich might never come.\n\n> > +     } else {\n> > +             fsync_or_die(fd, \"loose object file\");\n> > +     }\n> > +}\n> > +\n> >  int index_bulk_checkin(struct object_id *oid,\n> >                      int fd, size_t size, enum object_type type,\n> >                      const char *path, unsigned flags)\n> > @@ -288,6 +343,19 @@ int index_bulk_checkin(struct object_id *oid,\n> >  void plug_bulk_checkin(void)\n> >  {\n> >       assert(!bulk_checkin_plugged);\n> > +\n> > +     /*\n> > +      * A temporary object directory is used to hold the files\n> > +      * while they are not fsynced.\n> > +      */\n> > +     if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> > +             bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> > +             if (!bulk_fsync_objdir)\n> > +                     die(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n> > +\n> > +             tmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> > +     }\n> > +\n> >       bulk_checkin_plugged = 1;\n> >  }\n> >\n> > @@ -297,4 +365,6 @@ void unplug_bulk_checkin(void)\n> >       bulk_checkin_plugged = 0;\n> >       if (bulk_checkin_state.f)\n> >               finish_bulk_checkin(&bulk_checkin_state);\n> > +\n> > +     do_batch_fsync();\n> >  }\n> > diff --git a/bulk-checkin.h b/bulk-checkin.h\n> > index b26f3dc3b74..08f292379b6 100644\n> > --- a/bulk-checkin.h\n> > +++ b/bulk-checkin.h\n> > @@ -6,6 +6,8 @@\n> >\n> >  #include \"cache.h\"\n> >\n> > +void fsync_loose_object_bulk_checkin(int fd);\n> > +\n> >  int index_bulk_checkin(struct object_id *oid,\n> >                      int fd, size_t size, enum object_type type,\n> >                      const char *path, unsigned flags);\n> > diff --git a/cache.h b/cache.h\n> > index 3160bc1e489..d1ae51388c9 100644\n> > --- a/cache.h\n> > +++ b/cache.h\n> > @@ -1040,7 +1040,8 @@ extern int use_fsync;\n> >\n> >  enum fsync_method {\n> >       FSYNC_METHOD_FSYNC,\n> > -     FSYNC_METHOD_WRITEOUT_ONLY\n> > +     FSYNC_METHOD_WRITEOUT_ONLY,\n> > +     FSYNC_METHOD_BATCH\n> >  };\n>\n> Style.\n>\n> These days we allow trailing comma to enum definitions.  Perhaps\n> give a trailing comma after _BATCH so that the next update patch\n> will become less noisy?\n>\n\nFixed.\n\n> Thanks.\n\nThanks!\n-Neeraj\n"},{"id":"451767","messageId":"220321.86zgljnite.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"CANQDOdez2u4oTNeETM0zLQr7Xb6XXbEuoxXPhSqGuurwQWbkHA@mail.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T20:18:04Z","receivedAt":"2022-03-21T20:37:25Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 21 2022, Neeraj Singh wrote:\n\n[Don't have time for a full reply, sorry, just something quick]\n\n> On Mon, Mar 21, 2022 at 9:52 AM Ævar Arnfjörð Bjarmason\n> [...]\n>> So, my question is still why the temporary object dir migration part of\n>> this is needed.\n>>\n>> We are writing N loose object files, and we write those to temporary\n>> names already.\n>>\n>> AFAIKT we could do all of this by doing the same\n>> tmp/rename/sync_file_range dance on the main object store.\n>>\n>\n> Why not the main object store? We want to maintain the invariant that any\n> name in the main object store refers to a file that durably has the\n> correct contents.\n> If we do sync_file_range and then rename, and then crash, we now have a file\n> in the main object store with some SHA name, whose contents may or may not\n> match the SHA.  However, if we ensure an fsync happens before the rename,\n> a crash at any point will leave us either with no file in the main\n> object store or\n> with a file that is durable on the disk.\n\nAh, I see.\n\nWhy does that matter? If the \"bulk\" mode works as advertised we might\nhave such a corrupt loose or pack file, but we won't have anything\nreferring to it as far as reachability goes.\n\nI'm aware that the various code paths that handle OID writing don't deal\ntoo well with it in practice to say the least, which one can try with\nsay:\n\n    $ echo foo | git hash-object -w --stdin\n    45b983be36b73c0788dc9cbcb76cbb80fc7bb057\n    $ echo | sudo tee .git/objects/45/b983be36b73c0788dc9cbcb76cbb80fc7bb057\n\nI.e. \"fsck\", \"show\" etc. will all scream bloddy murder, and re-running\nthat hash-object again even returns successful (we see it's there\nalready, and think it's OK).\n\nBut in any case, I think it would me much easier to both review and\nreason about this code if these concerns were split up.\n\nI.e. things that want no fsync at all (I'd think especially so) might\nwant to have such updates serialized in this manner, and as Junio\npointed out making these things inseparable as you've done creates API\nconcerns & fallout that's got nothing to do with what we need for the\nperformance gains of the bulk checkin fsyncing technique,\ne.g. concurrent \"update-index\" consumers not being able to assume\nreported objects exist as soon as they're reported.\n\n>> Then instead of the \"bulk_fsync\" cookie file don't close() the last file\n>> object file we write until we issue the fsync on it.\n>>\n>> But maybe this is all needed, I just can't understand from the commit\n>> message why the \"bulk checkin\" part is being done.\n>>\n>> I think since we've been over this a few times without any success it\n>> would really help to have some example of the smallest set of syscalls\n>> to write a file like this safely. I.e. this is doing (pseudocode):\n>>\n>>     /* first the bulk path */\n>>     open(\"bulk/x.tmp\");\n>>     write(\"bulk/x.tmp\");\n>>     sync_file_range(\"bulk/x.tmp\");\n>>     close(\"bulk/x.tmp\");\n>>     rename(\"bulk/x.tmp\", \"bulk/x\");\n>>     open(\"bulk/y.tmp\");\n>>     write(\"bulk/y.tmp\");\n>>     sync_file_range(\"bulk/y.tmp\");\n>>     close(\"bulk/y.tmp\");\n>>     rename(\"bulk/y.tmp\", \"bulk/y\");\n>>     /* Rename to \"real\" */\n>>     rename(\"bulk/x\", x\");\n>>     rename(\"bulk/y\", y\");\n>>     /* sync a cookie */\n>>     fsync(\"cookie\");\n>>\n>\n> The '/* Rename to \"real\" */' and '/* sync a cookie */' steps are\n> reversed in your above sequence. It should be\n\nSorry.\n\n> 1: (for each file)\n>     a) open\n>     b) write\n>     c) sync_file_range\n>     d) close\n>     e) rename in tmp_objdir  -- The rename step is not required for\n> bulk-fsync. An earlier version of this series didn't do it, but\n> Jeff King pointed out that it was required for concurrency:\n> https://lore.kernel.org/all/YVOrikAl%2Fu5%2FVi61@coredump.intra.peff.net/\n\nYes we definitely need the rename, I was wondering about why we needed\nit 2x for each file, but that was answered above.\n\n>> And I'm asking why it's not:\n>>\n>>     /* Rename to \"real\" as we go */\n>>     open(\"x.tmp\");\n>>     write(\"x.tmp\");\n>>     sync_file_range(\"x.tmp\");\n>>     close(\"x.tmp\");\n>>     rename(\"x.tmp\", \"x\");\n>>     last_fd = open(\"y.tmp\"); /* don't close() the last one yet */\n>>     write(\"y.tmp\");\n>>     sync_file_range(\"y.tmp\");\n>>     rename(\"y.tmp\", \"y\");\n>>     /* sync a cookie */\n>>     fsync(last_fd);\n>>\n>> Which I guess is two questions:\n>>\n>>  A. do we need the cookie, or can we re-use the fd of the last thing we\n>>     write?\n>\n> We can re-use the FD of the last thing we write, but that results in a\n> tricker API which\n> is more intrusive on callers. I was originally using a lockfile, but\n> found a usage where\n> there was no lockfile in unpack-objects.\n\nOk, so it's something we could do, but passing down 2-3 functions to\nobject-file.c was a hassle.\n\nI tried to hack that up earlier and found that it wasn't *too\nbad*. I.e. we'd pass some \"flags\" about our intent, and amend various\nfunctions to take \"don't close this one\" and pass up the fd (or even do\nthat as a global).\n\nIn any case, having the commit message clearly document what's needed\nfor what & what's essential & just shortcut taken for the convenience of\nthe current implementation would be really useful.\n\nThen we can always e.g. change this later to just do the the fsync() on\nthe last of N we write.\n\n[Ran out of time, sorry]\n"},{"id":"451768","messageId":"xmqqv8w7vxng.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdfCJy68z0bNrhSmwo_uEVa6=y4V1dY0kZDq7JOTUD+6iQ@mail.gmail.com","subject":"Re: [PATCH v2 0/7] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T20:49:39Z","receivedAt":"2022-03-21T20:49:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> In any case, I've applied them on 0cac37f38f9 and then re-applied\n>> the result on top of fd008b1442 (i.e. the same base as the previous\n>> round was queued), which, with the magic of \"am -3\", applied\n>> cleanly.  Double checking the result was also simple (i.e. the tip of\n>> such an application on top of fd008b1442 can be merged with\n>> 0cac37f38f9 and the result should be identical to the result of\n>> applying them directly on top of 0cac37f38f9) and seems to have\n>> produced the right result.\n>>\n>> \\Thanks.\n>\n> Thanks Junio.  I was worried about how to properly represent the dependency\n> between these two in-flight branches without waiting for ns/core-fsyncmethod to\n> get into next.   Now ns/core-fsyncmethod appears to be there, so I'm assuming\n> that branch should have a stable OID until the end of the cycle.\n>\n> Should I base future versions of this series on the tip of\n> ns/core-fsyncmethod, or\n> on the merge point between that branch and 'next'?\n\nPlease base it on fd008b1442 (i.e. the same base as this and the\nprevious round was queued on), unless there is a strong reason to\nrebase elsewhere.\n\nThanks.\n"},{"id":"451772","messageId":"CANQDOdfa3SOee929pHmUBjQTgFa9hMHtX5kMZ45NaL+Xd6w+Rg@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqzgljyz34.fsf@gitster.g","subject":"Re: [PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T22:18:10Z","receivedAt":"2022-03-21T22:46:57Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 10:50 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > diff --git a/builtin/update-index.c b/builtin/update-index.c\n> > index 75d646377cc..38e9d7e88cb 100644\n> > --- a/builtin/update-index.c\n> > +++ b/builtin/update-index.c\n> > @@ -5,6 +5,7 @@\n> >   */\n> >  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n> >  #include \"cache.h\"\n> > +#include \"bulk-checkin.h\"\n> >  #include \"config.h\"\n> >  #include \"lockfile.h\"\n> >  #include \"quote.h\"\n> > @@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >\n> >       the_index.updated_skipworktree = 1;\n> >\n> > +     /* we might be adding many objects to the object database */\n> > +     plug_bulk_checkin();\n> > +\n> >       /*\n> >        * Custom copy of parse_options() because we want to handle\n> >        * filename arguments as they come.\n> > @@ -1190,6 +1194,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >               strbuf_release(&buf);\n> >       }\n> >\n> > +     /* by now we must have added all of the new objects */\n> > +     unplug_bulk_checkin();\n>\n> I understand read-from-stdin code path would be worth plugging, but\n> the list of paths on the command line?  How many of them would one\n> fit?\n>\n\ndo_reupdate could touch all the files in the index.  Also one can pass a\ndirectory, and re-add all files under the directory.\n\n> Of course, the feeder may be expecting for the objects to appear in\n> the object store as it feeds the paths and will be utterly broken by\n> this change, as you mentioned in the proposed log message.  The\n> existing plug/unplug will change the behaviour by making the objects\n> sent to the packfile available only after getting unplugged.  This\n> series makes it even worse by making loose objects also unavailable\n> until unplug is called.\n>\n> So, it probably is safer and more sensible approach to introduce a\n> new command line option to allow the bulk checkin, and those who do\n> not care about the intermediate state to opt into the new feature.\n>\n\nI don't believe this usage is likely today. How would the feeder know when\nit can expect to find an object in the object directory after passing something\non stdin?  When fed via stdin, git-update-index will asynchronously add that\nobject to the object database, leaving no indication to the feeder of when it\nactually happens, aside from it happening before the git-update-index process\nterminates.  I used to have a comment here about the feeder being able to\nparse the --verbose output to get feedback from git-update-index, which\nwould be quite tricky. I thought it was unnecessarily detailed.\n\nThanks,\nNeeraj\n"},{"id":"451774","messageId":"CANQDOdeox=Wox94SW+oUgbLDdLZ+KOdw6AWdBzwFkR18xmaNtQ@mail.gmail.com","threadId":"57568","inReplyTo":"220321.865yo79wkf.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T22:09:25Z","receivedAt":"2022-03-21T22:59:22Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 8:04 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > The update-index functionality is used internally by 'git stash push' to\n> > setup the internal stashed commit.\n> >\n> > This change enables bulk-checkin for update-index infrastructure to\n> > speed up adding new objects to the object database by leveraging the\n> > batch fsync functionality.\n> >\n> > There is some risk with this change, since under batch fsync, the object\n> > files will be in a tmp-objdir until update-index is complete.  This\n> > usage is unlikely, since any tool invoking update-index and expecting to\n> > see objects would have to synchronize with the update-index process\n> > after passing it a file path.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  builtin/update-index.c | 6 ++++++\n> >  1 file changed, 6 insertions(+)\n> >\n> > diff --git a/builtin/update-index.c b/builtin/update-index.c\n> > index 75d646377cc..38e9d7e88cb 100644\n> > --- a/builtin/update-index.c\n> > +++ b/builtin/update-index.c\n> > @@ -5,6 +5,7 @@\n> >   */\n> >  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n> >  #include \"cache.h\"\n> > +#include \"bulk-checkin.h\"\n> >  #include \"config.h\"\n> >  #include \"lockfile.h\"\n> >  #include \"quote.h\"\n> > @@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >\n> >       the_index.updated_skipworktree = 1;\n> >\n> > +     /* we might be adding many objects to the object database */\n> > +     plug_bulk_checkin();\n> > +\n>\n> Shouldn't this be after parse_options_start()?\n\nDoes it make a difference?  Especially if we do the object dir creation lazily?\n"},{"id":"451810","messageId":"CANQDOdeQP7b1MefkGZKSuFb-kc9F01RPypxaowFwG-A+wyZGog@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqv8w7yyua.fsf@gitster.g","subject":"Re: [PATCH v2 4/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-21T23:02:20Z","receivedAt":"2022-03-21T23:13:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 10:55 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > The unpack-objects functionality is used by fetch, push, and fast-import\n> > to turn the transfered data into object database entries when there are\n> > fewer objects than the 'unpacklimit' setting.\n> >\n> > By enabling bulk-checkin when unpacking objects, we can take advantage\n> > of batched fsyncs.\n>\n> This feels confused in that we dispatch to unpack-objects (instead\n> of index-objects) only when the number of loose objects should not\n> matter from performance point of view, and bulk-checkin should shine\n> from performance point of view only when there are enough objects to\n> batch.\n>\n> Also if we ever add \"too many small loose objects is wasteful, let's\n> send them into a single 'batch pack'\" optimization, it would create\n> a funny situation where the caller sends the contents of a small\n> incoming packfile to unpack-objects, but the command chooses to\n> bunch them all together in a packfile anyway ;-)\n>\n> So, I dunno.\n>\n\nI'd be happy to just drop this patch.  I originally added it to answer Avarab's\nquestion: how does batch mode compare to packfiles? [1] [2].\n\n[1] https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/\n[2] https://lore.kernel.org/git/pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com/\n"},{"id":"451819","messageId":"220322.86v8w6opyv.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"CANQDOdeox=Wox94SW+oUgbLDdLZ+KOdw6AWdBzwFkR18xmaNtQ@mail.gmail.com","subject":"Re: [PATCH v2 3/7] update-index: use the bulk-checkin infrastructure","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T23:16:31Z","receivedAt":"2022-03-21T23:22:37Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 21 2022, Neeraj Singh wrote:\n\n> On Mon, Mar 21, 2022 at 8:04 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n>>\n>> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >\n>> > The update-index functionality is used internally by 'git stash push' to\n>> > setup the internal stashed commit.\n>> >\n>> > This change enables bulk-checkin for update-index infrastructure to\n>> > speed up adding new objects to the object database by leveraging the\n>> > batch fsync functionality.\n>> >\n>> > There is some risk with this change, since under batch fsync, the object\n>> > files will be in a tmp-objdir until update-index is complete.  This\n>> > usage is unlikely, since any tool invoking update-index and expecting to\n>> > see objects would have to synchronize with the update-index process\n>> > after passing it a file path.\n>> >\n>> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> > ---\n>> >  builtin/update-index.c | 6 ++++++\n>> >  1 file changed, 6 insertions(+)\n>> >\n>> > diff --git a/builtin/update-index.c b/builtin/update-index.c\n>> > index 75d646377cc..38e9d7e88cb 100644\n>> > --- a/builtin/update-index.c\n>> > +++ b/builtin/update-index.c\n>> > @@ -5,6 +5,7 @@\n>> >   */\n>> >  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>> >  #include \"cache.h\"\n>> > +#include \"bulk-checkin.h\"\n>> >  #include \"config.h\"\n>> >  #include \"lockfile.h\"\n>> >  #include \"quote.h\"\n>> > @@ -1110,6 +1111,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>> >\n>> >       the_index.updated_skipworktree = 1;\n>> >\n>> > +     /* we might be adding many objects to the object database */\n>> > +     plug_bulk_checkin();\n>> > +\n>>\n>> Shouldn't this be after parse_options_start()?\n>\n> Does it make a difference?  Especially if we do the object dir creation lazily?\n\nI think it won't matter for the machine, but it helps with readability\nto keep code like this as close to where it's used as possible.\n\nClose enough and we'd also spot the other bug I mentioned here,\ni.e. that we're setting this up where we're not writing objects at all\n:)\n"},{"id":"451824","messageId":"CANQDOdc+ENfKLxpQ1HZJPgzgK26DRZi2-qNkkn7B6n9qV_B-gg@mail.gmail.com","threadId":"57568","inReplyTo":"220321.86zgljnite.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-22T00:13:36Z","receivedAt":"2022-03-22T00:16:27Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 1:37 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Mar 21 2022, Neeraj Singh wrote:\n>\n> [Don't have time for a full reply, sorry, just something quick]\n>\n> > On Mon, Mar 21, 2022 at 9:52 AM Ævar Arnfjörð Bjarmason\n> > [...]\n> >> So, my question is still why the temporary object dir migration part of\n> >> this is needed.\n> >>\n> >> We are writing N loose object files, and we write those to temporary\n> >> names already.\n> >>\n> >> AFAIKT we could do all of this by doing the same\n> >> tmp/rename/sync_file_range dance on the main object store.\n> >>\n> >\n> > Why not the main object store? We want to maintain the invariant that any\n> > name in the main object store refers to a file that durably has the\n> > correct contents.\n> > If we do sync_file_range and then rename, and then crash, we now have a file\n> > in the main object store with some SHA name, whose contents may or may not\n> > match the SHA.  However, if we ensure an fsync happens before the rename,\n> > a crash at any point will leave us either with no file in the main\n> > object store or\n> > with a file that is durable on the disk.\n>\n> Ah, I see.\n>\n> Why does that matter? If the \"bulk\" mode works as advertised we might\n> have such a corrupt loose or pack file, but we won't have anything\n> referring to it as far as reachability goes.\n>\n> I'm aware that the various code paths that handle OID writing don't deal\n> too well with it in practice to say the least, which one can try with\n> say:\n>\n>     $ echo foo | git hash-object -w --stdin\n>     45b983be36b73c0788dc9cbcb76cbb80fc7bb057\n>     $ echo | sudo tee .git/objects/45/b983be36b73c0788dc9cbcb76cbb80fc7bb057\n>\n> I.e. \"fsck\", \"show\" etc. will all scream bloddy murder, and re-running\n> that hash-object again even returns successful (we see it's there\n> already, and think it's OK).\n>\n\nI was under the impression that in-practice a corrupt loose-object can create\npersistent problems in the repo for future commands, since we might not\naggressively verify that an existing file with a certain OID really is\nvalid when\nadding a new instance of the data with the same OID.\n\nIf you don't have an fsync barrier before producing the final\ncontent-addressable\nname, you can't reason about \"this operation happened before that operation,\"\nso it wouldn't really be valid to say that \"we won't have anything\nreferring to it as far\nas reachability goes.\"\n\nIt's entirely possible that you'd have trees pointing to other trees\nor blobs that aren't\nvalid, since data writes can be durable in any order. At this point,\nfuture attempts add\nthe same blobs or trees might silently drop the updates.  I'm betting that's why\ncore.fsyncObjectFiles was added in the first place, since someone\nobserved severe\npersistent consequences for this form of corruption.\n\n> But in any case, I think it would me much easier to both review and\n> reason about this code if these concerns were split up.\n>\n> I.e. things that want no fsync at all (I'd think especially so) might\n> want to have such updates serialized in this manner, and as Junio\n> pointed out making these things inseparable as you've done creates API\n> concerns & fallout that's got nothing to do with what we need for the\n> performance gains of the bulk checkin fsyncing technique,\n> e.g. concurrent \"update-index\" consumers not being able to assume\n> reported objects exist as soon as they're reported.\n>\n\nI want to explicitly not respond to this concern. I don't believe this\n100 line patch\ncan be usefully split.\n\n> >> Then instead of the \"bulk_fsync\" cookie file don't close() the last file\n> >> object file we write until we issue the fsync on it.\n> >>\n> >> But maybe this is all needed, I just can't understand from the commit\n> >> message why the \"bulk checkin\" part is being done.\n> >>\n> >> I think since we've been over this a few times without any success it\n> >> would really help to have some example of the smallest set of syscalls\n> >> to write a file like this safely. I.e. this is doing (pseudocode):\n> >>\n> >>     /* first the bulk path */\n> >>     open(\"bulk/x.tmp\");\n> >>     write(\"bulk/x.tmp\");\n> >>     sync_file_range(\"bulk/x.tmp\");\n> >>     close(\"bulk/x.tmp\");\n> >>     rename(\"bulk/x.tmp\", \"bulk/x\");\n> >>     open(\"bulk/y.tmp\");\n> >>     write(\"bulk/y.tmp\");\n> >>     sync_file_range(\"bulk/y.tmp\");\n> >>     close(\"bulk/y.tmp\");\n> >>     rename(\"bulk/y.tmp\", \"bulk/y\");\n> >>     /* Rename to \"real\" */\n> >>     rename(\"bulk/x\", x\");\n> >>     rename(\"bulk/y\", y\");\n> >>     /* sync a cookie */\n> >>     fsync(\"cookie\");\n> >>\n> >\n> > The '/* Rename to \"real\" */' and '/* sync a cookie */' steps are\n> > reversed in your above sequence. It should be\n>\n> Sorry.\n>\n> > 1: (for each file)\n> >     a) open\n> >     b) write\n> >     c) sync_file_range\n> >     d) close\n> >     e) rename in tmp_objdir  -- The rename step is not required for\n> > bulk-fsync. An earlier version of this series didn't do it, but\n> > Jeff King pointed out that it was required for concurrency:\n> > https://lore.kernel.org/all/YVOrikAl%2Fu5%2FVi61@coredump.intra.peff.net/\n>\n> Yes we definitely need the rename, I was wondering about why we needed\n> it 2x for each file, but that was answered above.\n>\n> >> And I'm asking why it's not:\n> >>\n> >>     /* Rename to \"real\" as we go */\n> >>     open(\"x.tmp\");\n> >>     write(\"x.tmp\");\n> >>     sync_file_range(\"x.tmp\");\n> >>     close(\"x.tmp\");\n> >>     rename(\"x.tmp\", \"x\");\n> >>     last_fd = open(\"y.tmp\"); /* don't close() the last one yet */\n> >>     write(\"y.tmp\");\n> >>     sync_file_range(\"y.tmp\");\n> >>     rename(\"y.tmp\", \"y\");\n> >>     /* sync a cookie */\n> >>     fsync(last_fd);\n> >>\n> >> Which I guess is two questions:\n> >>\n> >>  A. do we need the cookie, or can we re-use the fd of the last thing we\n> >>     write?\n> >\n> > We can re-use the FD of the last thing we write, but that results in a\n> > tricker API which\n> > is more intrusive on callers. I was originally using a lockfile, but\n> > found a usage where\n> > there was no lockfile in unpack-objects.\n>\n> Ok, so it's something we could do, but passing down 2-3 functions to\n> object-file.c was a hassle.\n>\n> I tried to hack that up earlier and found that it wasn't *too\n> bad*. I.e. we'd pass some \"flags\" about our intent, and amend various\n> functions to take \"don't close this one\" and pass up the fd (or even do\n> that as a global).\n>\n> In any case, having the commit message clearly document what's needed\n> for what & what's essential & just shortcut taken for the convenience of\n> the current implementation would be really useful.\n>\n> Then we can always e.g. change this later to just do the the fsync() on\n> the last of N we write.\n>\n\nI left a comment in the (now very long) commit message that indicates the\ndummy file is there to make the API simpler.\n\nThanks,\nNeeraj\n"},{"id":"451827","messageId":"CANQDOddc7TPY=BJRPOBMJudFFeEv3j3nUd9+Ewkpg3bwFLPVUA@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqpmmfxigx.fsf@gitster.g","subject":"Re: [PATCH v2 6/7] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-22T05:54:57Z","receivedAt":"2022-03-22T05:55:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 11:34 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Add test cases to exercise batch mode for:\n> >  * 'git add'\n>\n> I was wondering why the obviously safe and good candidate 'git add' is\n> not gaining plug/unplug pair in this series.  It is obviously safe,\n> unlike 'update-index', that nobody can interact with it, observe its\n> intermediate output, and expect anything from it.\n>\n> I think the stupid reason of the lack of new plug/unplug is because\n> we already had them, which is good ;-).\n>\n> >  * 'git stash'\n> >  * 'git update-index'\n>\n> As I said, I suspect that we'd want to do this safely by adding a\n> new option to \"update-index\" and passing it from \"stash\" which knows\n> that it does not care about the intermediate state.\n>\n> > These tests ensure that the added data winds up in the object database.\n>\n> In other words, \"git add $path; git rev-parse :$path\" (and its\n> cousins) would be happy?  Like new object files not left hanging in\n> a tentative object store etc. _after_ the commands finish.\n>\n> Good.\n>\n> > In this change we introduce a new test helper lib-unique-files.sh. The\n> > goal of this library is to create a tree of files that have different\n> > oids from any other files that may have been created in the current test\n> > repo. This helps us avoid missing validation of an object being added due\n> > to it already being in the repo.\n>\n> More on this below.\n\nTo me the idea of putting the 'why' into the commit message and the 'what' into\ncode comments makes sense, since I'd assume people looking into the history\ncare about the why, but people making future changes would read the\ndocumentation\nin the comments for e.g. lib-unique-files.\n\n>\n> > We aren't actually issuing any fsyncs in these tests, since\n> > GIT_TEST_FSYNC is 0, but we still exercise all of the tmp_objdir logic\n> > in bulk-checkin.\n>\n> Shouldn't we manually override that, if it matters?\n> Not a suggestion but a question.\n\nI manually override it for the performance tests.  I think it's\nsensible to manually override\nthis variable for the small number of tests added in this commit so\nthat we can exercise\nthe underlying system calls, so I'll do that.\n\n> > +# Create multiple files with unique contents. Takes the number of\n> > +# directories, the number of files in each directory, and the base\n> > +# directory.\n>\n> This is more honest, compared to the claim made in the proposed log\n> message, in that the uniqueness guarantee is only among the files\n> created by this helper.  If we created other test contents without\n> using this helper, that may crash with the ones created here.\n>\n\nI've revised this comment to indicate that the files are only unique\nwithin this test run.\n\n> > +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> > +#                                     each in my_dir, all with unique\n> > +#                                     contents.\n> > +\n> > +test_create_unique_files() {\n>\n> Style.  SP on both sides of ().  I.e.\n>\n>         test_create_unique_files () {\n>\n\nFixed\n\n> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n> > +\n> > +     local dirs=$1\n> > +     local files=$2\n> > +     local basedir=$3\n> > +     local counter=0\n> > +     test_tick\n> > +     local basedata=$test_tick\n>\n> I am not sure if consumption and reliance on tick is a wise thing.\n> $basedir must be unique across all the other directories in this\n> test repository (there is no other $basedir)---can't we key\n> uniqueness off of it?\n\nIn the performance tests, we create sets of files with the same basedir\nbut we want the files to have different contents since we don't blow away\nthe repo between tests.  The current approach still generates uniqueness\nif someone simply copy/pastes a test_create_unique_files invocation, rather\nthan subtly failing to make new objects.\n\n> > +     rm -rf $basedir\n>\n> Can $basedir have any $IFS character in it?  We should \"$quote\" it.\n\nFixed.\n\n>\n> > +     for i in $(test_seq $dirs)\n> > +     do\n> > +             local dir=$basedir/dir$i\n> > +\n> > +             mkdir -p \"$dir\"\n> > +             for j in $(test_seq $files)\n> > +             do\n> > +                     counter=$((counter + 1))\n> > +                     echo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n>\n> An extra SP before \">\"?\n>\n\nFixed.\n\n> > +             done\n> > +     done\n> > +}\n>\n> There is no &&- cascade here, and we expect nothing in this to\n> fail.  Is that sensible?\n>\n\nI apologize, there's a lot of subtlety about UNIX shell scripting that\nI simply do not know.   I put an '&&' chain in, but I might still have it wrong.\n\n> > +test_expect_success 'git add: core.fsyncmethod=batch' \"\n> > +     test_create_unique_files 2 4 fsync-files &&\n> > +     git $BATCH_CONFIGURATION add -- ./fsync-files/ &&\n> > +     rm -f fsynced_files &&\n> > +     git ls-files --stage fsync-files/ > fsynced_files &&\n>\n> Style.  No SP between redirection operator and its target.  I.e.\n>\n>         git ls-files --stage fsync-files/ >fsynced_files &&\n>\n> Mixture of names-with-dash and name_with_understore looks somewhat\n> irritating.\n>\n\nWill switch to underscores for both. Also will rename the base dir to be more\ndistinct from the list of files that we expect to see.\n\n> > +     test_line_count = 8 fsynced_files &&\n>\n> The magic \"8\" matches \"2 4\" we saw earlier for create_unique_files?\n>\n\nWill comment to explain the 8.\n\n> > +     awk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n>\n> A test helper that takes the name of a file that has \"ls-files -s\" output\n> may prove to be useful.  I dunno.\n>\n> > diff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\n> > index a11d61206ad..8e2f73cc68f 100755\n> > --- a/t/t5300-pack-object.sh\n> > +++ b/t/t5300-pack-object.sh\n> > @@ -162,23 +162,25 @@ test_expect_success 'pack-objects with bogus arguments' '\n> >\n> >  check_unpack () {\n> >       test_when_finished \"rm -rf git2\" &&\n> > -     git init --bare git2 &&\n> > -     git -C git2 unpack-objects -n <\"$1\".pack &&\n> > -     git -C git2 unpack-objects <\"$1\".pack &&\n> > -     (cd .git && find objects -type f -print) |\n> > -     while read path\n> > -     do\n> > -             cmp git2/$path .git/$path || {\n> > -                     echo $path differs.\n> > -                     return 1\n> > -             }\n> > -     done\n> > +     git $2 init --bare git2 &&\n> > +     (\n> > +             git $2 -C git2 unpack-objects -n <\"$1\".pack &&\n> > +             git $2 -C git2 unpack-objects <\"$1\".pack &&\n> > +             git $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n> > +     ) <obj-list >current &&\n> > +     cmp obj-list current\n> >  }\n>\n> I think the change from the old \"the existence and the contents of\n> the object files must all match\" to the new \"cat-file should say\n> that the objects we expect to exist indeed do\" is not a bad thing.\n>\n> We used to only depend on the contents of the provided packfile but\n> now we assume that obj-list file gives us the list of objects.  Is\n> that sensible?  I somehow do not think so.  Don't we have the\n> corresponding \"$1.idx\" that we can feed to \"git show-index\", e.g.\n>\n\nI believe that \"obj-list\" is in some sense more authoritative,\nassuming arbitrary\nfuture bugs in the pack implementation.  We make sure that the pack we were\nhanded isn't missing any objects.  It makes sense to accept obj-list from\nthe outside so that check_unpack itself doesn't depend on that file name.\nBut the previous code wouldn't catch a bug where the pack code mistakenly drops\nan object.\n\n>         git show-index <\"$1.pack\" >expect.full &&\n>         cut -d\" \" -f2 >expect <expect.full &&\n>         ... your test in \"$2\", but feeding expect instead of obj-list ...\n>         test_cmp expect actual\n>\n> Also make sure you quote whatever is coming from outside, even if\n> you happen to call the helper with tokens that do not need quoting\n> in the current code.  It is a good discipline to help readers.\n>\n> Thanks.\n\nI'll take another pass over the shell code in the morning to make sure\nI'm following all the conventions and recommendations.  Apologize\nfor the set of mistakes. In terms of shell, I am a rookie.\n\nThanks,\n-Neeraj\n"},{"id":"451837","messageId":"220322.86mthinxnn.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"CANQDOdc+ENfKLxpQ1HZJPgzgK26DRZi2-qNkkn7B6n9qV_B-gg@mail.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-22T08:52:49Z","receivedAt":"2022-03-22T09:29:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 21 2022, Neeraj Singh wrote:\n\n> On Mon, Mar 21, 2022 at 1:37 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Mon, Mar 21 2022, Neeraj Singh wrote:\n>>\n>> [Don't have time for a full reply, sorry, just something quick]\n>>\n>> > On Mon, Mar 21, 2022 at 9:52 AM Ævar Arnfjörð Bjarmason\n>> > [...]\n>> >> So, my question is still why the temporary object dir migration part of\n>> >> this is needed.\n>> >>\n>> >> We are writing N loose object files, and we write those to temporary\n>> >> names already.\n>> >>\n>> >> AFAIKT we could do all of this by doing the same\n>> >> tmp/rename/sync_file_range dance on the main object store.\n>> >>\n>> >\n>> > Why not the main object store? We want to maintain the invariant that any\n>> > name in the main object store refers to a file that durably has the\n>> > correct contents.\n>> > If we do sync_file_range and then rename, and then crash, we now have a file\n>> > in the main object store with some SHA name, whose contents may or may not\n>> > match the SHA.  However, if we ensure an fsync happens before the rename,\n>> > a crash at any point will leave us either with no file in the main\n>> > object store or\n>> > with a file that is durable on the disk.\n>>\n>> Ah, I see.\n>>\n>> Why does that matter? If the \"bulk\" mode works as advertised we might\n>> have such a corrupt loose or pack file, but we won't have anything\n>> referring to it as far as reachability goes.\n>>\n>> I'm aware that the various code paths that handle OID writing don't deal\n>> too well with it in practice to say the least, which one can try with\n>> say:\n>>\n>>     $ echo foo | git hash-object -w --stdin\n>>     45b983be36b73c0788dc9cbcb76cbb80fc7bb057\n>>     $ echo | sudo tee .git/objects/45/b983be36b73c0788dc9cbcb76cbb80fc7bb057\n>>\n>> I.e. \"fsck\", \"show\" etc. will all scream bloddy murder, and re-running\n>> that hash-object again even returns successful (we see it's there\n>> already, and think it's OK).\n>>\n>\n> I was under the impression that in-practice a corrupt loose-object can create\n> persistent problems in the repo for future commands, since we might not\n> aggressively verify that an existing file with a certain OID really is\n> valid when\n> adding a new instance of the data with the same OID.\n\nYes, it can. As the hash-object case shows we don't even check at all.\n\nFor \"incoming push\" we *will* notice, but will just uselessly error\nout.\n\nI actually had some patches a while ago to turn off our own home-grown\nSHA-1 collision checking.\n\nIt had the nice side effect of making it easier to recover from loose\nobject corruption, since you could (re-)push the corrupted OID as a\nPACK, we wouldn't check (and die) on the bad loose object, and since we\ntake a PACK over LOOSE we'd recover:\nhttps://lore.kernel.org/git/20181028225023.26427-5-avarab@gmail.com/\n\n> If you don't have an fsync barrier before producing the final\n> content-addressable\n> name, you can't reason about \"this operation happened before that operation,\"\n> so it wouldn't really be valid to say that \"we won't have anything\n> referring to it as far\n> as reachability goes.\"\n\nThat's correct, but we're discussing a feature that *does have* that\nfsync barrier. So if we get an error while writing the loose objects\nbefore the \"cookie\" fsync we'll presumably error out. That'll then be\nfollowed by an fsync() of whatever makes the objects reachable.\n\n> It's entirely possible that you'd have trees pointing to other trees\n> or blobs that aren't\n> valid, since data writes can be durable in any order. At this point,\n> future attempts add\n> the same blobs or trees might silently drop the updates.  I'm betting that's why\n> core.fsyncObjectFiles was added in the first place, since someone\n> observed severe\n> persistent consequences for this form of corruption.\n\nWell, you can see Linus's original rant-as-documentation for why we\nadded it :) I.e. the original git implementation made some heavy\nlinux-FS assumption about the order of writes and an fsync() flushing\nany previous writes, which wasn't portable.\n\n>> But in any case, I think it would me much easier to both review and\n>> reason about this code if these concerns were split up.\n>>\n>> I.e. things that want no fsync at all (I'd think especially so) might\n>> want to have such updates serialized in this manner, and as Junio\n>> pointed out making these things inseparable as you've done creates API\n>> concerns & fallout that's got nothing to do with what we need for the\n>> performance gains of the bulk checkin fsyncing technique,\n>> e.g. concurrent \"update-index\" consumers not being able to assume\n>> reported objects exist as soon as they're reported.\n>>\n>\n> I want to explicitly not respond to this concern. I don't believe this\n> 100 line patch\n> can be usefully split.\n\nLeaving \"usefully\" aside for a second (since that's subjective), it\nclearly \"can\". I just tried this on top of \"seen\":\n\n\tdiff --git a/bulk-checkin.c b/bulk-checkin.c\n\tindex a702e0ff203..9e994c4d6ae 100644\n\t--- a/bulk-checkin.c\n\t+++ b/bulk-checkin.c\n\t@@ -9,15 +9,12 @@\n\t #include \"pack.h\"\n\t #include \"strbuf.h\"\n\t #include \"string-list.h\"\n\t-#include \"tmp-objdir.h\"\n\t #include \"packfile.h\"\n\t #include \"object-store.h\"\n\t \n\t static int bulk_checkin_plugged;\n\t static int needs_batch_fsync;\n\t \n\t-static struct tmp_objdir *bulk_fsync_objdir;\n\t-\n\t static struct bulk_checkin_state {\n\t \tchar *pack_tmp_name;\n\t \tstruct hashfile *f;\n\t@@ -110,11 +107,6 @@ static void do_batch_fsync(void)\n\t \t\tstrbuf_release(&temp_path);\n\t \t\tneeds_batch_fsync = 0;\n\t \t}\n\t-\n\t-\tif (bulk_fsync_objdir) {\n\t-\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n\t-\t\tbulk_fsync_objdir = NULL;\n\t-\t}\n\t }\n\t \n\t static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n\t@@ -321,7 +313,6 @@ void fsync_loose_object_bulk_checkin(int fd)\n\t \t */\n\t \tif (bulk_checkin_plugged &&\n\t \t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n\t-\t\tassert(bulk_fsync_objdir);\n\t \t\tif (!needs_batch_fsync)\n\t \t\t\tneeds_batch_fsync = 1;\n\t \t} else {\n\t@@ -343,19 +334,6 @@ int index_bulk_checkin(struct object_id *oid,\n\t void plug_bulk_checkin(void)\n\t {\n\t \tassert(!bulk_checkin_plugged);\n\t-\n\t-\t/*\n\t-\t * A temporary object directory is used to hold the files\n\t-\t * while they are not fsynced.\n\t-\t */\n\t-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n\t-\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n\t-\t\tif (!bulk_fsync_objdir)\n\t-\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n\t-\n\t-\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n\t-\t}\n\t-\n\t \tbulk_checkin_plugged = 1;\n\t }\n\nAnd then tried running:\n\n    $ GIT_PERF_MAKE_OPTS='CFLAGS=-O3' ./run HEAD~ HEAD -- p3900-stash.sh\t\n\nAnd got:\n    \n    Test                                              HEAD~              HEAD\n    --------------------------------------------------------------------------------------------\n    3900.2: stash 500 files (object_fsyncing=false)   0.56(0.08+0.09)    0.60(0.08+0.08) +7.1%\n    3900.4: stash 500 files (object_fsyncing=true)    14.50(0.07+0.15)   17.13(0.10+0.12) +18.1%\n    3900.6: stash 500 files (object_fsyncing=batch)   1.14(0.08+0.11)    1.03(0.08+0.10) -9.6%\n\nNow, I really don't trust that perf run to say anything except these\nbeing in the same ballpark, but it's clearly going to be a *bit* faster\nsince we'll be doing fewer IOps.\n\nAs to \"usefully\" I really do get what you're saying that you only find\nthese useful when you combine the two because you'd like to have 100%\nsafety, and that's fair enough.\n\nBut since we are going to have a knob to turn off fsyncing entirely, and\nwe have this \"bulk\" mode which requires you to carefully reason about\nyour FS semantics to ascertain safety the performance/safety trade-off\nis clearly something that's useful to have tweaks for.\n\nAnd with \"bulk\" the concern about leaving behind stray corrupt objects\nis entirely orthagonal to corcerns about losing a ref update, which is\nthe main thing we're worried about.\n\nI also don't see how even if you're arguing that nobody would want one\nwithout the other because everyone who cares about \"bulk\" cares about\nthis stray-corrupt-loose-but-no-ref-update case, how it has any business\nbeing tied up in the \"bulk\" mode as far as the implementation goes.\n\nThat's because the same edge case is exposed by\ncore.fsyncObjectFiles=false for those who are assuming the initial\n\"ordered\" semantics.\n\nI.e. if we're saying that before we write the ref we'd like to not\nexpose the WIP objects in the primary object store because they're not\nfsync'd yet, how is that mode different than \"bulk\" if we crash while\ndoing that operation (before the eventual fsync()).\n\nSo I really think it's much better to split these concerns up.\n\nI think even if you insist on the same end-state it makes the patch\nprogression much *easier* to reason about. We'd then solve one problem\nat a time, and start with a commit where just the semantics that are\nunique to \"bulk\" are implemented, with nothing else conflated with\nthose.\n\n> [...]\n>> Ok, so it's something we could do, but passing down 2-3 functions to\n>> object-file.c was a hassle.\n>>\n>> I tried to hack that up earlier and found that it wasn't *too\n>> bad*. I.e. we'd pass some \"flags\" about our intent, and amend various\n>> functions to take \"don't close this one\" and pass up the fd (or even do\n>> that as a global).\n>>\n>> In any case, having the commit message clearly document what's needed\n>> for what & what's essential & just shortcut taken for the convenience of\n>> the current implementation would be really useful.\n>>\n>> Then we can always e.g. change this later to just do the the fsync() on\n>> the last of N we write.\n>>\n>\n> I left a comment in the (now very long) commit message that indicates the\n> dummy file is there to make the API simpler.\n\nIn terms of more understandable progression I also think this series\nwould be much easier to understand if it converted one caller without\nneeding the \"cookie\" where doing so is easy, e.g. the unpack-objects.c\ncaller where we're processing nr_objects, so we can just pass down a\nflag to do the fsync() for i == nr_objects.\n\nThat'll then clearly show that the whole business of having the global\nstate on the side is just a replacement for passing down such a flag.\n"},{"id":"451924","messageId":"CANQDOde2OG8fVSM1hQE3FBmzWy5FkgQCWAUYhFztB8UGFyJELg@mail.gmail.com","threadId":"57568","inReplyTo":"220322.86mthinxnn.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-22T20:05:05Z","receivedAt":"2022-03-22T20:05:24Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 22, 2022 at 2:29 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Mar 21 2022, Neeraj Singh wrote:\n>\n> > On Mon, Mar 21, 2022 at 1:37 PM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >>\n> >> On Mon, Mar 21 2022, Neeraj Singh wrote:\n> >>\n> >> [Don't have time for a full reply, sorry, just something quick]\n> >>\n> >> > On Mon, Mar 21, 2022 at 9:52 AM Ævar Arnfjörð Bjarmason\n> >> > [...]\n> >> >> So, my question is still why the temporary object dir migration part of\n> >> >> this is needed.\n> >> >>\n> >> >> We are writing N loose object files, and we write those to temporary\n> >> >> names already.\n> >> >>\n> >> >> AFAIKT we could do all of this by doing the same\n> >> >> tmp/rename/sync_file_range dance on the main object store.\n> >> >>\n> >> >\n> >> > Why not the main object store? We want to maintain the invariant that any\n> >> > name in the main object store refers to a file that durably has the\n> >> > correct contents.\n> >> > If we do sync_file_range and then rename, and then crash, we now have a file\n> >> > in the main object store with some SHA name, whose contents may or may not\n> >> > match the SHA.  However, if we ensure an fsync happens before the rename,\n> >> > a crash at any point will leave us either with no file in the main\n> >> > object store or\n> >> > with a file that is durable on the disk.\n> >>\n> >> Ah, I see.\n> >>\n> >> Why does that matter? If the \"bulk\" mode works as advertised we might\n> >> have such a corrupt loose or pack file, but we won't have anything\n> >> referring to it as far as reachability goes.\n> >>\n> >> I'm aware that the various code paths that handle OID writing don't deal\n> >> too well with it in practice to say the least, which one can try with\n> >> say:\n> >>\n> >>     $ echo foo | git hash-object -w --stdin\n> >>     45b983be36b73c0788dc9cbcb76cbb80fc7bb057\n> >>     $ echo | sudo tee .git/objects/45/b983be36b73c0788dc9cbcb76cbb80fc7bb057\n> >>\n> >> I.e. \"fsck\", \"show\" etc. will all scream bloddy murder, and re-running\n> >> that hash-object again even returns successful (we see it's there\n> >> already, and think it's OK).\n> >>\n> >\n> > I was under the impression that in-practice a corrupt loose-object can create\n> > persistent problems in the repo for future commands, since we might not\n> > aggressively verify that an existing file with a certain OID really is\n> > valid when\n> > adding a new instance of the data with the same OID.\n>\n> Yes, it can. As the hash-object case shows we don't even check at all.\n>\n> For \"incoming push\" we *will* notice, but will just uselessly error\n> out.\n>\n> I actually had some patches a while ago to turn off our own home-grown\n> SHA-1 collision checking.\n>\n> It had the nice side effect of making it easier to recover from loose\n> object corruption, since you could (re-)push the corrupted OID as a\n> PACK, we wouldn't check (and die) on the bad loose object, and since we\n> take a PACK over LOOSE we'd recover:\n> https://lore.kernel.org/git/20181028225023.26427-5-avarab@gmail.com/\n>\n> > If you don't have an fsync barrier before producing the final\n> > content-addressable\n> > name, you can't reason about \"this operation happened before that operation,\"\n> > so it wouldn't really be valid to say that \"we won't have anything\n> > referring to it as far\n> > as reachability goes.\"\n>\n> That's correct, but we're discussing a feature that *does have* that\n> fsync barrier. So if we get an error while writing the loose objects\n> before the \"cookie\" fsync we'll presumably error out. That'll then be\n> followed by an fsync() of whatever makes the objects reachable.\n>\n\nBecause we have a content-addressable store which generally trusts\nits contents are valid (at least when adding new instances of the same\ncontent), the mere existence of a loose-object with a certain name is\nenough to make it \"reachable\" to future operations, even if there are\nno other immediate ways to get to that object.\n\n> > It's entirely possible that you'd have trees pointing to other trees\n> > or blobs that aren't\n> > valid, since data writes can be durable in any order. At this point,\n> > future attempts add\n> > the same blobs or trees might silently drop the updates.  I'm betting that's why\n> > core.fsyncObjectFiles was added in the first place, since someone\n> > observed severe\n> > persistent consequences for this form of corruption.\n>\n> Well, you can see Linus's original rant-as-documentation for why we\n> added it :) I.e. the original git implementation made some heavy\n> linux-FS assumption about the order of writes and an fsync() flushing\n> any previous writes, which wasn't portable.\n>\n> >> But in any case, I think it would me much easier to both review and\n> >> reason about this code if these concerns were split up.\n> >>\n> >> I.e. things that want no fsync at all (I'd think especially so) might\n> >> want to have such updates serialized in this manner, and as Junio\n> >> pointed out making these things inseparable as you've done creates API\n> >> concerns & fallout that's got nothing to do with what we need for the\n> >> performance gains of the bulk checkin fsyncing technique,\n> >> e.g. concurrent \"update-index\" consumers not being able to assume\n> >> reported objects exist as soon as they're reported.\n> >>\n> >\n> > I want to explicitly not respond to this concern. I don't believe this\n> > 100 line patch\n> > can be usefully split.\n>\n> Leaving \"usefully\" aside for a second (since that's subjective), it\n> clearly \"can\". I just tried this on top of \"seen\":\n>\n>         diff --git a/bulk-checkin.c b/bulk-checkin.c\n>         index a702e0ff203..9e994c4d6ae 100644\n>         --- a/bulk-checkin.c\n>         +++ b/bulk-checkin.c\n>         @@ -9,15 +9,12 @@\n>          #include \"pack.h\"\n>          #include \"strbuf.h\"\n>          #include \"string-list.h\"\n>         -#include \"tmp-objdir.h\"\n>          #include \"packfile.h\"\n>          #include \"object-store.h\"\n>\n>          static int bulk_checkin_plugged;\n>          static int needs_batch_fsync;\n>\n>         -static struct tmp_objdir *bulk_fsync_objdir;\n>         -\n>          static struct bulk_checkin_state {\n>                 char *pack_tmp_name;\n>                 struct hashfile *f;\n>         @@ -110,11 +107,6 @@ static void do_batch_fsync(void)\n>                         strbuf_release(&temp_path);\n>                         needs_batch_fsync = 0;\n>                 }\n>         -\n>         -       if (bulk_fsync_objdir) {\n>         -               tmp_objdir_migrate(bulk_fsync_objdir);\n>         -               bulk_fsync_objdir = NULL;\n>         -       }\n>          }\n>\n>          static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>         @@ -321,7 +313,6 @@ void fsync_loose_object_bulk_checkin(int fd)\n>                  */\n>                 if (bulk_checkin_plugged &&\n>                     git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n>         -               assert(bulk_fsync_objdir);\n>                         if (!needs_batch_fsync)\n>                                 needs_batch_fsync = 1;\n>                 } else {\n>         @@ -343,19 +334,6 @@ int index_bulk_checkin(struct object_id *oid,\n>          void plug_bulk_checkin(void)\n>          {\n>                 assert(!bulk_checkin_plugged);\n>         -\n>         -       /*\n>         -        * A temporary object directory is used to hold the files\n>         -        * while they are not fsynced.\n>         -        */\n>         -       if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n>         -               bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n>         -               if (!bulk_fsync_objdir)\n>         -                       die(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n>         -\n>         -               tmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n>         -       }\n>         -\n>                 bulk_checkin_plugged = 1;\n>          }\n>\n> And then tried running:\n>\n>     $ GIT_PERF_MAKE_OPTS='CFLAGS=-O3' ./run HEAD~ HEAD -- p3900-stash.sh\n>\n> And got:\n>\n>     Test                                              HEAD~              HEAD\n>     --------------------------------------------------------------------------------------------\n>     3900.2: stash 500 files (object_fsyncing=false)   0.56(0.08+0.09)    0.60(0.08+0.08) +7.1%\n>     3900.4: stash 500 files (object_fsyncing=true)    14.50(0.07+0.15)   17.13(0.10+0.12) +18.1%\n>     3900.6: stash 500 files (object_fsyncing=batch)   1.14(0.08+0.11)    1.03(0.08+0.10) -9.6%\n>\n> Now, I really don't trust that perf run to say anything except these\n> being in the same ballpark, but it's clearly going to be a *bit* faster\n> since we'll be doing fewer IOps.\n>\n> As to \"usefully\" I really do get what you're saying that you only find\n> these useful when you combine the two because you'd like to have 100%\n> safety, and that's fair enough.\n>\n> But since we are going to have a knob to turn off fsyncing entirely, and\n> we have this \"bulk\" mode which requires you to carefully reason about\n> your FS semantics to ascertain safety the performance/safety trade-off\n> is clearly something that's useful to have tweaks for.\n>\n> And with \"bulk\" the concern about leaving behind stray corrupt objects\n> is entirely orthagonal to corcerns about losing a ref update, which is\n> the main thing we're worried about.\n>\n> I also don't see how even if you're arguing that nobody would want one\n> without the other because everyone who cares about \"bulk\" cares about\n> this stray-corrupt-loose-but-no-ref-update case, how it has any business\n> being tied up in the \"bulk\" mode as far as the implementation goes.\n>\n> That's because the same edge case is exposed by\n> core.fsyncObjectFiles=false for those who are assuming the initial\n> \"ordered\" semantics.\n>\n> I.e. if we're saying that before we write the ref we'd like to not\n> expose the WIP objects in the primary object store because they're not\n> fsync'd yet, how is that mode different than \"bulk\" if we crash while\n> doing that operation (before the eventual fsync()).\n>\n> So I really think it's much better to split these concerns up.\n>\n> I think even if you insist on the same end-state it makes the patch\n> progression much *easier* to reason about. We'd then solve one problem\n> at a time, and start with a commit where just the semantics that are\n> unique to \"bulk\" are implemented, with nothing else conflated with\n> those.\n\nOn Windows, where we want to have a consistent ODB by default, I'm\nadding a faster way\nto achieve that safety. No user is asking for a world where we are\ndoing half the\nwork to make a consistent ODB but not the other half.\n\nThis one patch works holistically to provide the full batch safety\nfeature, and splitting it\ninto two patches (which in the new version wouldn't be as clean as\nyou've done it above)\ndoesn't make the correctness of the whole thing more reviewable.  In\nfact it's less reviewable\nsince the fsync and objdir migration are in two separate patches and a\nfuture historian\nwouldn't get as clear of a picture of the whole mechanism.\n\n> > [...]\n> >> Ok, so it's something we could do, but passing down 2-3 functions to\n> >> object-file.c was a hassle.\n> >>\n> >> I tried to hack that up earlier and found that it wasn't *too\n> >> bad*. I.e. we'd pass some \"flags\" about our intent, and amend various\n> >> functions to take \"don't close this one\" and pass up the fd (or even do\n> >> that as a global).\n> >>\n> >> In any case, having the commit message clearly document what's needed\n> >> for what & what's essential & just shortcut taken for the convenience of\n> >> the current implementation would be really useful.\n> >>\n> >> Then we can always e.g. change this later to just do the the fsync() on\n> >> the last of N we write.\n> >>\n> >\n> > I left a comment in the (now very long) commit message that indicates the\n> > dummy file is there to make the API simpler.\n>\n> In terms of more understandable progression I also think this series\n> would be much easier to understand if it converted one caller without\n> needing the \"cookie\" where doing so is easy, e.g. the unpack-objects.c\n> caller where we're processing nr_objects, so we can just pass down a\n> flag to do the fsync() for i == nr_objects.\n>\n> That'll then clearly show that the whole business of having the global\n> state on the side is just a replacement for passing down such a flag.\n\nThat seems appropriate for our mailing list discussion, but I don't see\nhow it helps the patch series, because we'd be doing work to fsync\nthe final object and then reversing that work when producing the final\nend state, which uses the dummy file.\n\nThanks,\nNeeraj\n"},{"id":"451928","messageId":"CANQDOdcpvv3p8N07q29GekaJ=hZgTjfp=w-Ky1PQCqqRmr8JyQ@mail.gmail.com","threadId":"57568","inReplyTo":"CANQDOdeQP7b1MefkGZKSuFb-kc9F01RPypxaowFwG-A+wyZGog@mail.gmail.com","subject":"Re: [PATCH v2 4/7] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-22T20:54:54Z","receivedAt":"2022-03-22T20:55:12Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 21, 2022 at 4:02 PM Neeraj Singh <nksingh85@gmail.com> wrote:\n>\n> On Mon, Mar 21, 2022 at 10:55 AM Junio C Hamano <gitster@pobox.com> wrote:\n> >\n> > \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >\n> > > From: Neeraj Singh <neerajsi@microsoft.com>\n> > >\n> > > The unpack-objects functionality is used by fetch, push, and fast-import\n> > > to turn the transfered data into object database entries when there are\n> > > fewer objects than the 'unpacklimit' setting.\n> > >\n> > > By enabling bulk-checkin when unpacking objects, we can take advantage\n> > > of batched fsyncs.\n> >\n> > This feels confused in that we dispatch to unpack-objects (instead\n> > of index-objects) only when the number of loose objects should not\n> > matter from performance point of view, and bulk-checkin should shine\n> > from performance point of view only when there are enough objects to\n> > batch.\n> >\n> > Also if we ever add \"too many small loose objects is wasteful, let's\n> > send them into a single 'batch pack'\" optimization, it would create\n> > a funny situation where the caller sends the contents of a small\n> > incoming packfile to unpack-objects, but the command chooses to\n> > bunch them all together in a packfile anyway ;-)\n> >\n> > So, I dunno.\n> >\n>\n> I'd be happy to just drop this patch.  I originally added it to answer Avarab's\n> question: how does batch mode compare to packfiles? [1] [2].\n>\n> [1] https://lore.kernel.org/git/87mtp5cwpn.fsf@evledraar.gmail.com/\n> [2] https://lore.kernel.org/git/pull.1076.v5.git.git.1632514331.gitgitgadget@gmail.com/\n\nWell looking back again at the spreadsheet [3], at 90 objects, which\nis below the\ndefault transfer.unpackLimit, we see a 3x difference in performance\nbetween batch\nmode and the default fsync mode.  That's a different interaction class\n(230 ms versus 760 ms).\n\nI'll include a small table in the commit description with these\nperformance numbers to\nhelp justify it.\n\n[3] https://docs.google.com/spreadsheets/d/1uxMBkEXFFnQ1Y3lXKqcKpw6Mq44BzhpCAcPex14T-QQ\n"},{"id":"451939","messageId":"RFC-patch-2.7-00dbffc2331-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 2/7] unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:31Z","receivedAt":"2022-03-23T03:48:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"In preparation for making the bulk-checkin.c logic operate from\nobject-file.c itself in some common cases let's add\nHASH_N_OBJECTS{,_{FIRST,LAST}} flags.\n\nThis will allow us to adjust for-loops that add N objects to just pass\ndown whether they have >1 objects (HASH_N_OBJECTS), as well as passing\ndown flags for whether we have the first or last object.\n\nWe'll thus be able to drive any sort of batch-object mechanism from\nwrite_object_file_flags() directly, which until now didn't know if it\nwas doing one object, or some arbitrary N.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c | 60 +++++++++++++++++++++++-----------------\n cache.h                  |  3 ++\n 2 files changed, 37 insertions(+), 26 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex c55b6616aed..ec40c6fd966 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -233,7 +233,8 @@ static void write_rest(void)\n }\n \n static void added_object(unsigned nr, enum object_type type,\n-\t\t\t void *data, unsigned long size);\n+\t\t\t void *data, unsigned long size,\n+\t\t\t unsigned oflags);\n \n /*\n  * Write out nr-th object from the list, now we know the contents\n@@ -241,21 +242,21 @@ static void added_object(unsigned nr, enum object_type type,\n  * to be checked at the end.\n  */\n static void write_object(unsigned nr, enum object_type type,\n-\t\t\t void *buf, unsigned long size)\n+\t\t\t void *buf, unsigned long size, unsigned oflags)\n {\n \tif (!strict) {\n-\t\tif (write_object_file(buf, size, type,\n-\t\t\t\t      &obj_list[nr].oid) < 0)\n+\t\tif (write_object_file_flags(buf, size, type,\n+\t\t\t\t      &obj_list[nr].oid, oflags) < 0)\n \t\t\tdie(\"failed to write object\");\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tfree(buf);\n \t\tobj_list[nr].obj = NULL;\n \t} else if (type == OBJ_BLOB) {\n \t\tstruct blob *blob;\n-\t\tif (write_object_file(buf, size, type,\n-\t\t\t\t      &obj_list[nr].oid) < 0)\n+\t\tif (write_object_file_flags(buf, size, type,\n+\t\t\t\t\t    &obj_list[nr].oid, oflags) < 0)\n \t\t\tdie(\"failed to write object\");\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tfree(buf);\n \n \t\tblob = lookup_blob(the_repository, &obj_list[nr].oid);\n@@ -269,7 +270,7 @@ static void write_object(unsigned nr, enum object_type type,\n \t\tint eaten;\n \t\thash_object_file(the_hash_algo, buf, size, type,\n \t\t\t\t &obj_list[nr].oid);\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tobj = parse_object_buffer(the_repository, &obj_list[nr].oid,\n \t\t\t\t\t  type, size, buf,\n \t\t\t\t\t  &eaten);\n@@ -283,7 +284,7 @@ static void write_object(unsigned nr, enum object_type type,\n \n static void resolve_delta(unsigned nr, enum object_type type,\n \t\t\t  void *base, unsigned long base_size,\n-\t\t\t  void *delta, unsigned long delta_size)\n+\t\t\t  void *delta, unsigned long delta_size, unsigned oflags)\n {\n \tvoid *result;\n \tunsigned long result_size;\n@@ -294,7 +295,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n \tif (!result)\n \t\tdie(\"failed to apply delta\");\n \tfree(delta);\n-\twrite_object(nr, type, result, result_size);\n+\twrite_object(nr, type, result, result_size, oflags);\n }\n \n /*\n@@ -302,7 +303,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n  * resolve all the deltified objects that are based on it.\n  */\n static void added_object(unsigned nr, enum object_type type,\n-\t\t\t void *data, unsigned long size)\n+\t\t\t void *data, unsigned long size, unsigned oflags)\n {\n \tstruct delta_info **p = &delta_list;\n \tstruct delta_info *info;\n@@ -313,7 +314,7 @@ static void added_object(unsigned nr, enum object_type type,\n \t\t\t*p = info->next;\n \t\t\tp = &delta_list;\n \t\t\tresolve_delta(info->nr, type, data, size,\n-\t\t\t\t      info->delta, info->size);\n+\t\t\t\t      info->delta, info->size, oflags);\n \t\t\tfree(info);\n \t\t\tcontinue;\n \t\t}\n@@ -322,18 +323,19 @@ static void added_object(unsigned nr, enum object_type type,\n }\n \n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n-\t\t\t\t   unsigned nr)\n+\t\t\t\t   unsigned nr, unsigned oflags)\n {\n \tvoid *buf = get_data(size);\n \n \tif (!dry_run && buf)\n-\t\twrite_object(nr, type, buf, size);\n+\t\twrite_object(nr, type, buf, size, oflags);\n \telse\n \t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n-\t\t\t\tvoid *delta_data, unsigned long delta_size)\n+\t\t\t\tvoid *delta_data, unsigned long delta_size,\n+\t\t\t\tunsigned oflags)\n {\n \tstruct object *obj;\n \tstruct obj_buffer *obj_buffer;\n@@ -344,12 +346,12 @@ static int resolve_against_held(unsigned nr, const struct object_id *base,\n \tif (!obj_buffer)\n \t\treturn 0;\n \tresolve_delta(nr, obj->type, obj_buffer->buffer,\n-\t\t      obj_buffer->size, delta_data, delta_size);\n+\t\t      obj_buffer->size, delta_data, delta_size, oflags);\n \treturn 1;\n }\n \n static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n-\t\t\t       unsigned nr)\n+\t\t\t       unsigned nr, unsigned oflags)\n {\n \tvoid *delta_data, *base;\n \tunsigned long base_size;\n@@ -366,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n-\t\t\t\t\t      delta_data, delta_size))\n+\t\t\t\t\t      delta_data, delta_size, oflags))\n \t\t\treturn; /* we are done */\n \t\telse {\n \t\t\t/* cannot resolve yet --- queue it */\n@@ -428,7 +430,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t}\n \t}\n \n-\tif (resolve_against_held(nr, &base_oid, delta_data, delta_size))\n+\tif (resolve_against_held(nr, &base_oid, delta_data, delta_size, oflags))\n \t\treturn;\n \n \tbase = read_object_file(&base_oid, &type, &base_size);\n@@ -440,11 +442,11 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\thas_errors = 1;\n \t\treturn;\n \t}\n-\tresolve_delta(nr, type, base, base_size, delta_data, delta_size);\n+\tresolve_delta(nr, type, base, base_size, delta_data, delta_size, oflags);\n \tfree(base);\n }\n \n-static void unpack_one(unsigned nr)\n+static void unpack_one(unsigned nr, unsigned oflags)\n {\n \tunsigned shift;\n \tunsigned char *pack;\n@@ -472,11 +474,11 @@ static void unpack_one(unsigned nr)\n \tcase OBJ_TREE:\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tunpack_non_delta_entry(type, size, nr);\n+\t\tunpack_non_delta_entry(type, size, nr, oflags);\n \t\treturn;\n \tcase OBJ_REF_DELTA:\n \tcase OBJ_OFS_DELTA:\n-\t\tunpack_delta_entry(type, size, nr);\n+\t\tunpack_delta_entry(type, size, nr, oflags);\n \t\treturn;\n \tdefault:\n \t\terror(\"bad object type %d\", type);\n@@ -491,6 +493,7 @@ static void unpack_all(void)\n {\n \tint i;\n \tstruct pack_header *hdr = fill(sizeof(struct pack_header));\n+\tunsigned oflags;\n \n \tnr_objects = ntohl(hdr->hdr_entries);\n \n@@ -505,9 +508,14 @@ static void unpack_all(void)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n \tplug_bulk_checkin();\n+\toflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n \tfor (i = 0; i < nr_objects; i++) {\n-\t\tunpack_one(i);\n-\t\tdisplay_progress(progress, i + 1);\n+\t\tint nth = i + 1;\n+\t\tunsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n+\t\t\tnr_objects == nth ? HASH_N_OBJECTS_LAST : 0;\n+\n+\t\tunpack_one(i, oflags | f);\n+\t\tdisplay_progress(progress, nth);\n \t}\n \tunplug_bulk_checkin();\n \tstop_progress(&progress);\ndiff --git a/cache.h b/cache.h\nindex 5d863f8c5e8..320248a54e0 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -896,6 +896,9 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n #define HASH_FORMAT_CHECK 2\n #define HASH_RENORMALIZE  4\n #define HASH_SILENT 8\n+#define HASH_N_OBJECTS 1<<4\n+#define HASH_N_OBJECTS_FIRST 1<<5\n+#define HASH_N_OBJECTS_LAST 1<<6\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451940","messageId":"RFC-patch-1.7-e03c119c784-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 1/7] write-or-die.c: remove unused fsync_component() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:30Z","receivedAt":"2022-03-23T03:48:12Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This function added in 020406eaa52 (core.fsync: introduce granular\nfsync control infrastructure, 2022-03-10) hasn't been used, and\nappears not to be used by the follow-up series either?\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n cache.h        | 1 -\n write-or-die.c | 7 -------\n 2 files changed, 8 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 84fafe2ed71..5d863f8c5e8 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1766,7 +1766,6 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n \n void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n-int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n static inline int batch_fsync_enabled(enum fsync_component component)\ndiff --git a/write-or-die.c b/write-or-die.c\nindex c4fd91b5b43..103698450c3 100644\n--- a/write-or-die.c\n+++ b/write-or-die.c\n@@ -76,13 +76,6 @@ void fsync_or_die(int fd, const char *msg)\n \t\tdie_errno(\"fsync error on '%s'\", msg);\n }\n \n-int fsync_component(enum fsync_component component, int fd)\n-{\n-\tif (fsync_components & component)\n-\t\treturn maybe_fsync(fd);\n-\treturn 0;\n-}\n-\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n {\n \tif (fsync_components & component)\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451941","messageId":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"CANQDOde2OG8fVSM1hQE3FBmzWy5FkgQCWAUYhFztB8UGFyJELg@mail.gmail.com","subject":"[RFC PATCH 0/7] bottom-up ns/batched-fsync & \"plugging\" in object-file.c","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:29Z","receivedAt":"2022-03-23T03:48:15Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This RFC series is a continuation of the thread at\nhttps://lore.kernel.org/git/CANQDOde2OG8fVSM1hQE3FBmzWy5FkgQCWAUYhFztB8UGFyJELg@mail.gmail.com/;\nMore details in individual commit messages.\n\nI'd suggested (upthread of) there pass new object flags down to the\nobject machinery instead of the {un,}plug_bulk_checkin() API\nroute. This has advantages described in more details in individual\npatches.\n\nThis also shows that the not-using tmpdir approach can be\nsignificantly faster than using it, and per my understanding just as\nsafe fsync-wise for those willing to deal with the caveat of possibly\nhaving truncated *unreachable* objects.\n\nI thought that showing some working code with what I was suggesting\nwas more productive than continuing the current back & forth :)\n\nÆvar Arnfjörð Bjarmason (7):\n  write-or-die.c: remove unused fsync_component() function\n  unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n  object-file: pass down unpack-objects.c flags for \"bulk\" checkin\n  update-index: use a utility function for stdin consumption\n  update-index: pass down an \"oflags\" argument\n  update-index: rename \"buf\" to \"line\"\n  update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n\n builtin/add.c            |  3 --\n builtin/unpack-objects.c | 62 ++++++++++++++------------\n builtin/update-index.c   | 96 ++++++++++++++++++++++++++--------------\n bulk-checkin.c           | 86 -----------------------------------\n bulk-checkin.h           |  6 ---\n cache.h                  |  9 ++--\n object-file.c            | 39 +++++++++++-----\n t/t1050-large.sh         |  3 ++\n write-or-die.c           |  7 ---\n 9 files changed, 131 insertions(+), 180 deletions(-)\n\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451942","messageId":"RFC-patch-4.7-2c5395a3716-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 4/7] update-index: use a utility function for stdin consumption","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:33Z","receivedAt":"2022-03-23T03:48:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c | 36 ++++++++++++++++++++++--------------\n 1 file changed, 22 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 95ed3c47b2e..80b96ec5721 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -971,6 +971,25 @@ static enum parse_opt_result reupdate_callback(\n \treturn 0;\n }\n \n+static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n+\t\t\t    const char *prefix, int prefix_length,\n+\t\t\t    const int nul_term_line, const int set_executable_bit)\n+{\n+\tchar *p;\n+\n+\tif (!nul_term_line && buf->buf[0] == '\"') {\n+\t\tstrbuf_reset(unquoted);\n+\t\tif (unquote_c_style(unquoted, buf->buf, NULL))\n+\t\t\tdie(\"line is badly quoted\");\n+\t\tstrbuf_swap(buf, unquoted);\n+\t}\n+\tp = prefix_path(prefix, prefix_length, buf->buf);\n+\tupdate_one(p);\n+\tif (set_executable_bit)\n+\t\tchmod_path(set_executable_bit, p);\n+\tfree(p);\n+}\n+\n int cmd_update_index(int argc, const char **argv, const char *prefix)\n {\n \tint newfd, entries, has_errors = 0, nul_term_line = 0;\n@@ -1174,20 +1193,9 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n-\t\twhile (getline_fn(&buf, stdin) != EOF) {\n-\t\t\tchar *p;\n-\t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n-\t\t\t\tstrbuf_reset(&unquoted);\n-\t\t\t\tif (unquote_c_style(&unquoted, buf.buf, NULL))\n-\t\t\t\t\tdie(\"line is badly quoted\");\n-\t\t\t\tstrbuf_swap(&buf, &unquoted);\n-\t\t\t}\n-\t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n-\t\t\tupdate_one(p);\n-\t\t\tif (set_executable_bit)\n-\t\t\t\tchmod_path(set_executable_bit, p);\n-\t\t\tfree(p);\n-\t\t}\n+\t\twhile (getline_fn(&buf, stdin) != EOF)\n+\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n+\t\t\t\t\tnul_term_line, set_executable_bit);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451943","messageId":"RFC-patch-3.7-beda9f99529-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 3/7] object-file: pass down unpack-objects.c flags for \"bulk\" checkin","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:32Z","receivedAt":"2022-03-23T03:48:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Remove much of this as a POC for exploring some of what I mentioned in\nhttps://lore.kernel.org/git/220322.86mthinxnn.gmgdl@evledraar.gmail.com/\n\nThis commit is obviously not what we *should* do as end-state, but\ndemonstrates what's needed (I think) for a bare-minimum implementation\nof just the \"bulk\" syncing method for loose objects without the part\nwhere we do the tmp-objdir.c dance.\n\nPerformance with this is already quite promising. Benchmarking with:\n\n\tgit hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' \\\n\t    \t-p 'rm -rf r.git && git init --bare r.git' \\\n\t\t'./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack'\n\nI.e. unpacking a small packfile (my dotfiles) yields, on a Linux\nramdisk:\n\n\tBenchmark 1: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync\n\t  Time (mean ± σ):     815.9 ms ±   8.2 ms    [User: 522.9 ms, System: 287.9 ms]\n\t  Range (min … max):   805.6 ms … 835.9 ms    10 runs\n\n\tBenchmark 2: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD\n\t  Time (mean ± σ):     779.4 ms ±  15.4 ms    [User: 505.7 ms, System: 270.2 ms]\n\t  Range (min … max):   763.1 ms … 813.9 ms    10 runs\n\n\tSummary\n\t  './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n\t    1.05 ± 0.02 times faster than './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n\nDoing the same with \"strace --summary-only\", which probably helps to\nemulate cases with slower syscalls is ~15% faster than using the\ntmp-objdir indirection:\n\n\tSummary\n\t  'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n\t    1.16 ± 0.01 times faster than 'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n\nWhich makes sense in terms of syscalls. In my case HEAD has ~101k\ncalls, and the parent topic is making ~129k calls, with around 2x the\nnumber of unlink(), link() as expected.\n\nOf course some users will want to use the tmp-objdir.c method. So a\nversion of this commit could be rewritten to come earlier in the\nseries, with the \"bulk\" on top being optional.\n\nIt seems to me that it's a much better strategy to do this whole thing\nin close_loose_object() after passing down the new HASH_N_OBJECTS /\nHASH_N_OBJECTS_FIRST / HASH_N_OBJECTS_LAST flags.\n\nDoing that for the \"builtin/add.c\" and \"builtin/unpack-objects.c\" code\nhaving its {un,}plug_bulk_checkin() removed here is then just a matter\nof passing down a similar set of flags indicating whether we're\ndealing with N objects, and if so if we're dealing with the last one\nor not.\n\nAs we'll see in subsequent commits doing it this way also effortlessly\nintegrates with other HASH_* flags. E.g. for \"update-index\" the code\nbeing rm'd here doesn't handle the interaction with\n\"HASH_WRITE_OBJECT\" properly, but once we've moved all this sync\nbootstrapping logic to close_loose_object() we'll never get to it if\nwe're not actually writing something.\n\nThis code currently doesn't use the HASH_N_OBJECTS_FIRST flag, but\nthat's what we'd use later to optionally call tmp_objdir_create().\n\nAside: This also changes logic that was a bit confusing and repetitive\nin close_loose_object(). Previously we'd first call\nbatch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT) which is just as\nshorthand for:\n\n\tfsync_components & FSYNC_COMPONENT_LOOSE_OBJECT &&\n\tfsync_method == FSYNC_METHOD_BATCH\n\nWe'd then proceed to call\nfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT) later in the same\nfunction, which is just a way of calling fsync_or_die() if:\n\n\tfsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\n\nNow we instead just define a local \"fsync_loose\" variable by checking\n\"fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\", which shows us that\nthe previous case of fsync_component_or_die(...)\" could just be added\nto the existing \"fsync_object_files > 0\" branch.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/add.c            |  3 --\n builtin/unpack-objects.c |  2 -\n builtin/update-index.c   |  4 --\n bulk-checkin.c           | 86 ----------------------------------------\n bulk-checkin.h           |  6 ---\n cache.h                  |  5 ---\n object-file.c            | 37 ++++++++++++-----\n t/t1050-large.sh         |  3 ++\n 8 files changed, 29 insertions(+), 117 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..a6d7f4dc1e1 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,8 +670,6 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\tstring_list_clear(&only_match_skip_worktree, 0);\n \t}\n \n-\tplug_bulk_checkin();\n-\n \tif (add_renormalize)\n \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n \telse\n@@ -682,7 +680,6 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex ec40c6fd966..93da436581b 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -507,7 +507,6 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n-\tplug_bulk_checkin();\n \toflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tint nth = i + 1;\n@@ -517,7 +516,6 @@ static void unpack_all(void)\n \t\tunpack_one(i, oflags | f);\n \t\tdisplay_progress(progress, nth);\n \t}\n-\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex cbd2b0d633b..95ed3c47b2e 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -1118,8 +1118,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n \n-\t/* optimize adding many objects to the object database */\n-\tplug_bulk_checkin();\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1194,8 +1192,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n-\t/* by now we must have added all of the new objects */\n-\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex a0dca79ba6a..4ffea87f44d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -9,14 +9,11 @@\n #include \"pack.h\"\n #include \"strbuf.h\"\n #include \"string-list.h\"\n-#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n-static struct tmp_objdir *bulk_fsync_objdir;\n-\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -85,40 +82,6 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \treprepare_packed_git(the_repository);\n }\n \n-/*\n- * Cleanup after batch-mode fsync_object_files.\n- */\n-static void do_batch_fsync(void)\n-{\n-\tstruct strbuf temp_path = STRBUF_INIT;\n-\tstruct tempfile *temp;\n-\n-\tif (!bulk_fsync_objdir)\n-\t\treturn;\n-\n-\t/*\n-\t * Issue a full hardware flush against a temporary file to ensure\n-\t * that all objects are durable before any renames occur. The code in\n-\t * fsync_loose_object_bulk_checkin has already issued a writeout\n-\t * request, but it has not flushed any writeback cache in the storage\n-\t * hardware or any filesystem logs. This fsync call acts as a barrier\n-\t * to ensure that the data in each new object file is durable before\n-\t * the final name is visible.\n-\t */\n-\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n-\ttemp = xmks_tempfile(temp_path.buf);\n-\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n-\tdelete_tempfile(&temp);\n-\tstrbuf_release(&temp_path);\n-\n-\t/*\n-\t * Make the object files visible in the primary ODB after their data is\n-\t * fully durable.\n-\t */\n-\ttmp_objdir_migrate(bulk_fsync_objdir);\n-\tbulk_fsync_objdir = NULL;\n-}\n-\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -313,26 +276,6 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n-void prepare_loose_object_bulk_checkin(void)\n-{\n-\tif (bulk_checkin_plugged && !bulk_fsync_objdir)\n-\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n-}\n-\n-void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n-{\n-\t/*\n-\t * If we have a plugged bulk checkin, we issue a call that\n-\t * cleans the filesystem page cache but avoids a hardware flush\n-\t * command. Later on we will issue a single hardware flush\n-\t * before as part of do_batch_fsync.\n-\t */\n-\tif (!bulk_fsync_objdir ||\n-\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n-\t\tfsync_or_die(fd, filename);\n-\t}\n-}\n-\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -343,32 +286,3 @@ int index_bulk_checkin(struct object_id *oid,\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n-\n-void plug_bulk_checkin(void)\n-{\n-\tassert(!bulk_checkin_plugged);\n-\n-\t/*\n-\t * A temporary object directory is used to hold the files\n-\t * while they are not fsynced.\n-\t */\n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n-\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n-\t\tif (!bulk_fsync_objdir)\n-\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncMethod=batch\"));\n-\n-\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n-\t}\n-\n-\tbulk_checkin_plugged = 1;\n-}\n-\n-void unplug_bulk_checkin(void)\n-{\n-\tassert(bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 0;\n-\tif (bulk_checkin_state.f)\n-\t\tfinish_bulk_checkin(&bulk_checkin_state);\n-\n-\tdo_batch_fsync();\n-}\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 181d3447ff9..76fc33e0c8f 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,14 +6,8 @@\n \n #include \"cache.h\"\n \n-void prepare_loose_object_bulk_checkin(void);\n-void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n-\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n-void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n-\n #endif\ndiff --git a/cache.h b/cache.h\nindex 320248a54e0..997bf2f57fd 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1771,11 +1771,6 @@ void write_or_die(int fd, const void *buf, size_t count);\n void fsync_or_die(int fd, const char *);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n-static inline int batch_fsync_enabled(enum fsync_component component)\n-{\n-\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n-}\n-\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/object-file.c b/object-file.c\nindex cd0ddb49e4b..dbeb3df502d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1886,19 +1886,37 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n }\n \n+static void sync_loose_object_batch(int fd, const char *filename,\n+\t\t\t\t    const unsigned oflags)\n+{\n+\tconst int last = oflags & HASH_N_OBJECTS_LAST;\n+\n+\t/*\n+\t * We're doing a sync_file_range() (or equivalent) for 1..N-1\n+\t * objects, and then a \"real\" fsync() for N. On some OS's\n+\t * enabling core.fsync=loose-object && core.fsyncMethod=batch\n+\t * improves the performance by a lot.\n+\t */\n+\tif (last || (!last && git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0))\n+\t\tfsync_or_die(fd, filename);\n+}\n+\n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd, const char *filename)\n+static void close_loose_object(int fd, const char *filename,\n+\t\t\t       const unsigned oflags)\n {\n+\tint fsync_loose;\n+\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n-\t\tfsync_loose_object_bulk_checkin(fd, filename);\n-\telse if (fsync_object_files > 0)\n+\tfsync_loose = fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT;\n+\n+\tif (oflags & HASH_N_OBJECTS && fsync_loose &&\n+\t    fsync_method == FSYNC_METHOD_BATCH)\n+\t\tsync_loose_object_batch(fd, filename, oflags);\n+\telse if (fsync_object_files > 0 || fsync_loose)\n \t\tfsync_or_die(fd, filename);\n-\telse\n-\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n-\t\t\t\t       filename);\n \n out:\n \tif (close(fd) != 0)\n@@ -1962,9 +1980,6 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n-\t\tprepare_loose_object_bulk_checkin();\n-\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n@@ -2015,7 +2030,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd, tmp_file.buf);\n+\tclose_loose_object(fd, tmp_file.buf, flags);\n \n \tif (mtime) {\n \t\tstruct utimbuf utb;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 4f3aa17c994..1baaa8024c8 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -5,6 +5,9 @@ test_description='adding and checking out large blobs'\n \n . ./test-lib.sh\n \n+skip_all='TODO: migrate the builtin/add.c code'\n+test_done\n+\n test_expect_success setup '\n \t# clone does not allow us to pass core.bigfilethreshold to\n \t# new repos, so set core.bigfilethreshold globally\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451944","messageId":"RFC-patch-5.7-a1474968991-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 5/7] update-index: pass down an \"oflags\" argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:34Z","receivedAt":"2022-03-23T03:48:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"We do nothing with this yet, but will soon.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c | 37 +++++++++++++++++++++----------------\n object-file.c          |  2 +-\n 2 files changed, 22 insertions(+), 17 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 80b96ec5721..1884124224c 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -267,10 +267,12 @@ static int process_lstat_error(const char *path, int err)\n \treturn error(\"lstat(\\\"%s\\\"): %s\", path, strerror(err));\n }\n \n-static int add_one_path(const struct cache_entry *old, const char *path, int len, struct stat *st)\n+static int add_one_path(const struct cache_entry *old, const char *path,\n+\t\t\tint len, struct stat *st, const unsigned oflags)\n {\n \tint option;\n \tstruct cache_entry *ce;\n+\tunsigned f;\n \n \t/* Was the old index entry already up-to-date? */\n \tif (old && !ce_stage(old) && !ce_match_stat(old, st, 0))\n@@ -283,8 +285,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n \tfill_stat_cache_info(&the_index, ce, st);\n \tce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n \n-\tif (index_path(&the_index, &ce->oid, path, st,\n-\t\t       info_only ? 0 : HASH_WRITE_OBJECT)) {\n+\tf = oflags | (info_only ? 0 : HASH_WRITE_OBJECT);\n+\tif (index_path(&the_index, &ce->oid, path, st, f)) {\n \t\tdiscard_cache_entry(ce);\n \t\treturn -1;\n \t}\n@@ -320,7 +322,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n  *  - it doesn't exist at all in the index, but it is a valid\n  *    git directory, and it should be *added* as a gitlink.\n  */\n-static int process_directory(const char *path, int len, struct stat *st)\n+static int process_directory(const char *path, int len, struct stat *st,\n+\t\t\t     const unsigned oflags)\n {\n \tstruct object_id oid;\n \tint pos = cache_name_pos(path, len);\n@@ -334,7 +337,7 @@ static int process_directory(const char *path, int len, struct stat *st)\n \t\t\tif (resolve_gitlink_ref(path, \"HEAD\", &oid) < 0)\n \t\t\t\treturn 0;\n \n-\t\t\treturn add_one_path(ce, path, len, st);\n+\t\t\treturn add_one_path(ce, path, len, st, oflags);\n \t\t}\n \t\t/* Should this be an unconditional error? */\n \t\treturn remove_one_path(path);\n@@ -358,13 +361,14 @@ static int process_directory(const char *path, int len, struct stat *st)\n \n \t/* No match - should we add it as a gitlink? */\n \tif (!resolve_gitlink_ref(path, \"HEAD\", &oid))\n-\t\treturn add_one_path(NULL, path, len, st);\n+\t\treturn add_one_path(NULL, path, len, st, oflags);\n \n \t/* Error out. */\n \treturn error(\"%s: is a directory - add files inside instead\", path);\n }\n \n-static int process_path(const char *path, struct stat *st, int stat_errno)\n+static int process_path(const char *path, struct stat *st, int stat_errno,\n+\t\t\tconst unsigned oflags)\n {\n \tint pos, len;\n \tconst struct cache_entry *ce;\n@@ -395,9 +399,9 @@ static int process_path(const char *path, struct stat *st, int stat_errno)\n \t\treturn process_lstat_error(path, stat_errno);\n \n \tif (S_ISDIR(st->st_mode))\n-\t\treturn process_directory(path, len, st);\n+\t\treturn process_directory(path, len, st, oflags);\n \n-\treturn add_one_path(ce, path, len, st);\n+\treturn add_one_path(ce, path, len, st, oflags);\n }\n \n static int add_cacheinfo(unsigned int mode, const struct object_id *oid,\n@@ -446,7 +450,7 @@ static void chmod_path(char flip, const char *path)\n \tdie(\"git update-index: cannot chmod %cx '%s'\", flip, path);\n }\n \n-static void update_one(const char *path)\n+static void update_one(const char *path, const unsigned oflags)\n {\n \tint stat_errno = 0;\n \tstruct stat st;\n@@ -485,7 +489,7 @@ static void update_one(const char *path)\n \t\treport(\"remove '%s'\", path);\n \t\treturn;\n \t}\n-\tif (process_path(path, &st, stat_errno))\n+\tif (process_path(path, &st, stat_errno, oflags))\n \t\tdie(\"Unable to process path %s\", path);\n \treport(\"add '%s'\", path);\n }\n@@ -776,7 +780,7 @@ static int do_reupdate(int ac, const char **av,\n \t\t */\n \t\tsave_nr = active_nr;\n \t\tpath = xstrdup(ce->name);\n-\t\tupdate_one(path);\n+\t\tupdate_one(path, 0);\n \t\tfree(path);\n \t\tdiscard_cache_entry(old);\n \t\tif (save_nr != active_nr)\n@@ -973,7 +977,8 @@ static enum parse_opt_result reupdate_callback(\n \n static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n \t\t\t    const char *prefix, int prefix_length,\n-\t\t\t    const int nul_term_line, const int set_executable_bit)\n+\t\t\t    const int nul_term_line, const int set_executable_bit,\n+\t\t\t    const unsigned oflags)\n {\n \tchar *p;\n \n@@ -984,7 +989,7 @@ static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n \t\tstrbuf_swap(buf, unquoted);\n \t}\n \tp = prefix_path(prefix, prefix_length, buf->buf);\n-\tupdate_one(p);\n+\tupdate_one(p, oflags);\n \tif (set_executable_bit)\n \t\tchmod_path(set_executable_bit, p);\n \tfree(p);\n@@ -1157,7 +1162,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \t\t\tsetup_work_tree();\n \t\t\tp = prefix_path(prefix, prefix_length, path);\n-\t\t\tupdate_one(p);\n+\t\t\tupdate_one(p, 0);\n \t\t\tif (set_executable_bit)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n@@ -1195,7 +1200,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tsetup_work_tree();\n \t\twhile (getline_fn(&buf, stdin) != EOF)\n \t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n-\t\t\t\t\tnul_term_line, set_executable_bit);\n+\t\t\t\t\tnul_term_line, set_executable_bit, 0);\n \t\tstrbuf_release(&unquoted);\n \t\tstrbuf_release(&buf);\n \t}\ndiff --git a/object-file.c b/object-file.c\nindex dbeb3df502d..8999fce2b15 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2211,7 +2211,7 @@ static int index_mem(struct index_state *istate,\n \t}\n \n \tif (write_object)\n-\t\tret = write_object_file(buf, size, type, oid);\n+\t\tret = write_object_file_flags(buf, size, type, oid, flags);\n \telse\n \t\thash_object_file(the_hash_algo, buf, size, type, oid);\n \tif (re_allocated)\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451945","messageId":"RFC-patch-6.7-4fad333e9a1-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 6/7] update-index: rename \"buf\" to \"line\"","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:35Z","receivedAt":"2022-03-23T03:48:32Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This variable renaming makes a subsequent more meaningful change\nsmaller.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 1884124224c..af02ff39756 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -1194,15 +1194,15 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t}\n \n \tif (read_from_stdin) {\n-\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\tstruct strbuf line = STRBUF_INIT;\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n \n \t\tsetup_work_tree();\n-\t\twhile (getline_fn(&buf, stdin) != EOF)\n-\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n+\t\twhile (getline_fn(&line, stdin) != EOF)\n+\t\t\tline_from_stdin(&line, &unquoted, prefix, prefix_length,\n \t\t\t\t\tnul_term_line, set_executable_bit, 0);\n \t\tstrbuf_release(&unquoted);\n-\t\tstrbuf_release(&buf);\n+\t\tstrbuf_release(&line);\n \t}\n \n \tif (split_index > 0) {\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451946","messageId":"RFC-patch-7.7-481f1d771cb-20220323T033928Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH 7/7] update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T03:47:36Z","receivedAt":"2022-03-23T03:48:35Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"As with unpack-objects in a preceding commit have update-index.c make\nuse of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\nmode again for \"update-index\".\n\nAdding the t/* directory from git.git on a Linux ramdisk is a bit\nfaster than with the tmp-objdir indirection:\n\n\tgit hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10\n\tBenchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync\n\t  Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]\n\t  Range (min … max):   285.6 ms … 297.0 ms    10 runs\n\n\tBenchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD\n\t  Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]\n\t  Range (min … max):   267.8 ms … 291.3 ms    10 runs\n\n\tSummary\n\t  'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n\t    1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n\nAnd as before running that with \"strace --summary-only\" slows things\ndown a bit (probably mimicking slower I/O a bit). I then get:\n\n\tSummary\n\t  'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n\t    1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n\nWe also go from ~51k syscalls to ~39k, with ~2x the number of link()\nand unlink() in ns/batched-fsync.\n\nIn the process of doing this conversion we lost the \"bulk\" mode for\nfiles added on the command-line. I don't think it's useful to optimize\nthat, but we could if anyone cared.\n\nWe've also converted this to a string_list, we could walk with\ngetline_fn() and get one line \"ahead\" to see what we have left, but I\nfound that state machine a bit painful, and at least in my testing\nbuffering this doesn't harm things. But we could also change this to\nstream again, at the cost of some getline_fn() twiddling.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c | 31 +++++++++++++++++++++++++++----\n 1 file changed, 27 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex af02ff39756..c7cbfe1123b 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -1194,15 +1194,38 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t}\n \n \tif (read_from_stdin) {\n+\t\tstruct string_list list = STRING_LIST_INIT_NODUP;\n \t\tstruct strbuf line = STRBUF_INIT;\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n+\t\tsize_t i, nr;\n+\t\tunsigned oflags;\n \n \t\tsetup_work_tree();\n-\t\twhile (getline_fn(&line, stdin) != EOF)\n-\t\t\tline_from_stdin(&line, &unquoted, prefix, prefix_length,\n-\t\t\t\t\tnul_term_line, set_executable_bit, 0);\n+\t\twhile (getline_fn(&line, stdin) != EOF) {\n+\t\t\tsize_t len = line.len;\n+\t\t\tchar *str = strbuf_detach(&line, NULL);\n+\n+\t\t\tstring_list_append_nodup(&list, str)->util = (void *)len;\n+\t\t}\n+\n+\t\tnr = list.nr;\n+\t\toflags = nr > 1 ? HASH_N_OBJECTS : 0;\n+\t\tfor (i = 0; i < nr; i++) {\n+\t\t\tsize_t nth = i + 1;\n+\t\t\tunsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n+\t\t\t\t  nr == nth ? HASH_N_OBJECTS_LAST : 0;\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\t\tstruct string_list_item *item = list.items + i;\n+\t\t\tconst size_t len = (size_t)item->util;\n+\n+\t\t\tstrbuf_attach(&buf, item->string, len, len);\n+\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n+\t\t\t\t\tnul_term_line, set_executable_bit,\n+\t\t\t\t\toflags | f);\n+\t\t\tstrbuf_release(&buf);\n+\t\t}\n \t\tstrbuf_release(&unquoted);\n-\t\tstrbuf_release(&line);\n+\t\tstring_list_clear(&list, 0);\n \t}\n \n \tif (split_index > 0) {\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451948","messageId":"CANQDOdd6w8d_1sx8Sob3EZ4+hy7m1DBdN4EYFJPExLmTgr5Kqw@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-1.7-e03c119c784-20220323T033928Z-avarab@gmail.com","subject":"Re: [RFC PATCH 1/7] write-or-die.c: remove unused fsync_component() function","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T05:27:34Z","receivedAt":"2022-03-23T05:27:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 22, 2022 at 8:48 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> This function added in 020406eaa52 (core.fsync: introduce granular\n> fsync control infrastructure, 2022-03-10) hasn't been used, and\n> appears not to be used by the follow-up series either?\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  cache.h        | 1 -\n>  write-or-die.c | 7 -------\n>  2 files changed, 8 deletions(-)\n>\n> diff --git a/cache.h b/cache.h\n> index 84fafe2ed71..5d863f8c5e8 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1766,7 +1766,6 @@ int copy_file_with_time(const char *dst, const char *src, int mode);\n>\n>  void write_or_die(int fd, const void *buf, size_t count);\n>  void fsync_or_die(int fd, const char *);\n> -int fsync_component(enum fsync_component component, int fd);\n>  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n>\n>  static inline int batch_fsync_enabled(enum fsync_component component)\n> diff --git a/write-or-die.c b/write-or-die.c\n> index c4fd91b5b43..103698450c3 100644\n> --- a/write-or-die.c\n> +++ b/write-or-die.c\n> @@ -76,13 +76,6 @@ void fsync_or_die(int fd, const char *msg)\n>                 die_errno(\"fsync error on '%s'\", msg);\n>  }\n>\n> -int fsync_component(enum fsync_component component, int fd)\n> -{\n> -       if (fsync_components & component)\n> -               return maybe_fsync(fd);\n> -       return 0;\n> -}\n> -\n>  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg)\n>  {\n>         if (fsync_components & component)\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nThis helper was put in for Patrick's patch at\nhttps://lore.kernel.org/git/f1e8a7bb3bf0f4c0414819cb1d5579dc08fd2a4f.1646905589.git.ps@pks.im/.\n\nThanks,\nNeeraj\n"},{"id":"451949","messageId":"CANQDOdcfoS-kLkaoXSKBhemDV13aWsoLSf=xUaU84B6_ajqbJA@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-7.7-481f1d771cb-20220323T033928Z-avarab@gmail.com","subject":"Re: [RFC PATCH 7/7] update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T05:51:35Z","receivedAt":"2022-03-23T05:51:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 22, 2022 at 8:48 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> As with unpack-objects in a preceding commit have update-index.c make\n> use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\n> mode again for \"update-index\".\n>\n> Adding the t/* directory from git.git on a Linux ramdisk is a bit\n> faster than with the tmp-objdir indirection:\n>\n>         git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10\n>         Benchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync\n>           Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]\n>           Range (min … max):   285.6 ms … 297.0 ms    10 runs\n>\n>         Benchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD\n>           Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]\n>           Range (min … max):   267.8 ms … 291.3 ms    10 runs\n>\n>         Summary\n>           'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n>             1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n>\n> And as before running that with \"strace --summary-only\" slows things\n> down a bit (probably mimicking slower I/O a bit). I then get:\n>\n>         Summary\n>           'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n>             1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n>\n> We also go from ~51k syscalls to ~39k, with ~2x the number of link()\n> and unlink() in ns/batched-fsync.\n>\n> In the process of doing this conversion we lost the \"bulk\" mode for\n> files added on the command-line. I don't think it's useful to optimize\n> that, but we could if anyone cared.\n>\n> We've also converted this to a string_list, we could walk with\n> getline_fn() and get one line \"ahead\" to see what we have left, but I\n> found that state machine a bit painful, and at least in my testing\n> buffering this doesn't harm things. But we could also change this to\n> stream again, at the cost of some getline_fn() twiddling.\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  builtin/update-index.c | 31 +++++++++++++++++++++++++++----\n>  1 file changed, 27 insertions(+), 4 deletions(-)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index af02ff39756..c7cbfe1123b 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -1194,15 +1194,38 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>         }\n>\n>         if (read_from_stdin) {\n> +               struct string_list list = STRING_LIST_INIT_NODUP;\n>                 struct strbuf line = STRBUF_INIT;\n>                 struct strbuf unquoted = STRBUF_INIT;\n> +               size_t i, nr;\n> +               unsigned oflags;\n>\n>                 setup_work_tree();\n> -               while (getline_fn(&line, stdin) != EOF)\n> -                       line_from_stdin(&line, &unquoted, prefix, prefix_length,\n> -                                       nul_term_line, set_executable_bit, 0);\n> +               while (getline_fn(&line, stdin) != EOF) {\n> +                       size_t len = line.len;\n> +                       char *str = strbuf_detach(&line, NULL);\n> +\n> +                       string_list_append_nodup(&list, str)->util = (void *)len;\n> +               }\n> +\n> +               nr = list.nr;\n> +               oflags = nr > 1 ? HASH_N_OBJECTS : 0;\n> +               for (i = 0; i < nr; i++) {\n> +                       size_t nth = i + 1;\n> +                       unsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n> +                                 nr == nth ? HASH_N_OBJECTS_LAST : 0;\n> +                       struct strbuf buf = STRBUF_INIT;\n> +                       struct string_list_item *item = list.items + i;\n> +                       const size_t len = (size_t)item->util;\n> +\n> +                       strbuf_attach(&buf, item->string, len, len);\n> +                       line_from_stdin(&buf, &unquoted, prefix, prefix_length,\n> +                                       nul_term_line, set_executable_bit,\n> +                                       oflags | f);\n> +                       strbuf_release(&buf);\n> +               }\n>                 strbuf_release(&unquoted);\n> -               strbuf_release(&line);\n> +               string_list_clear(&list, 0);\n>         }\n>\n>         if (split_index > 0) {\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nThis buffering introduces the same potential risk of the\n\"stdin-feeder\" process not being able to see objects right away as my\nversion had. I'm planning to mitigate the issue by unplugging the bulk\ncheckin when issuing a verbose report so that anyone who's using that\noutput to synchronize can still see what they're expecting.\n\nI think the code you've presented here is a lot of diff to accomplish\nthe same thing that my series does, where this specific update-index\ncaller has been roto-tilled to provide the needed\nbegin/end-transaction points.  And I think there will be a lot of\ncomplexity in supporting the same hints for command-line additions\n(which is roughly equivalent to the git-add workflow). Every caller\nthat wants batch treatment will have to either implement a state\nmachine or implement a buffering mechanism in order to figure out the\nbegin-end points. Having a separate plug/unplug call eliminates this\ncomplexity on each caller.\n\nBtw, I'm planning in a future series to reduce the system calls\ninvolved in renaming a file by taking advantage of the renameat2\nsystem call and equivalents on other platforms.  There's a pretty\nstrong motivation to do that on Windows.\n\nThanks for the concrete code,\n-Neeraj\n"},{"id":"451972","messageId":"220323.86sfr9ndpr.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"CANQDOdcfoS-kLkaoXSKBhemDV13aWsoLSf=xUaU84B6_ajqbJA@mail.gmail.com","subject":"Re: [RFC PATCH 7/7] update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T09:48:40Z","receivedAt":"2022-03-23T10:52:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Mar 22 2022, Neeraj Singh wrote:\n\n> On Tue, Mar 22, 2022 at 8:48 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>> As with unpack-objects in a preceding commit have update-index.c make\n>> use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\n>> mode again for \"update-index\".\n>>\n>> Adding the t/* directory from git.git on a Linux ramdisk is a bit\n>> faster than with the tmp-objdir indirection:\n>>\n>>         git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10\n>>         Benchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync\n>>           Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]\n>>           Range (min … max):   285.6 ms … 297.0 ms    10 runs\n>>\n>>         Benchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD\n>>           Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]\n>>           Range (min … max):   267.8 ms … 291.3 ms    10 runs\n>>\n>>         Summary\n>>           'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n>>             1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n>>\n>> And as before running that with \"strace --summary-only\" slows things\n>> down a bit (probably mimicking slower I/O a bit). I then get:\n>>\n>>         Summary\n>>           'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n>>             1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n>>\n>> We also go from ~51k syscalls to ~39k, with ~2x the number of link()\n>> and unlink() in ns/batched-fsync.\n>>\n>> In the process of doing this conversion we lost the \"bulk\" mode for\n>> files added on the command-line. I don't think it's useful to optimize\n>> that, but we could if anyone cared.\n>>\n>> We've also converted this to a string_list, we could walk with\n>> getline_fn() and get one line \"ahead\" to see what we have left, but I\n>> found that state machine a bit painful, and at least in my testing\n>> buffering this doesn't harm things. But we could also change this to\n>> stream again, at the cost of some getline_fn() twiddling.\n>>\n>> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>> ---\n>>  builtin/update-index.c | 31 +++++++++++++++++++++++++++----\n>>  1 file changed, 27 insertions(+), 4 deletions(-)\n>>\n>> diff --git a/builtin/update-index.c b/builtin/update-index.c\n>> index af02ff39756..c7cbfe1123b 100644\n>> --- a/builtin/update-index.c\n>> +++ b/builtin/update-index.c\n>> @@ -1194,15 +1194,38 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>>         }\n>>\n>>         if (read_from_stdin) {\n>> +               struct string_list list = STRING_LIST_INIT_NODUP;\n>>                 struct strbuf line = STRBUF_INIT;\n>>                 struct strbuf unquoted = STRBUF_INIT;\n>> +               size_t i, nr;\n>> +               unsigned oflags;\n>>\n>>                 setup_work_tree();\n>> -               while (getline_fn(&line, stdin) != EOF)\n>> -                       line_from_stdin(&line, &unquoted, prefix, prefix_length,\n>> -                                       nul_term_line, set_executable_bit, 0);\n>> +               while (getline_fn(&line, stdin) != EOF) {\n>> +                       size_t len = line.len;\n>> +                       char *str = strbuf_detach(&line, NULL);\n>> +\n>> +                       string_list_append_nodup(&list, str)->util = (void *)len;\n>> +               }\n>> +\n>> +               nr = list.nr;\n>> +               oflags = nr > 1 ? HASH_N_OBJECTS : 0;\n>> +               for (i = 0; i < nr; i++) {\n>> +                       size_t nth = i + 1;\n>> +                       unsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n>> +                                 nr == nth ? HASH_N_OBJECTS_LAST : 0;\n>> +                       struct strbuf buf = STRBUF_INIT;\n>> +                       struct string_list_item *item = list.items + i;\n>> +                       const size_t len = (size_t)item->util;\n>> +\n>> +                       strbuf_attach(&buf, item->string, len, len);\n>> +                       line_from_stdin(&buf, &unquoted, prefix, prefix_length,\n>> +                                       nul_term_line, set_executable_bit,\n>> +                                       oflags | f);\n>> +                       strbuf_release(&buf);\n>> +               }\n>>                 strbuf_release(&unquoted);\n>> -               strbuf_release(&line);\n>> +               string_list_clear(&list, 0);\n>>         }\n>>\n>>         if (split_index > 0) {\n>> --\n>> 2.35.1.1428.g1c1a0152d61\n>>\n>\n> This buffering introduces the same potential risk of the\n> \"stdin-feeder\" process not being able to see objects right away as my\n> version had. I'm planning to mitigate the issue by unplugging the bulk\n> checkin when issuing a verbose report so that anyone who's using that\n> output to synchronize can still see what they're expecting.\n\nI was rather terse in the commit message, I meant (but forgot some\nwords) \"doesn't harm thing for performance [in the above test]\", but\nconverting this to a string_list is clearly & regression that shouldn't\nbe kept.\n\nI just wanted to demonstrate method of doing this by passing down the\nHASH_* flags, and found that writing the state-machine to \"buffer ahead\"\nby one line so that we can eventually know in the loop if we're in the\n\"last\" line or not was tedious, so I came up with this POC. But we\nclearly shouldn't lose the \"streaming\" aspect.\n\nBut anyway, now that I look at this again the smart thing here (surely?)\nis to keep the simple getline() loop and not ever issue a\nHASH_N_OBJECTS_LAST for the Nth item, instead we should in this case do\nthe \"checkpoint fsync\" at the point that we write the actual index.\n\nBecause an existing redundancy in your series is that you'll do the\nfsync() the same way for \"git unpack-objects\" as for \"git\n{update-index,add}\".\n\nI.e. in the former case adding the N objects is all we're doing, so the\n\"last object\" is the point at which we need to flush the previous N to\ndisk.\n\nBut for \"update-index/add\" you'll do at least 2 fsync()'s in the bulk\nmode, when it should be one. I.e. the equivalent of (leaving aside the\ntmp-objdir migration part of it), if writing objects A && B:\n\n    ## METHOD ONE\n    # A\n    write(objects/A.tmp)\n    bulk_fsync(objects/A.tmp)\n    rename(objects/A.tmp, objects/A)\n    # B\n    write(objects/B.tmp)\n    bulk_fsync(objects/B.tmp)\n    rename(objects/B.tmp, objects/B)\n    # \"cookie\"\n    write(bulk_fsync_XXXXXX)\n    fsync(bulk_fsync_XXXXXX)\n    # ref\n    write(INDEX.tmp, $(git rev-parse B))\n    fsync(INDEX.tmp)\n    rename(INDEX.tmp, INDEX)\n\nThis series on top changes that so we know that we're doing N, so we\ndon't need the seperate \"cookie\", we can just use the B object as the\ncookie, as we know it comes last;\n\n    ## METHOD TWO\n    # A -- SAME as above\n    write(objects/A.tmp)\n    bulk_fsync(objects/A.tmp)\n    rename(objects/A.tmp, objects/A)\n    # B -- SAME as above, with s/bulk_fsync/fsync/\n    write(objects/B.tmp)\n    fsync(objects/B.tmp)\n    rename(objects/B.tmp, objects/B)\n    # \"cookie\" -- GONE!\n    # ref -- SAME\n    write(INDEX.tmp, $(git rev-parse B))\n    fsync(INDEX.tmp)\n    rename(INDEX.tmp, INDEX)\n\nBut really, we should instead realize that we're not doing\n\"unpack-objects\", but have a \"ref update\" at the end (whether that's a\nref, or an index etc.) and do:\n\n    ## METHOD THREE\n    # A -- SAME as above\n    write(objects/A.tmp)\n    bulk_fsync(objects/A.tmp)\n    rename(objects/A.tmp, objects/A)\n    # B -- SAME as the first\n    write(objects/B.tmp)\n    bulk_fsync(objects/B.tmp)\n    rename(objects/B.tmp, objects/B)\n    # ref -- SAME\n    write(INDEX.tmp, $(git rev-parse B))\n    fsync(INDEX.tmp)\n    rename(INDEX.tmp, INDEX)\n\nWhich cuts our number of fsync() operations down from 2 to 1, ina\naddition to removing the need for the \"cookie\", which is only there\nbecause we didn't keep track of where we were in the sequence as in my\n2/7 and 5/7.\n\nAnd it would be the same for tmp-objdir, the rename dance is a bit\ndifferent, but we'd do the \"full\" fsync() while on the INDEX.tmp, then\nmigrate() the tmp-objdir, and once that's done do the final:\n\n    rename(INDEX.tmp, INDEX)\n\nI.e. we'd fsync() the content once, and only have the renme() or link()\noperations left. For POSIX we'd need a few more fsync() for the\nmetadata, but this (i.e. your) series already makes the hard assumption\nthat we don't need to do that for rename().\n\n> I think the code you've presented here is a lot of diff to accomplish\n> the same thing that my series does, where this specific update-index\n> caller has been roto-tilled to provide the needed\n> begin/end-transaction points.\n\nAny caller of these APIs will need the \"unsigned oflags\" sooner than\nlater anyway, as they need to pass down e.g. HASH_WRITE_OBJECT. We just\ndo it slightly earlier.\n\nAnd because of that in the general case it's really not the same, I\nthink it's a better approach. You've already got the bug in yours of\nneedlessly setting up the bulk checkin for !HASH_WRITE_OBJECT in\nupdate-index, which this neatly solves by deferring the \"bulk\" mechanism\nuntil the codepath that's past that and into the \"real\" object writing.\n\nWe can also die() or error out in the object writing before ever getting\nto writing the object, in which case we'd do some setup that we'd need\nto tear down again, by deferring it until the last moment...\n\n> And I think there will be a lot of\n> complexity in supporting the same hints for command-line additions\n> (which is roughly equivalent to the git-add workflow).\n\nI left that out due to Junio's comment in\nhttps://lore.kernel.org/git/xmqqzgljyz34.fsf@gitster.g/; i.e. I don't\nsee why we'd find it worthwhile to optimize that case, but we easily\ncould (especially per the \"just sync the INDEX.tmp\" above).\n\nBut even if we don't do \"THREE\" above I think it's still easy, for \"TWO\"\nwe already have as parse_options() state machine to parse argv as it\ncomes in. Doing the fsync() on the last object is just a matter of\n\"looking ahead\" there).\n\n> Every caller\n> that wants batch treatment will have to either implement a state\n> machine or implement a buffering mechanism in order to figure out the\n> begin-end points. Having a separate plug/unplug call eliminates this\n> complexity on each caller.\n\nThis is subjective, but I really think that's rather easy to do, and\nmuch easier to reason about than the global state on the side via\nsingletons that your method of avoiding modifying these callers and\ninstead having them all consult global state via bulk-checkin.c and\ncache.h demands.\n\nThat API also currently assumes single-threaded writers, if we start\nwriting some of this in parallel in e.g. \"unpack-objects\" we'd need\nmutexes in bulk-object.[ch]. Isn't that a lot easier when the caller\nwould instead know something about the special nature of the transaction\nthey're interacting with, and that the 1st and last item are important\n(for a \"BEGIN\" and \"FLUSH\").\n\n> Btw, I'm planning in a future series to reduce the system calls\n> involved in renaming a file by taking advantage of the renameat2\n> system call and equivalents on other platforms.  There's a pretty\n> strong motivation to do that on Windows.\n\nWhat do you have in mind for renameat2() specifically?  I.e. which of\nthe 3x flags it implements will benefit us? RENAME_NOREPLACE to \"move\"\nthe tmp_OBJ to an eventual OBJ?\n\nGenerally: There's some low-hanging fruit there. E.g. for tmp-objdir we\nslavishly go through the motion of writing an tmp_OBJ, writing (and\npossibly syncing it), then renaming that tmp_OBJ to OBJ.\n\nWe could clearly just avoid that in some/all cases that use\ntmp-objdir. I.e. we're writing to a temporary store anyway, so why the\ntmp_OBJ files? We could just write to the final destinations instead,\nthey're not reachable (by ref or OID lookup) from anyone else yet.\n\nBut even then I don't see how you'd get away with reducing some classes\nof syscalls past the 2x increase for some (leading an overall increase,\nbut not a ~2x overall increase as noted in:\nhttps://lore.kernel.org/git/RFC-patch-7.7-481f1d771cb-20220323T033928Z-avarab@gmail.com/)\nas long as you use the tmp-objdir API. It's always going to have to\nwrite tmpdir/OBJ and link()/rename() that to OBJ.\n\nNow, I do think there's an easy way by extending the API use I've\nintroduced in this RFC to do it. I.e. we'd just do:\n\n    ## METHOD FOUR\n    # A -- SAME as THREE, except no rename()\n    write(objects/A.tmp)\n    bulk_fsync(objects/A.tmp)\n    # B -- SAME as THREE, except no rename()\n    write(objects/B.tmp)\n    bulk_fsync(objects/B.tmp)\n    # ref -- SAME\n    write(INDEX.tmp, $(git rev-parse B))\n    fsync(INDEX.tmp)\n    # NEW: do all the renames at the end:\n    rename(objects/A.tmp, objects/A)\n    rename(objects/B.tmp, objects/B)\n    rename(INDEX.tmp, INDEX)\n\nThat seems like an obvious win to me in any case. I.e. the tmp-objdir\nAPI isn't really a close fit for what we *really* want to do in this\ncase.\n\nI.e. the reason it does everything this way is because it was explicitly\ndesigned for 722ff7f876c (receive-pack: quarantine objects until\npre-receive accepts, 2016-10-03), where it's the right trade-off,\nbecause we'd like to cheaply \"rm -rf\" the whole thing if e.g. the\n\"pre-receive\" hook rejects the push.\n\n*AND* because it's made for the case of other things concurrently\nneeding access to those objects. So pedantically you would need it for\nsome modes of \"git update-index\", but not e.g. \"git unpack-objects\"\nwhere we really are expecting to keep all of them.\n\n> Thanks for the concrete code,\n\n..but no thanks? I.e. it would be useful to explicitly know if you're\ninterested or open to running with some of the approach in this RFC.\n"},{"id":"451978","messageId":"220323.86k0ckol2v.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"3ed1dcd9b9ba9b34f26b3012eaba8da0269ee842.1647760560.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T13:26:50Z","receivedAt":"2022-03-23T13:27:43Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n> [..\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 889522956e4..a3798dfc334 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -628,6 +628,13 @@ core.fsyncMethod::\n>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>    filesystem and storage hardware, data added to the repository may not be\n>    durable in the event of a system crash. This is the default mode on macOS.\n> +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> +  updates in the disk writeback cache and then does a single full fsync of\n> +  a dummy file to trigger the disk cache flush at the end of the operation.\n\nI think adding a \\n\\n here would help make this more readable & break\nthe flow a bit. I.e. just add a \"+\" on its own line, followed by\n\"Currently...\n\n> +  Currently `batch` mode only applies to loose-object files. Other repository\n> +  data is made durable as if `fsync` was specified. This mode is expected to\n> +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> +  and on Windows for repos stored on NTFS or ReFS filesystems.\n"},{"id":"451982","messageId":"RFC-patch-v2-1.7-98921aa2052-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 1/7] unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:25Z","receivedAt":"2022-03-23T14:18:50Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"In preparation for making the bulk-checkin.c logic operate from\nobject-file.c itself in some common cases let's add\nHASH_N_OBJECTS{,_{FIRST,LAST}} flags.\n\nThis will allow us to adjust for-loops that add N objects to just pass\ndown whether they have >1 objects (HASH_N_OBJECTS), as well as passing\ndown flags for whether we have the first or last object.\n\nWe'll thus be able to drive any sort of batch-object mechanism from\nwrite_object_file_flags() directly, which until now didn't know if it\nwas doing one object, or some arbitrary N.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c | 60 +++++++++++++++++++++++-----------------\n cache.h                  |  3 ++\n 2 files changed, 37 insertions(+), 26 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex c55b6616aed..ec40c6fd966 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -233,7 +233,8 @@ static void write_rest(void)\n }\n \n static void added_object(unsigned nr, enum object_type type,\n-\t\t\t void *data, unsigned long size);\n+\t\t\t void *data, unsigned long size,\n+\t\t\t unsigned oflags);\n \n /*\n  * Write out nr-th object from the list, now we know the contents\n@@ -241,21 +242,21 @@ static void added_object(unsigned nr, enum object_type type,\n  * to be checked at the end.\n  */\n static void write_object(unsigned nr, enum object_type type,\n-\t\t\t void *buf, unsigned long size)\n+\t\t\t void *buf, unsigned long size, unsigned oflags)\n {\n \tif (!strict) {\n-\t\tif (write_object_file(buf, size, type,\n-\t\t\t\t      &obj_list[nr].oid) < 0)\n+\t\tif (write_object_file_flags(buf, size, type,\n+\t\t\t\t      &obj_list[nr].oid, oflags) < 0)\n \t\t\tdie(\"failed to write object\");\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tfree(buf);\n \t\tobj_list[nr].obj = NULL;\n \t} else if (type == OBJ_BLOB) {\n \t\tstruct blob *blob;\n-\t\tif (write_object_file(buf, size, type,\n-\t\t\t\t      &obj_list[nr].oid) < 0)\n+\t\tif (write_object_file_flags(buf, size, type,\n+\t\t\t\t\t    &obj_list[nr].oid, oflags) < 0)\n \t\t\tdie(\"failed to write object\");\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tfree(buf);\n \n \t\tblob = lookup_blob(the_repository, &obj_list[nr].oid);\n@@ -269,7 +270,7 @@ static void write_object(unsigned nr, enum object_type type,\n \t\tint eaten;\n \t\thash_object_file(the_hash_algo, buf, size, type,\n \t\t\t\t &obj_list[nr].oid);\n-\t\tadded_object(nr, type, buf, size);\n+\t\tadded_object(nr, type, buf, size, oflags);\n \t\tobj = parse_object_buffer(the_repository, &obj_list[nr].oid,\n \t\t\t\t\t  type, size, buf,\n \t\t\t\t\t  &eaten);\n@@ -283,7 +284,7 @@ static void write_object(unsigned nr, enum object_type type,\n \n static void resolve_delta(unsigned nr, enum object_type type,\n \t\t\t  void *base, unsigned long base_size,\n-\t\t\t  void *delta, unsigned long delta_size)\n+\t\t\t  void *delta, unsigned long delta_size, unsigned oflags)\n {\n \tvoid *result;\n \tunsigned long result_size;\n@@ -294,7 +295,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n \tif (!result)\n \t\tdie(\"failed to apply delta\");\n \tfree(delta);\n-\twrite_object(nr, type, result, result_size);\n+\twrite_object(nr, type, result, result_size, oflags);\n }\n \n /*\n@@ -302,7 +303,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n  * resolve all the deltified objects that are based on it.\n  */\n static void added_object(unsigned nr, enum object_type type,\n-\t\t\t void *data, unsigned long size)\n+\t\t\t void *data, unsigned long size, unsigned oflags)\n {\n \tstruct delta_info **p = &delta_list;\n \tstruct delta_info *info;\n@@ -313,7 +314,7 @@ static void added_object(unsigned nr, enum object_type type,\n \t\t\t*p = info->next;\n \t\t\tp = &delta_list;\n \t\t\tresolve_delta(info->nr, type, data, size,\n-\t\t\t\t      info->delta, info->size);\n+\t\t\t\t      info->delta, info->size, oflags);\n \t\t\tfree(info);\n \t\t\tcontinue;\n \t\t}\n@@ -322,18 +323,19 @@ static void added_object(unsigned nr, enum object_type type,\n }\n \n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n-\t\t\t\t   unsigned nr)\n+\t\t\t\t   unsigned nr, unsigned oflags)\n {\n \tvoid *buf = get_data(size);\n \n \tif (!dry_run && buf)\n-\t\twrite_object(nr, type, buf, size);\n+\t\twrite_object(nr, type, buf, size, oflags);\n \telse\n \t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n-\t\t\t\tvoid *delta_data, unsigned long delta_size)\n+\t\t\t\tvoid *delta_data, unsigned long delta_size,\n+\t\t\t\tunsigned oflags)\n {\n \tstruct object *obj;\n \tstruct obj_buffer *obj_buffer;\n@@ -344,12 +346,12 @@ static int resolve_against_held(unsigned nr, const struct object_id *base,\n \tif (!obj_buffer)\n \t\treturn 0;\n \tresolve_delta(nr, obj->type, obj_buffer->buffer,\n-\t\t      obj_buffer->size, delta_data, delta_size);\n+\t\t      obj_buffer->size, delta_data, delta_size, oflags);\n \treturn 1;\n }\n \n static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n-\t\t\t       unsigned nr)\n+\t\t\t       unsigned nr, unsigned oflags)\n {\n \tvoid *delta_data, *base;\n \tunsigned long base_size;\n@@ -366,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n-\t\t\t\t\t      delta_data, delta_size))\n+\t\t\t\t\t      delta_data, delta_size, oflags))\n \t\t\treturn; /* we are done */\n \t\telse {\n \t\t\t/* cannot resolve yet --- queue it */\n@@ -428,7 +430,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t}\n \t}\n \n-\tif (resolve_against_held(nr, &base_oid, delta_data, delta_size))\n+\tif (resolve_against_held(nr, &base_oid, delta_data, delta_size, oflags))\n \t\treturn;\n \n \tbase = read_object_file(&base_oid, &type, &base_size);\n@@ -440,11 +442,11 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\thas_errors = 1;\n \t\treturn;\n \t}\n-\tresolve_delta(nr, type, base, base_size, delta_data, delta_size);\n+\tresolve_delta(nr, type, base, base_size, delta_data, delta_size, oflags);\n \tfree(base);\n }\n \n-static void unpack_one(unsigned nr)\n+static void unpack_one(unsigned nr, unsigned oflags)\n {\n \tunsigned shift;\n \tunsigned char *pack;\n@@ -472,11 +474,11 @@ static void unpack_one(unsigned nr)\n \tcase OBJ_TREE:\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tunpack_non_delta_entry(type, size, nr);\n+\t\tunpack_non_delta_entry(type, size, nr, oflags);\n \t\treturn;\n \tcase OBJ_REF_DELTA:\n \tcase OBJ_OFS_DELTA:\n-\t\tunpack_delta_entry(type, size, nr);\n+\t\tunpack_delta_entry(type, size, nr, oflags);\n \t\treturn;\n \tdefault:\n \t\terror(\"bad object type %d\", type);\n@@ -491,6 +493,7 @@ static void unpack_all(void)\n {\n \tint i;\n \tstruct pack_header *hdr = fill(sizeof(struct pack_header));\n+\tunsigned oflags;\n \n \tnr_objects = ntohl(hdr->hdr_entries);\n \n@@ -505,9 +508,14 @@ static void unpack_all(void)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n \tplug_bulk_checkin();\n+\toflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n \tfor (i = 0; i < nr_objects; i++) {\n-\t\tunpack_one(i);\n-\t\tdisplay_progress(progress, i + 1);\n+\t\tint nth = i + 1;\n+\t\tunsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n+\t\t\tnr_objects == nth ? HASH_N_OBJECTS_LAST : 0;\n+\n+\t\tunpack_one(i, oflags | f);\n+\t\tdisplay_progress(progress, nth);\n \t}\n \tunplug_bulk_checkin();\n \tstop_progress(&progress);\ndiff --git a/cache.h b/cache.h\nindex 84fafe2ed71..72c91c91286 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -896,6 +896,9 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n #define HASH_FORMAT_CHECK 2\n #define HASH_RENORMALIZE  4\n #define HASH_SILENT 8\n+#define HASH_N_OBJECTS 1<<4\n+#define HASH_N_OBJECTS_FIRST 1<<5\n+#define HASH_N_OBJECTS_LAST 1<<6\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451983","messageId":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-0.7-00000000000-20220323T033928Z-avarab@gmail.com","subject":"[RFC PATCH v2 0/7] bottom-up ns/batched-fsync & \"plugging\" in object-file.c","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:24Z","receivedAt":"2022-03-23T14:18:55Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Quite a bit less WIP-y but still RFC version of patcehs to\nintegrate/squash into some form in Neeraj's fsync() series at\nhttps://lore.kernel.org/git/pull.1134.v2.git.1647760560.gitgitgadget@gmail.com/\n\nAs noted in the v1 this starts (in 2/7) by removing the tmp-objdir\npart of the \"bulk checkin\" as a POC. Clearly Neeraj wants to keep it,\nso we should have it eventually. But this series argues that in both\npatch organization and configurability (see the new 7/7!) that the \"do\nquarantine\" should be split up and optional to \"do bulk fsync\".\n\nThe new documentation in 7/7 currently documents a vaporware setting,\nbut it's what we'd get if this were rebased early into Neeraj's\nseries, and we made the tmp-objdir part contingent on a configuration\nsetting.\n\nUnlike v1 (I overzealously ripped out some unrelated bulk-checkin.c\ncode then) this doesn't fail any tests.\n\nBut most importantly the whole fsync() schema here is *much better* in\nterms of semantics. We still do away with the \"cookie\" placeholder to\nforce an fsync, but now as can be seen in the 4/7 and 5/7 we'll\n\"fsync()\" by using the updated index at the end as our cookie.\n\nI.e. there's no need to introduce a \"bulk_fsync\" cookie file to force\nan fsync() if we can instead alter the relevant calling code to be\naware of the new \"fsync() transaction\". It can then do the \"flush\" by\ndoing the fsync() on the file it wanted to update anyway (now an\nindex, in the future a ref). So this implements the \"METHOD THREE\"\nnoted in [1].\n\n\n\n1. https://lore.kernel.org/git/220323.86sfr9ndpr.gmgdl@evledraar.gmail.com/\n\nÆvar Arnfjörð Bjarmason (7):\n  unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n  object-file: pass down unpack-objects.c flags for \"bulk\" checkin\n  update-index: pass down skeleton \"oflags\" argument\n  update-index: have the index fsync() flush the loose objects\n  add: use WLI_NEED_LOOSE_FSYNC for new \"only the index\" bulk fsync()\n  fsync docs: update for new syncing semantics\n  fsync docs: add new fsyncMethod.batch.quarantine, elaborate on old\n\n Documentation/config/core.txt | 101 +++++++++++++++++++++++++++++-----\n builtin/add.c                 |   6 +-\n builtin/unpack-objects.c      |  62 +++++++++++----------\n builtin/update-index.c        |  39 ++++++-------\n bulk-checkin.c                |  74 -------------------------\n bulk-checkin.h                |   3 -\n cache.h                       |  10 ++--\n object-file.c                 |  39 +++++++++----\n read-cache.c                  |  37 ++++++++++++-\n 9 files changed, 214 insertions(+), 157 deletions(-)\n\nRange-diff against v1:\n1:  e03c119c784 < -:  ----------- write-or-die.c: remove unused fsync_component() function\n2:  00dbffc2331 = 1:  98921aa2052 unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n3:  beda9f99529 ! 2:  c6f776fc2bc object-file: pass down unpack-objects.c flags for \"bulk\" checkin\n    @@ Commit message\n         the previous case of fsync_component_or_die(...)\" could just be added\n         to the existing \"fsync_object_files > 0\" branch.\n     \n    -    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n    +    Note: This commit reverts much of \"core.fsyncmethod: batched disk\n    +    flushes for loose-objects\". We'll set up new structures to bring what\n    +    it was doing back in a different way. I.e. to do the tmp-objdir\n    +    plug-in in object-file.c\n     \n    - ## builtin/add.c ##\n    -@@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n    - \t\tstring_list_clear(&only_match_skip_worktree, 0);\n    - \t}\n    - \n    --\tplug_bulk_checkin();\n    --\n    - \tif (add_renormalize)\n    - \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n    - \telse\n    -@@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n    - \n    - \tif (chmod_arg && pathspec.nr)\n    - \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n    --\tunplug_bulk_checkin();\n    - \n    - finish:\n    - \tif (write_locked_index(&the_index, &lock_file,\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void unpack_all(void)\n    @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n     \n      ## bulk-checkin.c ##\n     @@\n    +  */\n    + #include \"cache.h\"\n    + #include \"bulk-checkin.h\"\n    +-#include \"lockfile.h\"\n    + #include \"repository.h\"\n    + #include \"csum-file.h\"\n      #include \"pack.h\"\n      #include \"strbuf.h\"\n    - #include \"string-list.h\"\n    +-#include \"string-list.h\"\n     -#include \"tmp-objdir.h\"\n      #include \"packfile.h\"\n      #include \"object-store.h\"\n    @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      \t\t       int fd, size_t size, enum object_type type,\n      \t\t       const char *path, unsigned flags)\n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n    - \t\tfinish_bulk_checkin(&bulk_checkin_state);\n    - \treturn status;\n    - }\n    --\n    --void plug_bulk_checkin(void)\n    --{\n    --\tassert(!bulk_checkin_plugged);\n    + void plug_bulk_checkin(void)\n    + {\n    + \tassert(!bulk_checkin_plugged);\n     -\n     -\t/*\n     -\t * A temporary object directory is used to hold the files\n    @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n     -\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n     -\t}\n     -\n    --\tbulk_checkin_plugged = 1;\n    --}\n    --\n    --void unplug_bulk_checkin(void)\n    --{\n    --\tassert(bulk_checkin_plugged);\n    --\tbulk_checkin_plugged = 0;\n    --\tif (bulk_checkin_state.f)\n    --\t\tfinish_bulk_checkin(&bulk_checkin_state);\n    + \tbulk_checkin_plugged = 1;\n    + }\n    + \n    +@@ bulk-checkin.c: void unplug_bulk_checkin(void)\n    + \tbulk_checkin_plugged = 0;\n    + \tif (bulk_checkin_state.f)\n    + \t\tfinish_bulk_checkin(&bulk_checkin_state);\n     -\n     -\tdo_batch_fsync();\n    --}\n    + }\n     \n      ## bulk-checkin.h ##\n     @@\n    @@ bulk-checkin.h\n      int index_bulk_checkin(struct object_id *oid,\n      \t\t       int fd, size_t size, enum object_type type,\n      \t\t       const char *path, unsigned flags);\n    - \n    --void plug_bulk_checkin(void);\n    --void unplug_bulk_checkin(void);\n    --\n    - #endif\n     \n      ## cache.h ##\n    -@@ cache.h: void write_or_die(int fd, const void *buf, size_t count);\n    - void fsync_or_die(int fd, const char *);\n    +@@ cache.h: void fsync_or_die(int fd, const char *);\n    + int fsync_component(enum fsync_component component, int fd);\n      void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n      \n     -static inline int batch_fsync_enabled(enum fsync_component component)\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \n      \tif (mtime) {\n      \t\tstruct utimbuf utb;\n    -\n    - ## t/t1050-large.sh ##\n    -@@ t/t1050-large.sh: test_description='adding and checking out large blobs'\n    - \n    - . ./test-lib.sh\n    - \n    -+skip_all='TODO: migrate the builtin/add.c code'\n    -+test_done\n    -+\n    - test_expect_success setup '\n    - \t# clone does not allow us to pass core.bigfilethreshold to\n    - \t# new repos, so set core.bigfilethreshold globally\n5:  a1474968991 ! 3:  4df8012100a update-index: pass down an \"oflags\" argument\n    @@ Metadata\n     Author: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Commit message ##\n    -    update-index: pass down an \"oflags\" argument\n    +    update-index: pass down skeleton \"oflags\" argument\n     \n    -    We do nothing with this yet, but will soon.\n    +    As with a preceding change to \"unpack-objects\" add an \"oflags\" going\n    +    from cmd_update_index() all the way down to the code in\n    +    object-file.c. Note also how index_mem() will now call\n    +    write_object_file_flags().\n     \n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n    @@ builtin/update-index.c: static int do_reupdate(int ac, const char **av,\n      \t\tfree(path);\n      \t\tdiscard_cache_entry(old);\n      \t\tif (save_nr != active_nr)\n    -@@ builtin/update-index.c: static enum parse_opt_result reupdate_callback(\n    - \n    - static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n    - \t\t\t    const char *prefix, int prefix_length,\n    --\t\t\t    const int nul_term_line, const int set_executable_bit)\n    -+\t\t\t    const int nul_term_line, const int set_executable_bit,\n    -+\t\t\t    const unsigned oflags)\n    - {\n    - \tchar *p;\n    - \n    -@@ builtin/update-index.c: static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n    - \t\tstrbuf_swap(buf, unquoted);\n    - \t}\n    - \tp = prefix_path(prefix, prefix_length, buf->buf);\n    --\tupdate_one(p);\n    -+\tupdate_one(p, oflags);\n    - \tif (set_executable_bit)\n    - \t\tchmod_path(set_executable_bit, p);\n    - \tfree(p);\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n      \n      \t\t\tsetup_work_tree();\n    @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n      \t\t\t\tchmod_path(set_executable_bit, p);\n      \t\t\tfree(p);\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    - \t\tsetup_work_tree();\n    - \t\twhile (getline_fn(&buf, stdin) != EOF)\n    - \t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n    --\t\t\t\t\tnul_term_line, set_executable_bit);\n    -+\t\t\t\t\tnul_term_line, set_executable_bit, 0);\n    - \t\tstrbuf_release(&unquoted);\n    - \t\tstrbuf_release(&buf);\n    - \t}\n    + \t\t\t\tstrbuf_swap(&buf, &unquoted);\n    + \t\t\t}\n    + \t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n    +-\t\t\tupdate_one(p);\n    ++\t\t\tupdate_one(p, 0);\n    + \t\t\tif (set_executable_bit)\n    + \t\t\t\tchmod_path(set_executable_bit, p);\n    + \t\t\tfree(p);\n     \n      ## object-file.c ##\n     @@ object-file.c: static int index_mem(struct index_state *istate,\n7:  481f1d771cb ! 4:  61f4f3d7ef4 update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n    @@ Metadata\n     Author: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Commit message ##\n    -    update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags\n    +    update-index: have the index fsync() flush the loose objects\n     \n         As with unpack-objects in a preceding commit have update-index.c make\n         use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\n    @@ Commit message\n         Adding the t/* directory from git.git on a Linux ramdisk is a bit\n         faster than with the tmp-objdir indirection:\n     \n    -            git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10\n    -            Benchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync\n    -              Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]\n    -              Range (min … max):   285.6 ms … 297.0 ms    10 runs\n    +            $ git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt' -p 'rm -rf repo/.git/objects/* repo/.git/index' './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' --warmup 1 -r 10Benchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync\n    +              Time (mean ± σ):     281.1 ms ±   2.6 ms    [User: 186.2 ms, System: 92.3 ms]\n    +              Range (min … max):   278.3 ms … 287.0 ms    10 runs\n     \n    -            Benchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD\n    -              Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]\n    -              Range (min … max):   267.8 ms … 291.3 ms    10 runs\n    +            Benchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD\n    +              Time (mean ± σ):     265.9 ms ±   2.6 ms    [User: 181.7 ms, System: 82.1 ms]\n    +              Range (min … max):   262.0 ms … 270.3 ms    10 runs\n     \n                 Summary\n    -              'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n    -                1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n    +              './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n    +                1.06 ± 0.01 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n     \n         And as before running that with \"strace --summary-only\" slows things\n         down a bit (probably mimicking slower I/O a bit). I then get:\n     \n                 Summary\n    -              'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n    -                1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n    +              'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n    +                1.19 ± 0.03 times faster than 'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n    +\n    +    This one has a twist though, instead of fsync()-ing on the last object\n    +    we write we'll not do that, and instead defer the fsync() until we\n    +    write the index itself. This is outlined in [1] (as \"METHOD THREE\").\n    +\n    +    Because of this under FSYNC_METHOD_BATCH we'll do the N\n    +    objects (possibly only one, because we're lazy) as HASH_N_OBJECTS, and\n    +    we'll even now support doing this via N arguments on the command-line.\n    +\n    +    Then we won't fsync() any of it, but we will rename it\n    +    in-place (which, if we were still using the tmp-objdir, would leave it\n    +    \"staged\" in the tmp-objdir).\n    +\n    +    We'll then have the fsync() for the index update \"flush\" that out, and\n    +    thus avoid two fsync() calls when one will do.\n    +\n    +    Running this with the \"git hyperfine\" command mentioned in a preceding\n    +    commit with \"strace --summary-only\" shows that we do 1 fsync() now\n    +    instead of 2, and have one more sync_file_range(), as expected.\n     \n         We also go from ~51k syscalls to ~39k, with ~2x the number of link()\n    -    and unlink() in ns/batched-fsync.\n    +    and unlink() in ns/batched-fsync, and of course one fsync() instead of\n    +    two()>\n     \n    -    In the process of doing this conversion we lost the \"bulk\" mode for\n    -    files added on the command-line. I don't think it's useful to optimize\n    -    that, but we could if anyone cared.\n    +    The flow of this code isn't quite set up for re-plugging the\n    +    tmp-objdir back in. In particular we no longer pass\n    +    HASH_N_OBJECTS_FIRST (but doing so would be trivial)< and there's no\n    +    HASH_N_OBJECTS_LAST.\n     \n    -    We've also converted this to a string_list, we could walk with\n    -    getline_fn() and get one line \"ahead\" to see what we have left, but I\n    -    found that state machine a bit painful, and at least in my testing\n    -    buffering this doesn't harm things. But we could also change this to\n    -    stream again, at the cost of some getline_fn() twiddling.\n    +    So this and other callers would need some light transaction-y API, or\n    +    to otherwise pass down a \"yes, I'd like to flush it\" down to\n    +    finalize_hashfile(), but doing so will be trivial.\n    +\n    +    And since we've started structuring it this way it'll become easy to\n    +    do any arbitrary number of things down the line that would \"bulk\n    +    fsync\" before the final fsync(). Now we write some objects and fsync()\n    +    on the index, but between those two could do any number of other\n    +    things where we'd defer the fsync().\n    +\n    +    This sort of thing might be especially interesting for \"git repack\"\n    +    when it writes e.g. a *.bitmap, *.rev, *.pack and *.idx. In that case\n    +    we could skip the fsync() on all of those, and only do it on the *.idx\n    +    before we renamed it in-place. I *think* nothing cares about a *.pack\n    +    without an *.idx, but even then we could fsync *.idx, rename *.pack,\n    +    rename *.idx and still safely do only one fsync(). See \"git show\n    +    --first-parent\" on 62874602032 (Merge branch\n    +    'tb/pack-finalize-ordering' into maint, 2021-10-12) for a good\n    +    overview of the code involved in that.\n    +\n    +    1. https://lore.kernel.org/git/220323.86sfr9ndpr.gmgdl@evledraar.gmail.com/\n     \n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## builtin/update-index.c ##\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    + \n    + \t\t\tsetup_work_tree();\n    + \t\t\tp = prefix_path(prefix, prefix_length, path);\n    +-\t\t\tupdate_one(p, 0);\n    ++\t\t\tupdate_one(p, HASH_N_OBJECTS);\n    + \t\t\tif (set_executable_bit)\n    + \t\t\t\tchmod_path(set_executable_bit, p);\n    + \t\t\tfree(p);\n    +@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    + \t\t\t\tstrbuf_swap(&buf, &unquoted);\n    + \t\t\t}\n    + \t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n    +-\t\t\tupdate_one(p, 0);\n    ++\t\t\tupdate_one(p, HASH_N_OBJECTS);\n    + \t\t\tif (set_executable_bit)\n    + \t\t\t\tchmod_path(set_executable_bit, p);\n    + \t\t\tfree(p);\n    +@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    + \t\t\t\texit(128);\n    + \t\t\tunable_to_lock_die(get_index_file(), lock_error);\n    + \t\t}\n    +-\t\tif (write_locked_index(&the_index, &lock_file, COMMIT_LOCK))\n    ++\t\tif (write_locked_index(&the_index, &lock_file,\n    ++\t\t\t\t       COMMIT_LOCK | WLI_NEED_LOOSE_FSYNC))\n    + \t\t\tdie(\"Unable to write new index file\");\n      \t}\n      \n    - \tif (read_from_stdin) {\n    -+\t\tstruct string_list list = STRING_LIST_INIT_NODUP;\n    - \t\tstruct strbuf line = STRBUF_INIT;\n    - \t\tstruct strbuf unquoted = STRBUF_INIT;\n    -+\t\tsize_t i, nr;\n    -+\t\tunsigned oflags;\n    +\n    + ## cache.h ##\n    +@@ cache.h: void ensure_full_index(struct index_state *istate);\n    + /* For use with `write_locked_index()`. */\n    + #define COMMIT_LOCK\t\t(1 << 0)\n    + #define SKIP_IF_UNCHANGED\t(1 << 1)\n    ++#define WLI_NEED_LOOSE_FSYNC\t(1 << 2)\n      \n    - \t\tsetup_work_tree();\n    --\t\twhile (getline_fn(&line, stdin) != EOF)\n    --\t\t\tline_from_stdin(&line, &unquoted, prefix, prefix_length,\n    --\t\t\t\t\tnul_term_line, set_executable_bit, 0);\n    -+\t\twhile (getline_fn(&line, stdin) != EOF) {\n    -+\t\t\tsize_t len = line.len;\n    -+\t\t\tchar *str = strbuf_detach(&line, NULL);\n    -+\n    -+\t\t\tstring_list_append_nodup(&list, str)->util = (void *)len;\n    -+\t\t}\n    + /*\n    +  * Write the index while holding an already-taken lock. Close the lock,\n    +\n    + ## read-cache.c ##\n    +@@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n    + \tint ieot_entries = 1;\n    + \tstruct index_entry_offset_table *ieot = NULL;\n    + \tint nr, nr_threads;\n    ++\tunsigned int wflags = FSYNC_COMPONENT_INDEX;\n     +\n    -+\t\tnr = list.nr;\n    -+\t\toflags = nr > 1 ? HASH_N_OBJECTS : 0;\n    -+\t\tfor (i = 0; i < nr; i++) {\n    -+\t\t\tsize_t nth = i + 1;\n    -+\t\t\tunsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n    -+\t\t\t\t  nr == nth ? HASH_N_OBJECTS_LAST : 0;\n    -+\t\t\tstruct strbuf buf = STRBUF_INIT;\n    -+\t\t\tstruct string_list_item *item = list.items + i;\n    -+\t\t\tconst size_t len = (size_t)item->util;\n     +\n    -+\t\t\tstrbuf_attach(&buf, item->string, len, len);\n    -+\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n    -+\t\t\t\t\tnul_term_line, set_executable_bit,\n    -+\t\t\t\t\toflags | f);\n    -+\t\t\tstrbuf_release(&buf);\n    -+\t\t}\n    - \t\tstrbuf_release(&unquoted);\n    --\t\tstrbuf_release(&line);\n    -+\t\tstring_list_clear(&list, 0);\n    - \t}\n    ++\t/*\n    ++\t * TODO: This is abuse of the API recently modified\n    ++\t * finalize_hashfile() which reveals a shortcoming of its\n    ++\t * \"fsync\" design.\n    ++\t * \n    ++\t * I.e. It expects a \"enum fsync_component component\" label,\n    ++\t * but here we're passing it an OR of the two, knowing that\n    ++\t * it'll call fsync_component_or_die() which (in\n    ++\t * write-or-die.c) will do \"(fsync_components & wflags)\" (to\n    ++\t * our \"wflags\" here).\n    ++\t *\n    ++\t * But the API really should be changed to explicitly take\n    ++\t * such flags, because in this case we'd like to fsync() the\n    ++\t * index if we're in the bulk mode, *even if* our\n    ++\t * \"core.fsync=index\" isn't configured.\n    ++\t *\n    ++\t * That's because at this point we've been queuing up object\n    ++\t * writes that we didn't fsync(), and are going to use this\n    ++\t * fsync() to \"flush\" the whole thing. Doing it this way\n    ++\t * avoids redundantly calling fsync() twice when once will do.\n    ++\t */\n    ++\tif (fsync_method == FSYNC_METHOD_BATCH && \n    ++\t    flags & WLI_NEED_LOOSE_FSYNC)\n    ++\t\twflags |= FSYNC_COMPONENT_LOOSE_OBJECT;\n    + \n    + \tf = hashfd(tempfile->fd, tempfile->filename.buf);\n    + \n    +@@ read-cache.c: static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n    + \tif (!alternate_index_output && (flags & COMMIT_LOCK))\n    + \t\tcsum_fsync_flag = CSUM_FSYNC;\n    + \n    +-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n    ++\tfinalize_hashfile(f, istate->oid.hash, wflags,\n    + \t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n      \n    - \tif (split_index > 0) {\n    + \tif (close_tempfile_gently(tempfile)) {\n-:  ----------- > 5:  2bf14fd4946 add: use WLI_NEED_LOOSE_FSYNC for new \"only the index\" bulk fsync()\n6:  4fad333e9a1 ! 6:  c20301d7967 update-index: rename \"buf\" to \"line\"\n    @@ Metadata\n     Author: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Commit message ##\n    -    update-index: rename \"buf\" to \"line\"\n    -\n    -    This variable renaming makes a subsequent more meaningful change\n    -    smaller.\n    +    fsync docs: update for new syncing semantics\n     \n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n    - ## builtin/update-index.c ##\n    -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    - \t}\n    - \n    - \tif (read_from_stdin) {\n    --\t\tstruct strbuf buf = STRBUF_INIT;\n    -+\t\tstruct strbuf line = STRBUF_INIT;\n    - \t\tstruct strbuf unquoted = STRBUF_INIT;\n    - \n    - \t\tsetup_work_tree();\n    --\t\twhile (getline_fn(&buf, stdin) != EOF)\n    --\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n    -+\t\twhile (getline_fn(&line, stdin) != EOF)\n    -+\t\t\tline_from_stdin(&line, &unquoted, prefix, prefix_length,\n    - \t\t\t\t\tnul_term_line, set_executable_bit, 0);\n    - \t\tstrbuf_release(&unquoted);\n    --\t\tstrbuf_release(&buf);\n    -+\t\tstrbuf_release(&line);\n    - \t}\n    + ## Documentation/config/core.txt ##\n    +@@ Documentation/config/core.txt: core.fsyncMethod::\n    +   filesystem and storage hardware, data added to the repository may not be\n    +   durable in the event of a system crash. This is the default mode on macOS.\n    + * `batch` enables a mode that uses writeout-only flushes to stage multiple\n    +-  updates in the disk writeback cache and then does a single full fsync of\n    +-  a dummy file to trigger the disk cache flush at the end of the operation.\n    +-  Currently `batch` mode only applies to loose-object files. Other repository\n    +-  data is made durable as if `fsync` was specified. This mode is expected to\n    +-  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n    +-  and on Windows for repos stored on NTFS or ReFS filesystems.\n    ++  updates in the disk writeback cache and, before doing a full fsync() of\n    ++  on the \"last\" file that to trigger the disk cache flush at the end of the\n    ++  operation.\n    +++\n    ++Other repository data is made durable as if `fsync` was\n    ++specified. This mode is expected to be as safe as `fsync` on macOS for\n    ++repos stored on HFS+ or APFS filesystems and on Windows for repos\n    ++stored on NTFS or ReFS filesystems.\n    +++\n    ++The `batch` is currently only applies to loose-object files and will\n    ++kick in when using the linkgit:git-unpack-objects[1] and\n    ++linkgit:update-index[1] commands. Note that the \"last\" file to be\n    ++synced may be the last object, as in the case of\n    ++linkgit:git-unpack-objects[1], or relevant \"index\" (or in the future,\n    ++\"ref\") update, as in the case of linkgit:git-update-index[1]. I.e. the\n    ++batch syncing of the loose objects may be deferred until a subsequent\n    ++fsync() to a file that makes them \"active\".\n      \n    - \tif (split_index > 0) {\n    + core.fsyncObjectFiles::\n    + \tThis boolean will enable 'fsync()' when writing object files.\n4:  2c5395a3716 ! 7:  a5951366c6e update-index: use a utility function for stdin consumption\n    @@ Metadata\n     Author: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Commit message ##\n    -    update-index: use a utility function for stdin consumption\n    +    fsync docs: add new fsyncMethod.batch.quarantine, elaborate on old\n    +\n    +    Add a new fsyncMethod.batch.quarantine setting which defaults to\n    +    \"false\". Preceding (RFC, and not meant to flip-flop like that\n    +    eventually) commits ripped out the \"tmp-objdir\" part of the\n    +    core.fsyncMethod=batch.\n    +\n    +    This documentation proposes to keep that as the default for the\n    +    reasons discussed in it, while allowing users to set\n    +    \"fsyncMethod.batch.quarantine=true\".\n    +\n    +    Furthermore update the discussion of \"core.fsyncObjectFiles\" with\n    +    information about what it *really* does, why you probably shouldn't\n    +    use it, and how to safely emulate most of what it gave users in the\n    +    past in terms of performance benefit.\n     \n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n    - ## builtin/update-index.c ##\n    -@@ builtin/update-index.c: static enum parse_opt_result reupdate_callback(\n    - \treturn 0;\n    - }\n    + ## Documentation/config/core.txt ##\n    +@@ Documentation/config/core.txt: stored on NTFS or ReFS filesystems.\n    + +\n    + The `batch` is currently only applies to loose-object files and will\n    + kick in when using the linkgit:git-unpack-objects[1] and\n    +-linkgit:update-index[1] commands. Note that the \"last\" file to be\n    ++linkgit:git-update-index[1] commands. Note that the \"last\" file to be\n    + synced may be the last object, as in the case of\n    + linkgit:git-unpack-objects[1], or relevant \"index\" (or in the future,\n    + \"ref\") update, as in the case of linkgit:git-update-index[1]. I.e. the\n    + batch syncing of the loose objects may be deferred until a subsequent\n    + fsync() to a file that makes them \"active\".\n      \n    -+static void line_from_stdin(struct strbuf *buf, struct strbuf *unquoted,\n    -+\t\t\t    const char *prefix, int prefix_length,\n    -+\t\t\t    const int nul_term_line, const int set_executable_bit)\n    -+{\n    -+\tchar *p;\n    -+\n    -+\tif (!nul_term_line && buf->buf[0] == '\"') {\n    -+\t\tstrbuf_reset(unquoted);\n    -+\t\tif (unquote_c_style(unquoted, buf->buf, NULL))\n    -+\t\t\tdie(\"line is badly quoted\");\n    -+\t\tstrbuf_swap(buf, unquoted);\n    -+\t}\n    -+\tp = prefix_path(prefix, prefix_length, buf->buf);\n    -+\tupdate_one(p);\n    -+\tif (set_executable_bit)\n    -+\t\tchmod_path(set_executable_bit, p);\n    -+\tfree(p);\n    -+}\n    ++fsyncMethod.batch.quarantine::\n    ++\tA boolean which if set to `true` will cause \"batched\" writes\n    ++\tto objects to be \"quarantined\" if\n    ++\t`core.fsyncMethod=batch`. This is `false` by default.\n    +++\n    ++The primary object of these fsync() settings is to protect against\n    ++repository corruption of things which are reachable, i.e. \"reachable\",\n    ++via references, the index etc. Not merely objects that were present in\n    ++the object store.\n    +++\n    ++Historically setting `core.fsyncObjectFiles=false` assumed that on a\n    ++filesystem with where an fsync() would flush all preceding outstanding\n    ++I/O that we might end up with a corrupt loose object, but that was OK\n    ++as long as no reference referred to it. We'd eventually the corrupt\n    ++object with linkgit:git-gc[1], and linkgit:git-fsck[1] would only\n    ++report it as a minor annoyance\n    +++\n    ++Setting `fsyncMethod.batch.quarantine=true` takes the view that\n    ++something like a corrupt *unreferenced* loose object in the object\n    ++store is something we'd like to avoid, at the cost of reduced\n    ++performance when using `core.fsyncMethod=batch`.\n    +++\n    ++Currently this uses the same mechanism described in the \"QUARANTINE\n    ++ENVIRONMENT\" in the linkgit:git-receive-pack[1] documentation, but\n    ++that's subject to change. The performance loss is because we need to\n    ++\"stage\" the objects in that quarantine environment, fsync() it, and\n    ++once that's done rename() or link() it in-place into the main object\n    ++store, possibly with an fsync() of the index or ref at the end\n    +++\n    ++With `fsyncMethod.batch.quarantine=false` we'll \"stage\" things in the\n    ++main object store, and then do one fsync() at the very end, either on\n    ++the last object we write, or file (index or ref) that'll make it\n    ++\"reachable\".\n    +++\n    ++The bad thing about setting this to `true` is lost performance, as\n    ++well as not being able to access the objects as they're written (which\n    ++e.g. consumers of linkgit:git-update-index[1]'s `--verbose` mode might\n    ++want to do).\n    +++\n    ++The good thing is that you should be guaranteed not to get e.g. short\n    ++or otherwise corrupt loose objects if you pull your power cord, in\n    ++practice various git commands deal quite badly with discovering such a\n    ++stray corrupt object (including perhaps assuming it's valid based on\n    ++its existence, or hard dying on an error rather than replacing\n    ++it). Repairing such \"unreachable corruption\" can require manual\n    ++intervention.\n     +\n    - int cmd_update_index(int argc, const char **argv, const char *prefix)\n    - {\n    - \tint newfd, entries, has_errors = 0, nul_term_line = 0;\n    -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    - \t\tstruct strbuf unquoted = STRBUF_INIT;\n    + core.fsyncObjectFiles::\n    +-\tThis boolean will enable 'fsync()' when writing object files.\n    +-\tThis setting is deprecated. Use core.fsync instead.\n    +-+\n    +-This setting affects data added to the Git repository in loose-object\n    +-form. When set to true, Git will issue an fsync or similar system call\n    +-to flush caches so that loose-objects remain consistent in the face\n    +-of a unclean system shutdown.\n    ++\tThis boolean will enable 'fsync()' when writing loose object\n    ++\tfiles.\n    +++\n    ++This setting is the historical fsync configuration setting. It's now\n    ++*deprecated*, you should use `core.fsync` instead, perhaps in\n    ++combination with `core.fsyncMethod=batch`.\n    +++\n    ++The `core.fsyncObjectFiles` was initially added based on integrity\n    ++assumptions that early (pre-ext-4) versions of Linux's \"ext\"\n    ++filesystems provided.\n    +++\n    ++I.e. that a write of file A without an `fsync()` followed by a write\n    ++of file `B` with `fsync()` would implicitly guarantee that `A' would\n    ++be `fsync()`'d by calling `fsync()` on `B`. This asssumption is *not*\n    ++backed up by any standard (e.g. POSIX), but worked in practice on some\n    ++Linux setups.\n    +++\n    ++Nowadays you should almost certainly want to use\n    ++`core.fsync=loose-object` instead in combination with\n    ++`core.fsyncMethod=bulk`, and possibly with\n    ++`fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n    ++OSX, Windows) that gives you most of the performance benefit of\n    ++`core.fsyncObjectFiles=false` with all of the safety of the old\n    ++`core.fsyncObjectFiles=true`.\n      \n    - \t\tsetup_work_tree();\n    --\t\twhile (getline_fn(&buf, stdin) != EOF) {\n    --\t\t\tchar *p;\n    --\t\t\tif (!nul_term_line && buf.buf[0] == '\"') {\n    --\t\t\t\tstrbuf_reset(&unquoted);\n    --\t\t\t\tif (unquote_c_style(&unquoted, buf.buf, NULL))\n    --\t\t\t\t\tdie(\"line is badly quoted\");\n    --\t\t\t\tstrbuf_swap(&buf, &unquoted);\n    --\t\t\t}\n    --\t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n    --\t\t\tupdate_one(p);\n    --\t\t\tif (set_executable_bit)\n    --\t\t\t\tchmod_path(set_executable_bit, p);\n    --\t\t\tfree(p);\n    --\t\t}\n    -+\t\twhile (getline_fn(&buf, stdin) != EOF)\n    -+\t\t\tline_from_stdin(&buf, &unquoted, prefix, prefix_length,\n    -+\t\t\t\t\tnul_term_line, set_executable_bit);\n    - \t\tstrbuf_release(&unquoted);\n    - \t\tstrbuf_release(&buf);\n    - \t}\n    + core.preloadIndex::\n    + \tEnable parallel index preload for operations like 'git diff'\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451984","messageId":"RFC-patch-v2-2.7-c6f776fc2bc-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 2/7] object-file: pass down unpack-objects.c flags for \"bulk\" checkin","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:26Z","receivedAt":"2022-03-23T14:18:57Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Remove much of this as a POC for exploring some of what I mentioned in\nhttps://lore.kernel.org/git/220322.86mthinxnn.gmgdl@evledraar.gmail.com/\n\nThis commit is obviously not what we *should* do as end-state, but\ndemonstrates what's needed (I think) for a bare-minimum implementation\nof just the \"bulk\" syncing method for loose objects without the part\nwhere we do the tmp-objdir.c dance.\n\nPerformance with this is already quite promising. Benchmarking with:\n\n\tgit hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' \\\n\t    \t-p 'rm -rf r.git && git init --bare r.git' \\\n\t\t'./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack'\n\nI.e. unpacking a small packfile (my dotfiles) yields, on a Linux\nramdisk:\n\n\tBenchmark 1: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync\n\t  Time (mean ± σ):     815.9 ms ±   8.2 ms    [User: 522.9 ms, System: 287.9 ms]\n\t  Range (min … max):   805.6 ms … 835.9 ms    10 runs\n\n\tBenchmark 2: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD\n\t  Time (mean ± σ):     779.4 ms ±  15.4 ms    [User: 505.7 ms, System: 270.2 ms]\n\t  Range (min … max):   763.1 ms … 813.9 ms    10 runs\n\n\tSummary\n\t  './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n\t    1.05 ± 0.02 times faster than './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n\nDoing the same with \"strace --summary-only\", which probably helps to\nemulate cases with slower syscalls is ~15% faster than using the\ntmp-objdir indirection:\n\n\tSummary\n\t  'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n\t    1.16 ± 0.01 times faster than 'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n\nWhich makes sense in terms of syscalls. In my case HEAD has ~101k\ncalls, and the parent topic is making ~129k calls, with around 2x the\nnumber of unlink(), link() as expected.\n\nOf course some users will want to use the tmp-objdir.c method. So a\nversion of this commit could be rewritten to come earlier in the\nseries, with the \"bulk\" on top being optional.\n\nIt seems to me that it's a much better strategy to do this whole thing\nin close_loose_object() after passing down the new HASH_N_OBJECTS /\nHASH_N_OBJECTS_FIRST / HASH_N_OBJECTS_LAST flags.\n\nDoing that for the \"builtin/add.c\" and \"builtin/unpack-objects.c\" code\nhaving its {un,}plug_bulk_checkin() removed here is then just a matter\nof passing down a similar set of flags indicating whether we're\ndealing with N objects, and if so if we're dealing with the last one\nor not.\n\nAs we'll see in subsequent commits doing it this way also effortlessly\nintegrates with other HASH_* flags. E.g. for \"update-index\" the code\nbeing rm'd here doesn't handle the interaction with\n\"HASH_WRITE_OBJECT\" properly, but once we've moved all this sync\nbootstrapping logic to close_loose_object() we'll never get to it if\nwe're not actually writing something.\n\nThis code currently doesn't use the HASH_N_OBJECTS_FIRST flag, but\nthat's what we'd use later to optionally call tmp_objdir_create().\n\nAside: This also changes logic that was a bit confusing and repetitive\nin close_loose_object(). Previously we'd first call\nbatch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT) which is just as\nshorthand for:\n\n\tfsync_components & FSYNC_COMPONENT_LOOSE_OBJECT &&\n\tfsync_method == FSYNC_METHOD_BATCH\n\nWe'd then proceed to call\nfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT) later in the same\nfunction, which is just a way of calling fsync_or_die() if:\n\n\tfsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\n\nNow we instead just define a local \"fsync_loose\" variable by checking\n\"fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\", which shows us that\nthe previous case of fsync_component_or_die(...)\" could just be added\nto the existing \"fsync_object_files > 0\" branch.\n\nNote: This commit reverts much of \"core.fsyncmethod: batched disk\nflushes for loose-objects\". We'll set up new structures to bring what\nit was doing back in a different way. I.e. to do the tmp-objdir\nplug-in in object-file.c\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c |  2 --\n builtin/update-index.c   |  4 ---\n bulk-checkin.c           | 74 ----------------------------------------\n bulk-checkin.h           |  3 --\n cache.h                  |  5 ---\n object-file.c            | 37 ++++++++++++++------\n 6 files changed, 26 insertions(+), 99 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex ec40c6fd966..93da436581b 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -507,7 +507,6 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n-\tplug_bulk_checkin();\n \toflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tint nth = i + 1;\n@@ -517,7 +516,6 @@ static void unpack_all(void)\n \t\tunpack_one(i, oflags | f);\n \t\tdisplay_progress(progress, nth);\n \t}\n-\tunplug_bulk_checkin();\n \tstop_progress(&progress);\n \n \tif (delta_list)\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex cbd2b0d633b..95ed3c47b2e 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -1118,8 +1118,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n \n-\t/* optimize adding many objects to the object database */\n-\tplug_bulk_checkin();\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1194,8 +1192,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n-\t/* by now we must have added all of the new objects */\n-\tunplug_bulk_checkin();\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex a0dca79ba6a..577b135e39c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,20 +3,15 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n-#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n-#include \"string-list.h\"\n-#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n-static struct tmp_objdir *bulk_fsync_objdir;\n-\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -85,40 +80,6 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \treprepare_packed_git(the_repository);\n }\n \n-/*\n- * Cleanup after batch-mode fsync_object_files.\n- */\n-static void do_batch_fsync(void)\n-{\n-\tstruct strbuf temp_path = STRBUF_INIT;\n-\tstruct tempfile *temp;\n-\n-\tif (!bulk_fsync_objdir)\n-\t\treturn;\n-\n-\t/*\n-\t * Issue a full hardware flush against a temporary file to ensure\n-\t * that all objects are durable before any renames occur. The code in\n-\t * fsync_loose_object_bulk_checkin has already issued a writeout\n-\t * request, but it has not flushed any writeback cache in the storage\n-\t * hardware or any filesystem logs. This fsync call acts as a barrier\n-\t * to ensure that the data in each new object file is durable before\n-\t * the final name is visible.\n-\t */\n-\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n-\ttemp = xmks_tempfile(temp_path.buf);\n-\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n-\tdelete_tempfile(&temp);\n-\tstrbuf_release(&temp_path);\n-\n-\t/*\n-\t * Make the object files visible in the primary ODB after their data is\n-\t * fully durable.\n-\t */\n-\ttmp_objdir_migrate(bulk_fsync_objdir);\n-\tbulk_fsync_objdir = NULL;\n-}\n-\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -313,26 +274,6 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n-void prepare_loose_object_bulk_checkin(void)\n-{\n-\tif (bulk_checkin_plugged && !bulk_fsync_objdir)\n-\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n-}\n-\n-void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n-{\n-\t/*\n-\t * If we have a plugged bulk checkin, we issue a call that\n-\t * cleans the filesystem page cache but avoids a hardware flush\n-\t * command. Later on we will issue a single hardware flush\n-\t * before as part of do_batch_fsync.\n-\t */\n-\tif (!bulk_fsync_objdir ||\n-\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n-\t\tfsync_or_die(fd, filename);\n-\t}\n-}\n-\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -347,19 +288,6 @@ int index_bulk_checkin(struct object_id *oid,\n void plug_bulk_checkin(void)\n {\n \tassert(!bulk_checkin_plugged);\n-\n-\t/*\n-\t * A temporary object directory is used to hold the files\n-\t * while they are not fsynced.\n-\t */\n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n-\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n-\t\tif (!bulk_fsync_objdir)\n-\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncMethod=batch\"));\n-\n-\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n-\t}\n-\n \tbulk_checkin_plugged = 1;\n }\n \n@@ -369,6 +297,4 @@ void unplug_bulk_checkin(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n-\n-\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 181d3447ff9..b26f3dc3b74 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,9 +6,6 @@\n \n #include \"cache.h\"\n \n-void prepare_loose_object_bulk_checkin(void);\n-void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n-\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex 72c91c91286..2f3831fa853 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1772,11 +1772,6 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n-static inline int batch_fsync_enabled(enum fsync_component component)\n-{\n-\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n-}\n-\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/object-file.c b/object-file.c\nindex cd0ddb49e4b..dbeb3df502d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1886,19 +1886,37 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n }\n \n+static void sync_loose_object_batch(int fd, const char *filename,\n+\t\t\t\t    const unsigned oflags)\n+{\n+\tconst int last = oflags & HASH_N_OBJECTS_LAST;\n+\n+\t/*\n+\t * We're doing a sync_file_range() (or equivalent) for 1..N-1\n+\t * objects, and then a \"real\" fsync() for N. On some OS's\n+\t * enabling core.fsync=loose-object && core.fsyncMethod=batch\n+\t * improves the performance by a lot.\n+\t */\n+\tif (last || (!last && git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0))\n+\t\tfsync_or_die(fd, filename);\n+}\n+\n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd, const char *filename)\n+static void close_loose_object(int fd, const char *filename,\n+\t\t\t       const unsigned oflags)\n {\n+\tint fsync_loose;\n+\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n-\t\tfsync_loose_object_bulk_checkin(fd, filename);\n-\telse if (fsync_object_files > 0)\n+\tfsync_loose = fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT;\n+\n+\tif (oflags & HASH_N_OBJECTS && fsync_loose &&\n+\t    fsync_method == FSYNC_METHOD_BATCH)\n+\t\tsync_loose_object_batch(fd, filename, oflags);\n+\telse if (fsync_object_files > 0 || fsync_loose)\n \t\tfsync_or_die(fd, filename);\n-\telse\n-\t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n-\t\t\t\t       filename);\n \n out:\n \tif (close(fd) != 0)\n@@ -1962,9 +1980,6 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n-\t\tprepare_loose_object_bulk_checkin();\n-\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n@@ -2015,7 +2030,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd, tmp_file.buf);\n+\tclose_loose_object(fd, tmp_file.buf, flags);\n \n \tif (mtime) {\n \t\tstruct utimbuf utb;\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451985","messageId":"RFC-patch-v2-3.7-4df8012100a-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 3/7] update-index: pass down skeleton \"oflags\" argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:27Z","receivedAt":"2022-03-23T14:19:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"As with a preceding change to \"unpack-objects\" add an \"oflags\" going\nfrom cmd_update_index() all the way down to the code in\nobject-file.c. Note also how index_mem() will now call\nwrite_object_file_flags().\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c | 32 ++++++++++++++++++--------------\n object-file.c          |  2 +-\n 2 files changed, 19 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 95ed3c47b2e..34aaaa16c20 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -267,10 +267,12 @@ static int process_lstat_error(const char *path, int err)\n \treturn error(\"lstat(\\\"%s\\\"): %s\", path, strerror(err));\n }\n \n-static int add_one_path(const struct cache_entry *old, const char *path, int len, struct stat *st)\n+static int add_one_path(const struct cache_entry *old, const char *path,\n+\t\t\tint len, struct stat *st, const unsigned oflags)\n {\n \tint option;\n \tstruct cache_entry *ce;\n+\tunsigned f;\n \n \t/* Was the old index entry already up-to-date? */\n \tif (old && !ce_stage(old) && !ce_match_stat(old, st, 0))\n@@ -283,8 +285,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n \tfill_stat_cache_info(&the_index, ce, st);\n \tce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n \n-\tif (index_path(&the_index, &ce->oid, path, st,\n-\t\t       info_only ? 0 : HASH_WRITE_OBJECT)) {\n+\tf = oflags | (info_only ? 0 : HASH_WRITE_OBJECT);\n+\tif (index_path(&the_index, &ce->oid, path, st, f)) {\n \t\tdiscard_cache_entry(ce);\n \t\treturn -1;\n \t}\n@@ -320,7 +322,8 @@ static int add_one_path(const struct cache_entry *old, const char *path, int len\n  *  - it doesn't exist at all in the index, but it is a valid\n  *    git directory, and it should be *added* as a gitlink.\n  */\n-static int process_directory(const char *path, int len, struct stat *st)\n+static int process_directory(const char *path, int len, struct stat *st,\n+\t\t\t     const unsigned oflags)\n {\n \tstruct object_id oid;\n \tint pos = cache_name_pos(path, len);\n@@ -334,7 +337,7 @@ static int process_directory(const char *path, int len, struct stat *st)\n \t\t\tif (resolve_gitlink_ref(path, \"HEAD\", &oid) < 0)\n \t\t\t\treturn 0;\n \n-\t\t\treturn add_one_path(ce, path, len, st);\n+\t\t\treturn add_one_path(ce, path, len, st, oflags);\n \t\t}\n \t\t/* Should this be an unconditional error? */\n \t\treturn remove_one_path(path);\n@@ -358,13 +361,14 @@ static int process_directory(const char *path, int len, struct stat *st)\n \n \t/* No match - should we add it as a gitlink? */\n \tif (!resolve_gitlink_ref(path, \"HEAD\", &oid))\n-\t\treturn add_one_path(NULL, path, len, st);\n+\t\treturn add_one_path(NULL, path, len, st, oflags);\n \n \t/* Error out. */\n \treturn error(\"%s: is a directory - add files inside instead\", path);\n }\n \n-static int process_path(const char *path, struct stat *st, int stat_errno)\n+static int process_path(const char *path, struct stat *st, int stat_errno,\n+\t\t\tconst unsigned oflags)\n {\n \tint pos, len;\n \tconst struct cache_entry *ce;\n@@ -395,9 +399,9 @@ static int process_path(const char *path, struct stat *st, int stat_errno)\n \t\treturn process_lstat_error(path, stat_errno);\n \n \tif (S_ISDIR(st->st_mode))\n-\t\treturn process_directory(path, len, st);\n+\t\treturn process_directory(path, len, st, oflags);\n \n-\treturn add_one_path(ce, path, len, st);\n+\treturn add_one_path(ce, path, len, st, oflags);\n }\n \n static int add_cacheinfo(unsigned int mode, const struct object_id *oid,\n@@ -446,7 +450,7 @@ static void chmod_path(char flip, const char *path)\n \tdie(\"git update-index: cannot chmod %cx '%s'\", flip, path);\n }\n \n-static void update_one(const char *path)\n+static void update_one(const char *path, const unsigned oflags)\n {\n \tint stat_errno = 0;\n \tstruct stat st;\n@@ -485,7 +489,7 @@ static void update_one(const char *path)\n \t\treport(\"remove '%s'\", path);\n \t\treturn;\n \t}\n-\tif (process_path(path, &st, stat_errno))\n+\tif (process_path(path, &st, stat_errno, oflags))\n \t\tdie(\"Unable to process path %s\", path);\n \treport(\"add '%s'\", path);\n }\n@@ -776,7 +780,7 @@ static int do_reupdate(int ac, const char **av,\n \t\t */\n \t\tsave_nr = active_nr;\n \t\tpath = xstrdup(ce->name);\n-\t\tupdate_one(path);\n+\t\tupdate_one(path, 0);\n \t\tfree(path);\n \t\tdiscard_cache_entry(old);\n \t\tif (save_nr != active_nr)\n@@ -1138,7 +1142,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \t\t\tsetup_work_tree();\n \t\t\tp = prefix_path(prefix, prefix_length, path);\n-\t\t\tupdate_one(p);\n+\t\t\tupdate_one(p, 0);\n \t\t\tif (set_executable_bit)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n@@ -1183,7 +1187,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tstrbuf_swap(&buf, &unquoted);\n \t\t\t}\n \t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n-\t\t\tupdate_one(p);\n+\t\t\tupdate_one(p, 0);\n \t\t\tif (set_executable_bit)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\ndiff --git a/object-file.c b/object-file.c\nindex dbeb3df502d..8999fce2b15 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2211,7 +2211,7 @@ static int index_mem(struct index_state *istate,\n \t}\n \n \tif (write_object)\n-\t\tret = write_object_file(buf, size, type, oid);\n+\t\tret = write_object_file_flags(buf, size, type, oid, flags);\n \telse\n \t\thash_object_file(the_hash_algo, buf, size, type, oid);\n \tif (re_allocated)\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451986","messageId":"RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 4/7] update-index: have the index fsync() flush the loose objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:28Z","receivedAt":"2022-03-23T14:19:10Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"As with unpack-objects in a preceding commit have update-index.c make\nuse of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\nmode again for \"update-index\".\n\nAdding the t/* directory from git.git on a Linux ramdisk is a bit\nfaster than with the tmp-objdir indirection:\n\n\t$ git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt' -p 'rm -rf repo/.git/objects/* repo/.git/index' './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' --warmup 1 -r 10Benchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync\n\t  Time (mean ± σ):     281.1 ms ±   2.6 ms    [User: 186.2 ms, System: 92.3 ms]\n\t  Range (min … max):   278.3 ms … 287.0 ms    10 runs\n\n\tBenchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD\n\t  Time (mean ± σ):     265.9 ms ±   2.6 ms    [User: 181.7 ms, System: 82.1 ms]\n\t  Range (min … max):   262.0 ms … 270.3 ms    10 runs\n\n\tSummary\n\t  './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n\t    1.06 ± 0.01 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n\nAnd as before running that with \"strace --summary-only\" slows things\ndown a bit (probably mimicking slower I/O a bit). I then get:\n\n\tSummary\n\t  'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n\t    1.19 ± 0.03 times faster than 'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n\nThis one has a twist though, instead of fsync()-ing on the last object\nwe write we'll not do that, and instead defer the fsync() until we\nwrite the index itself. This is outlined in [1] (as \"METHOD THREE\").\n\nBecause of this under FSYNC_METHOD_BATCH we'll do the N\nobjects (possibly only one, because we're lazy) as HASH_N_OBJECTS, and\nwe'll even now support doing this via N arguments on the command-line.\n\nThen we won't fsync() any of it, but we will rename it\nin-place (which, if we were still using the tmp-objdir, would leave it\n\"staged\" in the tmp-objdir).\n\nWe'll then have the fsync() for the index update \"flush\" that out, and\nthus avoid two fsync() calls when one will do.\n\nRunning this with the \"git hyperfine\" command mentioned in a preceding\ncommit with \"strace --summary-only\" shows that we do 1 fsync() now\ninstead of 2, and have one more sync_file_range(), as expected.\n\nWe also go from ~51k syscalls to ~39k, with ~2x the number of link()\nand unlink() in ns/batched-fsync, and of course one fsync() instead of\ntwo()>\n\nThe flow of this code isn't quite set up for re-plugging the\ntmp-objdir back in. In particular we no longer pass\nHASH_N_OBJECTS_FIRST (but doing so would be trivial)< and there's no\nHASH_N_OBJECTS_LAST.\n\nSo this and other callers would need some light transaction-y API, or\nto otherwise pass down a \"yes, I'd like to flush it\" down to\nfinalize_hashfile(), but doing so will be trivial.\n\nAnd since we've started structuring it this way it'll become easy to\ndo any arbitrary number of things down the line that would \"bulk\nfsync\" before the final fsync(). Now we write some objects and fsync()\non the index, but between those two could do any number of other\nthings where we'd defer the fsync().\n\nThis sort of thing might be especially interesting for \"git repack\"\nwhen it writes e.g. a *.bitmap, *.rev, *.pack and *.idx. In that case\nwe could skip the fsync() on all of those, and only do it on the *.idx\nbefore we renamed it in-place. I *think* nothing cares about a *.pack\nwithout an *.idx, but even then we could fsync *.idx, rename *.pack,\nrename *.idx and still safely do only one fsync(). See \"git show\n--first-parent\" on 62874602032 (Merge branch\n'tb/pack-finalize-ordering' into maint, 2021-10-12) for a good\noverview of the code involved in that.\n\n1. https://lore.kernel.org/git/220323.86sfr9ndpr.gmgdl@evledraar.gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/update-index.c |  7 ++++---\n cache.h                |  1 +\n read-cache.c           | 29 ++++++++++++++++++++++++++++-\n 3 files changed, 33 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 34aaaa16c20..6cfec6efb38 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -1142,7 +1142,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \n \t\t\tsetup_work_tree();\n \t\t\tp = prefix_path(prefix, prefix_length, path);\n-\t\t\tupdate_one(p, 0);\n+\t\t\tupdate_one(p, HASH_N_OBJECTS);\n \t\t\tif (set_executable_bit)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n@@ -1187,7 +1187,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\tstrbuf_swap(&buf, &unquoted);\n \t\t\t}\n \t\t\tp = prefix_path(prefix, prefix_length, buf.buf);\n-\t\t\tupdate_one(p, 0);\n+\t\t\tupdate_one(p, HASH_N_OBJECTS);\n \t\t\tif (set_executable_bit)\n \t\t\t\tchmod_path(set_executable_bit, p);\n \t\t\tfree(p);\n@@ -1263,7 +1263,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\t\t\texit(128);\n \t\t\tunable_to_lock_die(get_index_file(), lock_error);\n \t\t}\n-\t\tif (write_locked_index(&the_index, &lock_file, COMMIT_LOCK))\n+\t\tif (write_locked_index(&the_index, &lock_file,\n+\t\t\t\t       COMMIT_LOCK | WLI_NEED_LOOSE_FSYNC))\n \t\t\tdie(\"Unable to write new index file\");\n \t}\n \ndiff --git a/cache.h b/cache.h\nindex 2f3831fa853..7542e009a34 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -751,6 +751,7 @@ void ensure_full_index(struct index_state *istate);\n /* For use with `write_locked_index()`. */\n #define COMMIT_LOCK\t\t(1 << 0)\n #define SKIP_IF_UNCHANGED\t(1 << 1)\n+#define WLI_NEED_LOOSE_FSYNC\t(1 << 2)\n \n /*\n  * Write the index while holding an already-taken lock. Close the lock,\ndiff --git a/read-cache.c b/read-cache.c\nindex 3e0e7d41837..275f6308c32 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2860,6 +2860,33 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n+\tunsigned int wflags = FSYNC_COMPONENT_INDEX;\n+\n+\n+\t/*\n+\t * TODO: This is abuse of the API recently modified\n+\t * finalize_hashfile() which reveals a shortcoming of its\n+\t * \"fsync\" design.\n+\t * \n+\t * I.e. It expects a \"enum fsync_component component\" label,\n+\t * but here we're passing it an OR of the two, knowing that\n+\t * it'll call fsync_component_or_die() which (in\n+\t * write-or-die.c) will do \"(fsync_components & wflags)\" (to\n+\t * our \"wflags\" here).\n+\t *\n+\t * But the API really should be changed to explicitly take\n+\t * such flags, because in this case we'd like to fsync() the\n+\t * index if we're in the bulk mode, *even if* our\n+\t * \"core.fsync=index\" isn't configured.\n+\t *\n+\t * That's because at this point we've been queuing up object\n+\t * writes that we didn't fsync(), and are going to use this\n+\t * fsync() to \"flush\" the whole thing. Doing it this way\n+\t * avoids redundantly calling fsync() twice when once will do.\n+\t */\n+\tif (fsync_method == FSYNC_METHOD_BATCH && \n+\t    flags & WLI_NEED_LOOSE_FSYNC)\n+\t\twflags |= FSYNC_COMPONENT_LOOSE_OBJECT;\n \n \tf = hashfd(tempfile->fd, tempfile->filename.buf);\n \n@@ -3094,7 +3121,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (!alternate_index_output && (flags & COMMIT_LOCK))\n \t\tcsum_fsync_flag = CSUM_FSYNC;\n \n-\tfinalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n+\tfinalize_hashfile(f, istate->oid.hash, wflags,\n \t\t\t  CSUM_HASH_IN_STREAM | csum_fsync_flag);\n \n \tif (close_tempfile_gently(tempfile)) {\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451987","messageId":"RFC-patch-v2-5.7-2bf14fd4946-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 5/7] add: use WLI_NEED_LOOSE_FSYNC for new \"only the index\" bulk fsync()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:29Z","receivedAt":"2022-03-23T14:19:17Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"We can now bring \"bulk\" syncing back to \"git add\" using a mechanism\ndiscussed in the preceding commit where we fsync() on the index, not\nthe last object we write.\n\nOn a ramdisk:\n\n\t$ git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/' -p 'rm -rf repo/.git/objects/* repo/.git/\n\tindex' './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' --warmup 1\n\tBenchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'ns/batched-fsync\n\t  Time (mean ± σ):     299.5 ms ±   1.6 ms    [User: 193.4 ms, System: 103.7 ms]\n\t  Range (min … max):   296.6 ms … 301.6 ms    10 runs\n\n\tBenchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'HEAD\n\t  Time (mean ± σ):     282.8 ms ±   2.1 ms    [User: 193.8 ms, System: 86.6 ms]\n\t  Range (min … max):   279.1 ms … 285.6 ms    10 runs\n\n\tSummary\n\t  './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'HEAD' ran\n\t    1.06 ± 0.01 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'ns/batched-fsync'\n\nMy times on my spinning disk are too fuzzy to quote with confidence,\nbut I have seen it go as well as 15-30% faster. FWIW doing \"strace\n--summary-only\" on the ramdisk is ~20% faster:\n\n\t$ git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/' -p 'rm -rf repo/.git/objects/* repo/.git/index' 'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' --warmup 1\n\tBenchmark 1: strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'ns/batched-fsync\n\t  Time (mean ± σ):     917.4 ms ±  18.8 ms    [User: 388.7 ms, System: 672.1 ms]\n\t  Range (min … max):   885.3 ms … 948.1 ms    10 runs\n\n\tBenchmark 2: strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'HEAD\n\t  Time (mean ± σ):     769.0 ms ±   9.2 ms    [User: 358.2 ms, System: 521.2 ms]\n\t  Range (min … max):   760.7 ms … 792.6 ms    10 runs\n\n\tSummary\n\t  'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'HEAD' ran\n\t    1.19 ± 0.03 times faster than 'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'ns/batched-fsync'\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/add.c | 6 ++++--\n cache.h       | 1 +\n read-cache.c  | 8 ++++++++\n 3 files changed, 13 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..6ef18b6246c 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -580,7 +580,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\t (intent_to_add ? ADD_CACHE_INTENT : 0) |\n \t\t (ignore_add_errors ? ADD_CACHE_IGNORE_ERRORS : 0) |\n \t\t (!(addremove || take_worktree_changes)\n-\t\t  ? ADD_CACHE_IGNORE_REMOVAL : 0));\n+\t\t  ? ADD_CACHE_IGNORE_REMOVAL : 0)) |\n+\t\tADD_CACHE_HASH_N_OBJECTS;\n \n \tif (read_cache_preload(&pathspec) < 0)\n \t\tdie(_(\"index file corrupt\"));\n@@ -686,7 +687,8 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\n-\t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n+\t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED |\n+\t\t\t       WLI_NEED_LOOSE_FSYNC))\n \t\tdie(_(\"Unable to write new index file\"));\n \n \tdir_clear(&dir);\ndiff --git a/cache.h b/cache.h\nindex 7542e009a34..d57af938cbc 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -857,6 +857,7 @@ int remove_file_from_index(struct index_state *, const char *path);\n #define ADD_CACHE_IGNORE_ERRORS\t4\n #define ADD_CACHE_IGNORE_REMOVAL 8\n #define ADD_CACHE_INTENT 16\n+#define ADD_CACHE_HASH_N_OBJECTS 32\n /*\n  * These two are used to add the contents of the file at path\n  * to the index, marking the working tree up-to-date by storing\ndiff --git a/read-cache.c b/read-cache.c\nindex 275f6308c32..788423b6dde 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -755,6 +755,14 @@ int add_to_index(struct index_state *istate, const char *path, struct stat *st,\n \tunsigned hash_flags = pretend ? 0 : HASH_WRITE_OBJECT;\n \tstruct object_id oid;\n \n+\t/*\n+\t * TODO: Can't we also set HASH_N_OBJECTS_FIRST as a function\n+\t * of !(ce->ce_flags & CE_ADDED) or something? I'm not too\n+\t * familiar with the cache API...\n+\t */\n+\tif (flags & ADD_CACHE_HASH_N_OBJECTS)\n+\t\thash_flags |= HASH_N_OBJECTS;\n+\n \tif (flags & ADD_CACHE_RENORMALIZE)\n \t\thash_flags |= HASH_RENORMALIZE;\n \n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451988","messageId":"RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 7/7] fsync docs: add new fsyncMethod.batch.quarantine, elaborate on old","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:31Z","receivedAt":"2022-03-23T14:19:19Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Add a new fsyncMethod.batch.quarantine setting which defaults to\n\"false\". Preceding (RFC, and not meant to flip-flop like that\neventually) commits ripped out the \"tmp-objdir\" part of the\ncore.fsyncMethod=batch.\n\nThis documentation proposes to keep that as the default for the\nreasons discussed in it, while allowing users to set\n\"fsyncMethod.batch.quarantine=true\".\n\nFurthermore update the discussion of \"core.fsyncObjectFiles\" with\ninformation about what it *really* does, why you probably shouldn't\nuse it, and how to safely emulate most of what it gave users in the\npast in terms of performance benefit.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 80 +++++++++++++++++++++++++++++++----\n 1 file changed, 72 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex f598925b597..365a12dc7ae 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -607,21 +607,85 @@ stored on NTFS or ReFS filesystems.\n +\n The `batch` is currently only applies to loose-object files and will\n kick in when using the linkgit:git-unpack-objects[1] and\n-linkgit:update-index[1] commands. Note that the \"last\" file to be\n+linkgit:git-update-index[1] commands. Note that the \"last\" file to be\n synced may be the last object, as in the case of\n linkgit:git-unpack-objects[1], or relevant \"index\" (or in the future,\n \"ref\") update, as in the case of linkgit:git-update-index[1]. I.e. the\n batch syncing of the loose objects may be deferred until a subsequent\n fsync() to a file that makes them \"active\".\n \n+fsyncMethod.batch.quarantine::\n+\tA boolean which if set to `true` will cause \"batched\" writes\n+\tto objects to be \"quarantined\" if\n+\t`core.fsyncMethod=batch`. This is `false` by default.\n++\n+The primary object of these fsync() settings is to protect against\n+repository corruption of things which are reachable, i.e. \"reachable\",\n+via references, the index etc. Not merely objects that were present in\n+the object store.\n++\n+Historically setting `core.fsyncObjectFiles=false` assumed that on a\n+filesystem with where an fsync() would flush all preceding outstanding\n+I/O that we might end up with a corrupt loose object, but that was OK\n+as long as no reference referred to it. We'd eventually the corrupt\n+object with linkgit:git-gc[1], and linkgit:git-fsck[1] would only\n+report it as a minor annoyance\n++\n+Setting `fsyncMethod.batch.quarantine=true` takes the view that\n+something like a corrupt *unreferenced* loose object in the object\n+store is something we'd like to avoid, at the cost of reduced\n+performance when using `core.fsyncMethod=batch`.\n++\n+Currently this uses the same mechanism described in the \"QUARANTINE\n+ENVIRONMENT\" in the linkgit:git-receive-pack[1] documentation, but\n+that's subject to change. The performance loss is because we need to\n+\"stage\" the objects in that quarantine environment, fsync() it, and\n+once that's done rename() or link() it in-place into the main object\n+store, possibly with an fsync() of the index or ref at the end\n++\n+With `fsyncMethod.batch.quarantine=false` we'll \"stage\" things in the\n+main object store, and then do one fsync() at the very end, either on\n+the last object we write, or file (index or ref) that'll make it\n+\"reachable\".\n++\n+The bad thing about setting this to `true` is lost performance, as\n+well as not being able to access the objects as they're written (which\n+e.g. consumers of linkgit:git-update-index[1]'s `--verbose` mode might\n+want to do).\n++\n+The good thing is that you should be guaranteed not to get e.g. short\n+or otherwise corrupt loose objects if you pull your power cord, in\n+practice various git commands deal quite badly with discovering such a\n+stray corrupt object (including perhaps assuming it's valid based on\n+its existence, or hard dying on an error rather than replacing\n+it). Repairing such \"unreachable corruption\" can require manual\n+intervention.\n+\n core.fsyncObjectFiles::\n-\tThis boolean will enable 'fsync()' when writing object files.\n-\tThis setting is deprecated. Use core.fsync instead.\n-+\n-This setting affects data added to the Git repository in loose-object\n-form. When set to true, Git will issue an fsync or similar system call\n-to flush caches so that loose-objects remain consistent in the face\n-of a unclean system shutdown.\n+\tThis boolean will enable 'fsync()' when writing loose object\n+\tfiles.\n++\n+This setting is the historical fsync configuration setting. It's now\n+*deprecated*, you should use `core.fsync` instead, perhaps in\n+combination with `core.fsyncMethod=batch`.\n++\n+The `core.fsyncObjectFiles` was initially added based on integrity\n+assumptions that early (pre-ext-4) versions of Linux's \"ext\"\n+filesystems provided.\n++\n+I.e. that a write of file A without an `fsync()` followed by a write\n+of file `B` with `fsync()` would implicitly guarantee that `A' would\n+be `fsync()`'d by calling `fsync()` on `B`. This asssumption is *not*\n+backed up by any standard (e.g. POSIX), but worked in practice on some\n+Linux setups.\n++\n+Nowadays you should almost certainly want to use\n+`core.fsync=loose-object` instead in combination with\n+`core.fsyncMethod=bulk`, and possibly with\n+`fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n+OSX, Windows) that gives you most of the performance benefit of\n+`core.fsyncObjectFiles=false` with all of the safety of the old\n+`core.fsyncObjectFiles=true`.\n \n core.preloadIndex::\n \tEnable parallel index preload for operations like 'git diff'\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"451989","messageId":"RFC-patch-v2-6.7-c20301d7967-20220323T140753Z-avarab@gmail.com","threadId":"57568","inReplyTo":"RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com","subject":"[RFC PATCH v2 6/7] fsync docs: update for new syncing semantics","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-23T14:18:30Z","receivedAt":"2022-03-23T14:19:21Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 23 +++++++++++++++++------\n 1 file changed, 17 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex cf0e9b8b088..f598925b597 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -596,12 +596,23 @@ core.fsyncMethod::\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n * `batch` enables a mode that uses writeout-only flushes to stage multiple\n-  updates in the disk writeback cache and then does a single full fsync of\n-  a dummy file to trigger the disk cache flush at the end of the operation.\n-  Currently `batch` mode only applies to loose-object files. Other repository\n-  data is made durable as if `fsync` was specified. This mode is expected to\n-  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n-  and on Windows for repos stored on NTFS or ReFS filesystems.\n+  updates in the disk writeback cache and, before doing a full fsync() of\n+  on the \"last\" file that to trigger the disk cache flush at the end of the\n+  operation.\n++\n+Other repository data is made durable as if `fsync` was\n+specified. This mode is expected to be as safe as `fsync` on macOS for\n+repos stored on HFS+ or APFS filesystems and on Windows for repos\n+stored on NTFS or ReFS filesystems.\n++\n+The `batch` is currently only applies to loose-object files and will\n+kick in when using the linkgit:git-unpack-objects[1] and\n+linkgit:update-index[1] commands. Note that the \"last\" file to be\n+synced may be the last object, as in the case of\n+linkgit:git-unpack-objects[1], or relevant \"index\" (or in the future,\n+\"ref\") update, as in the case of linkgit:git-update-index[1]. I.e. the\n+batch syncing of the loose objects may be deferred until a subsequent\n+fsync() to a file that makes them \"active\".\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\n-- \n2.35.1.1428.g1c1a0152d61\n\n"},{"id":"452025","messageId":"CANQDOddpo+a8r_0yghgy_1bHvfUe5XQaaaWc7D-OLqX6Anhgiw@mail.gmail.com","threadId":"57568","inReplyTo":"220323.86sfr9ndpr.gmgdl@evledraar.gmail.com","subject":"Re: [RFC PATCH 7/7] update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T20:19:32Z","receivedAt":"2022-03-23T20:19:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"I'm going to respond in more detail to your individual patches,\n(expect the last mail to contain a comment at the end \"LAST MAIL\").\n\nOn Wed, Mar 23, 2022 at 3:52 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Mar 22 2022, Neeraj Singh wrote:\n>\n> > On Tue, Mar 22, 2022 at 8:48 PM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >>\n> >> As with unpack-objects in a preceding commit have update-index.c make\n> >> use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\n> >> mode again for \"update-index\".\n> >>\n> >> Adding the t/* directory from git.git on a Linux ramdisk is a bit\n> >> faster than with the tmp-objdir indirection:\n> >>\n> >>         git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10\n> >>         Benchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync\n> >>           Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]\n> >>           Range (min … max):   285.6 ms … 297.0 ms    10 runs\n> >>\n> >>         Benchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD\n> >>           Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]\n> >>           Range (min … max):   267.8 ms … 291.3 ms    10 runs\n> >>\n> >>         Summary\n> >>           'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n> >>             1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n> >>\n> >> And as before running that with \"strace --summary-only\" slows things\n> >> down a bit (probably mimicking slower I/O a bit). I then get:\n> >>\n> >>         Summary\n> >>           'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran\n> >>             1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'\n> >>\n> >> We also go from ~51k syscalls to ~39k, with ~2x the number of link()\n> >> and unlink() in ns/batched-fsync.\n> >>\n> >> In the process of doing this conversion we lost the \"bulk\" mode for\n> >> files added on the command-line. I don't think it's useful to optimize\n> >> that, but we could if anyone cared.\n> >>\n> >> We've also converted this to a string_list, we could walk with\n> >> getline_fn() and get one line \"ahead\" to see what we have left, but I\n> >> found that state machine a bit painful, and at least in my testing\n> >> buffering this doesn't harm things. But we could also change this to\n> >> stream again, at the cost of some getline_fn() twiddling.\n> >>\n> >> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> >> ---\n> >>  builtin/update-index.c | 31 +++++++++++++++++++++++++++----\n> >>  1 file changed, 27 insertions(+), 4 deletions(-)\n> >>\n> >> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> >> index af02ff39756..c7cbfe1123b 100644\n> >> --- a/builtin/update-index.c\n> >> +++ b/builtin/update-index.c\n> >> @@ -1194,15 +1194,38 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >>         }\n> >>\n> >>         if (read_from_stdin) {\n> >> +               struct string_list list = STRING_LIST_INIT_NODUP;\n> >>                 struct strbuf line = STRBUF_INIT;\n> >>                 struct strbuf unquoted = STRBUF_INIT;\n> >> +               size_t i, nr;\n> >> +               unsigned oflags;\n> >>\n> >>                 setup_work_tree();\n> >> -               while (getline_fn(&line, stdin) != EOF)\n> >> -                       line_from_stdin(&line, &unquoted, prefix, prefix_length,\n> >> -                                       nul_term_line, set_executable_bit, 0);\n> >> +               while (getline_fn(&line, stdin) != EOF) {\n> >> +                       size_t len = line.len;\n> >> +                       char *str = strbuf_detach(&line, NULL);\n> >> +\n> >> +                       string_list_append_nodup(&list, str)->util = (void *)len;\n> >> +               }\n> >> +\n> >> +               nr = list.nr;\n> >> +               oflags = nr > 1 ? HASH_N_OBJECTS : 0;\n> >> +               for (i = 0; i < nr; i++) {\n> >> +                       size_t nth = i + 1;\n> >> +                       unsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n> >> +                                 nr == nth ? HASH_N_OBJECTS_LAST : 0;\n> >> +                       struct strbuf buf = STRBUF_INIT;\n> >> +                       struct string_list_item *item = list.items + i;\n> >> +                       const size_t len = (size_t)item->util;\n> >> +\n> >> +                       strbuf_attach(&buf, item->string, len, len);\n> >> +                       line_from_stdin(&buf, &unquoted, prefix, prefix_length,\n> >> +                                       nul_term_line, set_executable_bit,\n> >> +                                       oflags | f);\n> >> +                       strbuf_release(&buf);\n> >> +               }\n> >>                 strbuf_release(&unquoted);\n> >> -               strbuf_release(&line);\n> >> +               string_list_clear(&list, 0);\n> >>         }\n> >>\n> >>         if (split_index > 0) {\n> >> --\n> >> 2.35.1.1428.g1c1a0152d61\n> >>\n> >\n> > This buffering introduces the same potential risk of the\n> > \"stdin-feeder\" process not being able to see objects right away as my\n> > version had. I'm planning to mitigate the issue by unplugging the bulk\n> > checkin when issuing a verbose report so that anyone who's using that\n> > output to synchronize can still see what they're expecting.\n>\n> I was rather terse in the commit message, I meant (but forgot some\n> words) \"doesn't harm thing for performance [in the above test]\", but\n> converting this to a string_list is clearly & regression that shouldn't\n> be kept.\n>\n> I just wanted to demonstrate method of doing this by passing down the\n> HASH_* flags, and found that writing the state-machine to \"buffer ahead\"\n> by one line so that we can eventually know in the loop if we're in the\n> \"last\" line or not was tedious, so I came up with this POC. But we\n> clearly shouldn't lose the \"streaming\" aspect.\n>\n\nFrom my experience working on several state machines in the Windows\nOS, they are notoriously difficult to understand and extend.  I\nwouldn't want every top-level command that does something interesting\nto have to deal with that.\n\n> But anyway, now that I look at this again the smart thing here (surely?)\n> is to keep the simple getline() loop and not ever issue a\n> HASH_N_OBJECTS_LAST for the Nth item, instead we should in this case do\n> the \"checkpoint fsync\" at the point that we write the actual index.\n>\n> Because an existing redundancy in your series is that you'll do the\n> fsync() the same way for \"git unpack-objects\" as for \"git\n> {update-index,add}\".\n>\n> I.e. in the former case adding the N objects is all we're doing, so the\n> \"last object\" is the point at which we need to flush the previous N to\n> disk.\n>\n> But for \"update-index/add\" you'll do at least 2 fsync()'s in the bulk\n> mode, when it should be one. I.e. the equivalent of (leaving aside the\n> tmp-objdir migration part of it), if writing objects A && B:\n>\n>     ## METHOD ONE\n>     # A\n>     write(objects/A.tmp)\n>     bulk_fsync(objects/A.tmp)\n>     rename(objects/A.tmp, objects/A)\n>     # B\n>     write(objects/B.tmp)\n>     bulk_fsync(objects/B.tmp)\n>     rename(objects/B.tmp, objects/B)\n>     # \"cookie\"\n>     write(bulk_fsync_XXXXXX)\n>     fsync(bulk_fsync_XXXXXX)\n>     # ref\n>     write(INDEX.tmp, $(git rev-parse B))\n>     fsync(INDEX.tmp)\n>     rename(INDEX.tmp, INDEX)\n>\n> This series on top changes that so we know that we're doing N, so we\n> don't need the seperate \"cookie\", we can just use the B object as the\n> cookie, as we know it comes last;\n>\n>     ## METHOD TWO\n>     # A -- SAME as above\n>     write(objects/A.tmp)\n>     bulk_fsync(objects/A.tmp)\n>     rename(objects/A.tmp, objects/A)\n>     # B -- SAME as above, with s/bulk_fsync/fsync/\n>     write(objects/B.tmp)\n>     fsync(objects/B.tmp)\n>     rename(objects/B.tmp, objects/B)\n>     # \"cookie\" -- GONE!\n>     # ref -- SAME\n>     write(INDEX.tmp, $(git rev-parse B))\n>     fsync(INDEX.tmp)\n>     rename(INDEX.tmp, INDEX)\n>\n> But really, we should instead realize that we're not doing\n> \"unpack-objects\", but have a \"ref update\" at the end (whether that's a\n> ref, or an index etc.) and do:\n>\n>     ## METHOD THREE\n>     # A -- SAME as above\n>     write(objects/A.tmp)\n>     bulk_fsync(objects/A.tmp)\n>     rename(objects/A.tmp, objects/A)\n>     # B -- SAME as the first\n>     write(objects/B.tmp)\n>     bulk_fsync(objects/B.tmp)\n>     rename(objects/B.tmp, objects/B)\n>     # ref -- SAME\n>     write(INDEX.tmp, $(git rev-parse B))\n>     fsync(INDEX.tmp)\n>     rename(INDEX.tmp, INDEX)\n>\n> Which cuts our number of fsync() operations down from 2 to 1, ina\n> addition to removing the need for the \"cookie\", which is only there\n> because we didn't keep track of where we were in the sequence as in my\n> 2/7 and 5/7.\n>\n\nI agree that this is a great direction to go in as an extension to\nthis work (i.e. a subsequent patch).  I saw in one of your mails on v2\nof your rfc series that you mentioned a \"lightweight transaction-y\nthing\".  I've been thinking along the same lines myself, but wanted to\ntreat that as a separable concern.  In my ideal world, we'd just use a\nreal database for loose objects, the index, and refs and let that\nhandle the transaction management.  But in lieu of that, having a\ntransaction that looks across the ODB, index, and refs would let us\nbatch syncs optimally.\n\n> And it would be the same for tmp-objdir, the rename dance is a bit\n> different, but we'd do the \"full\" fsync() while on the INDEX.tmp, then\n> migrate() the tmp-objdir, and once that's done do the final:\n>\n>     rename(INDEX.tmp, INDEX)\n>\n> I.e. we'd fsync() the content once, and only have the renme() or link()\n> operations left. For POSIX we'd need a few more fsync() for the\n> metadata, but this (i.e. your) series already makes the hard assumption\n> that we don't need to do that for rename().\n>\n> > I think the code you've presented here is a lot of diff to accomplish\n> > the same thing that my series does, where this specific update-index\n> > caller has been roto-tilled to provide the needed\n> > begin/end-transaction points.\n>\n> Any caller of these APIs will need the \"unsigned oflags\" sooner than\n> later anyway, as they need to pass down e.g. HASH_WRITE_OBJECT. We just\n> do it slightly earlier.\n>\n> And because of that in the general case it's really not the same, I\n> think it's a better approach. You've already got the bug in yours of\n> needlessly setting up the bulk checkin for !HASH_WRITE_OBJECT in\n> update-index, which this neatly solves by deferring the \"bulk\" mechanism\n> until the codepath that's past that and into the \"real\" object writing.\n>\n> We can also die() or error out in the object writing before ever getting\n> to writing the object, in which case we'd do some setup that we'd need\n> to tear down again, by deferring it until the last moment...\n>\n\nI'll be submitting a new version to the list which sets up the tmp\nobjdir lazily on first actual write, so the concern about writing to\nthe ODB needlessly should go away.\n\n> > And I think there will be a lot of\n> > complexity in supporting the same hints for command-line additions\n> > (which is roughly equivalent to the git-add workflow).\n>\n> I left that out due to Junio's comment in\n> https://lore.kernel.org/git/xmqqzgljyz34.fsf@gitster.g/; i.e. I don't\n> see why we'd find it worthwhile to optimize that case, but we easily\n> could (especially per the \"just sync the INDEX.tmp\" above).\n>\n> But even if we don't do \"THREE\" above I think it's still easy, for \"TWO\"\n> we already have as parse_options() state machine to parse argv as it\n> comes in. Doing the fsync() on the last object is just a matter of\n> \"looking ahead\" there).\n>\n> > Every caller\n> > that wants batch treatment will have to either implement a state\n> > machine or implement a buffering mechanism in order to figure out the\n> > begin-end points. Having a separate plug/unplug call eliminates this\n> > complexity on each caller.\n>\n> This is subjective, but I really think that's rather easy to do, and\n> much easier to reason about than the global state on the side via\n> singletons that your method of avoiding modifying these callers and\n> instead having them all consult global state via bulk-checkin.c and\n> cache.h demands.\n\nThe nice thing about having the ODB handle the batch stuff internally\nis that it can present a nice minimal interface to all of the callers.\nYes, it has a complex implementation internally, but that complexity\nbacks a rather simple API surface:\n1. Begin/end transaction (plug/unplug checkin).\n2. Find-object by SHA\n3. Add object if it doesn't exist\n4. Get the SHA without adding anything.\n\nThe ODB work is implemented once and callers can easily adopt the\ntransaction API without having to implement their own stuff on the\nside.  Future series can make the transaction span nicely across the\nODB, index, and refs.\n\n> That API also currently assumes single-threaded writers, if we start\n> writing some of this in parallel in e.g. \"unpack-objects\" we'd need\n> mutexes in bulk-object.[ch]. Isn't that a lot easier when the caller\n> would instead know something about the special nature of the transaction\n> they're interacting with, and that the 1st and last item are important\n> (for a \"BEGIN\" and \"FLUSH\").\n>\n\nThe API as sketched above doesn't deeply assume single-threadedness\nfor the \"find object by SHA\" or \"add object if it doesn't exist\".\nThere is a single-threaded assumption for begin/end-transaction.  The\nimplementation can use pthread_once to handle anything that needs to\nbe done lazily when adding objects.\n\n> > Btw, I'm planning in a future series to reduce the system calls\n> > involved in renaming a file by taking advantage of the renameat2\n> > system call and equivalents on other platforms.  There's a pretty\n> > strong motivation to do that on Windows.\n>\n> What do you have in mind for renameat2() specifically?  I.e. which of\n> the 3x flags it implements will benefit us? RENAME_NOREPLACE to \"move\"\n> the tmp_OBJ to an eventual OBJ?\n>\n\nYes RENAME_NOREPLACE.  I'd want to introduce a helper called\ngit_rename_noreplace and use it instead of the link dance.\n\n> Generally: There's some low-hanging fruit there. E.g. for tmp-objdir we\n> slavishly go through the motion of writing an tmp_OBJ, writing (and\n> possibly syncing it), then renaming that tmp_OBJ to OBJ.\n>\n> We could clearly just avoid that in some/all cases that use\n> tmp-objdir. I.e. we're writing to a temporary store anyway, so why the\n> tmp_OBJ files? We could just write to the final destinations instead,\n> they're not reachable (by ref or OID lookup) from anyone else yet.\n>\n\nWe were thinking before that there could be some concurrency in the\ntmp_objdir, though I personally don't believe it's possible for the\ntypical bulk checkin case.  Using the final name in the tmp objdir\nwould be a nice optimization, but I think that it's a separable\nconcern that shouldn't block the bigger win from eliminating the cache\nflushes.\n\n> But even then I don't see how you'd get away with reducing some classes\n> of syscalls past the 2x increase for some (leading an overall increase,\n> but not a ~2x overall increase as noted in:\n> https://lore.kernel.org/git/RFC-patch-7.7-481f1d771cb-20220323T033928Z-avarab@gmail.com/)\n> as long as you use the tmp-objdir API. It's always going to have to\n> write tmpdir/OBJ and link()/rename() that to OBJ.\n>\n> Now, I do think there's an easy way by extending the API use I've\n> introduced in this RFC to do it. I.e. we'd just do:\n>\n>     ## METHOD FOUR\n>     # A -- SAME as THREE, except no rename()\n>     write(objects/A.tmp)\n>     bulk_fsync(objects/A.tmp)\n>     # B -- SAME as THREE, except no rename()\n>     write(objects/B.tmp)\n>     bulk_fsync(objects/B.tmp)\n>     # ref -- SAME\n>     write(INDEX.tmp, $(git rev-parse B))\n>     fsync(INDEX.tmp)\n>     # NEW: do all the renames at the end:\n>     rename(objects/A.tmp, objects/A)\n>     rename(objects/B.tmp, objects/B)\n>     rename(INDEX.tmp, INDEX)\n>\n> That seems like an obvious win to me in any case. I.e. the tmp-objdir\n> API isn't really a close fit for what we *really* want to do in this\n> case.\n>\n\nI think this is the right place to get to eventually.  I believe the\nbest way to get there is to keep the plug/unplug bulk checkin\nfunctionality (rebranding it as an 'ODB transaction') and then make\nthat a sub-transaction of a larger 'git repo transaction.'\n\n> I.e. the reason it does everything this way is because it was explicitly\n> designed for 722ff7f876c (receive-pack: quarantine objects until\n> pre-receive accepts, 2016-10-03), where it's the right trade-off,\n> because we'd like to cheaply \"rm -rf\" the whole thing if e.g. the\n> \"pre-receive\" hook rejects the push.\n>\n> *AND* because it's made for the case of other things concurrently\n> needing access to those objects. So pedantically you would need it for\n> some modes of \"git update-index\", but not e.g. \"git unpack-objects\"\n> where we really are expecting to keep all of them.\n>\n> > Thanks for the concrete code,\n>\n> ..but no thanks? I.e. it would be useful to explicitly know if you're\n> interested or open to running with some of the approach in this RFC.\n\nI'm still at the point of arguing with you about your RFC, but I'm\n_not_ currently leaning toward adopting your approach.  I think from a\nseparation-of-concerns perspective, we shouldn't change top-level git\ncommands to try hard to track first/last object.  The ODB should\nconceptually handle it internally as part of a higher-level\ntransaction.  Consider cmd_add, which does its interesting\nadd_file_to_index from the update_callback coming from the diff code:\nI believe it would be hopelessly complex/impossible to do the tracking\nrequired to pass the LAST_OF_N flag to a multiplexed write API.\n\nWe have a pretty clear example from the database world that\nbegin/end-transaction is the right way to design the API for the task\nwe want to accomplish.  It's also how many filesystems work\ninternally.  I don't want to reinvent the bicycle here.\n\nThanks,\nNeeraj\n"},{"id":"452026","messageId":"CANQDOdfoMztNW6o8=zM43um5+uQFnzvh-0gz85srZgr+4Bev2A@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-v2-1.7-98921aa2052-20220323T140753Z-avarab@gmail.com","subject":"Re: [RFC PATCH v2 1/7] unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flags","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T20:23:25Z","receivedAt":"2022-03-23T20:23:42Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 7:18 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> In preparation for making the bulk-checkin.c logic operate from\n> object-file.c itself in some common cases let's add\n> HASH_N_OBJECTS{,_{FIRST,LAST}} flags.\n>\n> This will allow us to adjust for-loops that add N objects to just pass\n> down whether they have >1 objects (HASH_N_OBJECTS), as well as passing\n> down flags for whether we have the first or last object.\n>\n> We'll thus be able to drive any sort of batch-object mechanism from\n> write_object_file_flags() directly, which until now didn't know if it\n> was doing one object, or some arbitrary N.\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  builtin/unpack-objects.c | 60 +++++++++++++++++++++++-----------------\n>  cache.h                  |  3 ++\n>  2 files changed, 37 insertions(+), 26 deletions(-)\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index c55b6616aed..ec40c6fd966 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -233,7 +233,8 @@ static void write_rest(void)\n>  }\n>\n>  static void added_object(unsigned nr, enum object_type type,\n> -                        void *data, unsigned long size);\n> +                        void *data, unsigned long size,\n> +                        unsigned oflags);\n>\n>  /*\n>   * Write out nr-th object from the list, now we know the contents\n> @@ -241,21 +242,21 @@ static void added_object(unsigned nr, enum object_type type,\n>   * to be checked at the end.\n>   */\n>  static void write_object(unsigned nr, enum object_type type,\n> -                        void *buf, unsigned long size)\n> +                        void *buf, unsigned long size, unsigned oflags)\n>  {\n>         if (!strict) {\n> -               if (write_object_file(buf, size, type,\n> -                                     &obj_list[nr].oid) < 0)\n> +               if (write_object_file_flags(buf, size, type,\n> +                                     &obj_list[nr].oid, oflags) < 0)\n>                         die(\"failed to write object\");\n> -               added_object(nr, type, buf, size);\n> +               added_object(nr, type, buf, size, oflags);\n>                 free(buf);\n>                 obj_list[nr].obj = NULL;\n>         } else if (type == OBJ_BLOB) {\n>                 struct blob *blob;\n> -               if (write_object_file(buf, size, type,\n> -                                     &obj_list[nr].oid) < 0)\n> +               if (write_object_file_flags(buf, size, type,\n> +                                           &obj_list[nr].oid, oflags) < 0)\n>                         die(\"failed to write object\");\n> -               added_object(nr, type, buf, size);\n> +               added_object(nr, type, buf, size, oflags);\n>                 free(buf);\n>\n>                 blob = lookup_blob(the_repository, &obj_list[nr].oid);\n> @@ -269,7 +270,7 @@ static void write_object(unsigned nr, enum object_type type,\n>                 int eaten;\n>                 hash_object_file(the_hash_algo, buf, size, type,\n>                                  &obj_list[nr].oid);\n> -               added_object(nr, type, buf, size);\n> +               added_object(nr, type, buf, size, oflags);\n>                 obj = parse_object_buffer(the_repository, &obj_list[nr].oid,\n>                                           type, size, buf,\n>                                           &eaten);\n> @@ -283,7 +284,7 @@ static void write_object(unsigned nr, enum object_type type,\n>\n>  static void resolve_delta(unsigned nr, enum object_type type,\n>                           void *base, unsigned long base_size,\n> -                         void *delta, unsigned long delta_size)\n> +                         void *delta, unsigned long delta_size, unsigned oflags)\n>  {\n>         void *result;\n>         unsigned long result_size;\n> @@ -294,7 +295,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n>         if (!result)\n>                 die(\"failed to apply delta\");\n>         free(delta);\n> -       write_object(nr, type, result, result_size);\n> +       write_object(nr, type, result, result_size, oflags);\n>  }\n>\n>  /*\n> @@ -302,7 +303,7 @@ static void resolve_delta(unsigned nr, enum object_type type,\n>   * resolve all the deltified objects that are based on it.\n>   */\n>  static void added_object(unsigned nr, enum object_type type,\n> -                        void *data, unsigned long size)\n> +                        void *data, unsigned long size, unsigned oflags)\n>  {\n>         struct delta_info **p = &delta_list;\n>         struct delta_info *info;\n> @@ -313,7 +314,7 @@ static void added_object(unsigned nr, enum object_type type,\n>                         *p = info->next;\n>                         p = &delta_list;\n>                         resolve_delta(info->nr, type, data, size,\n> -                                     info->delta, info->size);\n> +                                     info->delta, info->size, oflags);\n>                         free(info);\n>                         continue;\n>                 }\n> @@ -322,18 +323,19 @@ static void added_object(unsigned nr, enum object_type type,\n>  }\n>\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n> -                                  unsigned nr)\n> +                                  unsigned nr, unsigned oflags)\n>  {\n>         void *buf = get_data(size);\n>\n>         if (!dry_run && buf)\n> -               write_object(nr, type, buf, size);\n> +               write_object(nr, type, buf, size, oflags);\n>         else\n>                 free(buf);\n>  }\n>\n>  static int resolve_against_held(unsigned nr, const struct object_id *base,\n> -                               void *delta_data, unsigned long delta_size)\n> +                               void *delta_data, unsigned long delta_size,\n> +                               unsigned oflags)\n>  {\n>         struct object *obj;\n>         struct obj_buffer *obj_buffer;\n> @@ -344,12 +346,12 @@ static int resolve_against_held(unsigned nr, const struct object_id *base,\n>         if (!obj_buffer)\n>                 return 0;\n>         resolve_delta(nr, obj->type, obj_buffer->buffer,\n> -                     obj_buffer->size, delta_data, delta_size);\n> +                     obj_buffer->size, delta_data, delta_size, oflags);\n>         return 1;\n>  }\n>\n>  static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n> -                              unsigned nr)\n> +                              unsigned nr, unsigned oflags)\n>  {\n>         void *delta_data, *base;\n>         unsigned long base_size;\n> @@ -366,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>                 if (has_object_file(&base_oid))\n>                         ; /* Ok we have this one */\n>                 else if (resolve_against_held(nr, &base_oid,\n> -                                             delta_data, delta_size))\n> +                                             delta_data, delta_size, oflags))\n>                         return; /* we are done */\n>                 else {\n>                         /* cannot resolve yet --- queue it */\n> @@ -428,7 +430,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>                 }\n>         }\n>\n> -       if (resolve_against_held(nr, &base_oid, delta_data, delta_size))\n> +       if (resolve_against_held(nr, &base_oid, delta_data, delta_size, oflags))\n>                 return;\n>\n>         base = read_object_file(&base_oid, &type, &base_size);\n> @@ -440,11 +442,11 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>                 has_errors = 1;\n>                 return;\n>         }\n> -       resolve_delta(nr, type, base, base_size, delta_data, delta_size);\n> +       resolve_delta(nr, type, base, base_size, delta_data, delta_size, oflags);\n>         free(base);\n>  }\n>\n> -static void unpack_one(unsigned nr)\n> +static void unpack_one(unsigned nr, unsigned oflags)\n>  {\n>         unsigned shift;\n>         unsigned char *pack;\n> @@ -472,11 +474,11 @@ static void unpack_one(unsigned nr)\n>         case OBJ_TREE:\n>         case OBJ_BLOB:\n>         case OBJ_TAG:\n> -               unpack_non_delta_entry(type, size, nr);\n> +               unpack_non_delta_entry(type, size, nr, oflags);\n>                 return;\n>         case OBJ_REF_DELTA:\n>         case OBJ_OFS_DELTA:\n> -               unpack_delta_entry(type, size, nr);\n> +               unpack_delta_entry(type, size, nr, oflags);\n>                 return;\n>         default:\n>                 error(\"bad object type %d\", type);\n> @@ -491,6 +493,7 @@ static void unpack_all(void)\n>  {\n>         int i;\n>         struct pack_header *hdr = fill(sizeof(struct pack_header));\n> +       unsigned oflags;\n>\n>         nr_objects = ntohl(hdr->hdr_entries);\n>\n> @@ -505,9 +508,14 @@ static void unpack_all(void)\n>                 progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>         CALLOC_ARRAY(obj_list, nr_objects);\n>         plug_bulk_checkin();\n> +       oflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n>         for (i = 0; i < nr_objects; i++) {\n> -               unpack_one(i);\n> -               display_progress(progress, i + 1);\n> +               int nth = i + 1;\n> +               unsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :\n> +                       nr_objects == nth ? HASH_N_OBJECTS_LAST : 0;\n> +\n> +               unpack_one(i, oflags | f);\n> +               display_progress(progress, nth);\n>         }\n>         unplug_bulk_checkin();\n>         stop_progress(&progress);\n> diff --git a/cache.h b/cache.h\n> index 84fafe2ed71..72c91c91286 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -896,6 +896,9 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n>  #define HASH_FORMAT_CHECK 2\n>  #define HASH_RENORMALIZE  4\n>  #define HASH_SILENT 8\n> +#define HASH_N_OBJECTS 1<<4\n> +#define HASH_N_OBJECTS_FIRST 1<<5\n> +#define HASH_N_OBJECTS_LAST 1<<6\n>  int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n>  int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n>\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nThis patch works out okay because unpack-objects is the easy case.\nYou have a well defined number of objects.  I'd be fine with your\ndesign if all cases were like this.\n"},{"id":"452029","messageId":"CANQDOdd6m_xC6Jt41seLBnvjmM_xorFxmzNYCOs6uqGXcvpdqA@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-v2-2.7-c6f776fc2bc-20220323T140753Z-avarab@gmail.com","subject":"Re: [RFC PATCH v2 2/7] object-file: pass down unpack-objects.c flags for \"bulk\" checkin","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T20:25:43Z","receivedAt":"2022-03-23T20:26:00Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 7:18 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> Remove much of this as a POC for exploring some of what I mentioned in\n> https://lore.kernel.org/git/220322.86mthinxnn.gmgdl@evledraar.gmail.com/\n>\n> This commit is obviously not what we *should* do as end-state, but\n> demonstrates what's needed (I think) for a bare-minimum implementation\n> of just the \"bulk\" syncing method for loose objects without the part\n> where we do the tmp-objdir.c dance.\n>\n> Performance with this is already quite promising. Benchmarking with:\n>\n>         git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' \\\n>                 -p 'rm -rf r.git && git init --bare r.git' \\\n>                 './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack'\n>\n> I.e. unpacking a small packfile (my dotfiles) yields, on a Linux\n> ramdisk:\n>\n>         Benchmark 1: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync\n>           Time (mean ± σ):     815.9 ms ±   8.2 ms    [User: 522.9 ms, System: 287.9 ms]\n>           Range (min … max):   805.6 ms … 835.9 ms    10 runs\n>\n>         Benchmark 2: ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD\n>           Time (mean ± σ):     779.4 ms ±  15.4 ms    [User: 505.7 ms, System: 270.2 ms]\n>           Range (min … max):   763.1 ms … 813.9 ms    10 runs\n>\n>         Summary\n>           './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n>             1.05 ± 0.02 times faster than './git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n>\n> Doing the same with \"strace --summary-only\", which probably helps to\n> emulate cases with slower syscalls is ~15% faster than using the\n> tmp-objdir indirection:\n>\n>         Summary\n>           'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'HEAD' ran\n>             1.16 ± 0.01 times faster than 'strace --summary-only ./git -C r.git -c core.fsync=loose-object -c core.fsyncMethod=batch unpack-objects </tmp/pack-dotfiles.pack' in 'ns/batched-fsync'\n>\n> Which makes sense in terms of syscalls. In my case HEAD has ~101k\n> calls, and the parent topic is making ~129k calls, with around 2x the\n> number of unlink(), link() as expected.\n>\n> Of course some users will want to use the tmp-objdir.c method. So a\n> version of this commit could be rewritten to come earlier in the\n> series, with the \"bulk\" on top being optional.\n>\n> It seems to me that it's a much better strategy to do this whole thing\n> in close_loose_object() after passing down the new HASH_N_OBJECTS /\n> HASH_N_OBJECTS_FIRST / HASH_N_OBJECTS_LAST flags.\n>\n> Doing that for the \"builtin/add.c\" and \"builtin/unpack-objects.c\" code\n> having its {un,}plug_bulk_checkin() removed here is then just a matter\n> of passing down a similar set of flags indicating whether we're\n> dealing with N objects, and if so if we're dealing with the last one\n> or not.\n>\n> As we'll see in subsequent commits doing it this way also effortlessly\n> integrates with other HASH_* flags. E.g. for \"update-index\" the code\n> being rm'd here doesn't handle the interaction with\n> \"HASH_WRITE_OBJECT\" properly, but once we've moved all this sync\n> bootstrapping logic to close_loose_object() we'll never get to it if\n> we're not actually writing something.\n>\n> This code currently doesn't use the HASH_N_OBJECTS_FIRST flag, but\n> that's what we'd use later to optionally call tmp_objdir_create().\n>\n> Aside: This also changes logic that was a bit confusing and repetitive\n> in close_loose_object(). Previously we'd first call\n> batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT) which is just as\n> shorthand for:\n>\n>         fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT &&\n>         fsync_method == FSYNC_METHOD_BATCH\n>\n> We'd then proceed to call\n> fsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT) later in the same\n> function, which is just a way of calling fsync_or_die() if:\n>\n>         fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\n>\n> Now we instead just define a local \"fsync_loose\" variable by checking\n> \"fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT\", which shows us that\n> the previous case of fsync_component_or_die(...)\" could just be added\n> to the existing \"fsync_object_files > 0\" branch.\n>\n> Note: This commit reverts much of \"core.fsyncmethod: batched disk\n> flushes for loose-objects\". We'll set up new structures to bring what\n> it was doing back in a different way. I.e. to do the tmp-objdir\n> plug-in in object-file.c\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  builtin/unpack-objects.c |  2 --\n>  builtin/update-index.c   |  4 ---\n>  bulk-checkin.c           | 74 ----------------------------------------\n>  bulk-checkin.h           |  3 --\n>  cache.h                  |  5 ---\n>  object-file.c            | 37 ++++++++++++++------\n>  6 files changed, 26 insertions(+), 99 deletions(-)\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index ec40c6fd966..93da436581b 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -507,7 +507,6 @@ static void unpack_all(void)\n>         if (!quiet)\n>                 progress = start_progress(_(\"Unpacking objects\"), nr_objects);\n>         CALLOC_ARRAY(obj_list, nr_objects);\n> -       plug_bulk_checkin();\n>         oflags = nr_objects > 1 ? HASH_N_OBJECTS : 0;\n>         for (i = 0; i < nr_objects; i++) {\n>                 int nth = i + 1;\n> @@ -517,7 +516,6 @@ static void unpack_all(void)\n>                 unpack_one(i, oflags | f);\n>                 display_progress(progress, nth);\n>         }\n> -       unplug_bulk_checkin();\n>         stop_progress(&progress);\n>\n>         if (delta_list)\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index cbd2b0d633b..95ed3c47b2e 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -1118,8 +1118,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>         parse_options_start(&ctx, argc, argv, prefix,\n>                             options, PARSE_OPT_STOP_AT_NON_OPTION);\n>\n> -       /* optimize adding many objects to the object database */\n> -       plug_bulk_checkin();\n>         while (ctx.argc) {\n>                 if (parseopt_state != PARSE_OPT_DONE)\n>                         parseopt_state = parse_options_step(&ctx, options,\n> @@ -1194,8 +1192,6 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>                 strbuf_release(&buf);\n>         }\n>\n> -       /* by now we must have added all of the new objects */\n> -       unplug_bulk_checkin();\n>         if (split_index > 0) {\n>                 if (git_config_get_split_index() == 0)\n>                         warning(_(\"core.splitIndex is set to false; \"\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index a0dca79ba6a..577b135e39c 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -3,20 +3,15 @@\n>   */\n>  #include \"cache.h\"\n>  #include \"bulk-checkin.h\"\n> -#include \"lockfile.h\"\n>  #include \"repository.h\"\n>  #include \"csum-file.h\"\n>  #include \"pack.h\"\n>  #include \"strbuf.h\"\n> -#include \"string-list.h\"\n> -#include \"tmp-objdir.h\"\n>  #include \"packfile.h\"\n>  #include \"object-store.h\"\n>\n>  static int bulk_checkin_plugged;\n>\n> -static struct tmp_objdir *bulk_fsync_objdir;\n> -\n>  static struct bulk_checkin_state {\n>         char *pack_tmp_name;\n>         struct hashfile *f;\n> @@ -85,40 +80,6 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n>         reprepare_packed_git(the_repository);\n>  }\n>\n> -/*\n> - * Cleanup after batch-mode fsync_object_files.\n> - */\n> -static void do_batch_fsync(void)\n> -{\n> -       struct strbuf temp_path = STRBUF_INIT;\n> -       struct tempfile *temp;\n> -\n> -       if (!bulk_fsync_objdir)\n> -               return;\n> -\n> -       /*\n> -        * Issue a full hardware flush against a temporary file to ensure\n> -        * that all objects are durable before any renames occur. The code in\n> -        * fsync_loose_object_bulk_checkin has already issued a writeout\n> -        * request, but it has not flushed any writeback cache in the storage\n> -        * hardware or any filesystem logs. This fsync call acts as a barrier\n> -        * to ensure that the data in each new object file is durable before\n> -        * the final name is visible.\n> -        */\n> -       strbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> -       temp = xmks_tempfile(temp_path.buf);\n> -       fsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> -       delete_tempfile(&temp);\n> -       strbuf_release(&temp_path);\n> -\n> -       /*\n> -        * Make the object files visible in the primary ODB after their data is\n> -        * fully durable.\n> -        */\n> -       tmp_objdir_migrate(bulk_fsync_objdir);\n> -       bulk_fsync_objdir = NULL;\n> -}\n> -\n>  static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n>  {\n>         int i;\n> @@ -313,26 +274,6 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>         return 0;\n>  }\n>\n> -void prepare_loose_object_bulk_checkin(void)\n> -{\n> -       if (bulk_checkin_plugged && !bulk_fsync_objdir)\n> -               bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> -}\n> -\n> -void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n> -{\n> -       /*\n> -        * If we have a plugged bulk checkin, we issue a call that\n> -        * cleans the filesystem page cache but avoids a hardware flush\n> -        * command. Later on we will issue a single hardware flush\n> -        * before as part of do_batch_fsync.\n> -        */\n> -       if (!bulk_fsync_objdir ||\n> -           git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n> -               fsync_or_die(fd, filename);\n> -       }\n> -}\n> -\n>  int index_bulk_checkin(struct object_id *oid,\n>                        int fd, size_t size, enum object_type type,\n>                        const char *path, unsigned flags)\n> @@ -347,19 +288,6 @@ int index_bulk_checkin(struct object_id *oid,\n>  void plug_bulk_checkin(void)\n>  {\n>         assert(!bulk_checkin_plugged);\n> -\n> -       /*\n> -        * A temporary object directory is used to hold the files\n> -        * while they are not fsynced.\n> -        */\n> -       if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n> -               bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> -               if (!bulk_fsync_objdir)\n> -                       die(_(\"Could not create temporary object directory for core.fsyncMethod=batch\"));\n> -\n> -               tmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> -       }\n> -\n>         bulk_checkin_plugged = 1;\n>  }\n>\n> @@ -369,6 +297,4 @@ void unplug_bulk_checkin(void)\n>         bulk_checkin_plugged = 0;\n>         if (bulk_checkin_state.f)\n>                 finish_bulk_checkin(&bulk_checkin_state);\n> -\n> -       do_batch_fsync();\n>  }\n> diff --git a/bulk-checkin.h b/bulk-checkin.h\n> index 181d3447ff9..b26f3dc3b74 100644\n> --- a/bulk-checkin.h\n> +++ b/bulk-checkin.h\n> @@ -6,9 +6,6 @@\n>\n>  #include \"cache.h\"\n>\n> -void prepare_loose_object_bulk_checkin(void);\n> -void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n> -\n>  int index_bulk_checkin(struct object_id *oid,\n>                        int fd, size_t size, enum object_type type,\n>                        const char *path, unsigned flags);\n> diff --git a/cache.h b/cache.h\n> index 72c91c91286..2f3831fa853 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1772,11 +1772,6 @@ void fsync_or_die(int fd, const char *);\n>  int fsync_component(enum fsync_component component, int fd);\n>  void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n>\n> -static inline int batch_fsync_enabled(enum fsync_component component)\n> -{\n> -       return (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n> -}\n> -\n>  ssize_t read_in_full(int fd, void *buf, size_t count);\n>  ssize_t write_in_full(int fd, const void *buf, size_t count);\n>  ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\n> diff --git a/object-file.c b/object-file.c\n> index cd0ddb49e4b..dbeb3df502d 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1886,19 +1886,37 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>         hash_object_file_literally(algo, buf, len, type_name(type), oid);\n>  }\n>\n> +static void sync_loose_object_batch(int fd, const char *filename,\n> +                                   const unsigned oflags)\n> +{\n> +       const int last = oflags & HASH_N_OBJECTS_LAST;\n> +\n> +       /*\n> +        * We're doing a sync_file_range() (or equivalent) for 1..N-1\n> +        * objects, and then a \"real\" fsync() for N. On some OS's\n> +        * enabling core.fsync=loose-object && core.fsyncMethod=batch\n> +        * improves the performance by a lot.\n> +        */\n> +       if (last || (!last && git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0))\n> +               fsync_or_die(fd, filename);\n> +}\n> +\n>  /* Finalize a file on disk, and close it. */\n> -static void close_loose_object(int fd, const char *filename)\n> +static void close_loose_object(int fd, const char *filename,\n> +                              const unsigned oflags)\n>  {\n> +       int fsync_loose;\n> +\n>         if (the_repository->objects->odb->will_destroy)\n>                 goto out;\n>\n> -       if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> -               fsync_loose_object_bulk_checkin(fd, filename);\n> -       else if (fsync_object_files > 0)\n> +       fsync_loose = fsync_components & FSYNC_COMPONENT_LOOSE_OBJECT;\n> +\n> +       if (oflags & HASH_N_OBJECTS && fsync_loose &&\n> +           fsync_method == FSYNC_METHOD_BATCH)\n> +               sync_loose_object_batch(fd, filename, oflags);\n> +       else if (fsync_object_files > 0 || fsync_loose)\n>                 fsync_or_die(fd, filename);\n> -       else\n> -               fsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n> -                                      filename);\n>\n>  out:\n>         if (close(fd) != 0)\n> @@ -1962,9 +1980,6 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         static struct strbuf tmp_file = STRBUF_INIT;\n>         static struct strbuf filename = STRBUF_INIT;\n>\n> -       if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> -               prepare_loose_object_bulk_checkin();\n> -\n>         loose_object_path(the_repository, &filename, oid);\n>\n>         fd = create_tmpfile(&tmp_file, filename.buf);\n> @@ -2015,7 +2030,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>                 die(_(\"confused by unstable object source data for %s\"),\n>                     oid_to_hex(oid));\n>\n> -       close_loose_object(fd, tmp_file.buf);\n> +       close_loose_object(fd, tmp_file.buf, flags);\n>\n>         if (mtime) {\n>                 struct utimbuf utb;\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nFine. Doing this patch series as non-RFC, we could start from prior to\nmy fsyncMethod=batch series.\n"},{"id":"452030","messageId":"CANQDOdffANOvbTBAZs95PxQMgCu1Leww6+a7A960hcYi+4mNDQ@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com","subject":"Re: [RFC PATCH v2 4/7] update-index: have the index fsync() flush the loose objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T20:30:06Z","receivedAt":"2022-03-23T20:30:22Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 7:18 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> As with unpack-objects in a preceding commit have update-index.c make\n> use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a \"batch\"\n> mode again for \"update-index\".\n>\n> Adding the t/* directory from git.git on a Linux ramdisk is a bit\n> faster than with the tmp-objdir indirection:\n>\n>         $ git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt' -p 'rm -rf repo/.git/objects/* repo/.git/index' './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' --warmup 1 -r 10Benchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync\n>           Time (mean ± σ):     281.1 ms ±   2.6 ms    [User: 186.2 ms, System: 92.3 ms]\n>           Range (min … max):   278.3 ms … 287.0 ms    10 runs\n>\n>         Benchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD\n>           Time (mean ± σ):     265.9 ms ±   2.6 ms    [User: 181.7 ms, System: 82.1 ms]\n>           Range (min … max):   262.0 ms … 270.3 ms    10 runs\n>\n>         Summary\n>           './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n>             1.06 ± 0.01 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n>\n> And as before running that with \"strace --summary-only\" slows things\n> down a bit (probably mimicking slower I/O a bit). I then get:\n>\n>         Summary\n>           'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'HEAD' ran\n>             1.19 ± 0.03 times faster than 'strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'ns/batched-fsync'\n>\n> This one has a twist though, instead of fsync()-ing on the last object\n> we write we'll not do that, and instead defer the fsync() until we\n> write the index itself. This is outlined in [1] (as \"METHOD THREE\").\n>\n> Because of this under FSYNC_METHOD_BATCH we'll do the N\n> objects (possibly only one, because we're lazy) as HASH_N_OBJECTS, and\n> we'll even now support doing this via N arguments on the command-line.\n>\n> Then we won't fsync() any of it, but we will rename it\n> in-place (which, if we were still using the tmp-objdir, would leave it\n> \"staged\" in the tmp-objdir).\n>\n> We'll then have the fsync() for the index update \"flush\" that out, and\n> thus avoid two fsync() calls when one will do.\n>\n> Running this with the \"git hyperfine\" command mentioned in a preceding\n> commit with \"strace --summary-only\" shows that we do 1 fsync() now\n> instead of 2, and have one more sync_file_range(), as expected.\n>\n> We also go from ~51k syscalls to ~39k, with ~2x the number of link()\n> and unlink() in ns/batched-fsync, and of course one fsync() instead of\n> two()>\n>\n> The flow of this code isn't quite set up for re-plugging the\n> tmp-objdir back in. In particular we no longer pass\n> HASH_N_OBJECTS_FIRST (but doing so would be trivial)< and there's no\n> HASH_N_OBJECTS_LAST.\n>\n> So this and other callers would need some light transaction-y API, or\n> to otherwise pass down a \"yes, I'd like to flush it\" down to\n> finalize_hashfile(), but doing so will be trivial.\n>\n> And since we've started structuring it this way it'll become easy to\n> do any arbitrary number of things down the line that would \"bulk\n> fsync\" before the final fsync(). Now we write some objects and fsync()\n> on the index, but between those two could do any number of other\n> things where we'd defer the fsync().\n>\n> This sort of thing might be especially interesting for \"git repack\"\n> when it writes e.g. a *.bitmap, *.rev, *.pack and *.idx. In that case\n> we could skip the fsync() on all of those, and only do it on the *.idx\n> before we renamed it in-place. I *think* nothing cares about a *.pack\n> without an *.idx, but even then we could fsync *.idx, rename *.pack,\n> rename *.idx and still safely do only one fsync(). See \"git show\n> --first-parent\" on 62874602032 (Merge branch\n> 'tb/pack-finalize-ordering' into maint, 2021-10-12) for a good\n> overview of the code involved in that.\n>\n> 1. https://lore.kernel.org/git/220323.86sfr9ndpr.gmgdl@evledraar.gmail.com/\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  builtin/update-index.c |  7 ++++---\n>  cache.h                |  1 +\n>  read-cache.c           | 29 ++++++++++++++++++++++++++++-\n>  3 files changed, 33 insertions(+), 4 deletions(-)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index 34aaaa16c20..6cfec6efb38 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -1142,7 +1142,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>\n>                         setup_work_tree();\n>                         p = prefix_path(prefix, prefix_length, path);\n> -                       update_one(p, 0);\n> +                       update_one(p, HASH_N_OBJECTS);\n>                         if (set_executable_bit)\n>                                 chmod_path(set_executable_bit, p);\n>                         free(p);\n> @@ -1187,7 +1187,7 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>                                 strbuf_swap(&buf, &unquoted);\n>                         }\n>                         p = prefix_path(prefix, prefix_length, buf.buf);\n> -                       update_one(p, 0);\n> +                       update_one(p, HASH_N_OBJECTS);\n>                         if (set_executable_bit)\n>                                 chmod_path(set_executable_bit, p);\n>                         free(p);\n> @@ -1263,7 +1263,8 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>                                 exit(128);\n>                         unable_to_lock_die(get_index_file(), lock_error);\n>                 }\n> -               if (write_locked_index(&the_index, &lock_file, COMMIT_LOCK))\n> +               if (write_locked_index(&the_index, &lock_file,\n> +                                      COMMIT_LOCK | WLI_NEED_LOOSE_FSYNC))\n>                         die(\"Unable to write new index file\");\n>         }\n>\n> diff --git a/cache.h b/cache.h\n> index 2f3831fa853..7542e009a34 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -751,6 +751,7 @@ void ensure_full_index(struct index_state *istate);\n>  /* For use with `write_locked_index()`. */\n>  #define COMMIT_LOCK            (1 << 0)\n>  #define SKIP_IF_UNCHANGED      (1 << 1)\n> +#define WLI_NEED_LOOSE_FSYNC   (1 << 2)\n>\n>  /*\n>   * Write the index while holding an already-taken lock. Close the lock,\n> diff --git a/read-cache.c b/read-cache.c\n> index 3e0e7d41837..275f6308c32 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2860,6 +2860,33 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         int ieot_entries = 1;\n>         struct index_entry_offset_table *ieot = NULL;\n>         int nr, nr_threads;\n> +       unsigned int wflags = FSYNC_COMPONENT_INDEX;\n> +\n> +\n> +       /*\n> +        * TODO: This is abuse of the API recently modified\n> +        * finalize_hashfile() which reveals a shortcoming of its\n> +        * \"fsync\" design.\n> +        *\n> +        * I.e. It expects a \"enum fsync_component component\" label,\n> +        * but here we're passing it an OR of the two, knowing that\n> +        * it'll call fsync_component_or_die() which (in\n> +        * write-or-die.c) will do \"(fsync_components & wflags)\" (to\n> +        * our \"wflags\" here).\n> +        *\n> +        * But the API really should be changed to explicitly take\n> +        * such flags, because in this case we'd like to fsync() the\n> +        * index if we're in the bulk mode, *even if* our\n> +        * \"core.fsync=index\" isn't configured.\n> +        *\n> +        * That's because at this point we've been queuing up object\n> +        * writes that we didn't fsync(), and are going to use this\n> +        * fsync() to \"flush\" the whole thing. Doing it this way\n> +        * avoids redundantly calling fsync() twice when once will do.\n> +        */\n> +       if (fsync_method == FSYNC_METHOD_BATCH &&\n> +           flags & WLI_NEED_LOOSE_FSYNC)\n> +               wflags |= FSYNC_COMPONENT_LOOSE_OBJECT;\n>\n>         f = hashfd(tempfile->fd, tempfile->filename.buf);\n>\n> @@ -3094,7 +3121,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         if (!alternate_index_output && (flags & COMMIT_LOCK))\n>                 csum_fsync_flag = CSUM_FSYNC;\n>\n> -       finalize_hashfile(f, istate->oid.hash, FSYNC_COMPONENT_INDEX,\n> +       finalize_hashfile(f, istate->oid.hash, wflags,\n>                           CSUM_HASH_IN_STREAM | csum_fsync_flag);\n>\n>         if (close_tempfile_gently(tempfile)) {\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nIn the long run, we should attach the \"need to fsync the index\" to an\nongoing 'repo-transaction' so that we can composably sync at the best\npoint regardless of what the top-level git operation does.\n"},{"id":"452067","messageId":"CANQDOdcFN5GgOPZ3hqCsjHDTiRfRpqoAKxjF1n9D6S8oD9--_A@mail.gmail.com","threadId":"57568","inReplyTo":"RFC-patch-v2-7.7-a5951366c6e-20220323T140753Z-avarab@gmail.com","subject":"Re: [RFC PATCH v2 7/7] fsync docs: add new fsyncMethod.batch.quarantine, elaborate on old","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-23T21:08:37Z","receivedAt":"2022-03-23T21:08:53Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 7:18 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> Add a new fsyncMethod.batch.quarantine setting which defaults to\n> \"false\". Preceding (RFC, and not meant to flip-flop like that\n> eventually) commits ripped out the \"tmp-objdir\" part of the\n> core.fsyncMethod=batch.\n>\n> This documentation proposes to keep that as the default for the\n> reasons discussed in it, while allowing users to set\n> \"fsyncMethod.batch.quarantine=true\".\n>\n> Furthermore update the discussion of \"core.fsyncObjectFiles\" with\n> information about what it *really* does, why you probably shouldn't\n> use it, and how to safely emulate most of what it gave users in the\n> past in terms of performance benefit.\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  Documentation/config/core.txt | 80 +++++++++++++++++++++++++++++++----\n>  1 file changed, 72 insertions(+), 8 deletions(-)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index f598925b597..365a12dc7ae 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -607,21 +607,85 @@ stored on NTFS or ReFS filesystems.\n>  +\n>  The `batch` is currently only applies to loose-object files and will\n>  kick in when using the linkgit:git-unpack-objects[1] and\n> -linkgit:update-index[1] commands. Note that the \"last\" file to be\n> +linkgit:git-update-index[1] commands. Note that the \"last\" file to be\n>  synced may be the last object, as in the case of\n>  linkgit:git-unpack-objects[1], or relevant \"index\" (or in the future,\n>  \"ref\") update, as in the case of linkgit:git-update-index[1]. I.e. the\n>  batch syncing of the loose objects may be deferred until a subsequent\n>  fsync() to a file that makes them \"active\".\n>\n> +fsyncMethod.batch.quarantine::\n> +       A boolean which if set to `true` will cause \"batched\" writes\n> +       to objects to be \"quarantined\" if\n> +       `core.fsyncMethod=batch`. This is `false` by default.\n> ++\n> +The primary object of these fsync() settings is to protect against\n> +repository corruption of things which are reachable, i.e. \"reachable\",\n> +via references, the index etc. Not merely objects that were present in\n> +the object store.\n> ++\n> +Historically setting `core.fsyncObjectFiles=false` assumed that on a\n> +filesystem with where an fsync() would flush all preceding outstanding\n> +I/O that we might end up with a corrupt loose object, but that was OK\n> +as long as no reference referred to it. We'd eventually the corrupt\n> +object with linkgit:git-gc[1], and linkgit:git-fsck[1] would only\n> +report it as a minor annoyance\n> ++\n> +Setting `fsyncMethod.batch.quarantine=true` takes the view that\n> +something like a corrupt *unreferenced* loose object in the object\n> +store is something we'd like to avoid, at the cost of reduced\n> +performance when using `core.fsyncMethod=batch`.\n> ++\n> +Currently this uses the same mechanism described in the \"QUARANTINE\n> +ENVIRONMENT\" in the linkgit:git-receive-pack[1] documentation, but\n> +that's subject to change. The performance loss is because we need to\n> +\"stage\" the objects in that quarantine environment, fsync() it, and\n> +once that's done rename() or link() it in-place into the main object\n> +store, possibly with an fsync() of the index or ref at the end\n> ++\n> +With `fsyncMethod.batch.quarantine=false` we'll \"stage\" things in the\n> +main object store, and then do one fsync() at the very end, either on\n> +the last object we write, or file (index or ref) that'll make it\n> +\"reachable\".\n> ++\n> +The bad thing about setting this to `true` is lost performance, as\n> +well as not being able to access the objects as they're written (which\n> +e.g. consumers of linkgit:git-update-index[1]'s `--verbose` mode might\n> +want to do).\n\nI wasn't able to understand clearly from your performance numbers.\nWhat did you measure as the additional cost from quarantine=true\nversus quarantine=false? Just if you have the numbers handy...\n\n> ++\n> +The good thing is that you should be guaranteed not to get e.g. short\n> +or otherwise corrupt loose objects if you pull your power cord, in\n> +practice various git commands deal quite badly with discovering such a\n> +stray corrupt object (including perhaps assuming it's valid based on\n> +its existence, or hard dying on an error rather than replacing\n> +it). Repairing such \"unreachable corruption\" can require manual\n> +intervention.\n> +\n>  core.fsyncObjectFiles::\n> -       This boolean will enable 'fsync()' when writing object files.\n> -       This setting is deprecated. Use core.fsync instead.\n> -+\n> -This setting affects data added to the Git repository in loose-object\n> -form. When set to true, Git will issue an fsync or similar system call\n> -to flush caches so that loose-objects remain consistent in the face\n> -of a unclean system shutdown.\n> +       This boolean will enable 'fsync()' when writing loose object\n> +       files.\n> ++\n> +This setting is the historical fsync configuration setting. It's now\n> +*deprecated*, you should use `core.fsync` instead, perhaps in\n> +combination with `core.fsyncMethod=batch`.\n> ++\n> +The `core.fsyncObjectFiles` was initially added based on integrity\n> +assumptions that early (pre-ext-4) versions of Linux's \"ext\"\n> +filesystems provided.\n> ++\n> +I.e. that a write of file A without an `fsync()` followed by a write\n> +of file `B` with `fsync()` would implicitly guarantee that `A' would\n> +be `fsync()`'d by calling `fsync()` on `B`. This asssumption is *not*\n> +backed up by any standard (e.g. POSIX), but worked in practice on some\n> +Linux setups.\n> ++\n> +Nowadays you should almost certainly want to use\n> +`core.fsync=loose-object` instead in combination with\n> +`core.fsyncMethod=bulk`, and possibly with\n> +`fsyncMethod.batch.quarantine=true`, see above. On modern OS's (Linux,\n> +OSX, Windows) that gives you most of the performance benefit of\n> +`core.fsyncObjectFiles=false` with all of the safety of the old\n> +`core.fsyncObjectFiles=true`.\n>\n>  core.preloadIndex::\n>         Enable parallel index preload for operations like 'git diff'\n> --\n> 2.35.1.1428.g1c1a0152d61\n>\n\nI think the notion of minimizing fsyncs across the whole repository is\na great one.  However, your implementation isn't clean from an API\nperspective, since people modifying the top-level commands need to\nreason about the full set of operations to avoid silently breaking the\nfsync requirements.  I think we should phrase this as a \"transaction\"\nthat the top level command can begin and end. Subcomponents of the\nrepo can \"enlist\" in the transaction and do the right thing optimally\nwhen the overall transaction commits or aborts.\n\nIn the end, I think the optimal solution should be layered on top of\nthe final form of my current patch series as an incremental\nimprovement.  I'm going to start the rebranding of\nplug/unplug_bulk_checkin in V3 of the patch series.\n\nThanks,\nNeeraj\nLAST MAIL\n"},{"id":"452080","messageId":"CANQDOdekfyPEZTfzrJ3Una=9Xt-kCcGDtqPskKr=29cicxO-=g@mail.gmail.com","threadId":"57568","inReplyTo":"220323.86k0ckol2v.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v2 2/7] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T02:04:21Z","receivedAt":"2022-03-24T02:04:38Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 23, 2022 at 6:27 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Sun, Mar 20 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> > [..\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index 889522956e4..a3798dfc334 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -628,6 +628,13 @@ core.fsyncMethod::\n> >  * `writeout-only` issues pagecache writeback requests, but depending on the\n> >    filesystem and storage hardware, data added to the repository may not be\n> >    durable in the event of a system crash. This is the default mode on macOS.\n> > +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> > +  updates in the disk writeback cache and then does a single full fsync of\n> > +  a dummy file to trigger the disk cache flush at the end of the operation.\n>\n> I think adding a \\n\\n here would help make this more readable & break\n> the flow a bit. I.e. just add a \"+\" on its own line, followed by\n> \"Currently...\n>\n> > +  Currently `batch` mode only applies to loose-object files. Other repository\n> > +  data is made durable as if `fsync` was specified. This mode is expected to\n> > +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> > +  and on Windows for repos stored on NTFS or ReFS filesystems.\n\nThanks, will fix.\n"},{"id":"452089","messageId":"53261f0099d53524155464fe79d10f9605fe93aa.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 01/11] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:16Z","receivedAt":"2022-03-24T04:58:38Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nMake it clearer in the naming and documentation of the plug_bulk_checkin\nand unplug_bulk_checkin APIs that they can be thought of as\na \"transaction\" to optimize operations on the object database.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c  |  4 ++--\n bulk-checkin.c |  4 ++--\n bulk-checkin.h | 14 ++++++++++++--\n 3 files changed, 16 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..9bf37ceae8e 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\tstring_list_clear(&only_match_skip_worktree, 0);\n \t}\n \n-\tplug_bulk_checkin();\n+\tbegin_odb_transaction();\n \n \tif (add_renormalize)\n \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n@@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\tend_odb_transaction();\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c9..a16ae3c629d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -285,12 +285,12 @@ int index_bulk_checkin(struct object_id *oid,\n \treturn status;\n }\n \n-void plug_bulk_checkin(void)\n+void begin_odb_transaction(void)\n {\n \tstate.plugged = 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void end_odb_transaction(void)\n {\n \tstate.plugged = 0;\n \tif (state.f)\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..69a94422ac7 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -10,7 +10,17 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n-void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+/*\n+ * Tell the object database to optimize for adding\n+ * multiple objects. end_odb_transaction must be called\n+ * to make new objects visible.\n+ */\n+void begin_odb_transaction(void);\n+\n+/*\n+ * Tell the object database to make any objects from the\n+ * current transaction visible.\n+ */\n+void end_odb_transaction(void);\n \n #endif\n-- \ngitgitgadget\n\n"},{"id":"452090","messageId":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v2.git.1647760560.gitgitgadget@gmail.com","subject":"[PATCH v3 00/11] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:15Z","receivedAt":"2022-03-24T04:58:41Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"V3 changes:\n\n * Rebrand plug/unplug-bulk-checkin to \"begin_odb_transaction\" and\n   \"end_odb_transaction\"\n * Add a patch to pass filenames to fsync_or_die, rather than the string\n   \"loose object\"\n * Update the commit description for \"core.fsyncmethod to explain why we do\n   not directly expose objects until an fsync occurs.\n * Also explain in the commit description why we're using a dummy file for\n   the fsync.\n * Create the bulk-fsync tmp-objdir lazily the first time a loose object is\n   added. We now do fsync iff that objdir exists.\n * Do batch fsync if core.fsyncMethod=batch and core.fsync contains\n   loose-object, regardless of the core.fsyncObjectFiles setting.\n * Mitigate the risk in update-index of an object not being visible due to\n   bulk checkin.\n * Add a perf comment to justify the unpack-objects usage of bulk-checkin.\n * Add a new patch to create helpers for parsing OIDs from git commands.\n * Add a comment to the lib-unique-files.sh helper about uniqueness only\n   within a repo.\n * Fix style and add '&&' chaining to test helpers.\n * Comment on some magic numbers in tests.\n * Take the object list as an argument in\n   ./t5300-pack-object.sh:check_unpack ()\n * Drop accidental change to t/perf/perf-lib.sh\n\nV2 changes:\n\n * Change doc to indicate that only some repo updates are batched\n * Null and zero out control variables in do_batch_fsync under\n   unplug_bulk_checkin\n * Make batch mode default on Windows.\n * Update the description for the initial patch that cleans up the\n   bulk-checkin infrastructure.\n * Rebase onto 'seen' at 0cac37f38f9.\n\n--Original definition-- When core.fsync includes loose-object, we issue an\nfsync after every written object. For a 'git-add' or similar command that\nadds a lot of files to the repo, the costs of these fsyncs adds up. One\nmajor factor in this cost is the time it takes for the physical storage\ncontroller to flush its caches to durable media.\n\nThis series takes advantage of the writeout-only mode of git_fsync to issue\nOS cache writebacks for all of the objects being added to the repository\nfollowed by a single fsync to a dummy file, which should trigger a\nfilesystem log flush and storage controller cache flush. This mechanism is\nknown to be safe on common Windows filesystems and expected to be safe on\nmacOS. Some linux filesystems, such as XFS, will probably do the right thing\nas well. See [1] for previous discussion on the predecessor of this patch\nseries.\n\nThis series is important on Windows, where loose-objects are included in the\nfsync set by default in Git-For-Windows. In this series, I'm also setting\nthe default mode for Windows to turn on loose object fsyncing with batch\nmode, so that we can get CI coverage of the actual git-for-windows\nconfiguration upstream. We still don't actually issue fsyncs for the test\nsuite since GIT_TEST_FSYNC is set to 0, but we exercise all of the\nsurrounding batch mode code.\n\nThis work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n\n[1]\nhttps://lore.kernel.org/git/2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com/\n\nNeeraj Singh (11):\n  bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  object-file: pass filename to fsync_or_die\n  core.fsyncmethod: batched disk flushes for loose-objects\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  test-lib-functions: add parsing helpers for ls-files and ls-tree\n  core.fsyncmethod: tests for batch mode\n  core.fsyncmethod: performance tests for add and stash\n  core.fsyncmethod: correctly camel-case warning message\n\n Documentation/config/core.txt          |  8 +++\n builtin/add.c                          |  4 +-\n builtin/unpack-objects.c               |  3 +\n builtin/update-index.c                 | 33 +++++++++\n bulk-checkin.c                         | 97 ++++++++++++++++++++++----\n bulk-checkin.h                         | 17 ++++-\n cache.h                                | 12 +++-\n compat/mingw.h                         |  3 +\n config.c                               |  6 +-\n git-compat-util.h                      |  2 +\n object-file.c                          | 15 ++--\n t/lib-unique-files.sh                  | 32 +++++++++\n t/perf/p3700-add.sh                    | 59 ++++++++++++++++\n t/perf/p3900-stash.sh                  | 62 ++++++++++++++++\n t/t3700-add.sh                         | 28 ++++++++\n t/t3903-stash.sh                       | 20 ++++++\n t/t5300-pack-object.sh                 | 41 +++++++----\n t/t5317-pack-objects-filter-objects.sh | 91 ++++++++++++------------\n t/test-lib-functions.sh                | 10 +++\n 19 files changed, 458 insertions(+), 85 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\n\nbase-commit: c54b8eb302ffb72f31e73a26044c8a864e2cb307\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1134%2Fneerajsi-msft%2Fns%2Fbatched-fsync-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1134/neerajsi-msft/ns/batched-fsync-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/1134\n\nRange-diff vs v2:\n\n  -:  ----------- >  1:  53261f0099d bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  1:  9c2abd12bbb !  2:  b2d9766a662 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n       \treturn status;\n       }\n       \n     - void plug_bulk_checkin(void)\n     + void begin_odb_transaction(void)\n       {\n      -\tstate.plugged = 1;\n      +\tassert(!bulk_checkin_plugged);\n      +\tbulk_checkin_plugged = 1;\n       }\n       \n     - void unplug_bulk_checkin(void)\n     + void end_odb_transaction(void)\n       {\n      -\tstate.plugged = 0;\n      -\tif (state.f)\n  -:  ----------- >  3:  26ce5b8fdda object-file: pass filename to fsync_or_die\n  2:  3ed1dcd9b9b !  4:  52638326790 core.fsyncmethod: batched disk flushes for loose-objects\n     @@ Commit message\n          One major source of the cost of fsync is the implied flush of the\n          hardware writeback cache within the disk drive. This commit introduces\n          a new `core.fsyncMethod=batch` option that batches up hardware flushes.\n     -    It hooks into the bulk-checkin plugging and unplugging functionality,\n     -    takes advantage of tmp-objdir, and uses the writeout-only support code.\n     +    It hooks into the bulk-checkin odb-transaction functionality, takes\n     +    advantage of tmp-objdir, and uses the writeout-only support code.\n      \n          When the new mode is enabled, we do the following for each new object:\n     -    1. Create the object in a tmp-objdir.\n     -    2. Issue a pagecache writeback request and wait for it to complete.\n     +    1a. Create the object in a tmp-objdir.\n     +    2a. Issue a pagecache writeback request and wait for it to complete.\n      \n          At the end of the entire transaction when unplugging bulk checkin:\n     -    1. Issue an fsync against a dummy file to flush the hardware writeback\n     -       cache, which should by now have seen the tmp-objdir writes.\n     -    2. Rename all of the tmp-objdir files to their final names.\n     -    3. When updating the index and/or refs, we assume that Git will issue\n     +    1b. Issue an fsync against a dummy file to flush the log and hardware\n     +       writeback cache, which should by now have seen the tmp-objdir writes.\n     +    2b. Rename all of the tmp-objdir files to their final names.\n     +    3b. When updating the index and/or refs, we assume that Git will issue\n             another fsync internal to that operation. This is not the default\n             today, but the user now has the option of syncing the index and there\n             is a separate patch series to implement syncing of refs.\n     @@ Commit message\n          operations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\n          we would expect the fsync to trigger a journal writeout so that this\n          sequence is enough to ensure that the user's data is durable by the time\n     -    the git command returns.\n     +    the git command returns. This sequence also ensures that no object files\n     +    appear in the main object store unless they are fsync-durable.\n      \n     -    Batch mode is only enabled if core.fsyncObjectFiles is false or unset.\n     +    Batch mode is only enabled if core.fsync includes loose-objects. If\n     +    the legacy core.fsyncObjectFiles setting is enabled, but core.fsync does\n     +    not include loose-objects, we will use file-by-file fsyncing.\n     +\n     +    In step (1a) of the sequence, the tmp-objdir is created lazily to avoid\n     +    work if no loose objects are ever added to the ODB. We use a tmp-objdir\n     +    to maintain the invariant that no loose-objects are visible in the main\n     +    ODB unless they are properly fsync-durable. This is important since\n     +    future ODB operations that try to create an object with specific\n     +    contents will silently drop the new data if an object with the target\n     +    hash exists without checking that the loose-object contents match the\n     +    hash. Only a full git-fsck would restore the ODB to a functional state\n     +    where dataloss doesn't occur.\n     +\n     +    In step (1b) of the sequence, we issue a fsync against a dummy file\n     +    created specifically for the purpose. This method has a little higher\n     +    cost than using one of the input object files, but makes adding new\n     +    callers of this mechanism easier, since we don't need to figure out\n     +    which object file is \"last\" or risk sharing violations by caching the fd\n     +    of the last object file.\n      \n          _Performance numbers_:\n      \n     @@ Documentation/config/core.txt: core.fsyncMethod::\n      +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n      +  updates in the disk writeback cache and then does a single full fsync of\n      +  a dummy file to trigger the disk cache flush at the end of the operation.\n     +++\n      +  Currently `batch` mode only applies to loose-object files. Other repository\n      +  data is made durable as if `fsync` was specified. This mode is expected to\n      +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n     @@ bulk-checkin.c\n       #include \"object-store.h\"\n       \n       static int bulk_checkin_plugged;\n     -+static int needs_batch_fsync;\n     -+\n     -+static struct tmp_objdir *bulk_fsync_objdir;\n       \n     ++static struct tmp_objdir *bulk_fsync_objdir;\n     ++\n       static struct bulk_checkin_state {\n       \tchar *pack_tmp_name;\n     + \tstruct hashfile *f;\n      @@ bulk-checkin.c: clear_exit:\n       \treprepare_packed_git(the_repository);\n       }\n     @@ bulk-checkin.c: clear_exit:\n      + */\n      +static void do_batch_fsync(void)\n      +{\n     ++\tstruct strbuf temp_path = STRBUF_INIT;\n     ++\tstruct tempfile *temp;\n     ++\n     ++\tif (!bulk_fsync_objdir)\n     ++\t\treturn;\n     ++\n      +\t/*\n      +\t * Issue a full hardware flush against a temporary file to ensure\n     -+\t * that all objects are durable before any renames occur.  The code in\n     ++\t * that all objects are durable before any renames occur. The code in\n      +\t * fsync_loose_object_bulk_checkin has already issued a writeout\n      +\t * request, but it has not flushed any writeback cache in the storage\n     -+\t * hardware.\n     ++\t * hardware or any filesystem logs. This fsync call acts as a barrier\n     ++\t * to ensure that the data in each new object file is durable before\n     ++\t * the final name is visible.\n      +\t */\n     ++\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n     ++\ttemp = xmks_tempfile(temp_path.buf);\n     ++\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n     ++\tdelete_tempfile(&temp);\n     ++\tstrbuf_release(&temp_path);\n      +\n     -+\tif (needs_batch_fsync) {\n     -+\t\tstruct strbuf temp_path = STRBUF_INIT;\n     -+\t\tstruct tempfile *temp;\n     -+\n     -+\t\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n     -+\t\ttemp = xmks_tempfile(temp_path.buf);\n     -+\t\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n     -+\t\tdelete_tempfile(&temp);\n     -+\t\tstrbuf_release(&temp_path);\n     -+\t\tneeds_batch_fsync = 0;\n     -+\t}\n     -+\n     -+\tif (bulk_fsync_objdir) {\n     -+\t\ttmp_objdir_migrate(bulk_fsync_objdir);\n     -+\t\tbulk_fsync_objdir = NULL;\n     -+\t}\n     ++\t/*\n     ++\t * Make the object files visible in the primary ODB after their data is\n     ++\t * fully durable.\n     ++\t */\n     ++\ttmp_objdir_migrate(bulk_fsync_objdir);\n     ++\tbulk_fsync_objdir = NULL;\n      +}\n      +\n       static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n       \treturn 0;\n       }\n       \n     -+void fsync_loose_object_bulk_checkin(int fd)\n     ++void prepare_loose_object_bulk_checkin(void)\n     ++{\n     ++\t/*\n     ++\t * We lazily create the temporary object directory\n     ++\t * the first time an object might be added, since\n     ++\t * callers may not know whether any objects will be\n     ++\t * added at the time they call begin_odb_transaction.\n     ++\t */\n     ++\tif (!bulk_checkin_plugged || bulk_fsync_objdir)\n     ++\t\treturn;\n     ++\n     ++\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n     ++\tif (bulk_fsync_objdir)\n     ++\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n     ++}\n     ++\n     ++void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n      +{\n      +\t/*\n      +\t * If we have a plugged bulk checkin, we issue a call that\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +\t * command. Later on we will issue a single hardware flush\n      +\t * before as part of do_batch_fsync.\n      +\t */\n     -+\tif (bulk_checkin_plugged &&\n     -+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) >= 0) {\n     -+\t\tassert(bulk_fsync_objdir);\n     -+\t\tif (!needs_batch_fsync)\n     -+\t\t\tneeds_batch_fsync = 1;\n     -+\t} else {\n     -+\t\tfsync_or_die(fd, \"loose object file\");\n     ++\tif (!bulk_fsync_objdir ||\n     ++\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n     ++\t\tfsync_or_die(fd, filename);\n      +\t}\n      +}\n      +\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags)\n     -@@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n     - void plug_bulk_checkin(void)\n     - {\n     - \tassert(!bulk_checkin_plugged);\n     -+\n     -+\t/*\n     -+\t * A temporary object directory is used to hold the files\n     -+\t * while they are not fsynced.\n     -+\t */\n     -+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT)) {\n     -+\t\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n     -+\t\tif (!bulk_fsync_objdir)\n     -+\t\t\tdie(_(\"Could not create temporary object directory for core.fsyncobjectfiles=batch\"));\n     -+\n     -+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n     -+\t}\n     -+\n     - \tbulk_checkin_plugged = 1;\n     - }\n     - \n     -@@ bulk-checkin.c: void unplug_bulk_checkin(void)\n     +@@ bulk-checkin.c: void end_odb_transaction(void)\n       \tbulk_checkin_plugged = 0;\n       \tif (bulk_checkin_state.f)\n       \t\tfinish_bulk_checkin(&bulk_checkin_state);\n     @@ bulk-checkin.h\n       \n       #include \"cache.h\"\n       \n     -+void fsync_loose_object_bulk_checkin(int fd);\n     ++void prepare_loose_object_bulk_checkin(void);\n     ++void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n      +\n       int index_bulk_checkin(struct object_id *oid,\n       \t\t       int fd, size_t size, enum object_type type,\n     @@ cache.h: extern int use_fsync;\n       \tFSYNC_METHOD_FSYNC,\n      -\tFSYNC_METHOD_WRITEOUT_ONLY\n      +\tFSYNC_METHOD_WRITEOUT_ONLY,\n     -+\tFSYNC_METHOD_BATCH\n     ++\tFSYNC_METHOD_BATCH,\n       };\n       \n       extern enum fsync_method fsync_method;\n     @@ config.c: static int git_default_core_config(const char *var, const char *value,\n       \n      \n       ## object-file.c ##\n     -@@ object-file.c: static void close_loose_object(int fd)\n     +@@ object-file.c: static void close_loose_object(int fd, const char *filename)\n     + \tif (the_repository->objects->odb->will_destroy)\n     + \t\tgoto out;\n       \n     - \tif (fsync_object_files > 0)\n     - \t\tfsync_or_die(fd, \"loose object file\");\n     -+\telse if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n     -+\t\tfsync_loose_object_bulk_checkin(fd);\n     +-\tif (fsync_object_files > 0)\n     ++\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n     ++\t\tfsync_loose_object_bulk_checkin(fd, filename);\n     ++\telse if (fsync_object_files > 0)\n     + \t\tfsync_or_die(fd, filename);\n       \telse\n       \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n     - \t\t\t\t       \"loose object file\");\n     +@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n     + \tstatic struct strbuf tmp_file = STRBUF_INIT;\n     + \tstatic struct strbuf filename = STRBUF_INIT;\n     + \n     ++\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n     ++\t\tprepare_loose_object_bulk_checkin();\n     ++\n     + \tloose_object_path(the_repository, &filename, oid);\n     + \n     + \tfd = create_tmpfile(&tmp_file, filename.buf);\n  3:  54797dbc520 !  5:  913ce1b3df9 update-index: use the bulk-checkin infrastructure\n     @@ Commit message\n          The update-index functionality is used internally by 'git stash push' to\n          setup the internal stashed commit.\n      \n     -    This change enables bulk-checkin for update-index infrastructure to\n     +    This change enables odb-transactions for update-index infrastructure to\n          speed up adding new objects to the object database by leveraging the\n          batch fsync functionality.\n      \n          There is some risk with this change, since under batch fsync, the object\n     -    files will be in a tmp-objdir until update-index is complete.  This\n     -    usage is unlikely, since any tool invoking update-index and expecting to\n     -    see objects would have to synchronize with the update-index process\n     -    after passing it a file path.\n     +    files will be in a tmp-objdir until update-index is complete, so callers\n     +    using the --stdin option will not see them until update-index is done.\n     +    This risk is mitigated by unplugging the batch when reporting verbose\n     +    output, which is the only way a --stdin caller might synchronize with\n     +    the addition of an object.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ builtin/update-index.c\n       #include \"config.h\"\n       #include \"lockfile.h\"\n       #include \"quote.h\"\n     -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     +@@ builtin/update-index.c: static int allow_replace;\n     + static int info_only;\n     + static int force_remove;\n     + static int verbose;\n     ++static int odb_transaction_active;\n     + static int mark_valid_only;\n     + static int mark_skip_worktree_only;\n     + static int mark_fsmonitor_only;\n     +@@ builtin/update-index.c: enum uc_mode {\n     + \tUC_FORCE\n     + };\n       \n     - \tthe_index.updated_skipworktree = 1;\n     ++static void end_odb_transaction_if_active(void)\n     ++{\n     ++\tif (!odb_transaction_active)\n     ++\t\treturn;\n     ++\n     ++\tend_odb_transaction();\n     ++\todb_transaction_active = 0;\n     ++}\n     ++\n     + __attribute__((format (printf, 1, 2)))\n     + static void report(const char *fmt, ...)\n     + {\n     +@@ builtin/update-index.c: static void report(const char *fmt, ...)\n     + \tif (!verbose)\n     + \t\treturn;\n       \n     -+\t/* we might be adding many objects to the object database */\n     -+\tplug_bulk_checkin();\n     ++\t/*\n     ++\t * It is possible, though unlikely, that a caller\n     ++\t * could use the verbose output to synchronize with\n     ++\t * addition of objects to the object database, so\n     ++\t * unplug bulk checkin to make sure that future objects\n     ++\t * are immediately visible.\n     ++\t */\n     ++\n     ++\tend_odb_transaction_if_active();\n      +\n     - \t/*\n     - \t * Custom copy of parse_options() because we want to handle\n     - \t * filename arguments as they come.\n     + \tva_start(vp, fmt);\n     + \tvprintf(fmt, vp);\n     + \tputchar('\\n');\n     +@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     + \t */\n     + \tparse_options_start(&ctx, argc, argv, prefix,\n     + \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n     ++\n     ++\t/*\n     ++\t * Allow the object layer to optimize adding multiple objects in\n     ++\t * a batch.\n     ++\t */\n     ++\tbegin_odb_transaction();\n     ++\todb_transaction_active = 1;\n     + \twhile (ctx.argc) {\n     + \t\tif (parseopt_state != PARSE_OPT_DONE)\n     + \t\t\tparseopt_state = parse_options_step(&ctx, options,\n      @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n       \t\tstrbuf_release(&buf);\n       \t}\n       \n     -+\t/* by now we must have added all of the new objects */\n     -+\tunplug_bulk_checkin();\n     ++\t/*\n     ++\t * By now we have added all of the new objects\n     ++\t */\n     ++\tend_odb_transaction_if_active();\n     ++\n       \tif (split_index > 0) {\n       \t\tif (git_config_get_split_index() == 0)\n       \t\t\twarning(_(\"core.splitIndex is set to false; \"\n  4:  6662e2dae0f !  6:  84fd144ef18 unpack-objects: use the bulk-checkin infrastructure\n     @@ Commit message\n          to turn the transfered data into object database entries when there are\n          fewer objects than the 'unpacklimit' setting.\n      \n     -    By enabling bulk-checkin when unpacking objects, we can take advantage\n     +    By enabling an odb-transaction when unpacking objects, we can take advantage\n          of batched fsyncs.\n      \n     +    Here are some performance numbers to justify batch mode for\n     +    unpack-objects, collected on a WSL2 Ubuntu VM.\n     +\n     +    Fsync Mode | Time for 90 objects (ms)\n     +    -------------------------------------\n     +           Off | 170\n     +      On,fsync | 760\n     +      On,batch | 230\n     +\n     +    Note that the default unpackLimit is 100 objects, so there's a 3x\n     +    benefit in the worst case. The non-batch mode fsync scales linearly\n     +    with the number of objects, so there are significant benefits even with\n     +    smaller numbers of objects.\n     +\n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void unpack_all(void)\n       \tif (!quiet)\n       \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n       \tCALLOC_ARRAY(obj_list, nr_objects);\n     -+\tplug_bulk_checkin();\n     ++\tbegin_odb_transaction();\n       \tfor (i = 0; i < nr_objects; i++) {\n       \t\tunpack_one(i);\n       \t\tdisplay_progress(progress, i + 1);\n       \t}\n     -+\tunplug_bulk_checkin();\n     ++\tend_odb_transaction();\n       \tstop_progress(&progress);\n       \n       \tif (delta_list)\n  5:  03bf591742a !  7:  447263e8ef1 core.fsync: use batch mode and sync loose objects by default on Windows\n     @@ Commit message\n          in upstream Git so that we can get broad coverage of the new code\n          upstream.\n      \n     -    We don't actually do fsyncs in the test suite, since GIT_TEST_FSYNC is\n     -    set to 0. However, we do exercise all of the surrounding batch mode code\n     -    since GIT_TEST_FSYNC merely makes the maybe_fsync wrapper always appear\n     -    to succeed.\n     +    We don't actually do fsyncs in the most of the test suite, since\n     +    GIT_TEST_FSYNC is set to 0. However, we do exercise all of the\n     +    surrounding batch mode code since GIT_TEST_FSYNC merely makes the\n     +    maybe_fsync wrapper always appear to succeed.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n  -:  ----------- >  8:  8f1b01c9ca0 test-lib-functions: add parsing helpers for ls-files and ls-tree\n  6:  1937746df47 !  9:  b5f371e97fe core.fsyncmethod: tests for batch mode\n     @@ Commit message\n          In this change we introduce a new test helper lib-unique-files.sh. The\n          goal of this library is to create a tree of files that have different\n          oids from any other files that may have been created in the current test\n     -    repo. This helps us avoid missing validation of an object being added due\n     -    to it already being in the repo.\n     -\n     -    We aren't actually issuing any fsyncs in these tests, since\n     -    GIT_TEST_FSYNC is 0, but we still exercise all of the tmp_objdir logic\n     -    in bulk-checkin.\n     +    repo. This helps us avoid missing validation of an object being added\n     +    due to it already being in the repo.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ t/lib-unique-files.sh (new)\n      @@\n      +# Helper to create files with unique contents\n      +\n     -+\n     -+# Create multiple files with unique contents. Takes the number of\n     -+# directories, the number of files in each directory, and the base\n     ++# Create multiple files with unique contents within this test run. Takes the\n     ++# number of directories, the number of files in each directory, and the base\n      +# directory.\n      +#\n      +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n     -+#\t\t\t\t\t each in my_dir, all with unique\n     -+#\t\t\t\t\t contents.\n     ++#\t\t\t\t\t each in my_dir, all with contents\n     ++#\t\t\t\t\t different from previous invocations\n     ++#\t\t\t\t\t of this command in this run.\n      +\n     -+test_create_unique_files() {\n     ++test_create_unique_files () {\n      +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n      +\n     -+\tlocal dirs=$1\n     -+\tlocal files=$2\n     -+\tlocal basedir=$3\n     -+\tlocal counter=0\n     -+\ttest_tick\n     -+\tlocal basedata=$test_tick\n     -+\n     -+\n     -+\trm -rf $basedir\n     -+\n     ++\tlocal dirs=\"$1\" &&\n     ++\tlocal files=\"$2\" &&\n     ++\tlocal basedir=\"$3\" &&\n     ++\tlocal counter=0 &&\n     ++\ttest_tick &&\n     ++\tlocal basedata=$basedir$test_tick &&\n     ++\trm -rf \"$basedir\" &&\n      +\tfor i in $(test_seq $dirs)\n      +\tdo\n     -+\t\tlocal dir=$basedir/dir$i\n     -+\n     -+\t\tmkdir -p \"$dir\"\n     ++\t\tlocal dir=$basedir/dir$i &&\n     ++\t\tmkdir -p \"$dir\" &&\n      +\t\tfor j in $(test_seq $files)\n      +\t\tdo\n     -+\t\t\tcounter=$((counter + 1))\n     -+\t\t\techo \"$basedata.$counter\"  >\"$dir/file$j.txt\"\n     ++\t\t\tcounter=$((counter + 1)) &&\n     ++\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n      +\t\tdone\n      +\tdone\n      +}\n     @@ t/t3700-add.sh: test_expect_success \\\n      +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n      +\n      +test_expect_success 'git add: core.fsyncmethod=batch' \"\n     -+\ttest_create_unique_files 2 4 fsync-files &&\n     -+\tgit $BATCH_CONFIGURATION add -- ./fsync-files/ &&\n     -+\trm -f fsynced_files &&\n     -+\tgit ls-files --stage fsync-files/ > fsynced_files &&\n     -+\ttest_line_count = 8 fsynced_files &&\n     -+\tawk -- '{print \\$2}' fsynced_files | xargs -n1 git cat-file -e\n     ++\ttest_create_unique_files 2 4 files_base_dir1 &&\n     ++\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION add -- ./files_base_dir1/ &&\n     ++\tgit ls-files --stage files_base_dir1/ |\n     ++\ttest_parse_ls_files_stage_oids >added_files_oids &&\n     ++\n     ++\t# We created 2 subdirs with 4 files each (8 files total) above\n     ++\ttest_line_count = 8 added_files_oids &&\n     ++\tgit cat-file --batch-check='%(objectname)' <added_files_oids >added_files_actual &&\n     ++\ttest_cmp added_files_oids added_files_actual\n      +\"\n      +\n      +test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n     -+\ttest_create_unique_files 2 4 fsync-files2 &&\n     -+\tfind fsync-files2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n     -+\trm -f fsynced_files2 &&\n     -+\tgit ls-files --stage fsync-files2/ > fsynced_files2 &&\n     -+\ttest_line_count = 8 fsynced_files2 &&\n     -+\tawk -- '{print \\$2}' fsynced_files2 | xargs -n1 git cat-file -e\n     ++\ttest_create_unique_files 2 4 files_base_dir2 &&\n     ++\tfind files_base_dir2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n     ++\tgit ls-files --stage files_base_dir2 |\n     ++\ttest_parse_ls_files_stage_oids >added_files2_oids &&\n     ++\n     ++\t# We created 2 subdirs with 4 files each (8 files total) above\n     ++\ttest_line_count = 8 added_files2_oids &&\n     ++\tgit cat-file --batch-check='%(objectname)' <added_files2_oids >added_files2_actual &&\n     ++\ttest_cmp added_files2_oids added_files2_actual\n      +\"\n      +\n       test_expect_success \\\n     @@ t/t3903-stash.sh: test_expect_success 'stash handles skip-worktree entries nicel\n      +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n      +\n      +test_expect_success 'stash with core.fsyncmethod=batch' \"\n     -+\ttest_create_unique_files 2 4 fsync-files &&\n     -+\tgit $BATCH_CONFIGURATION stash push -u -- ./fsync-files/ &&\n     -+\trm -f fsynced_files &&\n     ++\ttest_create_unique_files 2 4 files_base_dir &&\n     ++\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION stash push -u -- ./files_base_dir/ &&\n      +\n      +\t# The files were untracked, so use the third parent,\n      +\t# which contains the untracked files\n     -+\tgit ls-tree -r stash^3 -- ./fsync-files/ > fsynced_files &&\n     -+\ttest_line_count = 8 fsynced_files &&\n     -+\tawk -- '{print \\$3}' fsynced_files | xargs -n1 git cat-file -e\n     ++\tgit ls-tree -r stash^3 -- ./files_base_dir/ |\n     ++\ttest_parse_ls_tree_oids >stashed_files_oids &&\n     ++\n     ++\t# We created 2 dirs with 4 files each (8 files total) above\n     ++\ttest_line_count = 8 stashed_files_oids &&\n     ++\tgit cat-file --batch-check='%(objectname)' <stashed_files_oids >stashed_files_actual &&\n     ++\ttest_cmp stashed_files_oids stashed_files_actual\n      +\"\n      +\n      +\n     @@ t/t3903-stash.sh: test_expect_success 'stash handles skip-worktree entries nicel\n      \n       ## t/t5300-pack-object.sh ##\n      @@ t/t5300-pack-object.sh: test_expect_success 'pack-objects with bogus arguments' '\n     + '\n       \n       check_unpack () {\n     ++\tlocal packname=\"$1\" &&\n     ++\tlocal object_list=\"$2\" &&\n     ++\tlocal git_config=\"$3\" &&\n       \ttest_when_finished \"rm -rf git2\" &&\n      -\tgit init --bare git2 &&\n      -\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n     @@ t/t5300-pack-object.sh: test_expect_success 'pack-objects with bogus arguments'\n      -\t\t\treturn 1\n      -\t\t}\n      -\tdone\n     -+\tgit $2 init --bare git2 &&\n     ++\tgit $git_config init --bare git2 &&\n      +\t(\n     -+\t\tgit $2 -C git2 unpack-objects -n <\"$1\".pack &&\n     -+\t\tgit $2 -C git2 unpack-objects <\"$1\".pack &&\n     -+\t\tgit $2 -C git2 cat-file --batch-check=\"%(objectname)\"\n     -+\t) <obj-list >current &&\n     -+\tcmp obj-list current\n     ++\t\tgit $git_config -C git2 unpack-objects -n <\"$packname\".pack &&\n     ++\t\tgit $git_config -C git2 unpack-objects <\"$packname\".pack &&\n     ++\t\tgit $git_config -C git2 cat-file --batch-check=\"%(objectname)\"\n     ++\t) <\"$object_list\" >current &&\n     ++\tcmp \"$object_list\" current\n       }\n       \n       test_expect_success 'unpack without delta' '\n     - \tcheck_unpack test-1-${packname_1}\n     - '\n     - \n     +-\tcheck_unpack test-1-${packname_1}\n     ++\tcheck_unpack test-1-${packname_1} obj-list\n     ++'\n     ++\n      +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n      +\n      +test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n     -+\tcheck_unpack test-1-${packname_1} \"$BATCH_CONFIGURATION\"\n     -+'\n     -+\n     ++\tcheck_unpack test-1-${packname_1} obj-list \"$BATCH_CONFIGURATION\"\n     + '\n     + \n       test_expect_success 'pack with REF_DELTA' '\n     - \tpackname_2=$(git pack-objects --progress test-2 <obj-list 2>stderr) &&\n     - \tcheck_deltas stderr -gt 0\n     -@@ t/t5300-pack-object.sh: test_expect_success 'unpack with REF_DELTA' '\n     - \tcheck_unpack test-2-${packname_2}\n     +@@ t/t5300-pack-object.sh: test_expect_success 'pack with REF_DELTA' '\n       '\n       \n     -+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n     -+       check_unpack test-2-${packname_2} \"$BATCH_CONFIGURATION\"\n     + test_expect_success 'unpack with REF_DELTA' '\n     +-\tcheck_unpack test-2-${packname_2}\n     ++\tcheck_unpack test-2-${packname_2} obj-list\n      +'\n      +\n     ++test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n     ++       check_unpack test-2-${packname_2} obj-list \"$BATCH_CONFIGURATION\"\n     + '\n     + \n       test_expect_success 'pack with OFS_DELTA' '\n     - \tpackname_3=$(git pack-objects --progress --delta-base-offset test-3 \\\n     - \t\t\t<obj-list 2>stderr) &&\n     -@@ t/t5300-pack-object.sh: test_expect_success 'unpack with OFS_DELTA' '\n     - \tcheck_unpack test-3-${packname_3}\n     +@@ t/t5300-pack-object.sh: test_expect_success 'pack with OFS_DELTA' '\n       '\n       \n     -+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n     -+       check_unpack test-3-${packname_3} \"$BATCH_CONFIGURATION\"\n     + test_expect_success 'unpack with OFS_DELTA' '\n     +-\tcheck_unpack test-3-${packname_3}\n     ++\tcheck_unpack test-3-${packname_3} obj-list\n      +'\n      +\n     ++test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n     ++       check_unpack test-3-${packname_3} obj-list \"$BATCH_CONFIGURATION\"\n     + '\n     + \n       test_expect_success 'compare delta flavors' '\n     - \tperl -e '\\''\n     - \t\tdefined($_ = -s $_) or die for @ARGV;\n  7:  624244078c7 ! 10:  b99b32a469c core.fsyncmethod: performance tests for add and stash\n     @@ Commit message\n      \n          Add basic performance tests for \"git add\" and \"git stash\" of a lot of\n          new objects with various fsync settings. This shows the benefit of batch\n     -    mode relative to an ordinary stash command.\n     +    mode relative to full fsync.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ t/perf/p3900-stash.sh (new)\n      +done\n      +\n      +test_done\n     -\n     - ## t/perf/perf-lib.sh ##\n     -@@ t/perf/perf-lib.sh: test_perf_create_repo_from () {\n     - \tmkdir -p \"$repo/.git\"\n     - \t(\n     - \t\tcd \"$source\" &&\n     --\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" 2>/dev/null ||\n     --\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\"; } &&\n     -+\t\t{ cp -Rl \"$objects_dir\" \"$repo/.git/\" ||\n     -+\t\t\tcp -R \"$objects_dir\" \"$repo/.git/\" 2>/dev/null;} &&\n     - \n     - \t\t# common_dir must come first here, since we want source_git to\n     - \t\t# take precedence and overwrite any overlapping files\n  -:  ----------- > 11:  6b832e89bc4 core.fsyncmethod: correctly camel-case warning message\n\n-- \ngitgitgadget\n"},{"id":"452091","messageId":"b2d9766a662894d177fa5b7f2a5de2a45351f47e.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 02/11] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:17Z","receivedAt":"2022-03-24T04:58:45Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit prepares for adding batch-fsync to the bulk-checkin\ninfrastructure.\n\nThe bulk-checkin infrastructure is currently used to batch up addition\nof large blobs to a packfile. When a blob is larger than\nbig_file_threshold, we unconditionally add it to a pack. If bulk\ncheckins are 'plugged', we allow multiple large blobs to be added to a\nsingle pack until we reach the packfile size limit; otherwise, we simply\nmake a new packfile for each large blob. The 'unplug' call tells us when\nthe series of blob additions is done so that we can finish the packfiles\nand make their objects available to subsequent operations.\n\nStated another way, bulk-checkin allows callers to define a transaction\nthat adds multiple objects to the object database, where the object\ndatabase can optimize its internal operations within the transaction\nboundary.\n\nBatched fsync will fit into bulk-checkin by taking advantage of the\nplug/unplug functionality to determine the appropriate time to fsync\nand make newly-added objects available in the primary object database.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex a16ae3c629d..ffe142841b2 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -278,21 +278,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void begin_odb_transaction(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void end_odb_transaction(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"452092","messageId":"26ce5b8fddaa507e09d3201c4c6492e4d1184d53.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 03/11] object-file: pass filename to fsync_or_die","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:18Z","receivedAt":"2022-03-24T04:58:46Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf we die while trying to fsync a loose object file, pass the actual\nfilename we're trying to sync. This is likely to be more helpful for a\nuser trying to diagnose the cause of the failure than the former\n'loose object file' string. It also sidesteps any concerns about\ntranslating the die message differently for loose objects versus\nsomething else that has a real path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n object-file.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b254bc50d70..5ffbf3d4fd4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1888,16 +1888,16 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n }\n \n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n+static void close_loose_object(int fd, const char *filename)\n {\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n \tif (fsync_object_files > 0)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n-\t\t\t\t       \"loose object file\");\n+\t\t\t\t       filename);\n \n out:\n \tif (close(fd) != 0)\n@@ -2011,7 +2011,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n+\tclose_loose_object(fd, tmp_file.buf);\n \n \tif (mtime) {\n \t\tstruct utimbuf utb;\n-- \ngitgitgadget\n\n"},{"id":"452093","messageId":"52638326790aaf5113ff64c74023100beb90760d.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 04/11] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:19Z","receivedAt":"2022-03-24T04:58:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin odb-transaction functionality, takes\nadvantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1a. Create the object in a tmp-objdir.\n2a. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1b. Issue an fsync against a dummy file to flush the log and hardware\n   writeback cache, which should by now have seen the tmp-objdir writes.\n2b. Rename all of the tmp-objdir files to their final names.\n3b. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns. This sequence also ensures that no object files\nappear in the main object store unless they are fsync-durable.\n\nBatch mode is only enabled if core.fsync includes loose-objects. If\nthe legacy core.fsyncObjectFiles setting is enabled, but core.fsync does\nnot include loose-objects, we will use file-by-file fsyncing.\n\nIn step (1a) of the sequence, the tmp-objdir is created lazily to avoid\nwork if no loose objects are ever added to the ODB. We use a tmp-objdir\nto maintain the invariant that no loose-objects are visible in the main\nODB unless they are properly fsync-durable. This is important since\nfuture ODB operations that try to create an object with specific\ncontents will silently drop the new data if an object with the target\nhash exists without checking that the loose-object contents match the\nhash. Only a full git-fsck would restore the ODB to a functional state\nwhere dataloss doesn't occur.\n\nIn step (1b) of the sequence, we issue a fsync against a dummy file\ncreated specifically for the purpose. This method has a little higher\ncost than using one of the input object files, but makes adding new\ncallers of this mechanism easier, since we don't need to figure out\nwhich object file is \"last\" or risk sharing violations by caching the fd\nof the last object file.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  8 ++++\n bulk-checkin.c                | 71 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  3 ++\n cache.h                       |  8 +++-\n config.c                      |  2 +\n object-file.c                 |  7 +++-\n 6 files changed, 97 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 9da3e5d88f6..3c90ba0b395 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -596,6 +596,14 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then does a single full fsync of\n+  a dummy file to trigger the disk cache flush at the end of the operation.\n++\n+  Currently `batch` mode only applies to loose-object files. Other repository\n+  data is made durable as if `fsync` was specified. This mode is expected to\n+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+  and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex ffe142841b2..a5c40a08b8d 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int bulk_checkin_plugged;\n \n+static struct tmp_objdir *bulk_fsync_objdir;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -80,6 +85,40 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\tstruct strbuf temp_path = STRBUF_INIT;\n+\tstruct tempfile *temp;\n+\n+\tif (!bulk_fsync_objdir)\n+\t\treturn;\n+\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur. The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware or any filesystem logs. This fsync call acts as a barrier\n+\t * to ensure that the data in each new object file is durable before\n+\t * the final name is visible.\n+\t */\n+\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\ttemp = xmks_tempfile(temp_path.buf);\n+\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\tdelete_tempfile(&temp);\n+\tstrbuf_release(&temp_path);\n+\n+\t/*\n+\t * Make the object files visible in the primary ODB after their data is\n+\t * fully durable.\n+\t */\n+\ttmp_objdir_migrate(bulk_fsync_objdir);\n+\tbulk_fsync_objdir = NULL;\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +313,36 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void prepare_loose_object_bulk_checkin(void)\n+{\n+\t/*\n+\t * We lazily create the temporary object directory\n+\t * the first time an object might be added, since\n+\t * callers may not know whether any objects will be\n+\t * added at the time they call begin_odb_transaction.\n+\t */\n+\tif (!bulk_checkin_plugged || bulk_fsync_objdir)\n+\t\treturn;\n+\n+\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+}\n+\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n+{\n+\t/*\n+\t * If we have a plugged bulk checkin, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (!bulk_fsync_objdir ||\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n+\t\tfsync_or_die(fd, filename);\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -297,4 +366,6 @@ void end_odb_transaction(void)\n \tbulk_checkin_plugged = 0;\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 69a94422ac7..70edf745be8 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,9 @@\n \n #include \"cache.h\"\n \n+void prepare_loose_object_bulk_checkin(void);\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex ef7d34b7a09..a5bf15a5131 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH,\n };\n \n extern enum fsync_method fsync_method;\n@@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex 3c9b6b589ab..511f4584eeb 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd4..d2e0c13198f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1893,7 +1893,9 @@ static void close_loose_object(int fd, const char *filename)\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (fsync_object_files > 0)\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd, filename);\n+\telse if (fsync_object_files > 0)\n \t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n@@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n-- \ngitgitgadget\n\n"},{"id":"452094","messageId":"913ce1b3df9cf273f1572c256dffad1cacc192a6.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 05/11] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:20Z","receivedAt":"2022-03-24T04:58:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables odb-transactions for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete, so callers\nusing the --stdin option will not see them until update-index is done.\nThis risk is mitigated by unplugging the batch when reporting verbose\noutput, which is the only way a --stdin caller might synchronize with\nthe addition of an object.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 33 +++++++++++++++++++++++++++++++++\n 1 file changed, 33 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex aafe7eeac2a..ae7887cfe37 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -32,6 +33,7 @@ static int allow_replace;\n static int info_only;\n static int force_remove;\n static int verbose;\n+static int odb_transaction_active;\n static int mark_valid_only;\n static int mark_skip_worktree_only;\n static int mark_fsmonitor_only;\n@@ -49,6 +51,15 @@ enum uc_mode {\n \tUC_FORCE\n };\n \n+static void end_odb_transaction_if_active(void)\n+{\n+\tif (!odb_transaction_active)\n+\t\treturn;\n+\n+\tend_odb_transaction();\n+\todb_transaction_active = 0;\n+}\n+\n __attribute__((format (printf, 1, 2)))\n static void report(const char *fmt, ...)\n {\n@@ -57,6 +68,16 @@ static void report(const char *fmt, ...)\n \tif (!verbose)\n \t\treturn;\n \n+\t/*\n+\t * It is possible, though unlikely, that a caller\n+\t * could use the verbose output to synchronize with\n+\t * addition of objects to the object database, so\n+\t * unplug bulk checkin to make sure that future objects\n+\t * are immediately visible.\n+\t */\n+\n+\tend_odb_transaction_if_active();\n+\n \tva_start(vp, fmt);\n \tvprintf(fmt, vp);\n \tputchar('\\n');\n@@ -1116,6 +1137,13 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t */\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n+\n+\t/*\n+\t * Allow the object layer to optimize adding multiple objects in\n+\t * a batch.\n+\t */\n+\tbegin_odb_transaction();\n+\todb_transaction_active = 1;\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1190,6 +1218,11 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/*\n+\t * By now we have added all of the new objects\n+\t */\n+\tend_odb_transaction_if_active();\n+\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"452095","messageId":"84fd144ef1889aaf4f88040a60cb8b156522e1b4.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 06/11] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:21Z","receivedAt":"2022-03-24T04:58:53Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling an odb-transaction when unpacking objects, we can take advantage\nof batched fsyncs.\n\nHere are some performance numbers to justify batch mode for\nunpack-objects, collected on a WSL2 Ubuntu VM.\n\nFsync Mode | Time for 90 objects (ms)\n-------------------------------------\n       Off | 170\n  On,fsync | 760\n  On,batch | 230\n\nNote that the default unpackLimit is 100 objects, so there's a 3x\nbenefit in the worst case. The non-batch mode fsync scales linearly\nwith the number of objects, so there are significant benefits even with\nsmaller numbers of objects.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..56d05e2725d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tbegin_odb_transaction();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tend_odb_transaction();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"452096","messageId":"447263e8ef14021714976e489248cdcc48ec0303.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 07/11] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:22Z","receivedAt":"2022-03-24T04:59:04Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the most of the test suite, since\nGIT_TEST_FSYNC is set to 0. However, we do exercise all of the\nsurrounding batch mode code since GIT_TEST_FSYNC merely makes the\nmaybe_fsync wrapper always appear to succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex a5bf15a5131..7f6cbb254b4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 511f4584eeb..e9cac5f4707 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 0892e209a2f..fffe42ce7c1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \ngitgitgadget\n\n"},{"id":"452097","messageId":"8f1b01c9ca0aff1412657d230f58fdcd7b3aeade.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 08/11] test-lib-functions: add parsing helpers for ls-files and ls-tree","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:23Z","receivedAt":"2022-03-24T04:59:16Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nSeveral tests use awk to parse OIDs from the output of 'git ls-files\n--stage' and 'git ls-tree'. Introduce helpers to centralize these uses\nof awk.\n\nUpdate t5317-pack-objects-filter-objects.sh to use the new ls-files\nhelper so that it has some usages to review. Other updates are left for\nthe future.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/t5317-pack-objects-filter-objects.sh | 91 +++++++++++++-------------\n t/test-lib-functions.sh                | 10 +++\n 2 files changed, 54 insertions(+), 47 deletions(-)\n\ndiff --git a/t/t5317-pack-objects-filter-objects.sh b/t/t5317-pack-objects-filter-objects.sh\nindex 33b740ce628..bb633c9b099 100755\n--- a/t/t5317-pack-objects-filter-objects.sh\n+++ b/t/t5317-pack-objects-filter-objects.sh\n@@ -10,9 +10,6 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n # Test blob:none filter.\n \n test_expect_success 'setup r1' '\n-\techo \"{print \\$1}\" >print_1.awk &&\n-\techo \"{print \\$2}\" >print_2.awk &&\n-\n \tgit init r1 &&\n \tfor n in 1 2 3 4 5\n \tdo\n@@ -22,10 +19,13 @@ test_expect_success 'setup r1' '\n \tdone\n '\n \n+parse_verify_pack_blob_oid () {\n+\tawk '{print $1}' -\n+}\n+\n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r1 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -35,7 +35,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -54,12 +54,12 @@ test_expect_success 'verify blob:none packfile has no blobs' '\n test_expect_success 'verify normal and blob:none packfiles have same commits/trees' '\n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r1 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -123,8 +123,8 @@ test_expect_success 'setup r2' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -134,7 +134,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -161,8 +161,8 @@ test_expect_success 'verify blob:limit=1000' '\n '\n \n test_expect_success 'verify blob:limit=1001' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1001 >filter.pack <<-EOF &&\n@@ -172,15 +172,15 @@ test_expect_success 'verify blob:limit=1001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=10001' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=10001 >filter.pack <<-EOF &&\n@@ -190,15 +190,15 @@ test_expect_success 'verify blob:limit=10001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1k' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1k >filter.pack <<-EOF &&\n@@ -208,15 +208,15 @@ test_expect_success 'verify blob:limit=1k' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify explicitly specifying oversized blob in input' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \techo HEAD >objects &&\n@@ -226,15 +226,15 @@ test_expect_success 'verify explicitly specifying oversized blob in input' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1m' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1m >filter.pack <<-EOF &&\n@@ -244,7 +244,7 @@ test_expect_success 'verify blob:limit=1m' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -253,12 +253,12 @@ test_expect_success 'verify blob:limit=1m' '\n test_expect_success 'verify normal and blob:limit packfiles have same commits/trees' '\n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -289,9 +289,8 @@ test_expect_success 'setup r3' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r3 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -301,7 +300,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r3 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -342,9 +341,8 @@ test_expect_success 'setup r4' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -354,19 +352,19 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r4 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=OID' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 ls-files -s pattern >staged &&\n-\toid=$(awk -f print_2.awk staged) &&\n+\toid=$(test_parse_ls_files_stage_oids <staged) &&\n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=$oid >filter.pack <<-EOF &&\n \tHEAD\n \tEOF\n@@ -374,15 +372,15 @@ test_expect_success 'verify sparse:oid=OID' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=oid-ish' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=main:pattern >filter.pack <<-EOF &&\n@@ -392,7 +390,7 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -402,9 +400,8 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n # This models previously omitted objects that we did not receive.\n \n test_expect_success 'setup r1 - delete loose blobs' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tfor id in `cat expected | sed \"s|..|&/|\"`\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex a027f0c409e..e6011409e2f 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1782,6 +1782,16 @@ test_oid_to_path () {\n \techo \"${1%$basename}/$basename\"\n }\n \n+# Parse oids from git ls-files --staged output\n+test_parse_ls_files_stage_oids () {\n+\tawk '{print $2}' -\n+}\n+\n+# Parse oids from git ls-tree output\n+test_parse_ls_tree_oids () {\n+\tawk '{print $3}' -\n+}\n+\n # Choose a port number based on the test script's number and store it in\n # the given variable name, unless that variable already contains a number.\n test_set_port () {\n-- \ngitgitgadget\n\n"},{"id":"452098","messageId":"b99b32a469c5bee6e1a4d4e2f374d69aff8db63e.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 10/11] core.fsyncmethod: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:25Z","receivedAt":"2022-03-24T04:59:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings. This shows the benefit of batch\nmode relative to full fsync.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh   | 59 ++++++++++++++++++++++++++++++++++++++++\n t/perf/p3900-stash.sh | 62 +++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 121 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n create mode 100755 t/perf/p3900-stash.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..2ea78c9449d\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,59 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of add\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\ttest_perf \"add $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG add files\n+\t\"\n+done\n+\n+test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..3526f06cef4\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,62 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of stash\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+# We need to create the files each time we run the perf test, but\n+# we do not want to measure the cost of creating the files, so run\n+# the test once.\n+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n+then\n+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n+\tGIT_PERF_REPEAT_COUNT=1\n+fi\n+\n+for m in false true batch\n+do\n+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n+\t\tgit reset --hard &&\n+\t\t# create files across directories\n+\t\ttest_create_unique_files $dir_count $files_per_dir files\n+\t'\n+\n+\tcase $m in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\t# We only stash files in the 'files' subdirectory since\n+\t# the perf test infrastructure creates files in the\n+\t# current working directory that need to be preserved\n+\ttest_perf \"stash $total_files files (object_fsyncing=$m)\" \"\n+\t\tgit $FSYNC_CONFIG stash push -u -- files\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"452099","messageId":"6b832e89bc47f00af4bd186b71577409914970f7.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 11/11] core.fsyncmethod: correctly camel-case warning message","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:26Z","receivedAt":"2022-03-24T04:59:21Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe warning for an unrecognized fsyncMethod was not\ncamel-cased.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n config.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/config.c b/config.c\nindex e9cac5f4707..ae819dee20b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1697,7 +1697,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tif (fsync_object_files < 0)\n-\t\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n+\t\t\twarning(_(\"core.fsyncObjectFiles is deprecated; use core.fsync instead\"));\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n-- \ngitgitgadget\n"},{"id":"452100","messageId":"b5f371e97fee69d87da1dccd3180de0691c15834.1648097906.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v3 09/11] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-24T04:58:24Z","receivedAt":"2022-03-24T04:59:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added\ndue to it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 32 ++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n t/t3903-stash.sh       | 20 ++++++++++++++++++++\n t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n 4 files changed, 107 insertions(+), 14 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..74efca91dd7\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,32 @@\n+# Helper to create files with unique contents\n+\n+# Create multiple files with unique contents within this test run. Takes the\n+# number of directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with contents\n+#\t\t\t\t\t different from previous invocations\n+#\t\t\t\t\t of this command in this run.\n+\n+test_create_unique_files () {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=\"$1\" &&\n+\tlocal files=\"$2\" &&\n+\tlocal basedir=\"$3\" &&\n+\tlocal counter=0 &&\n+\ttest_tick &&\n+\tlocal basedata=$basedir$test_tick &&\n+\trm -rf \"$basedir\" &&\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i &&\n+\t\tmkdir -p \"$dir\" &&\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1)) &&\n+\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..8979c8a5f03 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,32 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir1 &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION add -- ./files_base_dir1/ &&\n+\tgit ls-files --stage files_base_dir1/ |\n+\ttest_parse_ls_files_stage_oids >added_files_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files_oids >added_files_actual &&\n+\ttest_cmp added_files_oids added_files_actual\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir2 &&\n+\tfind files_base_dir2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\tgit ls-files --stage files_base_dir2 |\n+\ttest_parse_ls_files_stage_oids >added_files2_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files2_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files2_oids >added_files2_actual &&\n+\ttest_cmp added_files2_oids added_files2_actual\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 4abbc8fccae..20e94881964 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1410,6 +1411,25 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION stash push -u -- ./files_base_dir/ &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./files_base_dir/ |\n+\ttest_parse_ls_tree_oids >stashed_files_oids &&\n+\n+\t# We created 2 dirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 stashed_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <stashed_files_oids >stashed_files_actual &&\n+\ttest_cmp stashed_files_oids stashed_files_actual\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..f8a0f309e2d 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -161,22 +161,27 @@ test_expect_success 'pack-objects with bogus arguments' '\n '\n \n check_unpack () {\n+\tlocal packname=\"$1\" &&\n+\tlocal object_list=\"$2\" &&\n+\tlocal git_config=\"$3\" &&\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $git_config init --bare git2 &&\n+\t(\n+\t\tgit $git_config -C git2 unpack-objects -n <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 unpack-objects <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <\"$object_list\" >current &&\n+\tcmp \"$object_list\" current\n }\n \n test_expect_success 'unpack without delta' '\n-\tcheck_unpack test-1-${packname_1}\n+\tcheck_unpack test-1-${packname_1} obj-list\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with REF_DELTA' '\n@@ -185,7 +190,11 @@ test_expect_success 'pack with REF_DELTA' '\n '\n \n test_expect_success 'unpack with REF_DELTA' '\n-\tcheck_unpack test-2-${packname_2}\n+\tcheck_unpack test-2-${packname_2} obj-list\n+'\n+\n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with OFS_DELTA' '\n@@ -195,7 +204,11 @@ test_expect_success 'pack with OFS_DELTA' '\n '\n \n test_expect_success 'unpack with OFS_DELTA' '\n-\tcheck_unpack test-3-${packname_3}\n+\tcheck_unpack test-3-${packname_3} obj-list\n+'\n+\n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'compare delta flavors' '\n-- \ngitgitgadget\n\n"},{"id":"452126","messageId":"220324.86y20zmi84.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"53261f0099d53524155464fe79d10f9605fe93aa.1648097906.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 01/11] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-24T16:10:24Z","receivedAt":"2022-03-24T16:24:38Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Mar 24 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Make it clearer in the naming and documentation of the plug_bulk_checkin\n> and unplug_bulk_checkin APIs that they can be thought of as\n> a \"transaction\" to optimize operations on the object database.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/add.c  |  4 ++--\n>  bulk-checkin.c |  4 ++--\n>  bulk-checkin.h | 14 ++++++++++++--\n>  3 files changed, 16 insertions(+), 6 deletions(-)\n>\n> diff --git a/builtin/add.c b/builtin/add.c\n> index 3ffb86a4338..9bf37ceae8e 100644\n> --- a/builtin/add.c\n> +++ b/builtin/add.c\n> @@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n>  \t\tstring_list_clear(&only_match_skip_worktree, 0);\n>  \t}\n>  \n> -\tplug_bulk_checkin();\n> +\tbegin_odb_transaction();\n>  \n>  \tif (add_renormalize)\n>  \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n> @@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n>  \n>  \tif (chmod_arg && pathspec.nr)\n>  \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n> -\tunplug_bulk_checkin();\n> +\tend_odb_transaction();\n\nAside from anything else we've (dis)agreed on, I found this part really\nodd when hacking on my RFC-on-top, i.e. originally I (wrongly) thought\nthe plug_bulk_checkin() was something that originated with this series\nwhich adds the \"bulk\" mode.\n\nBut no, on second inspection it's a thing Junio added a long time ago so\nthat in this case we \"stream to N pack\" where we'd otherwise add N loose\nobjects.\n\nWhich, and I think Junio brought this up in an earlier round, but I\ndidn't fully understand that at the time makes this whole thing quite\nodd to me.\n\nSo first, shouldn't we add this begin_odb_transaction() as a new thing?\nI.e. surely wanting to do that object target redirection within a given\nbegin/end \"scope\" should be orthagonal to how fsync() happens within\nthat \"scope\", though in this case that happens to correspond.\n\nAnd secondly, per the commit message and comment when it was added in\n(568508e7657 (bulk-checkin: replace fast-import based implementation,\n2011-10-28)) is it something we need *for that purpose* with the series\nto unpack-objects without malloc()ing the size of the blob[1].\n\nAnd, if so and orthagonal to that: If we know how to either stream N\nobjects to a PACK (as fast-import does), *and* we now (or SOON) know how\nto stream loose objects without using size(blob) amounts of memory,\ndoesn't the \"optimize fsync()\" rather want to make use of the\nstream-to-pack approach?\n\nI.e. have you tried for the caseses where we create say 1k objects for\n\"git stash\" tried to stream those to a pack? How does that compare (both\nwith/without the fsync changes).\n\nI.e. I do worry (also per [2]) that while the whole \"bulk fsync\" is neat\n(and I think can use it in either case, to defer object syncs until the\n\"index\" or \"ref\" sync, as my RFC does) I worry that we're adding a bunch\nof configuration and complexity for something that:\n\n 1. Ultimately isn't all that important, as already for part of it we\n    can mostly configure it away. I.e. \"git-unpack-objects\" v.s. writing\n    a pack, cf. transfer.unpackLimit)\n 2. We don't have #1 for \"add\" and \"update-index\", but if we stream to\n    packs there is there any remaining benefit in practice?\n\n1. https://lore.kernel.org/git/cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com/\n2. https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n"},{"id":"452190","messageId":"220324.86tubnmgwk.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"b5f371e97fee69d87da1dccd3180de0691c15834.1648097906.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 09/11] core.fsyncmethod: tests for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-24T16:29:28Z","receivedAt":"2022-03-24T16:54:40Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Mar 24 2022, Neeraj Singh via GitGitGadget wrote:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Add test cases to exercise batch mode for:\n>  * 'git add'\n>  * 'git stash'\n>  * 'git update-index'\n>  * 'git unpack-objects'\n>\n> These tests ensure that the added data winds up in the object database.\n>\n> In this change we introduce a new test helper lib-unique-files.sh. The\n> goal of this library is to create a tree of files that have different\n> oids from any other files that may have been created in the current test\n> repo. This helps us avoid missing validation of an object being added\n> due to it already being in the repo.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  t/lib-unique-files.sh  | 32 ++++++++++++++++++++++++++++++++\n>  t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n>  t/t3903-stash.sh       | 20 ++++++++++++++++++++\n>  t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n>  4 files changed, 107 insertions(+), 14 deletions(-)\n>  create mode 100644 t/lib-unique-files.sh\n>\n> diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> new file mode 100644\n> index 00000000000..74efca91dd7\n> --- /dev/null\n> +++ b/t/lib-unique-files.sh\n> @@ -0,0 +1,32 @@\n> +# Helper to create files with unique contents\n> +\n> +# Create multiple files with unique contents within this test run. Takes the\n> +# number of directories, the number of files in each directory, and the base\n> +# directory.\n> +#\n> +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> +#\t\t\t\t\t each in my_dir, all with contents\n> +#\t\t\t\t\t different from previous invocations\n> +#\t\t\t\t\t of this command in this run.\n> +\n> +test_create_unique_files () {\n> +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n> +\n> +\tlocal dirs=\"$1\" &&\n> +\tlocal files=\"$2\" &&\n> +\tlocal basedir=\"$3\" &&\n> +\tlocal counter=0 &&\n> +\ttest_tick &&\n> +\tlocal basedata=$basedir$test_tick &&\n> +\trm -rf \"$basedir\" &&\n> +\tfor i in $(test_seq $dirs)\n> +\tdo\n> +\t\tlocal dir=$basedir/dir$i &&\n> +\t\tmkdir -p \"$dir\" &&\n> +\t\tfor j in $(test_seq $files)\n> +\t\tdo\n> +\t\t\tcounter=$((counter + 1)) &&\n> +\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n> +\t\tdone\n> +\tdone\n> +}\n\nHaving written my own perf tests for this series, I still don't get why\nthis is needed, at all.\n\ntl;dr: the below: I think this whole workaround is because you missed\nthat \"test_when_finished\" exists, and how it excludes perf timings.\n\nI.e. I get that if we ran this N times we'd want to wipe our repo\nbetween tests, as for e.g. \"git add\" you want it to actually add the\nobjects.\n\nIt's what I do with the \"hyperfine\" command in\nhttps://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\nwith the \"-p\" option.\n\nI.e. hyperfine has a way to say \"this is setup, but don't measure the\ntime\", which is 1/2 of what you're working around here and in 10/11.\n\nBut as 10/11 shows you're limited to one run with t/perf because you\nwant to not include those \"setup\" numbers, and \"test_perf\" has no easy\nway to avoid that (but more on that later).\n\nWhich b.t.w. I'm really skeptical of as an approach here in any case\n(even if we couldn't exclude it from the numbers).\n\nI.e. yes what \"hyperfine\" does would be preferrable, but in exchange for\navoiding that you're comparing samples of 1 runs.\n\nSurely we're better off with N run (even if noisy). Given enough of them\nthe difference will shake out, and our estimated +/- will narrow..\n\nBut aside from that, why isn't this just:\n\t\n\tfor cfg in true false blah\n\tdone\n\t\ttest_expect_success \"setup for $cfg\" '\n\t\t\tgit init repo-$cfg &&\n\t\t\tfor f in $(test_seq 1 100)\n\t\t\tdo\n\t\t\t\t>repo-$cfg/$f\n\t\t\tdone\n\t\t'\n\t\n\t\ttest_perf \"perf test for $cfg\" '\n\t\t\tgit -C repo-$cfg\n\t\t'\n\tdone\n\nWhich surely is going to be more accurate in the context of our limited\nt/perf environment because creating unique files is not sufficient at\nall to ensure that your tests don't interfere with each other.\n\nThat's because in the first iteration we'll create N objects in\n.git/objects/aa/* or whatever, which will *still be there* for your\nsecond test, which will impact performance.\n\nWhereas if you just make N repos you don't need unique files, and you\nwon't be introducing that as a conflating variable.\n\nBut anyway, reading perf-lib.sh again I haven't tested, but this whole\nworkaround seems truly unnecessary. I.e. in test_run_perf_ we do:\n\t\n\ttest_run_perf_ () {\n\t        test_cleanup=:\n\t        test_export_=\"test_cleanup\"\n\t        export test_cleanup test_export_\n\t        \"$GTIME\" -f \"%E %U %S\" -o test_time.$i \"$TEST_SHELL_PATH\" -c ' \n                \t[... code we run and time ...]\n\t\t'\n                [... later ...]\n                test_eval_ \"$test_cleanup\"\n\t}\n\nSo can't you just avoid this whole glorious workaround for the low low\ncost of approximately one shellscript string assignment? :)\n\nI.e. if you do:\n\n\tsetup_clean () {\n\t\trm -rf repo\n\t}\n\n\tsetup_first () {\n\t\tgit init repo &&\n\t\t[make a bunch of files or whatever in repo]\n\t}\n\n\tsetup_next () {\n\t\ttest_when_finished \"setup_clean\" &&\n\t\tsetup_first\n\t}\n\n\ttest_expect_success 'setup initial stuff' '\n\t\tsetup_first\n\t'\n\n\ttest_perf 'my perf test' '\n\t\ttest_when_finished \"setup_next\" &&\n\t\t[your perf test here]\n\t'\n\n\ttest_expect_success 'cleanup' '\n\t\t# Not really needed, but just for completeness, we are\n                # about to nuke the trash dir anyway...\n\t\tsetup_clean\n\t'\n\nI haven't tested (and need to run), but i'm pretty sure that does\nexactly what you want without these workarounds, i.e. you'll get\n\"trampoline setup\" without that setup being included in the perf\nnumbers.\n\nIs it pretty? No, but it's a lot less complex than this unique file\nbusiness & workarounds, and will give you just the numbers you want, and\nmost importantly you car run it N times now for better samples.\n\nI.e. \"what you want\" sans a *tiny* bit of noise that we use to just call\na function to do:\n\n    test_cleanup=setup_next\n\nWhich we'll then eval *after* we measure your numbers to setup the next\ntest.\n"},{"id":"452201","messageId":"xmqqo81vi6u7.fsf@gitster.g","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 00/11] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-24T17:44:00Z","receivedAt":"2022-03-24T17:44:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> V3 changes:\n>\n>  * Rebrand plug/unplug-bulk-checkin to \"begin_odb_transaction\" and\n>    \"end_odb_transaction\"\n\nOK.  Makes me wonder (not \"object\", more appropriate verb than\n\"object\" being \"be curious\") how well \"odb-transaction\" meshes with\nmechanisms to ensure that the bits hit the disk platter to protect\nthings outside the odb that you may or may not be covering in this\nseries (e.g. the index file, the refs, the working tree files).\n\n> This work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n\n\"at .\"???\n\n"},{"id":"452202","messageId":"CANQDOdfx91dbgvLQkEqjjZjx1=qOGvq_fZJQMQLww9SgqE157g@mail.gmail.com","threadId":"57568","inReplyTo":"220324.86y20zmi84.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v3 01/11] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T17:52:27Z","receivedAt":"2022-03-24T17:52:42Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 24, 2022 at 9:24 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Thu, Mar 24 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Make it clearer in the naming and documentation of the plug_bulk_checkin\n> > and unplug_bulk_checkin APIs that they can be thought of as\n> > a \"transaction\" to optimize operations on the object database.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  builtin/add.c  |  4 ++--\n> >  bulk-checkin.c |  4 ++--\n> >  bulk-checkin.h | 14 ++++++++++++--\n> >  3 files changed, 16 insertions(+), 6 deletions(-)\n> >\n> > diff --git a/builtin/add.c b/builtin/add.c\n> > index 3ffb86a4338..9bf37ceae8e 100644\n> > --- a/builtin/add.c\n> > +++ b/builtin/add.c\n> > @@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n> >               string_list_clear(&only_match_skip_worktree, 0);\n> >       }\n> >\n> > -     plug_bulk_checkin();\n> > +     begin_odb_transaction();\n> >\n> >       if (add_renormalize)\n> >               exit_status |= renormalize_tracked_files(&pathspec, flags);\n> > @@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n> >\n> >       if (chmod_arg && pathspec.nr)\n> >               exit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n> > -     unplug_bulk_checkin();\n> > +     end_odb_transaction();\n>\n> Aside from anything else we've (dis)agreed on, I found this part really\n> odd when hacking on my RFC-on-top, i.e. originally I (wrongly) thought\n> the plug_bulk_checkin() was something that originated with this series\n> which adds the \"bulk\" mode.\n>\n> But no, on second inspection it's a thing Junio added a long time ago so\n> that in this case we \"stream to N pack\" where we'd otherwise add N loose\n> objects.\n>\n> Which, and I think Junio brought this up in an earlier round, but I\n> didn't fully understand that at the time makes this whole thing quite\n> odd to me.\n>\n> So first, shouldn't we add this begin_odb_transaction() as a new thing?\n> I.e. surely wanting to do that object target redirection within a given\n> begin/end \"scope\" should be orthagonal to how fsync() happens within\n> that \"scope\", though in this case that happens to correspond.\n>\n> And secondly, per the commit message and comment when it was added in\n> (568508e7657 (bulk-checkin: replace fast-import based implementation,\n> 2011-10-28)) is it something we need *for that purpose* with the series\n> to unpack-objects without malloc()ing the size of the blob[1].\n>\n\nThe original change seems to be about optimizing addition of\nsuccessive large blobs to the ODB when we know we have a large batch.\nIt's a batch-mode optimization for the ODB, similar to my patch\nseries, just targeting large blobs rather than small blobs/trees.  It\nalso has the same property that the added data is \"invisible\" until\nthe transaction ends.\n\n> And, if so and orthagonal to that: If we know how to either stream N\n> objects to a PACK (as fast-import does), *and* we now (or SOON) know how\n> to stream loose objects without using size(blob) amounts of memory,\n> doesn't the \"optimize fsync()\" rather want to make use of the\n> stream-to-pack approach?\n>\n> I.e. have you tried for the caseses where we create say 1k objects for\n> \"git stash\" tried to stream those to a pack? How does that compare (both\n> with/without the fsync changes).\n>\n> I.e. I do worry (also per [2]) that while the whole \"bulk fsync\" is neat\n> (and I think can use it in either case, to defer object syncs until the\n> \"index\" or \"ref\" sync, as my RFC does) I worry that we're adding a bunch\n> of configuration and complexity for something that:\n>\n>  1. Ultimately isn't all that important, as already for part of it we\n>     can mostly configure it away. I.e. \"git-unpack-objects\" v.s. writing\n>     a pack, cf. transfer.unpackLimit)\n>  2. We don't have #1 for \"add\" and \"update-index\", but if we stream to\n>     packs there is there any remaining benefit in practice?\n>\n> 1. https://lore.kernel.org/git/cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com/\n> 2. https://lore.kernel.org/git/220323.86fsn8ohg8.gmgdl@evledraar.gmail.com/\n\nStream to pack is a good idea.  But I think we'd want a way to append\nto the most recent pack so that we don't explode the number of packs,\nwhich seems to impose a linear cost on ODB operations, at least to\nload up the indexes.  I think this is orthogonal and we can always\nchange the meaning of batch mode to use a pack mechanism when such a\nmechanism is ready.\n\nThanks,\nNeeraj\n"},{"id":"452205","messageId":"xmqqfsn7i594.fsf@gitster.g","threadId":"57568","inReplyTo":"913ce1b3df9cf273f1572c256dffad1cacc192a6.1648097906.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 05/11] update-index: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-24T18:18:15Z","receivedAt":"2022-03-24T18:18:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +static void end_odb_transaction_if_active(void)\n> +{\n> +\tif (!odb_transaction_active)\n> +\t\treturn;\n> +\n> +\tend_odb_transaction();\n> +\todb_transaction_active = 0;\n> +}\n\n>  __attribute__((format (printf, 1, 2)))\n>  static void report(const char *fmt, ...)\n>  {\n> @@ -57,6 +68,16 @@ static void report(const char *fmt, ...)\n>  \tif (!verbose)\n>  \t\treturn;\n>  \n> +\t/*\n> +\t * It is possible, though unlikely, that a caller\n> +\t * could use the verbose output to synchronize with\n> +\t * addition of objects to the object database, so\n> +\t * unplug bulk checkin to make sure that future objects\n> +\t * are immediately visible.\n> +\t */\n> +\n> +\tend_odb_transaction_if_active();\n> +\n>  \tva_start(vp, fmt);\n>  \tvprintf(fmt, vp);\n>  \tputchar('\\n');\n> @@ -1116,6 +1137,13 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t */\n>  \tparse_options_start(&ctx, argc, argv, prefix,\n>  \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n> +\n> +\t/*\n> +\t * Allow the object layer to optimize adding multiple objects in\n> +\t * a batch.\n> +\t */\n> +\tbegin_odb_transaction();\n> +\todb_transaction_active = 1;\n\nThis looks strange.  Shouldn't begin/end pair be responsible for\nknowing if there is a transaction active already?  For that matter,\ndidn't the original unplug in plug/unplug pair automatically turned\ninto no-op when it is already unplugged?\n\nIOW, I am not sure end_if_active() should exist in the first place.\nShouldn't end_transaction() do that instead?\n\n"},{"id":"452206","messageId":"CANQDOdfM_XyRa3e8Uo72yRdn6cmQxVSahb8J+7b2-cXogOg9pg@mail.gmail.com","threadId":"57568","inReplyTo":"220324.86tubnmgwk.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v3 09/11] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T18:23:11Z","receivedAt":"2022-03-24T18:23:28Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 24, 2022 at 9:53 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Thu, Mar 24 2022, Neeraj Singh via GitGitGadget wrote:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Add test cases to exercise batch mode for:\n> >  * 'git add'\n> >  * 'git stash'\n> >  * 'git update-index'\n> >  * 'git unpack-objects'\n> >\n> > These tests ensure that the added data winds up in the object database.\n> >\n> > In this change we introduce a new test helper lib-unique-files.sh. The\n> > goal of this library is to create a tree of files that have different\n> > oids from any other files that may have been created in the current test\n> > repo. This helps us avoid missing validation of an object being added\n> > due to it already being in the repo.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  t/lib-unique-files.sh  | 32 ++++++++++++++++++++++++++++++++\n> >  t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n> >  t/t3903-stash.sh       | 20 ++++++++++++++++++++\n> >  t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n> >  4 files changed, 107 insertions(+), 14 deletions(-)\n> >  create mode 100644 t/lib-unique-files.sh\n> >\n> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> > new file mode 100644\n> > index 00000000000..74efca91dd7\n> > --- /dev/null\n> > +++ b/t/lib-unique-files.sh\n> > @@ -0,0 +1,32 @@\n> > +# Helper to create files with unique contents\n> > +\n> > +# Create multiple files with unique contents within this test run. Takes the\n> > +# number of directories, the number of files in each directory, and the base\n> > +# directory.\n> > +#\n> > +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> > +#                                     each in my_dir, all with contents\n> > +#                                     different from previous invocations\n> > +#                                     of this command in this run.\n> > +\n> > +test_create_unique_files () {\n> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n> > +\n> > +     local dirs=\"$1\" &&\n> > +     local files=\"$2\" &&\n> > +     local basedir=\"$3\" &&\n> > +     local counter=0 &&\n> > +     test_tick &&\n> > +     local basedata=$basedir$test_tick &&\n> > +     rm -rf \"$basedir\" &&\n> > +     for i in $(test_seq $dirs)\n> > +     do\n> > +             local dir=$basedir/dir$i &&\n> > +             mkdir -p \"$dir\" &&\n> > +             for j in $(test_seq $files)\n> > +             do\n> > +                     counter=$((counter + 1)) &&\n> > +                     echo \"$basedata.$counter\">\"$dir/file$j.txt\"\n> > +             done\n> > +     done\n> > +}\n>\n> Having written my own perf tests for this series, I still don't get why\n> this is needed, at all.\n>\n> tl;dr: the below: I think this whole workaround is because you missed\n> that \"test_when_finished\" exists, and how it excludes perf timings.\n>\n\nI actually noticed test_when_finished, but I didn't think of your\n\"setup the next round on cleanup of last\" idea.  I was debating at the\ntime adding a \"test_perf_setup\" helper to do the setup work during\neach perf iteration.  How about I do that and just create a new repo\nin each test_perf_setup step?\n\n> I.e. I get that if we ran this N times we'd want to wipe our repo\n> between tests, as for e.g. \"git add\" you want it to actually add the\n> objects.\n>\n> It's what I do with the \"hyperfine\" command in\n> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n> with the \"-p\" option.\n>\n> I.e. hyperfine has a way to say \"this is setup, but don't measure the\n> time\", which is 1/2 of what you're working around here and in 10/11.\n>\n> But as 10/11 shows you're limited to one run with t/perf because you\n> want to not include those \"setup\" numbers, and \"test_perf\" has no easy\n> way to avoid that (but more on that later).\n>\n> Which b.t.w. I'm really skeptical of as an approach here in any case\n> (even if we couldn't exclude it from the numbers).\n>\n> I.e. yes what \"hyperfine\" does would be preferrable, but in exchange for\n> avoiding that you're comparing samples of 1 runs.\n>\n> Surely we're better off with N run (even if noisy). Given enough of them\n> the difference will shake out, and our estimated +/- will narrow..\n>\n> But aside from that, why isn't this just:\n>\n>         for cfg in true false blah\n>         done\n>                 test_expect_success \"setup for $cfg\" '\n>                         git init repo-$cfg &&\n>                         for f in $(test_seq 1 100)\n>                         do\n>                                 >repo-$cfg/$f\n>                         done\n>                 '\n>\n>                 test_perf \"perf test for $cfg\" '\n>                         git -C repo-$cfg\n>                 '\n>         done\n>\n> Which surely is going to be more accurate in the context of our limited\n> t/perf environment because creating unique files is not sufficient at\n> all to ensure that your tests don't interfere with each other.\n>\n> That's because in the first iteration we'll create N objects in\n> .git/objects/aa/* or whatever, which will *still be there* for your\n> second test, which will impact performance.\n>\n> Whereas if you just make N repos you don't need unique files, and you\n> won't be introducing that as a conflating variable.\n>\n> But anyway, reading perf-lib.sh again I haven't tested, but this whole\n> workaround seems truly unnecessary. I.e. in test_run_perf_ we do:\n>\n>         test_run_perf_ () {\n>                 test_cleanup=:\n>                 test_export_=\"test_cleanup\"\n>                 export test_cleanup test_export_\n>                 \"$GTIME\" -f \"%E %U %S\" -o test_time.$i \"$TEST_SHELL_PATH\" -c '\n>                         [... code we run and time ...]\n>                 '\n>                 [... later ...]\n>                 test_eval_ \"$test_cleanup\"\n>         }\n>\n> So can't you just avoid this whole glorious workaround for the low low\n> cost of approximately one shellscript string assignment? :)\n>\n> I.e. if you do:\n>\n>         setup_clean () {\n>                 rm -rf repo\n>         }\n>\n>         setup_first () {\n>                 git init repo &&\n>                 [make a bunch of files or whatever in repo]\n>         }\n>\n>         setup_next () {\n>                 test_when_finished \"setup_clean\" &&\n>                 setup_first\n>         }\n>\n>         test_expect_success 'setup initial stuff' '\n>                 setup_first\n>         '\n>\n>         test_perf 'my perf test' '\n>                 test_when_finished \"setup_next\" &&\n>                 [your perf test here]\n>         '\n>\n>         test_expect_success 'cleanup' '\n>                 # Not really needed, but just for completeness, we are\n>                 # about to nuke the trash dir anyway...\n>                 setup_clean\n>         '\n>\n> I haven't tested (and need to run), but i'm pretty sure that does\n> exactly what you want without these workarounds, i.e. you'll get\n> \"trampoline setup\" without that setup being included in the perf\n> numbers.\n>\n> Is it pretty? No, but it's a lot less complex than this unique file\n> business & workarounds, and will give you just the numbers you want, and\n> most importantly you car run it N times now for better samples.\n>\n> I.e. \"what you want\" sans a *tiny* bit of noise that we use to just call\n> a function to do:\n>\n>     test_cleanup=setup_next\n>\n> Which we'll then eval *after* we measure your numbers to setup the next\n> test.\n\nHow about I add a new test_perf_setup mechanism to make your idea work\nin a straightforward way?\n\nI still want the test_create_unique_files thing as a way to make\nmultiple files easily.  And for the non-perf tests it makes sense to\nhave differing contents within a test run.\n\nThanks,\nNeeraj\n"},{"id":"452220","messageId":"CANQDOdfcgaRUkygG9or7VS02uzo9vz1Ve=nkFUs_aYu3vVaCnA@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqo81vi6u7.fsf@gitster.g","subject":"Re: [PATCH v3 00/11] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T19:21:59Z","receivedAt":"2022-03-24T19:22:17Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 24, 2022 at 10:44 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj K. Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > V3 changes:\n> >\n> >  * Rebrand plug/unplug-bulk-checkin to \"begin_odb_transaction\" and\n> >    \"end_odb_transaction\"\n>\n> OK.  Makes me wonder (not \"object\", more appropriate verb than\n> \"object\" being \"be curious\") how well \"odb-transaction\" meshes with\n> mechanisms to ensure that the bits hit the disk platter to protect\n> things outside the odb that you may or may not be covering in this\n> series (e.g. the index file, the refs, the working tree files).\n>\n\nAs of this series, the odb-transaction will ensure that loose-objects\n(and trivially packs as well, since they're currently eagerly-synced)\nare efficiently made durable by the time the transaction ends.  Other\nparts of the repo (index, refs, etc) need to be updated and synced\nafter the odb transaction ends.  Patrick's original ref syncing work\nat [1] also contained a batch mode delimited by the existing ref\ntransactions.\n\nI think larger transactions would be interesting to have, but I'd\nargue that the current patch series is a worthwhile building block for\nthat world.  It solves the real-world multiplicative pain of adding N\nobjects to the ODB, where each one needs to be fsynced.  Patrick's\nbatch mode solves the real-world multiplicative pain of updating R\nrefs during a big mirror push.  Even talking just about the ODB, we\nstill have O(TreeSize) fsyncs for the updated trees and a few extra\nfsyncs for commits.  We can add odb transactions around those things\ntoo, which should be easy enough going forward.\n\n[1] https://lore.kernel.org/git/d9aa96913b1730f1d0c238d7d52e27c20bc55390.1636544377.git.ps@pks.im/\n\n> > This work is based on 'seen' at . It's dependent on ns/core-fsyncmethod.\n>\n> \"at .\"???\n>\n\nSorry, to make GGG/Github happy, I had to rebase onto b9f5d0358d2,\nwhich was the last non-merge commit that's present in next. Then I\ncould target next with the PR and get the right set of patches.\nBasing on fd008b1442 didn't work because GGG doesn't want to see a\nmerge commit in the set of changes not in the target branch.\n"},{"id":"452221","messageId":"CANQDOdcuBRvWx7iMYBvLYEEb6A_=SURLAGumk026ZyDODpfAsQ@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqfsn7i594.fsf@gitster.g","subject":"Re: [PATCH v3 05/11] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T20:25:38Z","receivedAt":"2022-03-24T20:25:55Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 24, 2022 at 11:18 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +static void end_odb_transaction_if_active(void)\n> > +{\n> > +     if (!odb_transaction_active)\n> > +             return;\n> > +\n> > +     end_odb_transaction();\n> > +     odb_transaction_active = 0;\n> > +}\n>\n> >  __attribute__((format (printf, 1, 2)))\n> >  static void report(const char *fmt, ...)\n> >  {\n> > @@ -57,6 +68,16 @@ static void report(const char *fmt, ...)\n> >       if (!verbose)\n> >               return;\n> >\n> > +     /*\n> > +      * It is possible, though unlikely, that a caller\n> > +      * could use the verbose output to synchronize with\n> > +      * addition of objects to the object database, so\n> > +      * unplug bulk checkin to make sure that future objects\n> > +      * are immediately visible.\n> > +      */\n> > +\n> > +     end_odb_transaction_if_active();\n> > +\n> >       va_start(vp, fmt);\n> >       vprintf(fmt, vp);\n> >       putchar('\\n');\n> > @@ -1116,6 +1137,13 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >        */\n> >       parse_options_start(&ctx, argc, argv, prefix,\n> >                           options, PARSE_OPT_STOP_AT_NON_OPTION);\n> > +\n> > +     /*\n> > +      * Allow the object layer to optimize adding multiple objects in\n> > +      * a batch.\n> > +      */\n> > +     begin_odb_transaction();\n> > +     odb_transaction_active = 1;\n>\n> This looks strange.  Shouldn't begin/end pair be responsible for\n> knowing if there is a transaction active already?  For that matter,\n> didn't the original unplug in plug/unplug pair automatically turned\n> into no-op when it is already unplugged?\n>\n> IOW, I am not sure end_if_active() should exist in the first place.\n> Shouldn't end_transaction() do that instead?\n>\n\nToday there's an \"assert(bulk_checkin_plugged)\" in\nend_odb_transaction. In principle we could just drop the assert and\nallow a transaction to be ended multiple times.  But maybe in the long\nrun for composability we'd like to have nested callers to begin/end\ntransaction (e.g. we could have a nested transaction around writing\nthe cache tree to the ODB to minimize fsyncs there).  In that world,\nhaving a subsystem not maintain a balanced pairing could be a problem.\nAn alternative API here could be to have an \"flush_odb_transaction\"\ncall to make the objects visible at this point.  Lastly, I could take\nyour original suggested approach of adding a new flag to update-index.\nI preferred the unplug-on-verbose approach since it would\nautomatically optimize most callers to update-index that might exist\nin the wild, without users having to change anything.\n\nThanks,\nNeeraj\n"},{"id":"452232","messageId":"xmqq8rszf31t.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdcuBRvWx7iMYBvLYEEb6A_=SURLAGumk026ZyDODpfAsQ@mail.gmail.com","subject":"Re: [PATCH v3 05/11] update-index: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-24T21:34:06Z","receivedAt":"2022-03-24T21:34:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> IOW, I am not sure end_if_active() should exist in the first place.\n>> Shouldn't end_transaction() do that instead?\n>>\n>\n> Today there's an \"assert(bulk_checkin_plugged)\" in\n> end_odb_transaction. In principle we could just drop the assert and\n> allow a transaction to be ended multiple times.  But maybe in the long\n> run for composability we'd like to have nested callers to begin/end\n> transaction (e.g. we could have a nested transaction around writing\n> the cache tree to the ODB to minimize fsyncs there).\n\nI am not convinced that \"transaction\" is a good mental model for\nthis mechanism to begin with, in the sense that the sense that it is\nnot a bug or failure of the implementation if two or more operations\nin the same <begin,end> bracket did not happen (or not happen)\natomically, or if 'begin' and 'end' were not properly nested.  With\nthe design getting more complex with things like tentative object\nstore that needs to be explicitly migrated after the outermost level\nof end-transaction, we may end up _requiring_ that sufficient number\nof 'end' must come once we issued 'begin', which I am not sure is\nnecessarily a good thing.\n\nIn any case, we aspire/envision to have a nested plug/unplug, I\nthink it is a good thing.  A helper for one subsystem may have its\nlarge batch of operations inside plug/unplug pair, another help may\ndo the same, and the caller of these two helpers may want to say\n\n\tplug\n\t\tcall helper A\n\t\t\tA does plug\n\t\t\tA does many things\n\t\t\tA does unplug\n\t\tcall helper B\n\t\t\tB does plug\n\t\t\tB does many things\n\t\t\tB does unplug\n\tunplug\n\nto \"cancel\" the unplug helper A and B has.\n\n> In that world,\n> having a subsystem not maintain a balanced pairing could be a problem.\n\nAnd in such a world, you never want to have end-if-active to\nimplement what you are doing here, as you may end up being not\nproperly nested:\n\n\tbegin\n\t\tbegin\n\t\t\tdo many things\n\t\t\tif some condtion\n\t\t\t\tend_if_active\n\t\t\tdo more things\n\t\tend\n\tend\n\n> An alternative API here could be to have an \"flush_odb_transaction\"\n> call to make the objects visible at this point.\n\nYes, what you want is a forced-flush instead, I think.\n\nSo I suspect you'd want these three primitives, perhaps?\n\n * begin increments the nesting level\n   - if outermost, you may have to do real \"setup\" things\n   - otherwise, you may not have anything other than just counting\n     the nesting level\n\n * flush implements unplug, fsync, etc. and does so immediately,\n   even when plugged.\n\n * end decrements the nesting level\n   - if outermost, you'd do \"flush\".\n   - otherwise, you may only count the nesting level and do nothing else,\n     but doing \"flush\" when you realize that you've queued too many\n     is not a bug or a crime.\n\n"},{"id":"452239","messageId":"CANQDOdeywLXa1UHFq7G51de1+6hDG2r7OWMvQ9tsJAZtajMdZg@mail.gmail.com","threadId":"57568","inReplyTo":"xmqq8rszf31t.fsf@gitster.g","subject":"Re: [PATCH v3 05/11] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-24T22:21:50Z","receivedAt":"2022-03-24T22:22:07Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 24, 2022 at 2:34 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> >> IOW, I am not sure end_if_active() should exist in the first place.\n> >> Shouldn't end_transaction() do that instead?\n> >>\n> >\n> > Today there's an \"assert(bulk_checkin_plugged)\" in\n> > end_odb_transaction. In principle we could just drop the assert and\n> > allow a transaction to be ended multiple times.  But maybe in the long\n> > run for composability we'd like to have nested callers to begin/end\n> > transaction (e.g. we could have a nested transaction around writing\n> > the cache tree to the ODB to minimize fsyncs there).\n>\n> I am not convinced that \"transaction\" is a good mental model for\n> this mechanism to begin with, in the sense that the sense that it is\n> not a bug or failure of the implementation if two or more operations\n> in the same <begin,end> bracket did not happen (or not happen)\n> atomically, or if 'begin' and 'end' were not properly nested.  With\n> the design getting more complex with things like tentative object\n> store that needs to be explicitly migrated after the outermost level\n> of end-transaction, we may end up _requiring_ that sufficient number\n> of 'end' must come once we issued 'begin', which I am not sure is\n> necessarily a good thing.\n\nI don't love the tentative object store that keeps things invisble,\nbut that was the safest way to maintain the invariant that no\nloose-object name appears in the ODB without durable contents.  I\nthink we want the \"durability/ordering boundary\" part of database\ntransactions without necessarily needing full abort/commit semantics.\nAs you say, we don't need full atomicity, but we do need ordering to\nensure that blobs are durable before trees pointing them, and so on up\nthe merkle chain.  The begin/end pairs help us defer the syncs\nrequired for ordering to the end rather than pessimistically assuming\nthat every object write is the end.\n\n> In any case, we aspire/envision to have a nested plug/unplug, I\n> think it is a good thing.  A helper for one subsystem may have its\n> large batch of operations inside plug/unplug pair, another help may\n> do the same, and the caller of these two helpers may want to say\n>\n>         plug\n>                 call helper A\n>                         A does plug\n>                         A does many things\n>                         A does unplug\n>                 call helper B\n>                         B does plug\n>                         B does many things\n>                         B does unplug\n>         unplug\n>\n> to \"cancel\" the unplug helper A and B has.\n>\n> > In that world,\n> > having a subsystem not maintain a balanced pairing could be a problem.\n>\n> And in such a world, you never want to have end-if-active to\n> implement what you are doing here, as you may end up being not\n> properly nested:\n>\n>         begin\n>                 begin\n>                         do many things\n>                         if some condtion\n>                                 end_if_active\n>                         do more things\n>                 end\n>         end\n>\n> > An alternative API here could be to have an \"flush_odb_transaction\"\n> > call to make the objects visible at this point.\n>\n> Yes, what you want is a forced-flush instead, I think.\n>\n> So I suspect you'd want these three primitives, perhaps?\n>\n>  * begin increments the nesting level\n>    - if outermost, you may have to do real \"setup\" things\n>    - otherwise, you may not have anything other than just counting\n>      the nesting level\n>\n>  * flush implements unplug, fsync, etc. and does so immediately,\n>    even when plugged.\n>\n>  * end decrements the nesting level\n>    - if outermost, you'd do \"flush\".\n>    - otherwise, you may only count the nesting level and do nothing else,\n>      but doing \"flush\" when you realize that you've queued too many\n>      is not a bug or a crime.\n>\n\nYes, I'll move in this direction. Thanks for the feedback.\n"},{"id":"452411","messageId":"220326.86o81sk9ao.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"CANQDOdfM_XyRa3e8Uo72yRdn6cmQxVSahb8J+7b2-cXogOg9pg@mail.gmail.com","subject":"Re: [PATCH v3 09/11] core.fsyncmethod: tests for batch mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-26T15:35:15Z","receivedAt":"2022-03-26T15:44:54Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Mar 24 2022, Neeraj Singh wrote:\n\n> On Thu, Mar 24, 2022 at 9:53 AM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>>\n>>\n>> On Thu, Mar 24 2022, Neeraj Singh via GitGitGadget wrote:\n>>\n>> > From: Neeraj Singh <neerajsi@microsoft.com>\n>> >\n>> > Add test cases to exercise batch mode for:\n>> >  * 'git add'\n>> >  * 'git stash'\n>> >  * 'git update-index'\n>> >  * 'git unpack-objects'\n>> >\n>> > These tests ensure that the added data winds up in the object database.\n>> >\n>> > In this change we introduce a new test helper lib-unique-files.sh. The\n>> > goal of this library is to create a tree of files that have different\n>> > oids from any other files that may have been created in the current test\n>> > repo. This helps us avoid missing validation of an object being added\n>> > due to it already being in the repo.\n>> >\n>> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n>> > ---\n>> >  t/lib-unique-files.sh  | 32 ++++++++++++++++++++++++++++++++\n>> >  t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n>> >  t/t3903-stash.sh       | 20 ++++++++++++++++++++\n>> >  t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n>> >  4 files changed, 107 insertions(+), 14 deletions(-)\n>> >  create mode 100644 t/lib-unique-files.sh\n>> >\n>> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n>> > new file mode 100644\n>> > index 00000000000..74efca91dd7\n>> > --- /dev/null\n>> > +++ b/t/lib-unique-files.sh\n>> > @@ -0,0 +1,32 @@\n>> > +# Helper to create files with unique contents\n>> > +\n>> > +# Create multiple files with unique contents within this test run. Takes the\n>> > +# number of directories, the number of files in each directory, and the base\n>> > +# directory.\n>> > +#\n>> > +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n>> > +#                                     each in my_dir, all with contents\n>> > +#                                     different from previous invocations\n>> > +#                                     of this command in this run.\n>> > +\n>> > +test_create_unique_files () {\n>> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n>> > +\n>> > +     local dirs=\"$1\" &&\n>> > +     local files=\"$2\" &&\n>> > +     local basedir=\"$3\" &&\n>> > +     local counter=0 &&\n>> > +     test_tick &&\n>> > +     local basedata=$basedir$test_tick &&\n>> > +     rm -rf \"$basedir\" &&\n>> > +     for i in $(test_seq $dirs)\n>> > +     do\n>> > +             local dir=$basedir/dir$i &&\n>> > +             mkdir -p \"$dir\" &&\n>> > +             for j in $(test_seq $files)\n>> > +             do\n>> > +                     counter=$((counter + 1)) &&\n>> > +                     echo \"$basedata.$counter\">\"$dir/file$j.txt\"\n>> > +             done\n>> > +     done\n>> > +}\n>>\n>> Having written my own perf tests for this series, I still don't get why\n>> this is needed, at all.\n>>\n>> tl;dr: the below: I think this whole workaround is because you missed\n>> that \"test_when_finished\" exists, and how it excludes perf timings.\n>>\n>\n> I actually noticed test_when_finished, but I didn't think of your\n> \"setup the next round on cleanup of last\" idea.  I was debating at the\n> time adding a \"test_perf_setup\" helper to do the setup work during\n> each perf iteration.  How about I do that and just create a new repo\n> in each test_perf_setup step?\n>\n>> I.e. I get that if we ran this N times we'd want to wipe our repo\n>> between tests, as for e.g. \"git add\" you want it to actually add the\n>> objects.\n>>\n>> It's what I do with the \"hyperfine\" command in\n>> https://lore.kernel.org/git/RFC-patch-v2-4.7-61f4f3d7ef4-20220323T140753Z-avarab@gmail.com/\n>> with the \"-p\" option.\n>>\n>> I.e. hyperfine has a way to say \"this is setup, but don't measure the\n>> time\", which is 1/2 of what you're working around here and in 10/11.\n>>\n>> But as 10/11 shows you're limited to one run with t/perf because you\n>> want to not include those \"setup\" numbers, and \"test_perf\" has no easy\n>> way to avoid that (but more on that later).\n>>\n>> Which b.t.w. I'm really skeptical of as an approach here in any case\n>> (even if we couldn't exclude it from the numbers).\n>>\n>> I.e. yes what \"hyperfine\" does would be preferrable, but in exchange for\n>> avoiding that you're comparing samples of 1 runs.\n>>\n>> Surely we're better off with N run (even if noisy). Given enough of them\n>> the difference will shake out, and our estimated +/- will narrow..\n>>\n>> But aside from that, why isn't this just:\n>>\n>>         for cfg in true false blah\n>>         done\n>>                 test_expect_success \"setup for $cfg\" '\n>>                         git init repo-$cfg &&\n>>                         for f in $(test_seq 1 100)\n>>                         do\n>>                                 >repo-$cfg/$f\n>>                         done\n>>                 '\n>>\n>>                 test_perf \"perf test for $cfg\" '\n>>                         git -C repo-$cfg\n>>                 '\n>>         done\n>>\n>> Which surely is going to be more accurate in the context of our limited\n>> t/perf environment because creating unique files is not sufficient at\n>> all to ensure that your tests don't interfere with each other.\n>>\n>> That's because in the first iteration we'll create N objects in\n>> .git/objects/aa/* or whatever, which will *still be there* for your\n>> second test, which will impact performance.\n>>\n>> Whereas if you just make N repos you don't need unique files, and you\n>> won't be introducing that as a conflating variable.\n>>\n>> But anyway, reading perf-lib.sh again I haven't tested, but this whole\n>> workaround seems truly unnecessary. I.e. in test_run_perf_ we do:\n>>\n>>         test_run_perf_ () {\n>>                 test_cleanup=:\n>>                 test_export_=\"test_cleanup\"\n>>                 export test_cleanup test_export_\n>>                 \"$GTIME\" -f \"%E %U %S\" -o test_time.$i \"$TEST_SHELL_PATH\" -c '\n>>                         [... code we run and time ...]\n>>                 '\n>>                 [... later ...]\n>>                 test_eval_ \"$test_cleanup\"\n>>         }\n>>\n>> So can't you just avoid this whole glorious workaround for the low low\n>> cost of approximately one shellscript string assignment? :)\n>>\n>> I.e. if you do:\n>>\n>>         setup_clean () {\n>>                 rm -rf repo\n>>         }\n>>\n>>         setup_first () {\n>>                 git init repo &&\n>>                 [make a bunch of files or whatever in repo]\n>>         }\n>>\n>>         setup_next () {\n>>                 test_when_finished \"setup_clean\" &&\n>>                 setup_first\n>>         }\n>>\n>>         test_expect_success 'setup initial stuff' '\n>>                 setup_first\n>>         '\n>>\n>>         test_perf 'my perf test' '\n>>                 test_when_finished \"setup_next\" &&\n>>                 [your perf test here]\n>>         '\n>>\n>>         test_expect_success 'cleanup' '\n>>                 # Not really needed, but just for completeness, we are\n>>                 # about to nuke the trash dir anyway...\n>>                 setup_clean\n>>         '\n>>\n>> I haven't tested (and need to run), but i'm pretty sure that does\n>> exactly what you want without these workarounds, i.e. you'll get\n>> \"trampoline setup\" without that setup being included in the perf\n>> numbers.\n>>\n>> Is it pretty? No, but it's a lot less complex than this unique file\n>> business & workarounds, and will give you just the numbers you want, and\n>> most importantly you car run it N times now for better samples.\n>>\n>> I.e. \"what you want\" sans a *tiny* bit of noise that we use to just call\n>> a function to do:\n>>\n>>     test_cleanup=setup_next\n>>\n>> Which we'll then eval *after* we measure your numbers to setup the next\n>> test.\n>\n> How about I add a new test_perf_setup mechanism to make your idea work\n> in a straightforward way?\n\nSure, that sounds great.\n\n> I still want the test_create_unique_files thing as a way to make\n> multiple files easily.  And for the non-perf tests it makes sense to\n> have differing contents within a test run.\n\nI think running your perf test on some generated data might still make\nsense, but I think given the above that the *method* really doesn't make\nany sense.\n\nI.e. pretty much the whole structure of t/perf is to write tests that\ncan be run on an arbitrary user-provided repo, some of them do make some\ncontent assumptions (or need no repo), but we've tried to have tests\nthere handle arbitrary repos.\n\nYou ended up with that \"generated random files\" to get around the X-Y\nproblem of not being able to reset the area without making that part of\nthe metrics, but as demo'd above we can use test_when_finished for that.\n\nAnd once that's resolved it would actually be much more handy to be able\nto run this on an arbitrary repo, as you can see in my \"git hyperfine\"\none-liner I grabbed the \"t\" directory, but we could just make our test\ndata all files in the dir (or specify a glob via an env var).\n\nI think it still sounds interesting to have a way to make arbitrary test\ndata, but surely that's then better as e.g.:\n\n\tcd t/perf\n \t./make-random-repo /tmp/random-repo &&\n\tGIT_PERF_REPO=/tmp/random-repo ./run p<your test>\n\nI.e. once we've resolved the metrics/play area issue needing to run this\non some very specific data is artificial limitation v.s. just being able\nto point it at a given repo.\n"},{"id":"452524","messageId":"c7a2a7efe6d532fc7fce1352b1dfce640cc9f2f6.1648514552.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 01/13] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:18Z","receivedAt":"2022-03-29T00:42:40Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit prepares for adding batch-fsync to the bulk-checkin\ninfrastructure.\n\nThe bulk-checkin infrastructure is currently used to batch up addition\nof large blobs to a packfile. When a blob is larger than\nbig_file_threshold, we unconditionally add it to a pack. If bulk\ncheckins are 'plugged', we allow multiple large blobs to be added to a\nsingle pack until we reach the packfile size limit; otherwise, we simply\nmake a new packfile for each large blob. The 'unplug' call tells us when\nthe series of blob additions is done so that we can finish the packfiles\nand make their objects available to subsequent operations.\n\nStated another way, bulk-checkin allows callers to define a transaction\nthat adds multiple objects to the object database, where the object\ndatabase can optimize its internal operations within the transaction\nboundary.\n\nBatched fsync will fit into bulk-checkin by taking advantage of the\nplug/unplug functionality to determine the appropriate time to fsync\nand make newly-added objects available in the primary object database.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c9..577b135e39c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -278,21 +278,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"452525","messageId":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v3.git.1648097906.gitgitgadget@gmail.com","subject":"[PATCH v4 00/13] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:17Z","receivedAt":"2022-03-29T00:42:42Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"V4 changes:\n\n * Make ODB transactions nestable.\n * Add an ODB transaction around writing out the cached tree.\n * Change update-index to use a more straightforward way of managing ODB\n   transactions.\n * Fix missing 'local's in lib-unique-files\n * Add a per-iteration setup mechanism to test_perf.\n * Fix camelCasing in warning message.\n\nV3 changes:\n\n * Rebrand plug/unplug-bulk-checkin to \"begin_odb_transaction\" and\n   \"end_odb_transaction\"\n * Add a patch to pass filenames to fsync_or_die, rather than the string\n   \"loose object\"\n * Update the commit description for \"core.fsyncmethod to explain why we do\n   not directly expose objects until an fsync occurs.\n * Also explain in the commit description why we're using a dummy file for\n   the fsync.\n * Create the bulk-fsync tmp-objdir lazily the first time a loose object is\n   added. We now do fsync iff that objdir exists.\n * Do batch fsync if core.fsyncMethod=batch and core.fsync contains\n   loose-object, regardless of the core.fsyncObjectFiles setting.\n * Mitigate the risk in update-index of an object not being visible due to\n   bulk checkin.\n * Add a perf comment to justify the unpack-objects usage of bulk-checkin.\n * Add a new patch to create helpers for parsing OIDs from git commands.\n * Add a comment to the lib-unique-files.sh helper about uniqueness only\n   within a repo.\n * Fix style and add '&&' chaining to test helpers.\n * Comment on some magic numbers in tests.\n * Take the object list as an argument in\n   ./t5300-pack-object.sh:check_unpack ()\n * Drop accidental change to t/perf/perf-lib.sh\n\nV2 changes:\n\n * Change doc to indicate that only some repo updates are batched\n * Null and zero out control variables in do_batch_fsync under\n   unplug_bulk_checkin\n * Make batch mode default on Windows.\n * Update the description for the initial patch that cleans up the\n   bulk-checkin infrastructure.\n * Rebase onto 'seen' at 0cac37f38f9.\n\n--Original definition-- When core.fsync includes loose-object, we issue an\nfsync after every written object. For a 'git-add' or similar command that\nadds a lot of files to the repo, the costs of these fsyncs adds up. One\nmajor factor in this cost is the time it takes for the physical storage\ncontroller to flush its caches to durable media.\n\nThis series takes advantage of the writeout-only mode of git_fsync to issue\nOS cache writebacks for all of the objects being added to the repository\nfollowed by a single fsync to a dummy file, which should trigger a\nfilesystem log flush and storage controller cache flush. This mechanism is\nknown to be safe on common Windows filesystems and expected to be safe on\nmacOS. Some linux filesystems, such as XFS, will probably do the right thing\nas well. See [1] for previous discussion on the predecessor of this patch\nseries.\n\nThis series is important on Windows, where loose-objects are included in the\nfsync set by default in Git-For-Windows. In this series, I'm also setting\nthe default mode for Windows to turn on loose object fsyncing with batch\nmode, so that we can get CI coverage of the actual git-for-windows\nconfiguration upstream. We still don't actually issue fsyncs for the test\nsuite since GIT_TEST_FSYNC is set to 0, but we exercise all of the\nsurrounding batch mode code.\n\nThis work is based on 'next' at c54b8eb302. It's dependent on\nns/core-fsyncmethod.\n\n[1]\nhttps://lore.kernel.org/git/2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com/\n\nNeeraj Singh (13):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  object-file: pass filename to fsync_or_die\n  core.fsyncmethod: batched disk flushes for loose-objects\n  cache-tree: use ODB transaction around writing a tree\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  test-lib-functions: add parsing helpers for ls-files and ls-tree\n  core.fsyncmethod: tests for batch mode\n  t/perf: add iteration setup mechanism to perf-lib\n  core.fsyncmethod: performance tests for add and stash\n  core.fsyncmethod: correctly camel-case warning message\n\n Documentation/config/core.txt          |   8 ++\n builtin/add.c                          |   4 +-\n builtin/unpack-objects.c               |   3 +\n builtin/update-index.c                 |  24 ++++++\n bulk-checkin.c                         | 101 ++++++++++++++++++++++---\n bulk-checkin.h                         |  17 ++++-\n cache-tree.c                           |   3 +\n cache.h                                |  12 ++-\n compat/mingw.h                         |   3 +\n config.c                               |   6 +-\n git-compat-util.h                      |   2 +\n object-file.c                          |  15 ++--\n t/lib-unique-files.sh                  |  34 +++++++++\n t/perf/p3700-add.sh                    |  59 +++++++++++++++\n t/perf/p4220-log-grep-engines.sh       |   3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |   3 +-\n t/perf/p5302-pack-index.sh             |  15 ++--\n t/perf/p7519-fsmonitor.sh              |  18 +----\n t/perf/p7820-grep-engines.sh           |   6 +-\n t/perf/perf-lib.sh                     |  62 +++++++++++++--\n t/t3700-add.sh                         |  28 +++++++\n t/t3903-stash.sh                       |  20 +++++\n t/t5300-pack-object.sh                 |  41 ++++++----\n t/t5317-pack-objects-filter-objects.sh |  91 +++++++++++-----------\n t/test-lib-functions.sh                |  10 +++\n 25 files changed, 469 insertions(+), 119 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p3700-add.sh\n\n\nbase-commit: c54b8eb302ffb72f31e73a26044c8a864e2cb307\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1134%2Fneerajsi-msft%2Fns%2Fbatched-fsync-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1134/neerajsi-msft/ns/batched-fsync-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/1134\n\nRange-diff vs v3:\n\n  2:  b2d9766a662 !  1:  c7a2a7efe6d bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n       \treturn status;\n       }\n       \n     - void begin_odb_transaction(void)\n     + void plug_bulk_checkin(void)\n       {\n      -\tstate.plugged = 1;\n      +\tassert(!bulk_checkin_plugged);\n      +\tbulk_checkin_plugged = 1;\n       }\n       \n     - void end_odb_transaction(void)\n     + void unplug_bulk_checkin(void)\n       {\n      -\tstate.plugged = 0;\n      -\tif (state.f)\n  1:  53261f0099d !  2:  d045b13795b bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n     @@ Commit message\n      \n          Make it clearer in the naming and documentation of the plug_bulk_checkin\n          and unplug_bulk_checkin APIs that they can be thought of as\n     -    a \"transaction\" to optimize operations on the object database.\n     +    a \"transaction\" to optimize operations on the object database. These\n     +    transactions may be nested so that subsystems like the cache-tree\n     +    writing code can optimize their operations without caring whether the\n     +    top-level code has a transaction active.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ builtin/add.c: int cmd_add(int argc, const char **argv, const char *prefix)\n       \tif (write_locked_index(&the_index, &lock_file,\n      \n       ## bulk-checkin.c ##\n     +@@\n     + #include \"packfile.h\"\n     + #include \"object-store.h\"\n     + \n     +-static int bulk_checkin_plugged;\n     ++static int odb_transaction_nesting;\n     + \n     + static struct bulk_checkin_state {\n     + \tchar *pack_tmp_name;\n      @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n     + {\n     + \tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n     + \t\t\t\t     path, flags);\n     +-\tif (!bulk_checkin_plugged)\n     ++\tif (!odb_transaction_nesting)\n     + \t\tfinish_bulk_checkin(&bulk_checkin_state);\n       \treturn status;\n       }\n       \n      -void plug_bulk_checkin(void)\n      +void begin_odb_transaction(void)\n       {\n     - \tstate.plugged = 1;\n     +-\tassert(!bulk_checkin_plugged);\n     +-\tbulk_checkin_plugged = 1;\n     ++\todb_transaction_nesting += 1;\n       }\n       \n      -void unplug_bulk_checkin(void)\n      +void end_odb_transaction(void)\n       {\n     - \tstate.plugged = 0;\n     - \tif (state.f)\n     +-\tassert(bulk_checkin_plugged);\n     +-\tbulk_checkin_plugged = 0;\n     ++\todb_transaction_nesting -= 1;\n     ++\tif (odb_transaction_nesting < 0)\n     ++\t\tBUG(\"Unbalanced ODB transaction nesting\");\n     ++\n     ++\tif (odb_transaction_nesting)\n     ++\t\treturn;\n     ++\n     + \tif (bulk_checkin_state.f)\n     + \t\tfinish_bulk_checkin(&bulk_checkin_state);\n     + }\n      \n       ## bulk-checkin.h ##\n      @@ bulk-checkin.h: int index_bulk_checkin(struct object_id *oid,\n  3:  26ce5b8fdda =  3:  2d1bc4568ac object-file: pass filename to fsync_or_die\n  4:  52638326790 !  4:  9e7ae22fa4a core.fsyncmethod: batched disk flushes for loose-objects\n     @@ bulk-checkin.c\n       #include \"packfile.h\"\n       #include \"object-store.h\"\n       \n     - static int bulk_checkin_plugged;\n     + static int odb_transaction_nesting;\n       \n      +static struct tmp_objdir *bulk_fsync_objdir;\n      +\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +\t * callers may not know whether any objects will be\n      +\t * added at the time they call begin_odb_transaction.\n      +\t */\n     -+\tif (!bulk_checkin_plugged || bulk_fsync_objdir)\n     ++\tif (!odb_transaction_nesting || bulk_fsync_objdir)\n      +\t\treturn;\n      +\n      +\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      +void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n      +{\n      +\t/*\n     -+\t * If we have a plugged bulk checkin, we issue a call that\n     ++\t * If we have an active ODB transaction, we issue a call that\n      +\t * cleans the filesystem page cache but avoids a hardware flush\n      +\t * command. Later on we will issue a single hardware flush\n      +\t * before as part of do_batch_fsync.\n     @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n       \t\t       int fd, size_t size, enum object_type type,\n       \t\t       const char *path, unsigned flags)\n      @@ bulk-checkin.c: void end_odb_transaction(void)\n     - \tbulk_checkin_plugged = 0;\n     + \n       \tif (bulk_checkin_state.f)\n       \t\tfinish_bulk_checkin(&bulk_checkin_state);\n      +\n  -:  ----------- >  5:  83fa4a5f3a5 cache-tree: use ODB transaction around writing a tree\n  5:  913ce1b3df9 !  6:  f03ebee695a update-index: use the bulk-checkin infrastructure\n     @@ Commit message\n          There is some risk with this change, since under batch fsync, the object\n          files will be in a tmp-objdir until update-index is complete, so callers\n          using the --stdin option will not see them until update-index is done.\n     -    This risk is mitigated by unplugging the batch when reporting verbose\n     -    output, which is the only way a --stdin caller might synchronize with\n     -    the addition of an object.\n     +    This risk is mitigated by not keeping an ODB transaction open around\n     +    --stdin processing if in --verbose mode. Without --verbose mode,\n     +    a caller feeding update-index via --stdin wouldn't know when\n     +    update-index adds an object, event without an ODB transaction.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     @@ builtin/update-index.c\n       #include \"config.h\"\n       #include \"lockfile.h\"\n       #include \"quote.h\"\n     -@@ builtin/update-index.c: static int allow_replace;\n     - static int info_only;\n     - static int force_remove;\n     - static int verbose;\n     -+static int odb_transaction_active;\n     - static int mark_valid_only;\n     - static int mark_skip_worktree_only;\n     - static int mark_fsmonitor_only;\n     -@@ builtin/update-index.c: enum uc_mode {\n     - \tUC_FORCE\n     - };\n     - \n     -+static void end_odb_transaction_if_active(void)\n     -+{\n     -+\tif (!odb_transaction_active)\n     -+\t\treturn;\n     -+\n     -+\tend_odb_transaction();\n     -+\todb_transaction_active = 0;\n     -+}\n     -+\n     - __attribute__((format (printf, 1, 2)))\n     - static void report(const char *fmt, ...)\n     - {\n     -@@ builtin/update-index.c: static void report(const char *fmt, ...)\n     - \tif (!verbose)\n     - \t\treturn;\n     - \n     -+\t/*\n     -+\t * It is possible, though unlikely, that a caller\n     -+\t * could use the verbose output to synchronize with\n     -+\t * addition of objects to the object database, so\n     -+\t * unplug bulk checkin to make sure that future objects\n     -+\t * are immediately visible.\n     -+\t */\n     -+\n     -+\tend_odb_transaction_if_active();\n     -+\n     - \tva_start(vp, fmt);\n     - \tvprintf(fmt, vp);\n     - \tputchar('\\n');\n      @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n       \t */\n       \tparse_options_start(&ctx, argc, argv, prefix,\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n      +\t * a batch.\n      +\t */\n      +\tbegin_odb_transaction();\n     -+\todb_transaction_active = 1;\n       \twhile (ctx.argc) {\n       \t\tif (parseopt_state != PARSE_OPT_DONE)\n       \t\t\tparseopt_state = parse_options_step(&ctx, options,\n     +@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n     + \t\tthe_index.version = preferred_index_format;\n     + \t}\n     + \n     ++\t/*\n     ++\t * It is possible, though unlikely, that a caller could use the verbose\n     ++\t * output to synchronize with addition of objects to the object\n     ++\t * database. The current implementation of ODB transactions leaves\n     ++\t * objects invisible while a transaction is active, so end the\n     ++\t * transaction here if verbose output is enabled.\n     ++\t */\n     ++\n     ++\tif (verbose)\n     ++\t\tend_odb_transaction();\n     ++\n     + \tif (read_from_stdin) {\n     + \t\tstruct strbuf buf = STRBUF_INIT;\n     + \t\tstruct strbuf unquoted = STRBUF_INIT;\n      @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n       \t\tstrbuf_release(&buf);\n       \t}\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n      +\t/*\n      +\t * By now we have added all of the new objects\n      +\t */\n     -+\tend_odb_transaction_if_active();\n     ++\tif (!verbose)\n     ++\t\tend_odb_transaction();\n      +\n       \tif (split_index > 0) {\n       \t\tif (git_config_get_split_index() == 0)\n  6:  84fd144ef18 =  7:  d85013f7d2c unpack-objects: use the bulk-checkin infrastructure\n  7:  447263e8ef1 =  8:  73e54f94c20 core.fsync: use batch mode and sync loose objects by default on Windows\n  8:  8f1b01c9ca0 =  9:  124450c86d9 test-lib-functions: add parsing helpers for ls-files and ls-tree\n  9:  b5f371e97fe ! 10:  282fbdef792 core.fsyncmethod: tests for batch mode\n     @@ t/lib-unique-files.sh (new)\n      +\tlocal files=\"$2\" &&\n      +\tlocal basedir=\"$3\" &&\n      +\tlocal counter=0 &&\n     ++\tlocal i &&\n     ++\tlocal j &&\n      +\ttest_tick &&\n      +\tlocal basedata=$basedir$test_tick &&\n      +\trm -rf \"$basedir\" &&\n  -:  ----------- > 11:  ee7ecf4cabe t/perf: add iteration setup mechanism to perf-lib\n 10:  b99b32a469c ! 12:  fdf90d45f52 core.fsyncmethod: performance tests for add and stash\n     @@ t/perf/p3700-add.sh (new)\n      +# core.fsyncMethod=batch mode, which is why we are testing different values\n      +# of that setting explicitly and creating a lot of unique objects.\n      +\n     -+test_description=\"Tests performance of add\"\n     ++test_description=\"Tests performance of adding things to the object database\"\n      +\n      +# Fsync is normally turned off for the test suite.\n      +GIT_TEST_FSYNC=1\n     @@ t/perf/p3700-add.sh (new)\n      +\n      +. $TEST_DIRECTORY/lib-unique-files.sh\n      +\n     -+test_perf_default_repo\n     ++test_perf_fresh_repo\n      +test_checkout_worktree\n      +\n      +dir_count=10\n      +files_per_dir=50\n      +total_files=$((dir_count * files_per_dir))\n      +\n     -+# We need to create the files each time we run the perf test, but\n     -+# we do not want to measure the cost of creating the files, so run\n     -+# the test once.\n     -+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n     -+then\n     -+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n     -+\tGIT_PERF_REPEAT_COUNT=1\n     -+fi\n     -+\n     -+for m in false true batch\n     ++for mode in false true batch\n      +do\n     -+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n     -+\t\tgit reset --hard &&\n     -+\t\t# create files across directories\n     -+\t\ttest_create_unique_files $dir_count $files_per_dir files\n     -+\t'\n     -+\n     -+\tcase $m in\n     ++\tcase $mode in\n      +\tfalse)\n      +\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n      +\t\t;;\n     @@ t/perf/p3700-add.sh (new)\n      +\t\t;;\n      +\tesac\n      +\n     -+\ttest_perf \"add $total_files files (object_fsyncing=$m)\" \"\n     -+\t\tgit $FSYNC_CONFIG add files\n     ++\ttest_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n     ++\t\t--setup \"\n     ++\t\t(rm -rf .git || 1) &&\n     ++\t\tgit init &&\n     ++\t\ttest_create_unique_files $dir_count $files_per_dir files_$mode\n     ++\t\" \"\n     ++\t\tgit $FSYNC_CONFIG add files_$mode\n      +\t\"\n     -+done\n     -+\n     -+test_done\n     -\n     - ## t/perf/p3900-stash.sh (new) ##\n     -@@\n     -+#!/bin/sh\n     -+#\n     -+# This test measures the performance of adding new files to the object database\n     -+# and index. The test was originally added to measure the effect of the\n     -+# core.fsyncMethod=batch mode, which is why we are testing different values\n     -+# of that setting explicitly and creating a lot of unique objects.\n     -+\n     -+test_description=\"Tests performance of stash\"\n     -+\n     -+# Fsync is normally turned off for the test suite.\n     -+GIT_TEST_FSYNC=1\n     -+export GIT_TEST_FSYNC\n     -+\n     -+. ./perf-lib.sh\n     -+\n     -+. $TEST_DIRECTORY/lib-unique-files.sh\n     -+\n     -+test_perf_default_repo\n     -+test_checkout_worktree\n     -+\n     -+dir_count=10\n     -+files_per_dir=50\n     -+total_files=$((dir_count * files_per_dir))\n     -+\n     -+# We need to create the files each time we run the perf test, but\n     -+# we do not want to measure the cost of creating the files, so run\n     -+# the test once.\n     -+if test \"${GIT_PERF_REPEAT_COUNT-1}\" -ne 1\n     -+then\n     -+\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n     -+\tGIT_PERF_REPEAT_COUNT=1\n     -+fi\n     -+\n     -+for m in false true batch\n     -+do\n     -+\ttest_expect_success \"create the files for object_fsyncing=$m\" '\n     -+\t\tgit reset --hard &&\n     -+\t\t# create files across directories\n     -+\t\ttest_create_unique_files $dir_count $files_per_dir files\n     -+\t'\n     -+\n     -+\tcase $m in\n     -+\tfalse)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n     -+\t\t;;\n     -+\ttrue)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n     -+\t\t;;\n     -+\tbatch)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n     -+\t\t;;\n     -+\tesac\n      +\n     -+\t# We only stash files in the 'files' subdirectory since\n     -+\t# the perf test infrastructure creates files in the\n     -+\t# current working directory that need to be preserved\n     -+\ttest_perf \"stash $total_files files (object_fsyncing=$m)\" \"\n     -+\t\tgit $FSYNC_CONFIG stash push -u -- files\n     ++\ttest_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n     ++\t\t--setup \"\n     ++\t\t(rm -rf .git || 1) &&\n     ++\t\tgit init &&\n     ++\t\ttest_commit first &&\n     ++\t\ttest_create_unique_files $dir_count $files_per_dir stash_files_$mode\n     ++\t\" \"\n     ++\t\tgit $FSYNC_CONFIG stash push -u -- stash_files_$mode\n      +\t\"\n      +done\n      +\n 11:  6b832e89bc4 = 13:  fb30bd02c8d core.fsyncmethod: correctly camel-case warning message\n\n-- \ngitgitgadget\n"},{"id":"452526","messageId":"d045b13795b38caa27f8e25340212f736b66bb05.1648514552.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 02/13] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:19Z","receivedAt":"2022-03-29T00:42:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nMake it clearer in the naming and documentation of the plug_bulk_checkin\nand unplug_bulk_checkin APIs that they can be thought of as\na \"transaction\" to optimize operations on the object database. These\ntransactions may be nested so that subsystems like the cache-tree\nwriting code can optimize their operations without caring whether the\ntop-level code has a transaction active.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c  |  4 ++--\n bulk-checkin.c | 20 ++++++++++++--------\n bulk-checkin.h | 14 ++++++++++++--\n 3 files changed, 26 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..9bf37ceae8e 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\tstring_list_clear(&only_match_skip_worktree, 0);\n \t}\n \n-\tplug_bulk_checkin();\n+\tbegin_odb_transaction();\n \n \tif (add_renormalize)\n \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n@@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\tend_odb_transaction();\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 577b135e39c..8b0fd5c7723 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,7 +10,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static int bulk_checkin_plugged;\n+static int odb_transaction_nesting;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -280,21 +280,25 @@ int index_bulk_checkin(struct object_id *oid,\n {\n \tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!bulk_checkin_plugged)\n+\tif (!odb_transaction_nesting)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n-void plug_bulk_checkin(void)\n+void begin_odb_transaction(void)\n {\n-\tassert(!bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 1;\n+\todb_transaction_nesting += 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void end_odb_transaction(void)\n {\n-\tassert(bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 0;\n+\todb_transaction_nesting -= 1;\n+\tif (odb_transaction_nesting < 0)\n+\t\tBUG(\"Unbalanced ODB transaction nesting\");\n+\n+\tif (odb_transaction_nesting)\n+\t\treturn;\n+\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..69a94422ac7 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -10,7 +10,17 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n-void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+/*\n+ * Tell the object database to optimize for adding\n+ * multiple objects. end_odb_transaction must be called\n+ * to make new objects visible.\n+ */\n+void begin_odb_transaction(void);\n+\n+/*\n+ * Tell the object database to make any objects from the\n+ * current transaction visible.\n+ */\n+void end_odb_transaction(void);\n \n #endif\n-- \ngitgitgadget\n\n"},{"id":"452527","messageId":"2d1bc4568ac744f11c886a5f964dbe563c04ce8b.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 03/13] object-file: pass filename to fsync_or_die","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:20Z","receivedAt":"2022-03-29T00:42:45Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf we die while trying to fsync a loose object file, pass the actual\nfilename we're trying to sync. This is likely to be more helpful for a\nuser trying to diagnose the cause of the failure than the former\n'loose object file' string. It also sidesteps any concerns about\ntranslating the die message differently for loose objects versus\nsomething else that has a real path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n object-file.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b254bc50d70..5ffbf3d4fd4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1888,16 +1888,16 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n }\n \n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n+static void close_loose_object(int fd, const char *filename)\n {\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n \tif (fsync_object_files > 0)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n-\t\t\t\t       \"loose object file\");\n+\t\t\t\t       filename);\n \n out:\n \tif (close(fd) != 0)\n@@ -2011,7 +2011,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n+\tclose_loose_object(fd, tmp_file.buf);\n \n \tif (mtime) {\n \t\tstruct utimbuf utb;\n-- \ngitgitgadget\n\n"},{"id":"452528","messageId":"9e7ae22fa4a2693fe26659f875dd780080c4cfb2.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 04/13] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:21Z","receivedAt":"2022-03-29T00:42:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin odb-transaction functionality, takes\nadvantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1a. Create the object in a tmp-objdir.\n2a. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1b. Issue an fsync against a dummy file to flush the log and hardware\n   writeback cache, which should by now have seen the tmp-objdir writes.\n2b. Rename all of the tmp-objdir files to their final names.\n3b. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns. This sequence also ensures that no object files\nappear in the main object store unless they are fsync-durable.\n\nBatch mode is only enabled if core.fsync includes loose-objects. If\nthe legacy core.fsyncObjectFiles setting is enabled, but core.fsync does\nnot include loose-objects, we will use file-by-file fsyncing.\n\nIn step (1a) of the sequence, the tmp-objdir is created lazily to avoid\nwork if no loose objects are ever added to the ODB. We use a tmp-objdir\nto maintain the invariant that no loose-objects are visible in the main\nODB unless they are properly fsync-durable. This is important since\nfuture ODB operations that try to create an object with specific\ncontents will silently drop the new data if an object with the target\nhash exists without checking that the loose-object contents match the\nhash. Only a full git-fsck would restore the ODB to a functional state\nwhere dataloss doesn't occur.\n\nIn step (1b) of the sequence, we issue a fsync against a dummy file\ncreated specifically for the purpose. This method has a little higher\ncost than using one of the input object files, but makes adding new\ncallers of this mechanism easier, since we don't need to figure out\nwhich object file is \"last\" or risk sharing violations by caching the fd\nof the last object file.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  8 ++++\n bulk-checkin.c                | 71 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  3 ++\n cache.h                       |  8 +++-\n config.c                      |  2 +\n object-file.c                 |  7 +++-\n 6 files changed, 97 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 9da3e5d88f6..3c90ba0b395 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -596,6 +596,14 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then does a single full fsync of\n+  a dummy file to trigger the disk cache flush at the end of the operation.\n++\n+  Currently `batch` mode only applies to loose-object files. Other repository\n+  data is made durable as if `fsync` was specified. This mode is expected to\n+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+  and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8b0fd5c7723..9799d247cad 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int odb_transaction_nesting;\n \n+static struct tmp_objdir *bulk_fsync_objdir;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -80,6 +85,40 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\tstruct strbuf temp_path = STRBUF_INIT;\n+\tstruct tempfile *temp;\n+\n+\tif (!bulk_fsync_objdir)\n+\t\treturn;\n+\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur. The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware or any filesystem logs. This fsync call acts as a barrier\n+\t * to ensure that the data in each new object file is durable before\n+\t * the final name is visible.\n+\t */\n+\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\ttemp = xmks_tempfile(temp_path.buf);\n+\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\tdelete_tempfile(&temp);\n+\tstrbuf_release(&temp_path);\n+\n+\t/*\n+\t * Make the object files visible in the primary ODB after their data is\n+\t * fully durable.\n+\t */\n+\ttmp_objdir_migrate(bulk_fsync_objdir);\n+\tbulk_fsync_objdir = NULL;\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +313,36 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void prepare_loose_object_bulk_checkin(void)\n+{\n+\t/*\n+\t * We lazily create the temporary object directory\n+\t * the first time an object might be added, since\n+\t * callers may not know whether any objects will be\n+\t * added at the time they call begin_odb_transaction.\n+\t */\n+\tif (!odb_transaction_nesting || bulk_fsync_objdir)\n+\t\treturn;\n+\n+\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+}\n+\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n+{\n+\t/*\n+\t * If we have an active ODB transaction, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (!bulk_fsync_objdir ||\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n+\t\tfsync_or_die(fd, filename);\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -301,4 +370,6 @@ void end_odb_transaction(void)\n \n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 69a94422ac7..70edf745be8 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,9 @@\n \n #include \"cache.h\"\n \n+void prepare_loose_object_bulk_checkin(void);\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex ef7d34b7a09..a5bf15a5131 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH,\n };\n \n extern enum fsync_method fsync_method;\n@@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex 3c9b6b589ab..511f4584eeb 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd4..d2e0c13198f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1893,7 +1893,9 @@ static void close_loose_object(int fd, const char *filename)\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (fsync_object_files > 0)\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd, filename);\n+\telse if (fsync_object_files > 0)\n \t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n@@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n-- \ngitgitgadget\n\n"},{"id":"452529","messageId":"83fa4a5f3a5c79fa814932c0705867ff16a584c7.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 05/13] cache-tree: use ODB transaction around writing a tree","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:22Z","receivedAt":"2022-03-29T00:42:49Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTake advantage of the odb transaction infrastructure around writing the\ncached tree to the object database.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache-tree.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6752f69d515..8c5e8822716 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -3,6 +3,7 @@\n #include \"tree.h\"\n #include \"tree-walk.h\"\n #include \"cache-tree.h\"\n+#include \"bulk-checkin.h\"\n #include \"object-store.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n@@ -474,8 +475,10 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \ttrace_performance_enter();\n \ttrace2_region_enter(\"cache_tree\", \"update\", the_repository);\n+\tbegin_odb_transaction();\n \ti = update_one(istate->cache_tree, istate->cache, istate->cache_nr,\n \t\t       \"\", 0, &skip, flags);\n+\tend_odb_transaction();\n \ttrace2_region_leave(\"cache_tree\", \"update\", the_repository);\n \ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n-- \ngitgitgadget\n\n"},{"id":"452530","messageId":"f03ebee695a5c21240d35aa5af12349e44931dff.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 06/13] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:23Z","receivedAt":"2022-03-29T00:42:58Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables odb-transactions for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete, so callers\nusing the --stdin option will not see them until update-index is done.\nThis risk is mitigated by not keeping an ODB transaction open around\n--stdin processing if in --verbose mode. Without --verbose mode,\na caller feeding update-index via --stdin wouldn't know when\nupdate-index adds an object, event without an ODB transaction.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 24 ++++++++++++++++++++++++\n 1 file changed, 24 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex aafe7eeac2a..50f9063e1c6 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1116,6 +1117,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t */\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n+\n+\t/*\n+\t * Allow the object layer to optimize adding multiple objects in\n+\t * a batch.\n+\t */\n+\tbegin_odb_transaction();\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1167,6 +1174,17 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tthe_index.version = preferred_index_format;\n \t}\n \n+\t/*\n+\t * It is possible, though unlikely, that a caller could use the verbose\n+\t * output to synchronize with addition of objects to the object\n+\t * database. The current implementation of ODB transactions leaves\n+\t * objects invisible while a transaction is active, so end the\n+\t * transaction here if verbose output is enabled.\n+\t */\n+\n+\tif (verbose)\n+\t\tend_odb_transaction();\n+\n \tif (read_from_stdin) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n@@ -1190,6 +1208,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/*\n+\t * By now we have added all of the new objects\n+\t */\n+\tif (!verbose)\n+\t\tend_odb_transaction();\n+\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"452531","messageId":"d85013f7d2cff17f279fa2d13569a65f42eebf60.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 07/13] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:24Z","receivedAt":"2022-03-29T00:42:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling an odb-transaction when unpacking objects, we can take advantage\nof batched fsyncs.\n\nHere are some performance numbers to justify batch mode for\nunpack-objects, collected on a WSL2 Ubuntu VM.\n\nFsync Mode | Time for 90 objects (ms)\n-------------------------------------\n       Off | 170\n  On,fsync | 760\n  On,batch | 230\n\nNote that the default unpackLimit is 100 objects, so there's a 3x\nbenefit in the worst case. The non-batch mode fsync scales linearly\nwith the number of objects, so there are significant benefits even with\nsmaller numbers of objects.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..56d05e2725d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tbegin_odb_transaction();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tend_odb_transaction();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"452532","messageId":"73e54f94c204759b0cf77e7b75501adb43b14994.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 08/13] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:25Z","receivedAt":"2022-03-29T00:43:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the most of the test suite, since\nGIT_TEST_FSYNC is set to 0. However, we do exercise all of the\nsurrounding batch mode code since GIT_TEST_FSYNC merely makes the\nmaybe_fsync wrapper always appear to succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex a5bf15a5131..7f6cbb254b4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 511f4584eeb..e9cac5f4707 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 0892e209a2f..fffe42ce7c1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \ngitgitgadget\n\n"},{"id":"452533","messageId":"fdf90d45f52d72cf2ff7fe6b620853da9fafc1b3.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 12/13] core.fsyncmethod: performance tests for add and stash","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:29Z","receivedAt":"2022-03-29T00:43:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for \"git add\" and \"git stash\" of a lot of\nnew objects with various fsync settings. This shows the benefit of batch\nmode relative to full fsync.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p3700-add.sh | 59 +++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 59 insertions(+)\n create mode 100755 t/perf/p3700-add.sh\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nnew file mode 100755\nindex 00000000000..ef6024f9897\n--- /dev/null\n+++ b/t/perf/p3700-add.sh\n@@ -0,0 +1,59 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object database\n+# and index. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values\n+# of that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of adding things to the object database\"\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_fresh_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+for mode in false true batch\n+do\n+\tcase $mode in\n+\tfalse)\n+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\ttrue)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n+\t\t;;\n+\tbatch)\n+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\t\t;;\n+\tesac\n+\n+\ttest_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n+\t\t--setup \"\n+\t\t(rm -rf .git || 1) &&\n+\t\tgit init &&\n+\t\ttest_create_unique_files $dir_count $files_per_dir files_$mode\n+\t\" \"\n+\t\tgit $FSYNC_CONFIG add files_$mode\n+\t\"\n+\n+\ttest_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n+\t\t--setup \"\n+\t\t(rm -rf .git || 1) &&\n+\t\tgit init &&\n+\t\ttest_commit first &&\n+\t\ttest_create_unique_files $dir_count $files_per_dir stash_files_$mode\n+\t\" \"\n+\t\tgit $FSYNC_CONFIG stash push -u -- stash_files_$mode\n+\t\"\n+done\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"452534","messageId":"282fbdef792b7157804a4139dbc61106f31bbaef.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 10/13] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:27Z","receivedAt":"2022-03-29T00:43:04Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added\ndue to it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 34 ++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n t/t3903-stash.sh       | 20 ++++++++++++++++++++\n t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n 4 files changed, 109 insertions(+), 14 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..34c01a65256\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,34 @@\n+# Helper to create files with unique contents\n+\n+# Create multiple files with unique contents within this test run. Takes the\n+# number of directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with contents\n+#\t\t\t\t\t different from previous invocations\n+#\t\t\t\t\t of this command in this run.\n+\n+test_create_unique_files () {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=\"$1\" &&\n+\tlocal files=\"$2\" &&\n+\tlocal basedir=\"$3\" &&\n+\tlocal counter=0 &&\n+\tlocal i &&\n+\tlocal j &&\n+\ttest_tick &&\n+\tlocal basedata=$basedir$test_tick &&\n+\trm -rf \"$basedir\" &&\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i &&\n+\t\tmkdir -p \"$dir\" &&\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1)) &&\n+\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..8979c8a5f03 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,32 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir1 &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION add -- ./files_base_dir1/ &&\n+\tgit ls-files --stage files_base_dir1/ |\n+\ttest_parse_ls_files_stage_oids >added_files_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files_oids >added_files_actual &&\n+\ttest_cmp added_files_oids added_files_actual\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir2 &&\n+\tfind files_base_dir2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\tgit ls-files --stage files_base_dir2 |\n+\ttest_parse_ls_files_stage_oids >added_files2_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files2_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files2_oids >added_files2_actual &&\n+\ttest_cmp added_files2_oids added_files2_actual\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 4abbc8fccae..20e94881964 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1410,6 +1411,25 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION stash push -u -- ./files_base_dir/ &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./files_base_dir/ |\n+\ttest_parse_ls_tree_oids >stashed_files_oids &&\n+\n+\t# We created 2 dirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 stashed_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <stashed_files_oids >stashed_files_actual &&\n+\ttest_cmp stashed_files_oids stashed_files_actual\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..f8a0f309e2d 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -161,22 +161,27 @@ test_expect_success 'pack-objects with bogus arguments' '\n '\n \n check_unpack () {\n+\tlocal packname=\"$1\" &&\n+\tlocal object_list=\"$2\" &&\n+\tlocal git_config=\"$3\" &&\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $git_config init --bare git2 &&\n+\t(\n+\t\tgit $git_config -C git2 unpack-objects -n <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 unpack-objects <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <\"$object_list\" >current &&\n+\tcmp \"$object_list\" current\n }\n \n test_expect_success 'unpack without delta' '\n-\tcheck_unpack test-1-${packname_1}\n+\tcheck_unpack test-1-${packname_1} obj-list\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with REF_DELTA' '\n@@ -185,7 +190,11 @@ test_expect_success 'pack with REF_DELTA' '\n '\n \n test_expect_success 'unpack with REF_DELTA' '\n-\tcheck_unpack test-2-${packname_2}\n+\tcheck_unpack test-2-${packname_2} obj-list\n+'\n+\n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with OFS_DELTA' '\n@@ -195,7 +204,11 @@ test_expect_success 'pack with OFS_DELTA' '\n '\n \n test_expect_success 'unpack with OFS_DELTA' '\n-\tcheck_unpack test-3-${packname_3}\n+\tcheck_unpack test-3-${packname_3} obj-list\n+'\n+\n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'compare delta flavors' '\n-- \ngitgitgadget\n\n"},{"id":"452535","messageId":"fb30bd02c8d240cbc8d50a335ac9b52884f06413.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 13/13] core.fsyncmethod: correctly camel-case warning message","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:30Z","receivedAt":"2022-03-29T00:43:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe warning for an unrecognized fsyncMethod was not\ncamel-cased.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n config.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/config.c b/config.c\nindex e9cac5f4707..ae819dee20b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1697,7 +1697,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tif (fsync_object_files < 0)\n-\t\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n+\t\t\twarning(_(\"core.fsyncObjectFiles is deprecated; use core.fsync instead\"));\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n-- \ngitgitgadget\n"},{"id":"452536","messageId":"ee7ecf4cabeff14cc64c979aa77fbb2597a9f986.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 11/13] t/perf: add iteration setup mechanism to perf-lib","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:28Z","receivedAt":"2022-03-29T00:43:08Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTests that affect the repo in stateful ways are easier to write if we\ncan run setup steps outside of the measured portion of perf iteration.\n\nThis change adds a \"--setup 'setup-script'\" parameter to test_perf. To\nmake invocations easier to understand, I also moved the prerequisites to\na new --prereq parameter.\n\nThe setup facility will be used in the upcoming perf tests for batch\nmode, but it already helps in some existing tests, like t5302 and t7820.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p4220-log-grep-engines.sh       |  3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |  3 +-\n t/perf/p5302-pack-index.sh             | 15 +++----\n t/perf/p7519-fsmonitor.sh              | 18 ++------\n t/perf/p7820-grep-engines.sh           |  6 ++-\n t/perf/perf-lib.sh                     | 62 +++++++++++++++++++++++---\n 6 files changed, 73 insertions(+), 34 deletions(-)\n\ndiff --git a/t/perf/p4220-log-grep-engines.sh b/t/perf/p4220-log-grep-engines.sh\nindex 2bc47ded4d1..03fbfbb85d3 100755\n--- a/t/perf/p4220-log-grep-engines.sh\n+++ b/t/perf/p4220-log-grep-engines.sh\n@@ -36,7 +36,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4220_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p4221-log-grep-engines-fixed.sh b/t/perf/p4221-log-grep-engines-fixed.sh\nindex 060971265a9..0a6d6dfc219 100755\n--- a/t/perf/p4221-log-grep-engines-fixed.sh\n+++ b/t/perf/p4221-log-grep-engines-fixed.sh\n@@ -26,7 +26,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4221_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p5302-pack-index.sh b/t/perf/p5302-pack-index.sh\nindex c16f6a3ff69..14c601bbf86 100755\n--- a/t/perf/p5302-pack-index.sh\n+++ b/t/perf/p5302-pack-index.sh\n@@ -26,9 +26,8 @@ test_expect_success 'set up thread-counting tests' '\n \tdone\n '\n \n-test_perf PERF_EXTRA 'index-pack 0 threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack 0 threads' --prereq PERF_EXTRA \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --threads=1 --stdin < $PACK\n '\n \n@@ -36,17 +35,15 @@ for t in $threads\n do\n \tTHREADS=$t\n \texport THREADS\n-\ttest_perf PERF_EXTRA \"index-pack $t threads\" '\n-\t\trm -rf repo.git &&\n-\t\tgit init --bare repo.git &&\n+\ttest_perf \"index-pack $t threads\" --prereq PERF_EXTRA \\\n+\t\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \t\tGIT_DIR=repo.git GIT_FORCE_THREADS=1 \\\n \t\tgit index-pack --threads=$THREADS --stdin <$PACK\n \t'\n done\n \n-test_perf 'index-pack default number of threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack default number of threads' \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --stdin < $PACK\n '\n \ndiff --git a/t/perf/p7519-fsmonitor.sh b/t/perf/p7519-fsmonitor.sh\nindex c8be58f3c76..5b489c968b8 100755\n--- a/t/perf/p7519-fsmonitor.sh\n+++ b/t/perf/p7519-fsmonitor.sh\n@@ -60,18 +60,6 @@ then\n \tesac\n fi\n \n-if test -n \"$GIT_PERF_7519_DROP_CACHE\"\n-then\n-\t# When using GIT_PERF_7519_DROP_CACHE, GIT_PERF_REPEAT_COUNT must be 1 to\n-\t# generate valid results. Otherwise the caching that happens for the nth\n-\t# run will negate the validity of the comparisons.\n-\tif test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n-\tthen\n-\t\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n-\t\tGIT_PERF_REPEAT_COUNT=1\n-\tfi\n-fi\n-\n trace_start() {\n \tif test -n \"$GIT_PERF_7519_TRACE\"\n \tthen\n@@ -167,10 +155,10 @@ setup_for_fsmonitor() {\n \n test_perf_w_drop_caches () {\n \tif test -n \"$GIT_PERF_7519_DROP_CACHE\"; then\n-\t\ttest-tool drop-caches\n+\t\ttest_perf \"$1\" --setup \"test-tool drop-caches\" \"$2\"\n+\telse\n+\t\ttest_perf \"$@\"\n \tfi\n-\n-\ttest_perf \"$@\"\n }\n \n test_fsmonitor_suite() {\ndiff --git a/t/perf/p7820-grep-engines.sh b/t/perf/p7820-grep-engines.sh\nindex 8b09c5bf328..9bfb86842a9 100755\n--- a/t/perf/p7820-grep-engines.sh\n+++ b/t/perf/p7820-grep-engines.sh\n@@ -49,13 +49,15 @@ do\n \t\tfi\n \t\tif ! test_have_prereq PERF_GREP_ENGINES_THREADS\n \t\tthen\n-\t\t\ttest_perf $prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \"\n+\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \\\n+\t\t\t\t--prereq \"$prereq\" \"\n \t\t\t\tgit -c grep.patternType=$engine grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine' || :\n \t\t\t\"\n \t\telse\n \t\t\tfor threads in $GIT_PERF_GREP_THREADS\n \t\t\tdo\n-\t\t\t\ttest_perf PTHREADS,$prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\" \"\n+\t\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\"\n+\t\t\t\t\t--prereq PTHREADS,$prereq \"\n \t\t\t\t\tgit -c grep.patternType=$engine -c grep.threads=$threads grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine.$threads' || :\n \t\t\t\t\"\n \t\t\tdone\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex 407252bac70..a935ad622d3 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -189,19 +189,38 @@ exit $ret' >&3 2>&4\n }\n \n test_wrapper_ () {\n-\ttest_wrapper_func_=$1; shift\n+\tlocal test_wrapper_func_=$1; shift\n+\tlocal test_title_=$1; shift\n \ttest_start_\n-\ttest \"$#\" = 3 && { test_prereq=$1; shift; } || test_prereq=\n-\ttest \"$#\" = 2 ||\n-\tBUG \"not 2 or 3 parameters to test-expect-success\"\n+\ttest_prereq=\n+\ttest_perf_setup_=\n+\twhile test $# != 0\n+\tdo\n+\t\tcase $1 in\n+\t\t--prereq)\n+\t\t\ttest_prereq=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t--setup)\n+\t\t\ttest_perf_setup_=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t*)\n+\t\t\tbreak\n+\t\t\t;;\n+\t\tesac\n+\t\tshift\n+\tdone\n+\ttest \"$#\" = 1 || BUG \"test_wrapper_ needs 2 positional parameters\"\n \texport test_prereq\n-\tif ! test_skip \"$@\"\n+\texport test_perf_setup_\n+\tif ! test_skip \"$test_title_\" \"$@\"\n \tthen\n \t\tbase=$(basename \"$0\" .sh)\n \t\techo \"$test_count\" >>\"$perf_results_dir\"/$base.subtests\n \t\techo \"$1\" >\"$perf_results_dir\"/$base.$test_count.descr\n \t\tbase=\"$perf_results_dir\"/\"$PERF_RESULTS_PREFIX$(basename \"$0\" .sh)\".\"$test_count\"\n-\t\t\"$test_wrapper_func_\" \"$@\"\n+\t\t\"$test_wrapper_func_\" \"$test_title_\" \"$@\"\n \tfi\n \n \ttest_finish_\n@@ -214,6 +233,16 @@ test_perf_ () {\n \t\techo \"perf $test_count - $1:\"\n \tfi\n \tfor i in $(test_seq 1 $GIT_PERF_REPEAT_COUNT); do\n+\t\tif test -n \"$test_perf_setup_\"\n+\t\tthen\n+\t\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\t\tif ! test_eval_ $test_perf_setup_\n+\t\t\tthen\n+\t\t\t\ttest_failure_ \"$test_perf_setup_\"\n+\t\t\t\tbreak\n+\t\t\tfi\n+\n+\t\tfi\n \t\tsay >&3 \"running: $2\"\n \t\tif test_run_perf_ \"$2\"\n \t\tthen\n@@ -237,11 +266,24 @@ test_perf_ () {\n \trm test_time.*\n }\n \n+# Usage: test_perf 'title' [options] 'perf-test'\n+#\tRun the performance test script specified in perf-test with\n+#\toptional prerequisite and setup steps.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to each measured iteration\n+#\n test_perf () {\n \ttest_wrapper_ test_perf_ \"$@\"\n }\n \n test_size_ () {\n+\tif test -n \"$test_perf_setup_\"\n+\tthen\n+\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\ttest_eval_ $test_perf_setup_\n+\tfi\n+\n \tsay >&3 \"running: $2\"\n \tif test_eval_ \"$2\" 3>\"$base\".result; then\n \t\ttest_ok_ \"$1\"\n@@ -250,6 +292,14 @@ test_size_ () {\n \tfi\n }\n \n+# Usage: test_size 'title' [options] 'size-test'\n+#\tRun the size test script specified in size-test with optional\n+#\tprerequisites and setup steps. Returns the numeric value\n+#\treturned by size-test.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to the size measurement\n+\n test_size () {\n \ttest_wrapper_ test_size_ \"$@\"\n }\n-- \ngitgitgadget\n\n"},{"id":"452537","messageId":"124450c86d9f703dde0b5c4fa32e0bd08d4df009.1648514553.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v4 09/13] test-lib-functions: add parsing helpers for ls-files and ls-tree","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-29T00:42:26Z","receivedAt":"2022-03-29T00:43:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nSeveral tests use awk to parse OIDs from the output of 'git ls-files\n--stage' and 'git ls-tree'. Introduce helpers to centralize these uses\nof awk.\n\nUpdate t5317-pack-objects-filter-objects.sh to use the new ls-files\nhelper so that it has some usages to review. Other updates are left for\nthe future.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/t5317-pack-objects-filter-objects.sh | 91 +++++++++++++-------------\n t/test-lib-functions.sh                | 10 +++\n 2 files changed, 54 insertions(+), 47 deletions(-)\n\ndiff --git a/t/t5317-pack-objects-filter-objects.sh b/t/t5317-pack-objects-filter-objects.sh\nindex 33b740ce628..bb633c9b099 100755\n--- a/t/t5317-pack-objects-filter-objects.sh\n+++ b/t/t5317-pack-objects-filter-objects.sh\n@@ -10,9 +10,6 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n # Test blob:none filter.\n \n test_expect_success 'setup r1' '\n-\techo \"{print \\$1}\" >print_1.awk &&\n-\techo \"{print \\$2}\" >print_2.awk &&\n-\n \tgit init r1 &&\n \tfor n in 1 2 3 4 5\n \tdo\n@@ -22,10 +19,13 @@ test_expect_success 'setup r1' '\n \tdone\n '\n \n+parse_verify_pack_blob_oid () {\n+\tawk '{print $1}' -\n+}\n+\n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r1 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -35,7 +35,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -54,12 +54,12 @@ test_expect_success 'verify blob:none packfile has no blobs' '\n test_expect_success 'verify normal and blob:none packfiles have same commits/trees' '\n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r1 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -123,8 +123,8 @@ test_expect_success 'setup r2' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -134,7 +134,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -161,8 +161,8 @@ test_expect_success 'verify blob:limit=1000' '\n '\n \n test_expect_success 'verify blob:limit=1001' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1001 >filter.pack <<-EOF &&\n@@ -172,15 +172,15 @@ test_expect_success 'verify blob:limit=1001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=10001' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=10001 >filter.pack <<-EOF &&\n@@ -190,15 +190,15 @@ test_expect_success 'verify blob:limit=10001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1k' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1k >filter.pack <<-EOF &&\n@@ -208,15 +208,15 @@ test_expect_success 'verify blob:limit=1k' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify explicitly specifying oversized blob in input' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \techo HEAD >objects &&\n@@ -226,15 +226,15 @@ test_expect_success 'verify explicitly specifying oversized blob in input' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1m' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1m >filter.pack <<-EOF &&\n@@ -244,7 +244,7 @@ test_expect_success 'verify blob:limit=1m' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -253,12 +253,12 @@ test_expect_success 'verify blob:limit=1m' '\n test_expect_success 'verify normal and blob:limit packfiles have same commits/trees' '\n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -289,9 +289,8 @@ test_expect_success 'setup r3' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r3 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -301,7 +300,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r3 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -342,9 +341,8 @@ test_expect_success 'setup r4' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -354,19 +352,19 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r4 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=OID' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 ls-files -s pattern >staged &&\n-\toid=$(awk -f print_2.awk staged) &&\n+\toid=$(test_parse_ls_files_stage_oids <staged) &&\n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=$oid >filter.pack <<-EOF &&\n \tHEAD\n \tEOF\n@@ -374,15 +372,15 @@ test_expect_success 'verify sparse:oid=OID' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=oid-ish' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=main:pattern >filter.pack <<-EOF &&\n@@ -392,7 +390,7 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -402,9 +400,8 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n # This models previously omitted objects that we did not receive.\n \n test_expect_success 'setup r1 - delete loose blobs' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tfor id in `cat expected | sed \"s|..|&/|\"`\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex a027f0c409e..e6011409e2f 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1782,6 +1782,16 @@ test_oid_to_path () {\n \techo \"${1%$basename}/$basename\"\n }\n \n+# Parse oids from git ls-files --staged output\n+test_parse_ls_files_stage_oids () {\n+\tawk '{print $2}' -\n+}\n+\n+# Parse oids from git ls-tree output\n+test_parse_ls_tree_oids () {\n+\tawk '{print $3}' -\n+}\n+\n # Choose a port number based on the test script's number and store it in\n # the given variable name, unless that variable already contains a number.\n test_set_port () {\n-- \ngitgitgadget\n\n"},{"id":"452550","messageId":"220329.86czi52ekn.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 00/13] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T10:47:27Z","receivedAt":"2022-03-29T11:17:20Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Mar 29 2022, Neeraj K. Singh via GitGitGadget wrote:\n\n> V4 changes:\n>\n>  * Make ODB transactions nestable.\n>  * Add an ODB transaction around writing out the cached tree.\n>  * Change update-index to use a more straightforward way of managing ODB\n>    transactions.\n>  * Fix missing 'local's in lib-unique-files\n>  * Add a per-iteration setup mechanism to test_perf.\n>  * Fix camelCasing in warning message.\n\nI haven't looked at the bulk of this in any detail, but:\n\n>  10:  b99b32a469c ! 12:  fdf90d45f52 core.fsyncmethod: performance tests for add and stash\n>      @@ t/perf/p3700-add.sh (new)\n>       +# core.fsyncMethod=batch mode, which is why we are testing different values\n>       +# of that setting explicitly and creating a lot of unique objects.\n>       +\n>      -+test_description=\"Tests performance of add\"\n>      ++test_description=\"Tests performance of adding things to the object database\"\n\nNow having both tests for \"add\" and \"stash\" in a test named p3700-add.sh\nisn't better, the rest of the perf tests are split up by command,\nperhaps just add a helper library and have both use it?\n\nAnd re the unaddressed feedback I ad of \"why the random data\"\ninhttps://lore.kernel.org/git/220326.86o81sk9ao.gmgdl@evledraar.gmail.com/\nI tried patching it on top to do what I suggested there, allowing us to\nrun these against any arbitrary repository and came up with this:\n\ndiff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\nindex ef6024f9897..60abd5ee076 100755\n--- a/t/perf/p3700-add.sh\n+++ b/t/perf/p3700-add.sh\n@@ -13,47 +13,26 @@ export GIT_TEST_FSYNC\n \n . ./perf-lib.sh\n \n-. $TEST_DIRECTORY/lib-unique-files.sh\n-\n-test_perf_fresh_repo\n+test_perf_default_repo\n test_checkout_worktree\n \n-dir_count=10\n-files_per_dir=50\n-total_files=$((dir_count * files_per_dir))\n-\n-for mode in false true batch\n+for cfg in \\\n+\t'-c core.fsync=-loose-object -c core.fsyncmethod=fsync' \\\n+\t'-c core.fsync=loose-object -c core.fsyncmethod=fsync' \\\n+\t'-c core.fsync=loose-object -c core.fsyncmethod=batch' \\\n+\t'-c core.fsyncmethod=batch'\n do\n-\tcase $mode in\n-\tfalse)\n-\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n-\t\t;;\n-\ttrue)\n-\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n-\t\t;;\n-\tbatch)\n-\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n-\t\t;;\n-\tesac\n-\n-\ttest_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n-\t\t--setup \"\n-\t\t(rm -rf .git || 1) &&\n-\t\tgit init &&\n-\t\ttest_create_unique_files $dir_count $files_per_dir files_$mode\n-\t\" \"\n-\t\tgit $FSYNC_CONFIG add files_$mode\n-\t\"\n-\n-\ttest_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n-\t\t--setup \"\n-\t\t(rm -rf .git || 1) &&\n-\t\tgit init &&\n-\t\ttest_commit first &&\n-\t\ttest_create_unique_files $dir_count $files_per_dir stash_files_$mode\n-\t\" \"\n-\t\tgit $FSYNC_CONFIG stash push -u -- stash_files_$mode\n-\t\"\n+\ttest_perf \"'git add' with '$cfg'\" \\\n+\t\t--setup '\n+\t\t\tmv -v .git .git.old &&\n+\t\t\tgit init .\n+\t\t' \\\n+\t\t--cleanup '\n+\t\t\trm -rf .git &&\n+\t\t\tmv .git.old .git\n+\t\t' '\n+\t\tgit $cfg add -f -- \":!.git.old/\"\n+\t'\n done\n \n test_done\ndiff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\nnew file mode 100755\nindex 00000000000..12c489069ba\n--- /dev/null\n+++ b/t/perf/p3900-stash.sh\n@@ -0,0 +1,34 @@\n+#!/bin/sh\n+\n+test_description='performance of \"git stash\" with different fsync settings'\n+\n+# Fsync is normally turned off for the test suite.\n+GIT_TEST_FSYNC=1\n+export GIT_TEST_FSYNC\n+\n+. ./perf-lib.sh\n+\n+test_perf_default_repo\n+test_checkout_worktree\n+\n+for cfg in \\\n+\t'-c core.fsync=-loose-object -c core.fsyncmethod=fsync' \\\n+\t'-c core.fsync=loose-object -c core.fsyncmethod=fsync' \\\n+\t'-c core.fsync=loose-object -c core.fsyncmethod=batch' \\\n+\t'-c core.fsyncmethod=batch'\n+do\n+\ttest_perf \"'stash push -u' with '$cfg'\" \\\n+\t\t--setup '\n+\t\t\tmv -v .git .git.old &&\n+\t\t\tgit init . &&\n+\t\t\ttest_commit dummy\n+\t\t' \\\n+\t\t--cleanup '\n+\t\t\trm -rf .git &&\n+\t\t\tmv .git.old .git\n+\t\t' '\n+\t\tgit $cfg stash push -a -u \":!.git.old/\" \":!test*\" \".\"\n+\t'\n+done\n+\n+test_done\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex a935ad622d3..24a5108f234 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -194,6 +194,7 @@ test_wrapper_ () {\n \ttest_start_\n \ttest_prereq=\n \ttest_perf_setup_=\n+\ttest_perf_cleanup_=\n \twhile test $# != 0\n \tdo\n \t\tcase $1 in\n@@ -205,6 +206,10 @@ test_wrapper_ () {\n \t\t\ttest_perf_setup_=$2\n \t\t\tshift\n \t\t\t;;\n+\t\t--cleanup)\n+\t\t\ttest_perf_cleanup_=$2\n+\t\t\tshift\n+\t\t\t;;\n \t\t*)\n \t\t\tbreak\n \t\t\t;;\n@@ -214,6 +219,7 @@ test_wrapper_ () {\n \ttest \"$#\" = 1 || BUG \"test_wrapper_ needs 2 positional parameters\"\n \texport test_prereq\n \texport test_perf_setup_\n+\texport test_perf_cleanup_\n \tif ! test_skip \"$test_title_\" \"$@\"\n \tthen\n \t\tbase=$(basename \"$0\" .sh)\n@@ -256,6 +262,16 @@ test_perf_ () {\n \t\t\ttest_failure_ \"$@\"\n \t\t\tbreak\n \t\tfi\n+\t\tif test -n \"$test_perf_cleanup_\"\n+\t\tthen\n+\t\t\tsay >&3 \"cleanup: $test_perf_cleanup_\"\n+\t\t\tif ! test_eval_ $test_perf_cleanup_\n+\t\t\tthen\n+\t\t\t\ttest_failure_ \"$test_perf_cleanup_\"\n+\t\t\t\tbreak\n+\t\t\tfi\n+\n+\t\tfi\n \tdone\n \tif test -z \"$verbose\"; then\n \t\techo \" ok\"\n\n\nHere it is against Cor.git (a random small-ish repo I had laying around):\n\t\n\t$ GIT_SKIP_TESTS='p3[79]00.[12]' GIT_PERF_MAKE_OPTS='CFLAGS=-O3' GIT_PERF_REPO=~/g/Cor/ ./run origin/master HEAD -- p3900-stash.sh\n\t=== Building abf474a5dd901f28013c52155411a48fd4c09922 (origin/master) ===\n\t    GEN git-add--interactive\n\t    GEN git-archimport\n\t    GEN git-cvsexportcommit\n\t    GEN git-cvsimport\n\t    GEN git-cvsserver\n\t    GEN git-send-email\n\t    GEN git-svn\n\t    GEN git-p4\n\t    SUBDIR templates\n\t=== Running 1 tests in /home/avar/g/git/t/perf/build/abf474a5dd901f28013c52155411a48fd4c09922/bin-wrappers ===\n\tok 1 # skip 'stash push -u' with '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n\tok 2 # skip 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n\tperf 3 - 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=batch': 1 2 3 ok\n\tperf 4 - 'stash push -u' with '-c core.fsyncmethod=batch': 1 2 3 ok\n\t# passed all 4 test(s)\n\t1..4\n\t=== Building ecda9c2b029e35d239e369b875b245f45fd2a097 (HEAD) ===\n\t    GEN git-add--interactive\n\t    GEN git-archimport\n\t    GEN git-cvsexportcommit\n\t    GEN git-cvsimport\n\t    GEN git-cvsserver\n\t    GEN git-send-email\n\t    GEN git-svn\n\t    GEN git-p4\n\t    SUBDIR templates\n\t=== Running 1 tests in /home/avar/g/git/t/perf/build/ecda9c2b029e35d239e369b875b245f45fd2a097/bin-wrappers ===\n\tok 1 # skip 'stash push -u' with '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n\tok 2 # skip 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n\tperf 3 - 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=batch': 1 2 3 ok\n\tperf 4 - 'stash push -u' with '-c core.fsyncmethod=batch': 1 2 3 ok\n\t# passed all 4 test(s)\n\t1..4\n\tTest       origin/master     HEAD\n\t---------------------------------------------------\n\t3900.3:    0.03(0.00+0.00)   0.02(0.00+0.00) -33.3%\n\t3900.4:    0.02(0.00+0.00)   0.03(0.00+0.00) +50.0%\n\t\n"},{"id":"452557","messageId":"220329.868rst2cei.gmgdl@evledraar.gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 00/13] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T11:45:17Z","receivedAt":"2022-03-29T12:04:15Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Mar 29 2022, Neeraj K. Singh via GitGitGadget wrote:\n\n> V4 changes:\n>\n>  * Make ODB transactions nestable.\n>  * Add an ODB transaction around writing out the cached tree.\n>  * Change update-index to use a more straightforward way of managing ODB\n>    transactions.\n>  * Fix missing 'local's in lib-unique-files\n>  * Add a per-iteration setup mechanism to test_perf.\n>  * Fix camelCasing in warning message.\n\nDespite my\nhttps://lore.kernel.org/git/220329.86czi52ekn.gmgdl@evledraar.gmail.com/\nI eventually gave up on trying to extract meaningful numbers from\nt/perf, I can never quite find out if they're because of its\nshellscripts shenanigans or actual code.\n\n(And also; I realize I didn't follow-up on\nhttps://lore.kernel.org/git/CANQDOdcFN5GgOPZ3hqCsjHDTiRfRpqoAKxjF1n9D6S8oD9--_A@mail.gmail.com/,\nsorry):\n\nBut I came up with this (uses my thin\nhttps://gitlab.com/avar/git-hyperfine/ wrapper, and you should be able\nto apt get hyperfine):\n\t\n\t#!/bin/sh\n\tset -xe\n\t\n\tif ! test -d /tmp/scalar.git\n\tthen\n\t\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git\n\t\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack\n\tfi\n\tgit hyperfine \\\n\t        --warmup 1 -r 3 \\\n\t\t-L rev neeraj-v4,avar-RFC \\\n\t\t-s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt' \\\n\t\t-p 'rm -rf repo/.git/objects/* repo/.git/index' \\\n\t\t$@'./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt'\n\t\n\tgit hyperfine \\\n\t        --warmup 1 -r 3 \\\n\t\t-L rev neeraj-v4,avar-RFC \\\n\t\t-s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/' \\\n\t\t-p 'rm -rf repo/.git/objects/* repo/.git/index' \\\n\t\t$@'./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .'\n\t\n\tgit hyperfine \\\n\t        --warmup 1 -r 3 \\\n\t\t-L rev neeraj-v4,avar-RFC \\\n\t        -s 'make CFLAGS=-O3' \\\n\t        -p 'git init --bare dest.git' \\\n\t        -c 'rm -rf dest.git' \\\n\t        $@'./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack'\n\nThose tags are your v4 here & the v2 of the RFC I sent at\nhttps://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/\n\nWhich shows my RFC v2 is ~20% faster with:\n\n    $ PFX='strace' ~/g/git.meta/benchmark.sh \"strace \"\n\n    1.22 ± 0.02 times faster than 'strace ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n    1.22 ± 0.01 times faster than 'strace ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n    1.00 ± 0.01 times faster than 'strace ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n\nBut only for add/update-index, is the unpack-objects not using the\ntmp-objdir? (presumably yes).\n\nAs noted before I've found \"strace\" to be a handy way to \"simulate\"\nslower FS ops on a ramdisk (I get about the same numbers sometimes on\nthe actual non-SSD disk, but due to load on the system (that I'm not in\nfull control of[1]) I can't get hyperfine to be happy with the\nnon-fuzzyness:\n\n    1.06 ± 0.02 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n    1.06 ± 0.03 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n    1.01 ± 0.01 times faster than './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n\nFWIW these are my actual non-fuzzy-with-strace numbers on the\nnot-ramdisk, as you can see the intervals overlap, but for the first two\nthe \"min\" time is never close to the RFC v2:\n\t\n\t$ XDG_RUNTIME_DIR=/tmp/ghf ~/g/git.meta/benchmark.sh\n\t+ test -d /tmp/scalar.git\n\t+ git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt -p rm -rf repo/.git/objects/* repo/.git/index ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt\n\tBenchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4\n\t  Time (mean ± σ):      1.043 s ±  0.143 s    [User: 0.184 s, System: 0.193 s]\n\t  Range (min … max):    0.943 s …  1.207 s    3 runs\n\t\n\tBenchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'avar-RFC\n\t  Time (mean ± σ):     877.6 ms ± 183.4 ms    [User: 197.9 ms, System: 149.4 ms]\n\t  Range (min … max):   697.8 ms … 1064.4 ms    3 runs\n\t\n\tSummary\n\t  './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'avar-RFC' ran\n\t    1.19 ± 0.30 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n\t+ git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ -p rm -rf repo/.git/objects/* repo/.git/index ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .\n\tBenchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4\n\t  Time (mean ± σ):      1.019 s ±  0.057 s    [User: 0.213 s, System: 0.194 s]\n\t  Range (min … max):    0.963 s …  1.076 s    3 runs\n\t\n\tBenchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'avar-RFC\n\t  Time (mean ± σ):     918.6 ms ±  34.4 ms    [User: 207.8 ms, System: 164.1 ms]\n\t  Range (min … max):   880.6 ms … 947.5 ms    3 runs\n\t\n\tSummary\n\t  './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'avar-RFC' ran\n\t    1.11 ± 0.07 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n\t+ git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 -p git init --bare dest.git -c rm -rf dest.git ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack\n\tBenchmark 1: ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4\n\t  Time (mean ± σ):      1.362 s ±  0.285 s    [User: 1.021 s, System: 0.186 s]\n\t  Range (min … max):    1.192 s …  1.691 s    3 runs\n\t\n\t  Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' options.\n\t\n\tBenchmark 2: ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'avar-RFC\n\t ⠏ Performing warmup runs         ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ⠙ Performing warmup runs         ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  Time (mean ± σ):      1.188 s ±  0.009 s    [User: 1.025 s, System: 0.161 s]\n\t  Range (min … max):    1.180 s …  1.199 s    3 runs\n\t \n\tSummary\n\t  './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'avar-RFC' ran\n\t    1.15 ± 0.24 times faster than './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n\n1. I do my git hacking on a bare metal box I rent with some friends, and\n   one of them is running one those persistent video game daemons\n   written in Java. So I think all my non-RAM I/O numbers are\n   continually fuzzed by what players are doing in Minecraft or whatever\n   that thing is...\n"},{"id":"452585","messageId":"CANQDOdcK=2WakB0o2PyGGOHgMMLzYrAeW8tTnpQgG5H0Y5UB-g@mail.gmail.com","threadId":"57568","inReplyTo":"220329.868rst2cei.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 00/13] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-29T16:51:33Z","receivedAt":"2022-03-29T16:51:53Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 29, 2022 at 5:04 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Mar 29 2022, Neeraj K. Singh via GitGitGadget wrote:\n>\n> > V4 changes:\n> >\n> >  * Make ODB transactions nestable.\n> >  * Add an ODB transaction around writing out the cached tree.\n> >  * Change update-index to use a more straightforward way of managing ODB\n> >    transactions.\n> >  * Fix missing 'local's in lib-unique-files\n> >  * Add a per-iteration setup mechanism to test_perf.\n> >  * Fix camelCasing in warning message.\n>\n> Despite my\n> https://lore.kernel.org/git/220329.86czi52ekn.gmgdl@evledraar.gmail.com/\n> I eventually gave up on trying to extract meaningful numbers from\n> t/perf, I can never quite find out if they're because of its\n> shellscripts shenanigans or actual code.\n>\n> (And also; I realize I didn't follow-up on\n> https://lore.kernel.org/git/CANQDOdcFN5GgOPZ3hqCsjHDTiRfRpqoAKxjF1n9D6S8oD9--_A@mail.gmail.com/,\n> sorry):\n>\n\nLooks like we aren't actually hitting fsync in the numbers you\nexpressed there, if they're down in the 20ms range.  Or we simply\naren't adding enough files.  Or if that's against a ramdisk, the\nramdisk doesn't have enough cost to represent real disk hardware.\n\n> But I came up with this (uses my thin\n> https://gitlab.com/avar/git-hyperfine/ wrapper, and you should be able\n> to apt get hyperfine):\n>\n>         #!/bin/sh\n>         set -xe\n>\n>         if ! test -d /tmp/scalar.git\n>         then\n>                 git clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git\n>                 mv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack\n>         fi\n>         git hyperfine \\\n>                 --warmup 1 -r 3 \\\n>                 -L rev neeraj-v4,avar-RFC \\\n>                 -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt' \\\n>                 -p 'rm -rf repo/.git/objects/* repo/.git/index' \\\n>                 $@'./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt'\n>\n>         git hyperfine \\\n>                 --warmup 1 -r 3 \\\n>                 -L rev neeraj-v4,avar-RFC \\\n>                 -s 'make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/' \\\n>                 -p 'rm -rf repo/.git/objects/* repo/.git/index' \\\n>                 $@'./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .'\n>\n>         git hyperfine \\\n>                 --warmup 1 -r 3 \\\n>                 -L rev neeraj-v4,avar-RFC \\\n>                 -s 'make CFLAGS=-O3' \\\n>                 -p 'git init --bare dest.git' \\\n>                 -c 'rm -rf dest.git' \\\n>                 $@'./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack'\n>\n> Those tags are your v4 here & the v2 of the RFC I sent at\n> https://lore.kernel.org/git/RFC-cover-v2-0.7-00000000000-20220323T140753Z-avarab@gmail.com/\n>\n> Which shows my RFC v2 is ~20% faster with:\n>\n>     $ PFX='strace' ~/g/git.meta/benchmark.sh \"strace \"\n>\n>     1.22 ± 0.02 times faster than 'strace ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n>     1.22 ± 0.01 times faster than 'strace ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n>     1.00 ± 0.01 times faster than 'strace ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n>\n> But only for add/update-index, is the unpack-objects not using the\n> tmp-objdir? (presumably yes).\n>\n> As noted before I've found \"strace\" to be a handy way to \"simulate\"\n> slower FS ops on a ramdisk (I get about the same numbers sometimes on\n> the actual non-SSD disk, but due to load on the system (that I'm not in\n> full control of[1]) I can't get hyperfine to be happy with the\n> non-fuzzyness:\n>\n\nAt least in this case, I don't think 'strace' is representative of\nwhat a real disk would behave like.  Unless you can somehow make\nstrace of sync_file_range cost less than strace of fsync.\n\n>     1.06 ± 0.02 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n>     1.06 ± 0.03 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n>     1.01 ± 0.01 times faster than './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n>\n> FWIW these are my actual non-fuzzy-with-strace numbers on the\n> not-ramdisk, as you can see the intervals overlap, but for the first two\n> the \"min\" time is never close to the RFC v2:\n>\n>         $ XDG_RUNTIME_DIR=/tmp/ghf ~/g/git.meta/benchmark.sh\n>         + test -d /tmp/scalar.git\n>         + git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ && git ls-files -- t >repo/.git/to-add.txt -p rm -rf repo/.git/objects/* repo/.git/index ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt\n>         Benchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4\n>           Time (mean ± σ):      1.043 s ±  0.143 s    [User: 0.184 s, System: 0.193 s]\n>           Range (min … max):    0.943 s …  1.207 s    3 runs\n>\n>         Benchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'avar-RFC\n>           Time (mean ± σ):     877.6 ms ± 183.4 ms    [User: 197.9 ms, System: 149.4 ms]\n>           Range (min … max):   697.8 ms … 1064.4 ms    3 runs\n>\n>         Summary\n>           './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'avar-RFC' ran\n>             1.19 ± 0.30 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin <repo/.git/to-add.txt' in 'neeraj-v4'\n>         + git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 && rm -rf repo && git init repo && cp -R t repo/ -p rm -rf repo/.git/objects/* repo/.git/index ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .\n>         Benchmark 1: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4\n>           Time (mean ± σ):      1.019 s ±  0.057 s    [User: 0.213 s, System: 0.194 s]\n>           Range (min … max):    0.963 s …  1.076 s    3 runs\n>\n>         Benchmark 2: ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'avar-RFC\n>           Time (mean ± σ):     918.6 ms ±  34.4 ms    [User: 207.8 ms, System: 164.1 ms]\n>           Range (min … max):   880.6 ms … 947.5 ms    3 runs\n>\n>         Summary\n>           './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'avar-RFC' ran\n>             1.11 ± 0.07 times faster than './git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo add .' in 'neeraj-v4'\n>         + git hyperfine --warmup 1 -r 3 -L rev neeraj-v4,avar-RFC -s make CFLAGS=-O3 -p git init --bare dest.git -c rm -rf dest.git ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack\n>         Benchmark 1: ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4\n>           Time (mean ± σ):      1.362 s ±  0.285 s    [User: 1.021 s, System: 0.186 s]\n>           Range (min … max):    1.192 s …  1.691 s    3 runs\n>\n>           Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' options.\n>\n>         Benchmark 2: ./git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'avar-RFC\n>          ⠏ Performing warmup runs         ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ⠙ Performing warmup runs         ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  Time (mean ± σ):      1.188 s ±  0.009 s    [User: 1.025 s, System: 0.161 s]\n>           Range (min … max):    1.180 s …  1.199 s    3 runs\n>\n>         Summary\n>           './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'avar-RFC' ran\n>             1.15 ± 0.24 times faster than './git -C dest.git -c core.fsyncMethod=batch unpack-objects </tmp/scalar.git/my.pack' in 'neeraj-v4'\n>\n> 1. I do my git hacking on a bare metal box I rent with some friends, and\n>    one of them is running one those persistent video game daemons\n>    written in Java. So I think all my non-RAM I/O numbers are\n>    continually fuzzed by what players are doing in Minecraft or whatever\n>    that thing is...\n\nThanks for the numbers.  So if I'm understanding correctly, the\ndifference on a real disk between quarantine and non-quarantine is 20%\nor so on your system?\n\nI did my own experiment by adding a 'batch-no-quarantine' method. No\nquarantine was slightly faster.\n*  For 'git add' I found a very small difference (.29s vs 30s).\n* For 'git stash' it was a bit bigger (.35s vs.55s).\n\nThis is with perf-lib, so we're just looking at min-times.  On the\nother hand, classic fsync is 1.04s for 'git add' and 1.21s for 'git\nstash', all with 500 tiny blobs.  FYI, this is measured on my laptop\nrunning Ubuntu in WSL.\n\n I don't think it's worth having a knob for no-quarantine for this\nsmall delta.  I believe a better use of time for a follow-on change\nwould be to implement an appendable pack format for newly-added\nobjects.\n\nThanks,\nNeeraj\n"},{"id":"452587","messageId":"CANQDOdf2tXM6f2U5=68Htw+AGyrAM0C6UAVPmfxCk8Ga7pek2g@mail.gmail.com","threadId":"57568","inReplyTo":"220329.86czi52ekn.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 00/13] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-29T17:09:26Z","receivedAt":"2022-03-29T17:09:45Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 29, 2022 at 4:17 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Mar 29 2022, Neeraj K. Singh via GitGitGadget wrote:\n>\n> > V4 changes:\n> >\n> >  * Make ODB transactions nestable.\n> >  * Add an ODB transaction around writing out the cached tree.\n> >  * Change update-index to use a more straightforward way of managing ODB\n> >    transactions.\n> >  * Fix missing 'local's in lib-unique-files\n> >  * Add a per-iteration setup mechanism to test_perf.\n> >  * Fix camelCasing in warning message.\n>\n> I haven't looked at the bulk of this in any detail, but:\n>\n> >  10:  b99b32a469c ! 12:  fdf90d45f52 core.fsyncmethod: performance tests for add and stash\n> >      @@ t/perf/p3700-add.sh (new)\n> >       +# core.fsyncMethod=batch mode, which is why we are testing different values\n> >       +# of that setting explicitly and creating a lot of unique objects.\n> >       +\n> >      -+test_description=\"Tests performance of add\"\n> >      ++test_description=\"Tests performance of adding things to the object database\"\n>\n> Now having both tests for \"add\" and \"stash\" in a test named p3700-add.sh\n> isn't better, the rest of the perf tests are split up by command,\n> perhaps just add a helper library and have both use it?\n>\n\nI was getting tired of editing two files that were nearly identical\nand thought that reviewers would be tired of reading them.  At least\nin the main test suite, t/t3700-add.sh covers update-index in addition\nto git-add.\n\n> And re the unaddressed feedback I ad of \"why the random data\"\n> inhttps://lore.kernel.org/git/220326.86o81sk9ao.gmgdl@evledraar.gmail.com/\n> I tried patching it on top to do what I suggested there, allowing us to\n> run these against any arbitrary repository and came up with this:\n>\n\nThe advantage of the random data is that it's easy to scale the number\nof files and change the tree shape.  I wanted the\ntest_create_unique_files helper anyway, so using it here made sense.\nAlso, I'm quite confident that I'm really getting new objects added to\nthe repo with this test scheme.\n\n> diff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\n> index ef6024f9897..60abd5ee076 100755\n> --- a/t/perf/p3700-add.sh\n> +++ b/t/perf/p3700-add.sh\n> @@ -13,47 +13,26 @@ export GIT_TEST_FSYNC\n>\n>  . ./perf-lib.sh\n>\n> -. $TEST_DIRECTORY/lib-unique-files.sh\n> -\n> -test_perf_fresh_repo\n> +test_perf_default_repo\n>  test_checkout_worktree\n>\n> -dir_count=10\n> -files_per_dir=50\n> -total_files=$((dir_count * files_per_dir))\n> -\n> -for mode in false true batch\n> +for cfg in \\\n> +       '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' \\\n> +       '-c core.fsync=loose-object -c core.fsyncmethod=fsync' \\\n> +       '-c core.fsync=loose-object -c core.fsyncmethod=batch' \\\n> +       '-c core.fsyncmethod=batch'\n>  do\n> -       case $mode in\n> -       false)\n> -               FSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n> -               ;;\n> -       true)\n> -               FSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n> -               ;;\n> -       batch)\n> -               FSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> -               ;;\n> -       esac\n> -\n> -       test_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n> -               --setup \"\n> -               (rm -rf .git || 1) &&\n> -               git init &&\n> -               test_create_unique_files $dir_count $files_per_dir files_$mode\n> -       \" \"\n> -               git $FSYNC_CONFIG add files_$mode\n> -       \"\n> -\n> -       test_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n> -               --setup \"\n> -               (rm -rf .git || 1) &&\n> -               git init &&\n> -               test_commit first &&\n> -               test_create_unique_files $dir_count $files_per_dir stash_files_$mode\n> -       \" \"\n> -               git $FSYNC_CONFIG stash push -u -- stash_files_$mode\n> -       \"\n> +       test_perf \"'git add' with '$cfg'\" \\\n> +               --setup '\n> +                       mv -v .git .git.old &&\n> +                       git init .\n> +               ' \\\n> +               --cleanup '\n> +                       rm -rf .git &&\n> +                       mv .git.old .git\n> +               ' '\n> +               git $cfg add -f -- \":!.git.old/\"\n> +       '\n>  done\n>\n>  test_done\n> diff --git a/t/perf/p3900-stash.sh b/t/perf/p3900-stash.sh\n> new file mode 100755\n> index 00000000000..12c489069ba\n> --- /dev/null\n> +++ b/t/perf/p3900-stash.sh\n> @@ -0,0 +1,34 @@\n> +#!/bin/sh\n> +\n> +test_description='performance of \"git stash\" with different fsync settings'\n> +\n> +# Fsync is normally turned off for the test suite.\n> +GIT_TEST_FSYNC=1\n> +export GIT_TEST_FSYNC\n> +\n> +. ./perf-lib.sh\n> +\n> +test_perf_default_repo\n> +test_checkout_worktree\n> +\n> +for cfg in \\\n> +       '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' \\\n> +       '-c core.fsync=loose-object -c core.fsyncmethod=fsync' \\\n> +       '-c core.fsync=loose-object -c core.fsyncmethod=batch' \\\n> +       '-c core.fsyncmethod=batch'\n> +do\n> +       test_perf \"'stash push -u' with '$cfg'\" \\\n> +               --setup '\n> +                       mv -v .git .git.old &&\n> +                       git init . &&\n> +                       test_commit dummy\n> +               ' \\\n> +               --cleanup '\n> +                       rm -rf .git &&\n> +                       mv .git.old .git\n> +               ' '\n> +               git $cfg stash push -a -u \":!.git.old/\" \":!test*\" \".\"\n> +       '\n> +done\n> +\n> +test_done\n> diff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\n> index a935ad622d3..24a5108f234 100644\n> --- a/t/perf/perf-lib.sh\n> +++ b/t/perf/perf-lib.sh\n> @@ -194,6 +194,7 @@ test_wrapper_ () {\n>         test_start_\n>         test_prereq=\n>         test_perf_setup_=\n> +       test_perf_cleanup_=\n>         while test $# != 0\n>         do\n>                 case $1 in\n> @@ -205,6 +206,10 @@ test_wrapper_ () {\n>                         test_perf_setup_=$2\n>                         shift\n>                         ;;\n> +               --cleanup)\n> +                       test_perf_cleanup_=$2\n> +                       shift\n> +                       ;;\n>                 *)\n>                         break\n>                         ;;\n> @@ -214,6 +219,7 @@ test_wrapper_ () {\n>         test \"$#\" = 1 || BUG \"test_wrapper_ needs 2 positional parameters\"\n>         export test_prereq\n>         export test_perf_setup_\n> +       export test_perf_cleanup_\n>         if ! test_skip \"$test_title_\" \"$@\"\n>         then\n>                 base=$(basename \"$0\" .sh)\n> @@ -256,6 +262,16 @@ test_perf_ () {\n>                         test_failure_ \"$@\"\n>                         break\n>                 fi\n> +               if test -n \"$test_perf_cleanup_\"\n> +               then\n> +                       say >&3 \"cleanup: $test_perf_cleanup_\"\n> +                       if ! test_eval_ $test_perf_cleanup_\n> +                       then\n> +                               test_failure_ \"$test_perf_cleanup_\"\n> +                               break\n> +                       fi\n> +\n> +               fi\n>         done\n>         if test -z \"$verbose\"; then\n>                 echo \" ok\"\n>\n>\n> Here it is against Cor.git (a random small-ish repo I had laying around):\n>\n>         $ GIT_SKIP_TESTS='p3[79]00.[12]' GIT_PERF_MAKE_OPTS='CFLAGS=-O3' GIT_PERF_REPO=~/g/Cor/ ./run origin/master HEAD -- p3900-stash.sh\n>         === Building abf474a5dd901f28013c52155411a48fd4c09922 (origin/master) ===\n>             GEN git-add--interactive\n>             GEN git-archimport\n>             GEN git-cvsexportcommit\n>             GEN git-cvsimport\n>             GEN git-cvsserver\n>             GEN git-send-email\n>             GEN git-svn\n>             GEN git-p4\n>             SUBDIR templates\n>         === Running 1 tests in /home/avar/g/git/t/perf/build/abf474a5dd901f28013c52155411a48fd4c09922/bin-wrappers ===\n>         ok 1 # skip 'stash push -u' with '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n>         ok 2 # skip 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n>         perf 3 - 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=batch': 1 2 3 ok\n>         perf 4 - 'stash push -u' with '-c core.fsyncmethod=batch': 1 2 3 ok\n>         # passed all 4 test(s)\n>         1..4\n>         === Building ecda9c2b029e35d239e369b875b245f45fd2a097 (HEAD) ===\n>             GEN git-add--interactive\n>             GEN git-archimport\n>             GEN git-cvsexportcommit\n>             GEN git-cvsimport\n>             GEN git-cvsserver\n>             GEN git-send-email\n>             GEN git-svn\n>             GEN git-p4\n>             SUBDIR templates\n>         === Running 1 tests in /home/avar/g/git/t/perf/build/ecda9c2b029e35d239e369b875b245f45fd2a097/bin-wrappers ===\n>         ok 1 # skip 'stash push -u' with '-c core.fsync=-loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n>         ok 2 # skip 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=fsync' (GIT_SKIP_TESTS)\n>         perf 3 - 'stash push -u' with '-c core.fsync=loose-object -c core.fsyncmethod=batch': 1 2 3 ok\n>         perf 4 - 'stash push -u' with '-c core.fsyncmethod=batch': 1 2 3 ok\n>         # passed all 4 test(s)\n>         1..4\n>         Test       origin/master     HEAD\n>         ---------------------------------------------------\n>         3900.3:    0.03(0.00+0.00)   0.02(0.00+0.00) -33.3%\n>         3900.4:    0.02(0.00+0.00)   0.03(0.00+0.00) +50.0%\n>\n\nSomething is wrong with your data here.  Or your repo is too small to\nhighlight the differences.\n\nI'd suggest that if you want to write a different perf test for this\nfeature, that it be a follow-on change.\n\nThanks,\nNeeraj\n"},{"id":"452588","messageId":"CANQDOdfjRcTB4=mq869NZrNAk+gapR5i8XHFnBhfH1eONzP0BA@mail.gmail.com","threadId":"57568","inReplyTo":"ee7ecf4cabeff14cc64c979aa77fbb2597a9f986.1648514553.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 11/13] t/perf: add iteration setup mechanism to perf-lib","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-29T17:14:45Z","receivedAt":"2022-03-29T17:15:02Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 28, 2022 at 5:42 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n>  test_wrapper_ () {\n> -       test_wrapper_func_=$1; shift\n> +       local test_wrapper_func_=$1; shift\n> +       local test_title_=$1; shift\n\nlocal with assignment seems to be a bash-ism.  I'll fix this.\n"},{"id":"452593","messageId":"CANQDOdfm9cuP=_+rErsBj97hw6QSyaJ6oQUoVkcaDSqZObGdiQ@mail.gmail.com","threadId":"57568","inReplyTo":"fdf90d45f52d72cf2ff7fe6b620853da9fafc1b3.1648514553.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 12/13] core.fsyncmethod: performance tests for add and stash","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-29T17:38:22Z","receivedAt":"2022-03-29T17:38:41Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Mon, Mar 28, 2022 at 5:42 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Add basic performance tests for \"git add\" and \"git stash\" of a lot of\n> new objects with various fsync settings. This shows the benefit of batch\n> mode relative to full fsync.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  t/perf/p3700-add.sh | 59 +++++++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 59 insertions(+)\n>  create mode 100755 t/perf/p3700-add.sh\n>\n> diff --git a/t/perf/p3700-add.sh b/t/perf/p3700-add.sh\n> new file mode 100755\n> index 00000000000..ef6024f9897\n> --- /dev/null\n> +++ b/t/perf/p3700-add.sh\n> @@ -0,0 +1,59 @@\n> +#!/bin/sh\n> +#\n> +# This test measures the performance of adding new files to the object database\n> +# and index. The test was originally added to measure the effect of the\n> +# core.fsyncMethod=batch mode, which is why we are testing different values\n> +# of that setting explicitly and creating a lot of unique objects.\n> +\n> +test_description=\"Tests performance of adding things to the object database\"\n> +\n> +# Fsync is normally turned off for the test suite.\n> +GIT_TEST_FSYNC=1\n> +export GIT_TEST_FSYNC\n> +\n> +. ./perf-lib.sh\n> +\n> +. $TEST_DIRECTORY/lib-unique-files.sh\n> +\n> +test_perf_fresh_repo\n> +test_checkout_worktree\n> +\n> +dir_count=10\n> +files_per_dir=50\n> +total_files=$((dir_count * files_per_dir))\n> +\n> +for mode in false true batch\n> +do\n> +       case $mode in\n> +       false)\n> +               FSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n> +               ;;\n> +       true)\n> +               FSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n> +               ;;\n> +       batch)\n> +               FSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> +               ;;\n> +       esac\n> +\n> +       test_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n> +               --setup \"\n> +               (rm -rf .git || 1) &&\n> +               git init &&\n> +               test_create_unique_files $dir_count $files_per_dir files_$mode\n> +       \" \"\n> +               git $FSYNC_CONFIG add files_$mode\n> +       \"\n> +\n> +       test_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n> +               --setup \"\n> +               (rm -rf .git || 1) &&\n> +               git init &&\n> +               test_commit first &&\n> +               test_create_unique_files $dir_count $files_per_dir stash_files_$mode\n> +       \" \"\n> +               git $FSYNC_CONFIG stash push -u -- stash_files_$mode\n> +       \"\n> +done\n> +\n> +test_done\n> --\n> gitgitgadget\n>\n\nÆvar suggested that it's a bit funny to have a stash test in a file\nnamed 'p3700-add.sh'.  So I'm thinking that I'll move these tests to a\np0008-odb-fsync.sh instead, and add a test for unpack-object and\ngit-commit as well.  I'm also going to adopt  Ævar's scheme for\niterating over the configs.\n"},{"id":"452597","messageId":"xmqqilrwtwx9.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdfjRcTB4=mq869NZrNAk+gapR5i8XHFnBhfH1eONzP0BA@mail.gmail.com","subject":"Re: [PATCH v4 11/13] t/perf: add iteration setup mechanism to perf-lib","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-29T18:50:58Z","receivedAt":"2022-03-29T18:51:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> On Mon, Mar 28, 2022 at 5:42 PM Neeraj Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>>\n>>  test_wrapper_ () {\n>> -       test_wrapper_func_=$1; shift\n>> +       local test_wrapper_func_=$1; shift\n>> +       local test_title_=$1; shift\n>\n> local with assignment seems to be a bash-ism.  I'll fix this.\n\nThanks.\n"},{"id":"452641","messageId":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v4.git.1648514552.gitgitgadget@gmail.com","subject":"[PATCH v5 00/14] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj K. Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:18Z","receivedAt":"2022-03-30T05:05:42Z","isPatch":true,"sender":{"key":"name:Neeraj K. Singh","avatar":null},"body":"V5 changes:\n\n * Remove 'local-with-assignment' from perf-lib.sh\n * Add a new patch to put an ODB transaction around add_files_to_cache.\n * Move the perf tests to t/perf/p0008-odb-fsync.sh and add more perf tests.\n\nV4 changes:\n\n * Make ODB transactions nestable.\n * Add an ODB transaction around writing out the cached tree.\n * Change update-index to use a more straightforward way of managing ODB\n   transactions.\n * Fix missing 'local's in lib-unique-files\n * Add a per-iteration setup mechanism to test_perf.\n * Fix camelCasing in warning message.\n\nV3 changes:\n\n * Rebrand plug/unplug-bulk-checkin to \"begin_odb_transaction\" and\n   \"end_odb_transaction\"\n * Add a patch to pass filenames to fsync_or_die, rather than the string\n   \"loose object\"\n * Update the commit description for \"core.fsyncmethod to explain why we do\n   not directly expose objects until an fsync occurs.\n * Also explain in the commit description why we're using a dummy file for\n   the fsync.\n * Create the bulk-fsync tmp-objdir lazily the first time a loose object is\n   added. We now do fsync iff that objdir exists.\n * Do batch fsync if core.fsyncMethod=batch and core.fsync contains\n   loose-object, regardless of the core.fsyncObjectFiles setting.\n * Mitigate the risk in update-index of an object not being visible due to\n   bulk checkin.\n * Add a perf comment to justify the unpack-objects usage of bulk-checkin.\n * Add a new patch to create helpers for parsing OIDs from git commands.\n * Add a comment to the lib-unique-files.sh helper about uniqueness only\n   within a repo.\n * Fix style and add '&&' chaining to test helpers.\n * Comment on some magic numbers in tests.\n * Take the object list as an argument in\n   ./t5300-pack-object.sh:check_unpack ()\n * Drop accidental change to t/perf/perf-lib.sh\n\nV2 changes:\n\n * Change doc to indicate that only some repo updates are batched\n * Null and zero out control variables in do_batch_fsync under\n   unplug_bulk_checkin\n * Make batch mode default on Windows.\n * Update the description for the initial patch that cleans up the\n   bulk-checkin infrastructure.\n * Rebase onto 'seen' at 0cac37f38f9.\n\n--Original definition-- When core.fsync includes loose-object, we issue an\nfsync after every written object. For a 'git-add' or similar command that\nadds a lot of files to the repo, the costs of these fsyncs adds up. One\nmajor factor in this cost is the time it takes for the physical storage\ncontroller to flush its caches to durable media.\n\nThis series takes advantage of the writeout-only mode of git_fsync to issue\nOS cache writebacks for all of the objects being added to the repository\nfollowed by a single fsync to a dummy file, which should trigger a\nfilesystem log flush and storage controller cache flush. This mechanism is\nknown to be safe on common Windows filesystems and expected to be safe on\nmacOS. Some linux filesystems, such as XFS, will probably do the right thing\nas well. See [1] for previous discussion on the predecessor of this patch\nseries.\n\nThis series is important on Windows, where loose-objects are included in the\nfsync set by default in Git-For-Windows. In this series, I'm also setting\nthe default mode for Windows to turn on loose object fsyncing with batch\nmode, so that we can get CI coverage of the actual git-for-windows\nconfiguration upstream. We still don't actually issue fsyncs for the test\nsuite since GIT_TEST_FSYNC is set to 0, but we exercise all of the\nsurrounding batch mode code.\n\nThis work is based on 'next' at c54b8eb302. It's dependent on\nns/core-fsyncmethod.\n\n[1]\nhttps://lore.kernel.org/git/2c1ddef6057157d85da74a7274e03eacf0374e45.1629856293.git.gitgitgadget@gmail.com/\n\nNeeraj Singh (14):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  object-file: pass filename to fsync_or_die\n  core.fsyncmethod: batched disk flushes for loose-objects\n  cache-tree: use ODB transaction around writing a tree\n  builtin/add: add ODB transaction around add_files_to_cache\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  test-lib-functions: add parsing helpers for ls-files and ls-tree\n  core.fsyncmethod: tests for batch mode\n  t/perf: add iteration setup mechanism to perf-lib\n  core.fsyncmethod: performance tests for batch mode\n  core.fsyncmethod: correctly camel-case warning message\n\n Documentation/config/core.txt          |   8 ++\n builtin/add.c                          |  13 +++-\n builtin/unpack-objects.c               |   3 +\n builtin/update-index.c                 |  24 ++++++\n bulk-checkin.c                         | 101 ++++++++++++++++++++++---\n bulk-checkin.h                         |  17 ++++-\n cache-tree.c                           |   3 +\n cache.h                                |  12 ++-\n compat/mingw.h                         |   3 +\n config.c                               |   6 +-\n git-compat-util.h                      |   2 +\n object-file.c                          |  15 ++--\n t/lib-unique-files.sh                  |  34 +++++++++\n t/perf/p0008-odb-fsync.sh              |  81 ++++++++++++++++++++\n t/perf/p4220-log-grep-engines.sh       |   3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |   3 +-\n t/perf/p5302-pack-index.sh             |  15 ++--\n t/perf/p7519-fsmonitor.sh              |  18 +----\n t/perf/p7820-grep-engines.sh           |   6 +-\n t/perf/perf-lib.sh                     |  62 +++++++++++++--\n t/t3700-add.sh                         |  28 +++++++\n t/t3903-stash.sh                       |  20 +++++\n t/t5300-pack-object.sh                 |  41 ++++++----\n t/t5317-pack-objects-filter-objects.sh |  91 +++++++++++-----------\n t/test-lib-functions.sh                |  10 +++\n 25 files changed, 501 insertions(+), 118 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p0008-odb-fsync.sh\n\n\nbase-commit: c54b8eb302ffb72f31e73a26044c8a864e2cb307\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1134%2Fneerajsi-msft%2Fns%2Fbatched-fsync-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1134/neerajsi-msft/ns/batched-fsync-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/1134\n\nRange-diff vs v4:\n\n  1:  c7a2a7efe6d =  1:  c7a2a7efe6d bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  2:  d045b13795b =  2:  d045b13795b bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  3:  2d1bc4568ac =  3:  2d1bc4568ac object-file: pass filename to fsync_or_die\n  4:  9e7ae22fa4a =  4:  9e7ae22fa4a core.fsyncmethod: batched disk flushes for loose-objects\n  5:  83fa4a5f3a5 =  5:  83fa4a5f3a5 cache-tree: use ODB transaction around writing a tree\n  -:  ----------- >  6:  d514842ad49 builtin/add: add ODB transaction around add_files_to_cache\n  6:  f03ebee695a =  7:  8cac94598a5 update-index: use the bulk-checkin infrastructure\n  7:  d85013f7d2c =  8:  523e5fbd63e unpack-objects: use the bulk-checkin infrastructure\n  8:  73e54f94c20 =  9:  faacc19aab2 core.fsync: use batch mode and sync loose objects by default on Windows\n  9:  124450c86d9 = 10:  4de7300a7b0 test-lib-functions: add parsing helpers for ls-files and ls-tree\n 10:  282fbdef792 = 11:  1a4aff8c350 core.fsyncmethod: tests for batch mode\n 11:  ee7ecf4cabe ! 12:  47cc63e1dda t/perf: add iteration setup mechanism to perf-lib\n     @@ t/perf/perf-lib.sh: exit $ret' >&3 2>&4\n       }\n       \n       test_wrapper_ () {\n     --\ttest_wrapper_func_=$1; shift\n     -+\tlocal test_wrapper_func_=$1; shift\n     -+\tlocal test_title_=$1; shift\n     ++\tlocal test_wrapper_func_ test_title_\n     + \ttest_wrapper_func_=$1; shift\n     ++\ttest_title_=$1; shift\n       \ttest_start_\n      -\ttest \"$#\" = 3 && { test_prereq=$1; shift; } || test_prereq=\n      -\ttest \"$#\" = 2 ||\n     @@ t/perf/perf-lib.sh: exit $ret' >&3 2>&4\n       \texport test_prereq\n      -\tif ! test_skip \"$@\"\n      +\texport test_perf_setup_\n     ++\n      +\tif ! test_skip \"$test_title_\" \"$@\"\n       \tthen\n       \t\tbase=$(basename \"$0\" .sh)\n 12:  fdf90d45f52 ! 13:  26be6ecb28b core.fsyncmethod: performance tests for add and stash\n     @@ Metadata\n      Author: Neeraj Singh <neerajsi@microsoft.com>\n      \n       ## Commit message ##\n     -    core.fsyncmethod: performance tests for add and stash\n     +    core.fsyncmethod: performance tests for batch mode\n      \n     -    Add basic performance tests for \"git add\" and \"git stash\" of a lot of\n     -    new objects with various fsync settings. This shows the benefit of batch\n     -    mode relative to full fsync.\n     +    Add basic performance tests for git commands that can add data to the\n     +    object database. We cover:\n     +    * git add\n     +    * git stash\n     +    * git update-index (via git stash)\n     +    * git unpack-objects\n     +    * git commit --all\n     +\n     +    We cover all currently available fsync methods as well.\n      \n          Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n      \n     - ## t/perf/p3700-add.sh (new) ##\n     + ## t/perf/p0008-odb-fsync.sh (new) ##\n      @@\n      +#!/bin/sh\n      +#\n     -+# This test measures the performance of adding new files to the object database\n     -+# and index. The test was originally added to measure the effect of the\n     -+# core.fsyncMethod=batch mode, which is why we are testing different values\n     -+# of that setting explicitly and creating a lot of unique objects.\n     ++# This test measures the performance of adding new files to the object\n     ++# database. The test was originally added to measure the effect of the\n     ++# core.fsyncMethod=batch mode, which is why we are testing different values of\n     ++# that setting explicitly and creating a lot of unique objects.\n      +\n      +test_description=\"Tests performance of adding things to the object database\"\n      +\n     -+# Fsync is normally turned off for the test suite.\n     -+GIT_TEST_FSYNC=1\n     -+export GIT_TEST_FSYNC\n     -+\n      +. ./perf-lib.sh\n      +\n      +. $TEST_DIRECTORY/lib-unique-files.sh\n     @@ t/perf/p3700-add.sh (new)\n      +files_per_dir=50\n      +total_files=$((dir_count * files_per_dir))\n      +\n     -+for mode in false true batch\n     -+do\n     -+\tcase $mode in\n     -+\tfalse)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=-loose-object -c core.fsyncmethod=fsync'\n     -+\t\t;;\n     -+\ttrue)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=fsync'\n     -+\t\t;;\n     -+\tbatch)\n     -+\t\tFSYNC_CONFIG='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n     -+\t\t;;\n     -+\tesac\n     ++populate_files () {\n     ++\ttest_create_unique_files $dir_count $files_per_dir files\n     ++}\n     ++\n     ++setup_repo () {\n     ++\t(rm -rf .git || 1) &&\n     ++\tgit init &&\n     ++\ttest_commit first &&\n     ++\tpopulate_files\n     ++}\n     ++\n     ++test_perf_fsync_cfgs () {\n     ++\tlocal method cfg &&\n     ++\tfor method in none fsync batch writeout-only\n     ++\tdo\n     ++\t\tcase $method in\n     ++\t\tnone)\n     ++\t\t\tcfg=\"-c core.fsync=none\"\n     ++\t\t\t;;\n     ++\t\t*)\n     ++\t\t\tcfg=\"-c core.fsync=loose-object -c core.fsyncMethod=$method\"\n     ++\t\tesac &&\n     ++\n     ++\t\t# Set GIT_TEST_FSYNC=1 explicitly since fsync is normally\n     ++\t\t# disabled by t/test-lib.sh.\n     ++\t\tif ! test_perf \"$1 (fsyncMethod=$method)\" \\\n     ++\t\t\t\t\t\t--setup \"$2\" \\\n     ++\t\t\t\t\t\t\"GIT_TEST_FSYNC=1 git $cfg $3\"\n     ++\t\tthen\n     ++\t\t\tbreak\n     ++\t\tfi\n     ++\tdone\n     ++}\n     ++\n     ++test_perf_fsync_cfgs \"add $total_files files\" \\\n     ++\t\"setup_repo\" \\\n     ++\t\"add -- files\"\n     ++\n     ++test_perf_fsync_cfgs \"stash $total_files files\" \\\n     ++\t\"setup_repo\" \\\n     ++\t\"stash push -u -- files\"\n      +\n     -+\ttest_perf \"add $total_files files (object_fsyncing=$mode)\" \\\n     -+\t\t--setup \"\n     -+\t\t(rm -rf .git || 1) &&\n     -+\t\tgit init &&\n     -+\t\ttest_create_unique_files $dir_count $files_per_dir files_$mode\n     -+\t\" \"\n     -+\t\tgit $FSYNC_CONFIG add files_$mode\n     ++test_perf_fsync_cfgs \"unpack $total_files files\" \\\n      +\t\"\n     ++\tsetup_repo &&\n     ++\tgit -c core.fsync=none add -- files &&\n     ++\tgit -c core.fsync=none commit -q -m second &&\n     ++\techo HEAD | git pack-objects -q --stdout --revs >test_pack.pack &&\n     ++\tsetup_repo\n     ++\t\" \\\n     ++\t\"unpack-objects -q <test_pack.pack\"\n      +\n     -+\ttest_perf \"stash $total_files files (object_fsyncing=$mode)\" \\\n     -+\t\t--setup \"\n     -+\t\t(rm -rf .git || 1) &&\n     -+\t\tgit init &&\n     -+\t\ttest_commit first &&\n     -+\t\ttest_create_unique_files $dir_count $files_per_dir stash_files_$mode\n     -+\t\" \"\n     -+\t\tgit $FSYNC_CONFIG stash push -u -- stash_files_$mode\n     ++test_perf_fsync_cfgs \"commit $total_files files\" \\\n      +\t\"\n     -+done\n     ++\tsetup_repo &&\n     ++\tgit -c core.fsync=none add -- files &&\n     ++\tpopulate_files\n     ++\t\" \\\n     ++\t\"commit -q -a -m test\"\n      +\n      +test_done\n 13:  fb30bd02c8d = 14:  88c1f84d4c3 core.fsyncmethod: correctly camel-case warning message\n\n-- \ngitgitgadget\n"},{"id":"452642","messageId":"d045b13795b38caa27f8e25340212f736b66bb05.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 02/14] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:20Z","receivedAt":"2022-03-30T05:05:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nMake it clearer in the naming and documentation of the plug_bulk_checkin\nand unplug_bulk_checkin APIs that they can be thought of as\na \"transaction\" to optimize operations on the object database. These\ntransactions may be nested so that subsystems like the cache-tree\nwriting code can optimize their operations without caring whether the\ntop-level code has a transaction active.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c  |  4 ++--\n bulk-checkin.c | 20 ++++++++++++--------\n bulk-checkin.h | 14 ++++++++++++--\n 3 files changed, 26 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..9bf37ceae8e 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\tstring_list_clear(&only_match_skip_worktree, 0);\n \t}\n \n-\tplug_bulk_checkin();\n+\tbegin_odb_transaction();\n \n \tif (add_renormalize)\n \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n@@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\tend_odb_transaction();\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 577b135e39c..8b0fd5c7723 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,7 +10,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static int bulk_checkin_plugged;\n+static int odb_transaction_nesting;\n \n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n@@ -280,21 +280,25 @@ int index_bulk_checkin(struct object_id *oid,\n {\n \tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!bulk_checkin_plugged)\n+\tif (!odb_transaction_nesting)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n-void plug_bulk_checkin(void)\n+void begin_odb_transaction(void)\n {\n-\tassert(!bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 1;\n+\todb_transaction_nesting += 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void end_odb_transaction(void)\n {\n-\tassert(bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 0;\n+\todb_transaction_nesting -= 1;\n+\tif (odb_transaction_nesting < 0)\n+\t\tBUG(\"Unbalanced ODB transaction nesting\");\n+\n+\tif (odb_transaction_nesting)\n+\t\treturn;\n+\n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..69a94422ac7 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -10,7 +10,17 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n-void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+/*\n+ * Tell the object database to optimize for adding\n+ * multiple objects. end_odb_transaction must be called\n+ * to make new objects visible.\n+ */\n+void begin_odb_transaction(void);\n+\n+/*\n+ * Tell the object database to make any objects from the\n+ * current transaction visible.\n+ */\n+void end_odb_transaction(void);\n \n #endif\n-- \ngitgitgadget\n\n"},{"id":"452643","messageId":"c7a2a7efe6d532fc7fce1352b1dfce640cc9f2f6.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:19Z","receivedAt":"2022-03-30T05:05:48Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit prepares for adding batch-fsync to the bulk-checkin\ninfrastructure.\n\nThe bulk-checkin infrastructure is currently used to batch up addition\nof large blobs to a packfile. When a blob is larger than\nbig_file_threshold, we unconditionally add it to a pack. If bulk\ncheckins are 'plugged', we allow multiple large blobs to be added to a\nsingle pack until we reach the packfile size limit; otherwise, we simply\nmake a new packfile for each large blob. The 'unplug' call tells us when\nthe series of blob additions is done so that we can finish the packfiles\nand make their objects available to subsequent operations.\n\nStated another way, bulk-checkin allows callers to define a transaction\nthat adds multiple objects to the object database, where the object\ndatabase can optimize its internal operations within the transaction\nboundary.\n\nBatched fsync will fit into bulk-checkin by taking advantage of the\nplug/unplug functionality to determine the appropriate time to fsync\nand make newly-added objects available in the primary object database.\n\n* Rename 'state' variable to 'bulk_checkin_state', since we will later\n  be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n  find in the debugger, since the name is more unique.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 22 ++++++++++++----------\n 1 file changed, 12 insertions(+), 10 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c9..577b135e39c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_state;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -278,21 +278,23 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tif (bulk_checkin_state.f)\n+\t\tfinish_bulk_checkin(&bulk_checkin_state);\n }\n-- \ngitgitgadget\n\n"},{"id":"452644","messageId":"2d1bc4568ac744f11c886a5f964dbe563c04ce8b.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 03/14] object-file: pass filename to fsync_or_die","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:21Z","receivedAt":"2022-03-30T05:05:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nIf we die while trying to fsync a loose object file, pass the actual\nfilename we're trying to sync. This is likely to be more helpful for a\nuser trying to diagnose the cause of the failure than the former\n'loose object file' string. It also sidesteps any concerns about\ntranslating the die message differently for loose objects versus\nsomething else that has a real path.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n object-file.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b254bc50d70..5ffbf3d4fd4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1888,16 +1888,16 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n }\n \n /* Finalize a file on disk, and close it. */\n-static void close_loose_object(int fd)\n+static void close_loose_object(int fd, const char *filename)\n {\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n \tif (fsync_object_files > 0)\n-\t\tfsync_or_die(fd, \"loose object file\");\n+\t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n-\t\t\t\t       \"loose object file\");\n+\t\t\t\t       filename);\n \n out:\n \tif (close(fd) != 0)\n@@ -2011,7 +2011,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n+\tclose_loose_object(fd, tmp_file.buf);\n \n \tif (mtime) {\n \t\tstruct utimbuf utb;\n-- \ngitgitgadget\n\n"},{"id":"452646","messageId":"9e7ae22fa4a2693fe26659f875dd780080c4cfb2.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:22Z","receivedAt":"2022-03-30T05:05:55Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin odb-transaction functionality, takes\nadvantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1a. Create the object in a tmp-objdir.\n2a. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1b. Issue an fsync against a dummy file to flush the log and hardware\n   writeback cache, which should by now have seen the tmp-objdir writes.\n2b. Rename all of the tmp-objdir files to their final names.\n3b. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns. This sequence also ensures that no object files\nappear in the main object store unless they are fsync-durable.\n\nBatch mode is only enabled if core.fsync includes loose-objects. If\nthe legacy core.fsyncObjectFiles setting is enabled, but core.fsync does\nnot include loose-objects, we will use file-by-file fsyncing.\n\nIn step (1a) of the sequence, the tmp-objdir is created lazily to avoid\nwork if no loose objects are ever added to the ODB. We use a tmp-objdir\nto maintain the invariant that no loose-objects are visible in the main\nODB unless they are properly fsync-durable. This is important since\nfuture ODB operations that try to create an object with specific\ncontents will silently drop the new data if an object with the target\nhash exists without checking that the loose-object contents match the\nhash. Only a full git-fsck would restore the ODB to a functional state\nwhere dataloss doesn't occur.\n\nIn step (1b) of the sequence, we issue a fsync against a dummy file\ncreated specifically for the purpose. This method has a little higher\ncost than using one of the input object files, but makes adding new\ncallers of this mechanism easier, since we don't need to figure out\nwhich object file is \"last\" or risk sharing violations by caching the fd\nof the last object file.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  8 ++++\n bulk-checkin.c                | 71 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  3 ++\n cache.h                       |  8 +++-\n config.c                      |  2 +\n object-file.c                 |  7 +++-\n 6 files changed, 97 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 9da3e5d88f6..3c90ba0b395 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -596,6 +596,14 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then does a single full fsync of\n+  a dummy file to trigger the disk cache flush at the end of the operation.\n++\n+  Currently `batch` mode only applies to loose-object files. Other repository\n+  data is made durable as if `fsync` was specified. This mode is expected to\n+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+  and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8b0fd5c7723..9799d247cad 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int odb_transaction_nesting;\n \n+static struct tmp_objdir *bulk_fsync_objdir;\n+\n static struct bulk_checkin_state {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -80,6 +85,40 @@ clear_exit:\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void do_batch_fsync(void)\n+{\n+\tstruct strbuf temp_path = STRBUF_INIT;\n+\tstruct tempfile *temp;\n+\n+\tif (!bulk_fsync_objdir)\n+\t\treturn;\n+\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur. The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware or any filesystem logs. This fsync call acts as a barrier\n+\t * to ensure that the data in each new object file is durable before\n+\t * the final name is visible.\n+\t */\n+\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\ttemp = xmks_tempfile(temp_path.buf);\n+\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\tdelete_tempfile(&temp);\n+\tstrbuf_release(&temp_path);\n+\n+\t/*\n+\t * Make the object files visible in the primary ODB after their data is\n+\t * fully durable.\n+\t */\n+\ttmp_objdir_migrate(bulk_fsync_objdir);\n+\tbulk_fsync_objdir = NULL;\n+}\n+\n static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +313,36 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \treturn 0;\n }\n \n+void prepare_loose_object_bulk_checkin(void)\n+{\n+\t/*\n+\t * We lazily create the temporary object directory\n+\t * the first time an object might be added, since\n+\t * callers may not know whether any objects will be\n+\t * added at the time they call begin_odb_transaction.\n+\t */\n+\tif (!odb_transaction_nesting || bulk_fsync_objdir)\n+\t\treturn;\n+\n+\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+}\n+\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n+{\n+\t/*\n+\t * If we have an active ODB transaction, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before as part of do_batch_fsync.\n+\t */\n+\tif (!bulk_fsync_objdir ||\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n+\t\tfsync_or_die(fd, filename);\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -301,4 +370,6 @@ void end_odb_transaction(void)\n \n \tif (bulk_checkin_state.f)\n \t\tfinish_bulk_checkin(&bulk_checkin_state);\n+\n+\tdo_batch_fsync();\n }\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 69a94422ac7..70edf745be8 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,9 @@\n \n #include \"cache.h\"\n \n+void prepare_loose_object_bulk_checkin(void);\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex ef7d34b7a09..a5bf15a5131 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH,\n };\n \n extern enum fsync_method fsync_method;\n@@ -1767,6 +1768,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex 3c9b6b589ab..511f4584eeb 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd4..d2e0c13198f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1893,7 +1893,9 @@ static void close_loose_object(int fd, const char *filename)\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (fsync_object_files > 0)\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd, filename);\n+\telse if (fsync_object_files > 0)\n \t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n@@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n-- \ngitgitgadget\n\n"},{"id":"452645","messageId":"83fa4a5f3a5c79fa814932c0705867ff16a584c7.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 05/14] cache-tree: use ODB transaction around writing a tree","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:23Z","receivedAt":"2022-03-30T05:05:56Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTake advantage of the odb transaction infrastructure around writing the\ncached tree to the object database.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache-tree.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6752f69d515..8c5e8822716 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -3,6 +3,7 @@\n #include \"tree.h\"\n #include \"tree-walk.h\"\n #include \"cache-tree.h\"\n+#include \"bulk-checkin.h\"\n #include \"object-store.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n@@ -474,8 +475,10 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \ttrace_performance_enter();\n \ttrace2_region_enter(\"cache_tree\", \"update\", the_repository);\n+\tbegin_odb_transaction();\n \ti = update_one(istate->cache_tree, istate->cache, istate->cache_nr,\n \t\t       \"\", 0, &skip, flags);\n+\tend_odb_transaction();\n \ttrace2_region_leave(\"cache_tree\", \"update\", the_repository);\n \ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n-- \ngitgitgadget\n\n"},{"id":"452647","messageId":"8cac94598a58704d9b625a9d8a593779f7adc30f.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 07/14] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:25Z","receivedAt":"2022-03-30T05:06:06Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables odb-transactions for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete, so callers\nusing the --stdin option will not see them until update-index is done.\nThis risk is mitigated by not keeping an ODB transaction open around\n--stdin processing if in --verbose mode. Without --verbose mode,\na caller feeding update-index via --stdin wouldn't know when\nupdate-index adds an object, event without an ODB transaction.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 24 ++++++++++++++++++++++++\n 1 file changed, 24 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex aafe7eeac2a..50f9063e1c6 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -1116,6 +1117,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t */\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n+\n+\t/*\n+\t * Allow the object layer to optimize adding multiple objects in\n+\t * a batch.\n+\t */\n+\tbegin_odb_transaction();\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1167,6 +1174,17 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tthe_index.version = preferred_index_format;\n \t}\n \n+\t/*\n+\t * It is possible, though unlikely, that a caller could use the verbose\n+\t * output to synchronize with addition of objects to the object\n+\t * database. The current implementation of ODB transactions leaves\n+\t * objects invisible while a transaction is active, so end the\n+\t * transaction here if verbose output is enabled.\n+\t */\n+\n+\tif (verbose)\n+\t\tend_odb_transaction();\n+\n \tif (read_from_stdin) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tstruct strbuf unquoted = STRBUF_INIT;\n@@ -1190,6 +1208,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/*\n+\t * By now we have added all of the new objects\n+\t */\n+\tif (!verbose)\n+\t\tend_odb_transaction();\n+\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \ngitgitgadget\n\n"},{"id":"452648","messageId":"faacc19aab2ecde1eb4134d1514b65a1f8ea6791.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 09/14] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:27Z","receivedAt":"2022-03-30T05:06:08Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the most of the test suite, since\nGIT_TEST_FSYNC is set to 0. However, we do exercise all of the\nsurrounding batch mode code since GIT_TEST_FSYNC merely makes the\nmaybe_fsync wrapper always appear to succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex a5bf15a5131..7f6cbb254b4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 511f4584eeb..e9cac5f4707 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 0892e209a2f..fffe42ce7c1 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \ngitgitgadget\n\n"},{"id":"452649","messageId":"4de7300a7b0061c1399738c66ce05bfbbe2db1d0.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 10/14] test-lib-functions: add parsing helpers for ls-files and ls-tree","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:28Z","receivedAt":"2022-03-30T05:06:10Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nSeveral tests use awk to parse OIDs from the output of 'git ls-files\n--stage' and 'git ls-tree'. Introduce helpers to centralize these uses\nof awk.\n\nUpdate t5317-pack-objects-filter-objects.sh to use the new ls-files\nhelper so that it has some usages to review. Other updates are left for\nthe future.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/t5317-pack-objects-filter-objects.sh | 91 +++++++++++++-------------\n t/test-lib-functions.sh                | 10 +++\n 2 files changed, 54 insertions(+), 47 deletions(-)\n\ndiff --git a/t/t5317-pack-objects-filter-objects.sh b/t/t5317-pack-objects-filter-objects.sh\nindex 33b740ce628..bb633c9b099 100755\n--- a/t/t5317-pack-objects-filter-objects.sh\n+++ b/t/t5317-pack-objects-filter-objects.sh\n@@ -10,9 +10,6 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n # Test blob:none filter.\n \n test_expect_success 'setup r1' '\n-\techo \"{print \\$1}\" >print_1.awk &&\n-\techo \"{print \\$2}\" >print_2.awk &&\n-\n \tgit init r1 &&\n \tfor n in 1 2 3 4 5\n \tdo\n@@ -22,10 +19,13 @@ test_expect_success 'setup r1' '\n \tdone\n '\n \n+parse_verify_pack_blob_oid () {\n+\tawk '{print $1}' -\n+}\n+\n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r1 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -35,7 +35,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -54,12 +54,12 @@ test_expect_success 'verify blob:none packfile has no blobs' '\n test_expect_success 'verify normal and blob:none packfiles have same commits/trees' '\n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r1 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -123,8 +123,8 @@ test_expect_success 'setup r2' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -134,7 +134,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -161,8 +161,8 @@ test_expect_success 'verify blob:limit=1000' '\n '\n \n test_expect_success 'verify blob:limit=1001' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1001 >filter.pack <<-EOF &&\n@@ -172,15 +172,15 @@ test_expect_success 'verify blob:limit=1001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=10001' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=10001 >filter.pack <<-EOF &&\n@@ -190,15 +190,15 @@ test_expect_success 'verify blob:limit=10001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1k' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1k >filter.pack <<-EOF &&\n@@ -208,15 +208,15 @@ test_expect_success 'verify blob:limit=1k' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify explicitly specifying oversized blob in input' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \techo HEAD >objects &&\n@@ -226,15 +226,15 @@ test_expect_success 'verify explicitly specifying oversized blob in input' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1m' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1m >filter.pack <<-EOF &&\n@@ -244,7 +244,7 @@ test_expect_success 'verify blob:limit=1m' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -253,12 +253,12 @@ test_expect_success 'verify blob:limit=1m' '\n test_expect_success 'verify normal and blob:limit packfiles have same commits/trees' '\n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -289,9 +289,8 @@ test_expect_success 'setup r3' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r3 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -301,7 +300,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r3 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -342,9 +341,8 @@ test_expect_success 'setup r4' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -354,19 +352,19 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r4 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=OID' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 ls-files -s pattern >staged &&\n-\toid=$(awk -f print_2.awk staged) &&\n+\toid=$(test_parse_ls_files_stage_oids <staged) &&\n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=$oid >filter.pack <<-EOF &&\n \tHEAD\n \tEOF\n@@ -374,15 +372,15 @@ test_expect_success 'verify sparse:oid=OID' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=oid-ish' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=main:pattern >filter.pack <<-EOF &&\n@@ -392,7 +390,7 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -402,9 +400,8 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n # This models previously omitted objects that we did not receive.\n \n test_expect_success 'setup r1 - delete loose blobs' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tfor id in `cat expected | sed \"s|..|&/|\"`\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex a027f0c409e..e6011409e2f 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1782,6 +1782,16 @@ test_oid_to_path () {\n \techo \"${1%$basename}/$basename\"\n }\n \n+# Parse oids from git ls-files --staged output\n+test_parse_ls_files_stage_oids () {\n+\tawk '{print $2}' -\n+}\n+\n+# Parse oids from git ls-tree output\n+test_parse_ls_tree_oids () {\n+\tawk '{print $3}' -\n+}\n+\n # Choose a port number based on the test script's number and store it in\n # the given variable name, unless that variable already contains a number.\n test_set_port () {\n-- \ngitgitgadget\n\n"},{"id":"452650","messageId":"523e5fbd63ef85035131bd4cec7565707c290e84.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 08/14] unpack-objects: use the bulk-checkin infrastructure","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:26Z","receivedAt":"2022-03-30T05:06:11Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling an odb-transaction when unpacking objects, we can take advantage\nof batched fsyncs.\n\nHere are some performance numbers to justify batch mode for\nunpack-objects, collected on a WSL2 Ubuntu VM.\n\nFsync Mode | Time for 90 objects (ms)\n-------------------------------------\n       Off | 170\n  On,fsync | 760\n  On,batch | 230\n\nNote that the default unpackLimit is 100 objects, so there's a 3x\nbenefit in the worst case. The non-batch mode fsync scales linearly\nwith the number of objects, so there are significant benefits even with\nsmaller numbers of objects.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..56d05e2725d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tbegin_odb_transaction();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tend_odb_transaction();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \ngitgitgadget\n\n"},{"id":"452651","messageId":"d514842ad493a819e3640ecb658f702e530d6e85.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 06/14] builtin/add: add ODB transaction around add_files_to_cache","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:24Z","receivedAt":"2022-03-30T05:06:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe add_files_to_cache function is invoked internally by\nbuiltin/commit.c and builtin/checkout.c for their flags that stage\nmodified files before doing the larger operation. These commands\ncan benefit from batched fsyncing.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c | 9 +++++++++\n 1 file changed, 9 insertions(+)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 9bf37ceae8e..e39770e4746 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -141,7 +141,16 @@ int add_files_to_cache(const char *prefix,\n \trev.diffopt.format_callback_data = &data;\n \trev.diffopt.flags.override_submodule_config = 1;\n \trev.max_count = 0; /* do not compare unmerged paths with stage #2 */\n+\n+\t/*\n+\t * Use an ODB transaction to optimize adding multiple objects.\n+\t * This function is invoked from commands other than 'add', which\n+\t * may not have their own transaction active.\n+\t */\n+\tbegin_odb_transaction();\n \trun_diff_files(&rev, DIFF_RACY_IS_MODIFIED);\n+\tend_odb_transaction();\n+\n \tclear_pathspec(&rev.prune_data);\n \treturn !!data.add_errors;\n }\n-- \ngitgitgadget\n\n"},{"id":"452652","messageId":"1a4aff8c350b5ffe3c7760faa4accc88c83ce11c.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 11/14] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:29Z","receivedAt":"2022-03-30T05:06:17Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added\ndue to it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 34 ++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n t/t3903-stash.sh       | 20 ++++++++++++++++++++\n t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n 4 files changed, 109 insertions(+), 14 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..34c01a65256\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,34 @@\n+# Helper to create files with unique contents\n+\n+# Create multiple files with unique contents within this test run. Takes the\n+# number of directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with contents\n+#\t\t\t\t\t different from previous invocations\n+#\t\t\t\t\t of this command in this run.\n+\n+test_create_unique_files () {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=\"$1\" &&\n+\tlocal files=\"$2\" &&\n+\tlocal basedir=\"$3\" &&\n+\tlocal counter=0 &&\n+\tlocal i &&\n+\tlocal j &&\n+\ttest_tick &&\n+\tlocal basedata=$basedir$test_tick &&\n+\trm -rf \"$basedir\" &&\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=$basedir/dir$i &&\n+\t\tmkdir -p \"$dir\" &&\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1)) &&\n+\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..8979c8a5f03 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,32 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir1 &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION add -- ./files_base_dir1/ &&\n+\tgit ls-files --stage files_base_dir1/ |\n+\ttest_parse_ls_files_stage_oids >added_files_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files_oids >added_files_actual &&\n+\ttest_cmp added_files_oids added_files_actual\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir2 &&\n+\tfind files_base_dir2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\tgit ls-files --stage files_base_dir2 |\n+\ttest_parse_ls_files_stage_oids >added_files2_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files2_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files2_oids >added_files2_actual &&\n+\ttest_cmp added_files2_oids added_files2_actual\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 4abbc8fccae..20e94881964 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1410,6 +1411,25 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION stash push -u -- ./files_base_dir/ &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./files_base_dir/ |\n+\ttest_parse_ls_tree_oids >stashed_files_oids &&\n+\n+\t# We created 2 dirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 stashed_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <stashed_files_oids >stashed_files_actual &&\n+\ttest_cmp stashed_files_oids stashed_files_actual\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..f8a0f309e2d 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -161,22 +161,27 @@ test_expect_success 'pack-objects with bogus arguments' '\n '\n \n check_unpack () {\n+\tlocal packname=\"$1\" &&\n+\tlocal object_list=\"$2\" &&\n+\tlocal git_config=\"$3\" &&\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $git_config init --bare git2 &&\n+\t(\n+\t\tgit $git_config -C git2 unpack-objects -n <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 unpack-objects <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <\"$object_list\" >current &&\n+\tcmp \"$object_list\" current\n }\n \n test_expect_success 'unpack without delta' '\n-\tcheck_unpack test-1-${packname_1}\n+\tcheck_unpack test-1-${packname_1} obj-list\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with REF_DELTA' '\n@@ -185,7 +190,11 @@ test_expect_success 'pack with REF_DELTA' '\n '\n \n test_expect_success 'unpack with REF_DELTA' '\n-\tcheck_unpack test-2-${packname_2}\n+\tcheck_unpack test-2-${packname_2} obj-list\n+'\n+\n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with OFS_DELTA' '\n@@ -195,7 +204,11 @@ test_expect_success 'pack with OFS_DELTA' '\n '\n \n test_expect_success 'unpack with OFS_DELTA' '\n-\tcheck_unpack test-3-${packname_3}\n+\tcheck_unpack test-3-${packname_3} obj-list\n+'\n+\n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'compare delta flavors' '\n-- \ngitgitgadget\n\n"},{"id":"452653","messageId":"47cc63e1dda6fa6417ff6bace4e9e371aaae0d7d.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 12/14] t/perf: add iteration setup mechanism to perf-lib","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:30Z","receivedAt":"2022-03-30T05:06:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTests that affect the repo in stateful ways are easier to write if we\ncan run setup steps outside of the measured portion of perf iteration.\n\nThis change adds a \"--setup 'setup-script'\" parameter to test_perf. To\nmake invocations easier to understand, I also moved the prerequisites to\na new --prereq parameter.\n\nThe setup facility will be used in the upcoming perf tests for batch\nmode, but it already helps in some existing tests, like t5302 and t7820.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p4220-log-grep-engines.sh       |  3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |  3 +-\n t/perf/p5302-pack-index.sh             | 15 +++----\n t/perf/p7519-fsmonitor.sh              | 18 ++------\n t/perf/p7820-grep-engines.sh           |  6 ++-\n t/perf/perf-lib.sh                     | 62 +++++++++++++++++++++++---\n 6 files changed, 74 insertions(+), 33 deletions(-)\n\ndiff --git a/t/perf/p4220-log-grep-engines.sh b/t/perf/p4220-log-grep-engines.sh\nindex 2bc47ded4d1..03fbfbb85d3 100755\n--- a/t/perf/p4220-log-grep-engines.sh\n+++ b/t/perf/p4220-log-grep-engines.sh\n@@ -36,7 +36,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4220_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p4221-log-grep-engines-fixed.sh b/t/perf/p4221-log-grep-engines-fixed.sh\nindex 060971265a9..0a6d6dfc219 100755\n--- a/t/perf/p4221-log-grep-engines-fixed.sh\n+++ b/t/perf/p4221-log-grep-engines-fixed.sh\n@@ -26,7 +26,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4221_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p5302-pack-index.sh b/t/perf/p5302-pack-index.sh\nindex c16f6a3ff69..14c601bbf86 100755\n--- a/t/perf/p5302-pack-index.sh\n+++ b/t/perf/p5302-pack-index.sh\n@@ -26,9 +26,8 @@ test_expect_success 'set up thread-counting tests' '\n \tdone\n '\n \n-test_perf PERF_EXTRA 'index-pack 0 threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack 0 threads' --prereq PERF_EXTRA \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --threads=1 --stdin < $PACK\n '\n \n@@ -36,17 +35,15 @@ for t in $threads\n do\n \tTHREADS=$t\n \texport THREADS\n-\ttest_perf PERF_EXTRA \"index-pack $t threads\" '\n-\t\trm -rf repo.git &&\n-\t\tgit init --bare repo.git &&\n+\ttest_perf \"index-pack $t threads\" --prereq PERF_EXTRA \\\n+\t\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \t\tGIT_DIR=repo.git GIT_FORCE_THREADS=1 \\\n \t\tgit index-pack --threads=$THREADS --stdin <$PACK\n \t'\n done\n \n-test_perf 'index-pack default number of threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack default number of threads' \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --stdin < $PACK\n '\n \ndiff --git a/t/perf/p7519-fsmonitor.sh b/t/perf/p7519-fsmonitor.sh\nindex c8be58f3c76..5b489c968b8 100755\n--- a/t/perf/p7519-fsmonitor.sh\n+++ b/t/perf/p7519-fsmonitor.sh\n@@ -60,18 +60,6 @@ then\n \tesac\n fi\n \n-if test -n \"$GIT_PERF_7519_DROP_CACHE\"\n-then\n-\t# When using GIT_PERF_7519_DROP_CACHE, GIT_PERF_REPEAT_COUNT must be 1 to\n-\t# generate valid results. Otherwise the caching that happens for the nth\n-\t# run will negate the validity of the comparisons.\n-\tif test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n-\tthen\n-\t\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n-\t\tGIT_PERF_REPEAT_COUNT=1\n-\tfi\n-fi\n-\n trace_start() {\n \tif test -n \"$GIT_PERF_7519_TRACE\"\n \tthen\n@@ -167,10 +155,10 @@ setup_for_fsmonitor() {\n \n test_perf_w_drop_caches () {\n \tif test -n \"$GIT_PERF_7519_DROP_CACHE\"; then\n-\t\ttest-tool drop-caches\n+\t\ttest_perf \"$1\" --setup \"test-tool drop-caches\" \"$2\"\n+\telse\n+\t\ttest_perf \"$@\"\n \tfi\n-\n-\ttest_perf \"$@\"\n }\n \n test_fsmonitor_suite() {\ndiff --git a/t/perf/p7820-grep-engines.sh b/t/perf/p7820-grep-engines.sh\nindex 8b09c5bf328..9bfb86842a9 100755\n--- a/t/perf/p7820-grep-engines.sh\n+++ b/t/perf/p7820-grep-engines.sh\n@@ -49,13 +49,15 @@ do\n \t\tfi\n \t\tif ! test_have_prereq PERF_GREP_ENGINES_THREADS\n \t\tthen\n-\t\t\ttest_perf $prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \"\n+\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \\\n+\t\t\t\t--prereq \"$prereq\" \"\n \t\t\t\tgit -c grep.patternType=$engine grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine' || :\n \t\t\t\"\n \t\telse\n \t\t\tfor threads in $GIT_PERF_GREP_THREADS\n \t\t\tdo\n-\t\t\t\ttest_perf PTHREADS,$prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\" \"\n+\t\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\"\n+\t\t\t\t\t--prereq PTHREADS,$prereq \"\n \t\t\t\t\tgit -c grep.patternType=$engine -c grep.threads=$threads grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine.$threads' || :\n \t\t\t\t\"\n \t\t\tdone\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex 407252bac70..e9cc16d045a 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -189,19 +189,40 @@ exit $ret' >&3 2>&4\n }\n \n test_wrapper_ () {\n+\tlocal test_wrapper_func_ test_title_\n \ttest_wrapper_func_=$1; shift\n+\ttest_title_=$1; shift\n \ttest_start_\n-\ttest \"$#\" = 3 && { test_prereq=$1; shift; } || test_prereq=\n-\ttest \"$#\" = 2 ||\n-\tBUG \"not 2 or 3 parameters to test-expect-success\"\n+\ttest_prereq=\n+\ttest_perf_setup_=\n+\twhile test $# != 0\n+\tdo\n+\t\tcase $1 in\n+\t\t--prereq)\n+\t\t\ttest_prereq=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t--setup)\n+\t\t\ttest_perf_setup_=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t*)\n+\t\t\tbreak\n+\t\t\t;;\n+\t\tesac\n+\t\tshift\n+\tdone\n+\ttest \"$#\" = 1 || BUG \"test_wrapper_ needs 2 positional parameters\"\n \texport test_prereq\n-\tif ! test_skip \"$@\"\n+\texport test_perf_setup_\n+\n+\tif ! test_skip \"$test_title_\" \"$@\"\n \tthen\n \t\tbase=$(basename \"$0\" .sh)\n \t\techo \"$test_count\" >>\"$perf_results_dir\"/$base.subtests\n \t\techo \"$1\" >\"$perf_results_dir\"/$base.$test_count.descr\n \t\tbase=\"$perf_results_dir\"/\"$PERF_RESULTS_PREFIX$(basename \"$0\" .sh)\".\"$test_count\"\n-\t\t\"$test_wrapper_func_\" \"$@\"\n+\t\t\"$test_wrapper_func_\" \"$test_title_\" \"$@\"\n \tfi\n \n \ttest_finish_\n@@ -214,6 +235,16 @@ test_perf_ () {\n \t\techo \"perf $test_count - $1:\"\n \tfi\n \tfor i in $(test_seq 1 $GIT_PERF_REPEAT_COUNT); do\n+\t\tif test -n \"$test_perf_setup_\"\n+\t\tthen\n+\t\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\t\tif ! test_eval_ $test_perf_setup_\n+\t\t\tthen\n+\t\t\t\ttest_failure_ \"$test_perf_setup_\"\n+\t\t\t\tbreak\n+\t\t\tfi\n+\n+\t\tfi\n \t\tsay >&3 \"running: $2\"\n \t\tif test_run_perf_ \"$2\"\n \t\tthen\n@@ -237,11 +268,24 @@ test_perf_ () {\n \trm test_time.*\n }\n \n+# Usage: test_perf 'title' [options] 'perf-test'\n+#\tRun the performance test script specified in perf-test with\n+#\toptional prerequisite and setup steps.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to each measured iteration\n+#\n test_perf () {\n \ttest_wrapper_ test_perf_ \"$@\"\n }\n \n test_size_ () {\n+\tif test -n \"$test_perf_setup_\"\n+\tthen\n+\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\ttest_eval_ $test_perf_setup_\n+\tfi\n+\n \tsay >&3 \"running: $2\"\n \tif test_eval_ \"$2\" 3>\"$base\".result; then\n \t\ttest_ok_ \"$1\"\n@@ -250,6 +294,14 @@ test_size_ () {\n \tfi\n }\n \n+# Usage: test_size 'title' [options] 'size-test'\n+#\tRun the size test script specified in size-test with optional\n+#\tprerequisites and setup steps. Returns the numeric value\n+#\treturned by size-test.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to the size measurement\n+\n test_size () {\n \ttest_wrapper_ test_size_ \"$@\"\n }\n-- \ngitgitgadget\n\n"},{"id":"452654","messageId":"26be6ecb28bc1f76fba380fdd10acf59820df997.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 13/14] core.fsyncmethod: performance tests for batch mode","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:31Z","receivedAt":"2022-03-30T05:06:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for git commands that can add data to the\nobject database. We cover:\n* git add\n* git stash\n* git update-index (via git stash)\n* git unpack-objects\n* git commit --all\n\nWe cover all currently available fsync methods as well.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p0008-odb-fsync.sh | 81 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 81 insertions(+)\n create mode 100755 t/perf/p0008-odb-fsync.sh\n\ndiff --git a/t/perf/p0008-odb-fsync.sh b/t/perf/p0008-odb-fsync.sh\nnew file mode 100755\nindex 00000000000..87092c2627e\n--- /dev/null\n+++ b/t/perf/p0008-odb-fsync.sh\n@@ -0,0 +1,81 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object\n+# database. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values of\n+# that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of adding things to the object database\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_fresh_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+populate_files () {\n+\ttest_create_unique_files $dir_count $files_per_dir files\n+}\n+\n+setup_repo () {\n+\t(rm -rf .git || 1) &&\n+\tgit init &&\n+\ttest_commit first &&\n+\tpopulate_files\n+}\n+\n+test_perf_fsync_cfgs () {\n+\tlocal method cfg &&\n+\tfor method in none fsync batch writeout-only\n+\tdo\n+\t\tcase $method in\n+\t\tnone)\n+\t\t\tcfg=\"-c core.fsync=none\"\n+\t\t\t;;\n+\t\t*)\n+\t\t\tcfg=\"-c core.fsync=loose-object -c core.fsyncMethod=$method\"\n+\t\tesac &&\n+\n+\t\t# Set GIT_TEST_FSYNC=1 explicitly since fsync is normally\n+\t\t# disabled by t/test-lib.sh.\n+\t\tif ! test_perf \"$1 (fsyncMethod=$method)\" \\\n+\t\t\t\t\t\t--setup \"$2\" \\\n+\t\t\t\t\t\t\"GIT_TEST_FSYNC=1 git $cfg $3\"\n+\t\tthen\n+\t\t\tbreak\n+\t\tfi\n+\tdone\n+}\n+\n+test_perf_fsync_cfgs \"add $total_files files\" \\\n+\t\"setup_repo\" \\\n+\t\"add -- files\"\n+\n+test_perf_fsync_cfgs \"stash $total_files files\" \\\n+\t\"setup_repo\" \\\n+\t\"stash push -u -- files\"\n+\n+test_perf_fsync_cfgs \"unpack $total_files files\" \\\n+\t\"\n+\tsetup_repo &&\n+\tgit -c core.fsync=none add -- files &&\n+\tgit -c core.fsync=none commit -q -m second &&\n+\techo HEAD | git pack-objects -q --stdout --revs >test_pack.pack &&\n+\tsetup_repo\n+\t\" \\\n+\t\"unpack-objects -q <test_pack.pack\"\n+\n+test_perf_fsync_cfgs \"commit $total_files files\" \\\n+\t\"\n+\tsetup_repo &&\n+\tgit -c core.fsync=none add -- files &&\n+\tpopulate_files\n+\t\" \\\n+\t\"commit -q -a -m test\"\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"452655","messageId":"88c1f84d4c3f71bb3cbd6e016771196a56f90bd9.1648616734.git.gitgitgadget@gmail.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v5 14/14] core.fsyncmethod: correctly camel-case warning message","fromName":"Neeraj Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-03-30T05:05:32Z","receivedAt":"2022-03-30T05:06:29Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe warning for an unrecognized fsyncMethod was not\ncamel-cased.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n config.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/config.c b/config.c\nindex e9cac5f4707..ae819dee20b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1697,7 +1697,7 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \n \tif (!strcmp(var, \"core.fsyncobjectfiles\")) {\n \t\tif (fsync_object_files < 0)\n-\t\t\twarning(_(\"core.fsyncobjectfiles is deprecated; use core.fsync instead\"));\n+\t\t\twarning(_(\"core.fsyncObjectFiles is deprecated; use core.fsync instead\"));\n \t\tfsync_object_files = git_config_bool(var, value);\n \t\treturn 0;\n \t}\n-- \ngitgitgadget\n"},{"id":"452671","messageId":"xmqqpmm39xhx.fsf@gitster.g","threadId":"57568","inReplyTo":"c7a2a7efe6d532fc7fce1352b1dfce640cc9f2f6.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:11:06Z","receivedAt":"2022-03-30T17:11:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Batched fsync will fit into bulk-checkin by taking advantage of the\n> plug/unplug functionality to determine the appropriate time to fsync\n> and make newly-added objects available in the primary object database.\n>\n> * Rename 'state' variable to 'bulk_checkin_state', since we will later\n>   be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n>   find in the debugger, since the name is more unique.\n>\n> * Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n>   static variable. Doing this avoids resetting the variable in\n>   finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n>   seem to unintentionally disable the plugging functionality the first\n>   time a new packfile must be created due to packfile size limits. While\n>   disabling the plugging state only results in suboptimal behavior for\n>   the current code, it would be fatal for the bulk-fsync functionality\n>   later in this patch series.\n\nParaphrasing to make sure I understand your reasoning here...\n\nIn the \"plug and then perform as many changes to the repository and\nfinally unplug\" flow, before or after this series, the \"perform\"\nstep in the middle is unaware of which \"bulk_checkin_state\" instance\nis being used to keep track of what is done to optimize by deferring\nsome operations until the \"unplug\" time.  So bulk_checkin_state is\nnot there to allow us to create multiple instances of it, pass them\naround to different sequences of \"plug, perform, unplug\".  Each of\nits members is inherently a singleton, so in the extreme, we could\nturn these members into separate file-scope global variables if we\nwanted to.  The \"plugged\" bit happens to be the only one getting\nejected by this patch, because it is inconvenient to \"clear\" other\nmembers otherwise.\n\nIs that what is going on?\n\nIf it is, I am mildly opposed to the flow of thought, from at least\ntwo reasons.  It makes it hard for the next developer to decide if\nthe new members they are adding should be in or out of the struct.\n\nMore importantly, I think the call of finish_bulk_checkin() we make\nin deflate_to_pack() you found (and there may possibly other places\nthat we do so; I didn't check) may not appear to be a bug in the\noriginal context, but it already is a bug.  And when we change the\nsemantics of plug-unplug to be more \"transaction-like\", it becomes a\nmore serious bug, as you said.\n\nThere is NO reason to end the ongoing transaction there inside the\nwhile() loop that tries to limit the size of the packfile being\nused.  We may want to flush the \"packfile part\", which may have been\nalmost synonymous to the entirety of bulk_checkin_state, but as you\nfound out, the \"plugged\" bit is *outside* the \"packfile part\", and\nthat makes it a bug to call finish_bulk_checkin() from there.\n\nWe should add a new function, flush_bulk_checking_packfile(), to\nflush only the packfile part of the bulk_checkin_state without\naffecting other things---the \"plugged\" bit is the only one in the\ncurrent code before this series, but it does not have to stay to be\nso.  When you start plugging the loose ref transactions, you may\nfind it handy (this is me handwaving) to have a list of refs that\nyou may have to do something at \"unplug\" time kept in the struct,\nand you do not want deflate_to_pack() affecting the ongoing\n\"plugged\" ref operations by calling finish_bulk_checkin() and\nreinitializing that list, for example.\n\nAnd then we should examine existing calls to finish_bulk_checkin()\nand replace the ones that should not be finishing, i.e. the ones\nthat wanted \"flush\" but called \"finish\".\n\n"},{"id":"452673","messageId":"xmqqk0cb9x78.fsf@gitster.g","threadId":"57568","inReplyTo":"d045b13795b38caa27f8e25340212f736b66bb05.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 02/14] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:17:31Z","receivedAt":"2022-03-30T17:17:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Make it clearer in the naming and documentation of the plug_bulk_checkin\n> and unplug_bulk_checkin APIs that they can be thought of as\n> a \"transaction\" to optimize operations on the object database. These\n> transactions may be nested so that subsystems like the cache-tree\n> writing code can optimize their operations without caring whether the\n> top-level code has a transaction active.\n\nI can see that \"checkin\" part of the name is too limiting (you may\nwant to do more than optimize checkin, e.g. fsync), and that you may\nprefer \"begin/end\" over \"plug/unplug\", but I am not sure if we want\nto limit ourselves to \"odb\".  If we find our code doing things on\nmany instances of something that are not objects (e.g. refs, config\nvariables), don't we want to give them the same chance to be optimized\nby batching them?\n\n{begin,end}_bulk_transaction perhaps?  I dunno.\n"},{"id":"452674","messageId":"xmqqfsmz9x4u.fsf@gitster.g","threadId":"57568","inReplyTo":"2d1bc4568ac744f11c886a5f964dbe563c04ce8b.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 03/14] object-file: pass filename to fsync_or_die","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:18:57Z","receivedAt":"2022-03-30T17:19:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> If we die while trying to fsync a loose object file, pass the actual\n> filename we're trying to sync. This is likely to be more helpful for a\n> user trying to diagnose the cause of the failure than the former\n> 'loose object file' string. It also sidesteps any concerns about\n> translating the die message differently for loose objects versus\n> something else that has a real path.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  object-file.c | 8 ++++----\n>  1 file changed, 4 insertions(+), 4 deletions(-)\n\nLooks good and obviously can be done totally outside the series.\nPerhaps while the rest of the topic is still cooking, extract this\n(we may find others) and have them graduate sooner?\n\n> diff --git a/object-file.c b/object-file.c\n> index b254bc50d70..5ffbf3d4fd4 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1888,16 +1888,16 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  }\n>  \n>  /* Finalize a file on disk, and close it. */\n> -static void close_loose_object(int fd)\n> +static void close_loose_object(int fd, const char *filename)\n>  {\n>  \tif (the_repository->objects->odb->will_destroy)\n>  \t\tgoto out;\n>  \n>  \tif (fsync_object_files > 0)\n> -\t\tfsync_or_die(fd, \"loose object file\");\n> +\t\tfsync_or_die(fd, filename);\n>  \telse\n>  \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n> -\t\t\t\t       \"loose object file\");\n> +\t\t\t\t       filename);\n>  \n>  out:\n>  \tif (close(fd) != 0)\n> @@ -2011,7 +2011,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n> -\tclose_loose_object(fd);\n> +\tclose_loose_object(fd, tmp_file.buf);\n>  \n>  \tif (mtime) {\n>  \t\tstruct utimbuf utb;\n"},{"id":"452678","messageId":"xmqq4k3f9w9s.fsf@gitster.g","threadId":"57568","inReplyTo":"9e7ae22fa4a2693fe26659f875dd780080c4cfb2.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:37:35Z","receivedAt":"2022-03-30T17:37:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index 9da3e5d88f6..3c90ba0b395 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -596,6 +596,14 @@ core.fsyncMethod::\n>  * `writeout-only` issues pagecache writeback requests, but depending on the\n>    filesystem and storage hardware, data added to the repository may not be\n>    durable in the event of a system crash. This is the default mode on macOS.\n> +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> +  updates in the disk writeback cache and then does a single full fsync of\n> +  a dummy file to trigger the disk cache flush at the end of the operation.\n> ++\n> +  Currently `batch` mode only applies to loose-object files. Other repository\n> +  data is made durable as if `fsync` was specified. This mode is expected to\n> +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> +  and on Windows for repos stored on NTFS or ReFS filesystems.\n\nDoes this format correctly?  I had an impression that the second and\nsubsequent paragraphs, connected with a line with a single \"+\" on\nit, has to be flushed left without indentation.\n\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index 8b0fd5c7723..9799d247cad 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -3,15 +3,20 @@\n>   */\n>  #include \"cache.h\"\n>  #include \"bulk-checkin.h\"\n> +#include \"lockfile.h\"\n>  #include \"repository.h\"\n>  #include \"csum-file.h\"\n>  #include \"pack.h\"\n>  #include \"strbuf.h\"\n> +#include \"string-list.h\"\n> +#include \"tmp-objdir.h\"\n>  #include \"packfile.h\"\n>  #include \"object-store.h\"\n>  \n>  static int odb_transaction_nesting;\n>  \n> +static struct tmp_objdir *bulk_fsync_objdir;\n\nI wonder if this should be added to the bulk_checkin_state structure\nas a new member, especially if we fix the erroneous call to\nfinish_bulk_checkin() as a preliminary fix-up of a bug that existed\neven before this series.\n\n> +/*\n> + * Cleanup after batch-mode fsync_object_files.\n> + */\n> +static void do_batch_fsync(void)\n> +{\n> +\tstruct strbuf temp_path = STRBUF_INIT;\n> +\tstruct tempfile *temp;\n> +\n> +\tif (!bulk_fsync_objdir)\n> +\t\treturn;\n> +\n> +\t/*\n> +\t * Issue a full hardware flush against a temporary file to ensure\n> +\t * that all objects are durable before any renames occur. The code in\n> +\t * fsync_loose_object_bulk_checkin has already issued a writeout\n> +\t * request, but it has not flushed any writeback cache in the storage\n> +\t * hardware or any filesystem logs. This fsync call acts as a barrier\n> +\t * to ensure that the data in each new object file is durable before\n> +\t * the final name is visible.\n> +\t */\n> +\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> +\ttemp = xmks_tempfile(temp_path.buf);\n> +\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> +\tdelete_tempfile(&temp);\n> +\tstrbuf_release(&temp_path);\n> +\n> +\t/*\n> +\t * Make the object files visible in the primary ODB after their data is\n> +\t * fully durable.\n> +\t */\n> +\ttmp_objdir_migrate(bulk_fsync_objdir);\n> +\tbulk_fsync_objdir = NULL;\n> +}\n\nOK.\n\n> +void prepare_loose_object_bulk_checkin(void)\n> +{\n> +\t/*\n> +\t * We lazily create the temporary object directory\n> +\t * the first time an object might be added, since\n> +\t * callers may not know whether any objects will be\n> +\t * added at the time they call begin_odb_transaction.\n> +\t */\n> +\tif (!odb_transaction_nesting || bulk_fsync_objdir)\n> +\t\treturn;\n> +\n> +\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> +\tif (bulk_fsync_objdir)\n> +\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> +}\n\nOK.  If we got a failure from tmp_objdir_create(), then we don't\nswap and end up creating a new loose object file in the primary\nobject store.  I wonder if we at least want to note that fact for\nlater use at \"unplug\" time.  We may create a few loose objects in\nthe primary object store without any fsync, then a later call may\nsuccessfully create a temporary object directory and we'd create\nmore loose objects in the temporary one, which are flushed with the\n\"create a dummy and fsync\" trick and migrated, but do we need to do\nsomething to the ones we created in the primary object store before\nall that happens?\n\n> +void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n> +{\n> +\t/*\n> +\t * If we have an active ODB transaction, we issue a call that\n> +\t * cleans the filesystem page cache but avoids a hardware flush\n> +\t * command. Later on we will issue a single hardware flush\n> +\t * before as part of do_batch_fsync.\n> +\t */\n> +\tif (!bulk_fsync_objdir ||\n> +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n> +\t\tfsync_or_die(fd, filename);\n> +\t}\n> +}\n\nAh, if we have successfully created the temporary directory, we\ndon't do full fsync but just writeout-only one, so there is no need\nfor the worry I mentioned in the previous paragraph.  OK.\n\n> @@ -301,4 +370,6 @@ void end_odb_transaction(void)\n>  \n>  \tif (bulk_checkin_state.f)\n>  \t\tfinish_bulk_checkin(&bulk_checkin_state);\n> +\n> +\tdo_batch_fsync();\n>  }\n\nOK.\n\n> @@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tstatic struct strbuf tmp_file = STRBUF_INIT;\n>  \tstatic struct strbuf filename = STRBUF_INIT;\n>  \n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> +\t\tprepare_loose_object_bulk_checkin();\n> +\n>  \tloose_object_path(the_repository, &filename, oid);\n>  \n>  \tfd = create_tmpfile(&tmp_file, filename.buf);\n\nThe necessary change to the \"workhorse\" code path is surprisingly\nsmall, which is pleasing to see.\n\nThanks.\n"},{"id":"452679","messageId":"xmqqy20r8h9o.fsf@gitster.g","threadId":"57568","inReplyTo":"83fa4a5f3a5c79fa814932c0705867ff16a584c7.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 05/14] cache-tree: use ODB transaction around writing a tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:46:59Z","receivedAt":"2022-03-30T17:47:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Take advantage of the odb transaction infrastructure around writing the\n> cached tree to the object database.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  cache-tree.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/cache-tree.c b/cache-tree.c\n> index 6752f69d515..8c5e8822716 100644\n> --- a/cache-tree.c\n> +++ b/cache-tree.c\n> @@ -3,6 +3,7 @@\n>  #include \"tree.h\"\n>  #include \"tree-walk.h\"\n>  #include \"cache-tree.h\"\n> +#include \"bulk-checkin.h\"\n>  #include \"object-store.h\"\n>  #include \"replace-object.h\"\n>  #include \"promisor-remote.h\"\n> @@ -474,8 +475,10 @@ int cache_tree_update(struct index_state *istate, int flags)\n>  \n>  \ttrace_performance_enter();\n>  \ttrace2_region_enter(\"cache_tree\", \"update\", the_repository);\n\nThere is no I/O in update_one() when the WRITE_TREE_DRY_RUN bit is\nset, so we _could_ optimize the begin/end away with\n\n\tif (!(flags & WRITE_TREE_DRY_RUN))\n\t\tbegin_odb_transaction()\n\n> +\tbegin_odb_transaction();\n>  \ti = update_one(istate->cache_tree, istate->cache, istate->cache_nr,\n>  \t\t       \"\", 0, &skip, flags);\n> +\tend_odb_transaction();\n>  \ttrace2_region_leave(\"cache_tree\", \"update\", the_repository);\n>  \ttrace_performance_leave(\"cache_tree_update\");\n>  \tif (i < 0)\n\nI do not know if that is worth it.  If we do not do any object\ncreation inside begin/end, we don't even create the temporary object\ndirectory and there is nothing we need to do when we \"unplug\".  So\nthis would be fine as-is, but I may be overlooking something, so I\nthought I'd mention it for completeness.\n\nThanks.\n"},{"id":"452680","messageId":"xmqqtubf8h8g.fsf@gitster.g","threadId":"57568","inReplyTo":"d514842ad493a819e3640ecb658f702e530d6e85.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 06/14] builtin/add: add ODB transaction around add_files_to_cache","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:47:43Z","receivedAt":"2022-03-30T17:48:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> The add_files_to_cache function is invoked internally by\n> builtin/commit.c and builtin/checkout.c for their flags that stage\n> modified files before doing the larger operation. These commands\n> can benefit from batched fsyncing.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  builtin/add.c | 9 +++++++++\n>  1 file changed, 9 insertions(+)\n>\n> diff --git a/builtin/add.c b/builtin/add.c\n> index 9bf37ceae8e..e39770e4746 100644\n> --- a/builtin/add.c\n> +++ b/builtin/add.c\n> @@ -141,7 +141,16 @@ int add_files_to_cache(const char *prefix,\n>  \trev.diffopt.format_callback_data = &data;\n>  \trev.diffopt.flags.override_submodule_config = 1;\n>  \trev.max_count = 0; /* do not compare unmerged paths with stage #2 */\n> +\n> +\t/*\n> +\t * Use an ODB transaction to optimize adding multiple objects.\n> +\t * This function is invoked from commands other than 'add', which\n> +\t * may not have their own transaction active.\n> +\t */\n> +\tbegin_odb_transaction();\n>  \trun_diff_files(&rev, DIFF_RACY_IS_MODIFIED);\n> +\tend_odb_transaction();\n> +\n\nThis one clearly is \"bulk\".  Makes sense.\n"},{"id":"452682","messageId":"xmqqpmm38h01.fsf@gitster.g","threadId":"57568","inReplyTo":"8cac94598a58704d9b625a9d8a593779f7adc30f.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 07/14] update-index: use the bulk-checkin infrastructure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T17:52:46Z","receivedAt":"2022-03-30T17:52:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +\t/*\n> +\t * Allow the object layer to optimize adding multiple objects in\n> +\t * a batch.\n> +\t */\n> +\tbegin_odb_transaction();\n>  \twhile (ctx.argc) {\n>  \t\tif (parseopt_state != PARSE_OPT_DONE)\n>  \t\t\tparseopt_state = parse_options_step(&ctx, options,\n> @@ -1167,6 +1174,17 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\tthe_index.version = preferred_index_format;\n>  \t}\n>  \n> +\t/*\n> +\t * It is possible, though unlikely, that a caller could use the verbose\n> +\t * output to synchronize with addition of objects to the object\n> +\t * database. The current implementation of ODB transactions leaves\n> +\t * objects invisible while a transaction is active, so end the\n> +\t * transaction here if verbose output is enabled.\n> +\t */\n> +\n> +\tif (verbose)\n> +\t\tend_odb_transaction();\n> +\n>  \tif (read_from_stdin) {\n>  \t\tstruct strbuf buf = STRBUF_INIT;\n>  \t\tstruct strbuf unquoted = STRBUF_INIT;\n> @@ -1190,6 +1208,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n>  \t\tstrbuf_release(&buf);\n>  \t}\n>  \n> +\t/*\n> +\t * By now we have added all of the new objects\n> +\t */\n> +\tif (!verbose)\n> +\t\tend_odb_transaction();\n\nIf we had \"flush\" in addition to \"begin\" and \"end\", then we could,\ninstead of the above\n\n    begin_transaction\n\tdo things\n    if condition:\n\tend_transaction\n    loop:\n\tdo thing\n    if !condition:\n\tend_transaction\n\n\nwhich is somewhat hard to follow and maintain, consider using a\ndifferent flow, which is\n\n    begin_transaction\n\tdo things\n    loop:\n\tdo thing\n\tif condition:\n\t    flush\n    end_transaction\n\nand that might make it easier to follow and maintain.  I am not 100%\nsure if it is worth it, but I am leaning to say it would be.\n\nThanks.\n"},{"id":"452683","messageId":"CANQDOdeFbMreRh_5w5fwfJzRfrhk+A9HBoqpkACx1PrNOpyq0Q@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqfsmz9x4u.fsf@gitster.g","subject":"Re: [PATCH v5 03/14] object-file: pass filename to fsync_or_die","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-30T17:54:06Z","receivedAt":"2022-03-30T17:54:25Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:18 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > If we die while trying to fsync a loose object file, pass the actual\n> > filename we're trying to sync. This is likely to be more helpful for a\n> > user trying to diagnose the cause of the failure than the former\n> > 'loose object file' string. It also sidesteps any concerns about\n> > translating the die message differently for loose objects versus\n> > something else that has a real path.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  object-file.c | 8 ++++----\n> >  1 file changed, 4 insertions(+), 4 deletions(-)\n>\n> Looks good and obviously can be done totally outside the series.\n> Perhaps while the rest of the topic is still cooking, extract this\n> (we may find others) and have them graduate sooner?\n>\n> > diff --git a/object-file.c b/object-file.c\n> > index b254bc50d70..5ffbf3d4fd4 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1888,16 +1888,16 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n> >  }\n> >\n> >  /* Finalize a file on disk, and close it. */\n> > -static void close_loose_object(int fd)\n> > +static void close_loose_object(int fd, const char *filename)\n> >  {\n> >       if (the_repository->objects->odb->will_destroy)\n> >               goto out;\n> >\n> >       if (fsync_object_files > 0)\n> > -             fsync_or_die(fd, \"loose object file\");\n> > +             fsync_or_die(fd, filename);\n> >       else\n> >               fsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n> > -                                    \"loose object file\");\n> > +                                    filename);\n> >\n> >  out:\n> >       if (close(fd) != 0)\n> > @@ -2011,7 +2011,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >               die(_(\"confused by unstable object source data for %s\"),\n> >                   oid_to_hex(oid));\n> >\n> > -     close_loose_object(fd);\n> > +     close_loose_object(fd, tmp_file.buf);\n> >\n> >       if (mtime) {\n> >               struct utimbuf utb;\n\n\nOk. I'll create two new single-patch series with changes that are independent.\n"},{"id":"452685","messageId":"xmqqbkxn8g0r.fsf@gitster.g","threadId":"57568","inReplyTo":"1a4aff8c350b5ffe3c7760faa4accc88c83ce11c.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 11/14] core.fsyncmethod: tests for batch mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T18:13:56Z","receivedAt":"2022-03-30T18:14:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> new file mode 100644\n> index 00000000000..34c01a65256\n> --- /dev/null\n> +++ b/t/lib-unique-files.sh\n> @@ -0,0 +1,34 @@\n> +# Helper to create files with unique contents\n> +\n> +# Create multiple files with unique contents within this test run. Takes the\n> +# number of directories, the number of files in each directory, and the base\n> +# directory.\n> +#\n> +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> +#\t\t\t\t\t each in my_dir, all with contents\n> +#\t\t\t\t\t different from previous invocations\n> +#\t\t\t\t\t of this command in this run.\n> +\n> +test_create_unique_files () {\n> +\ttest \"$#\" -ne 3 && BUG \"3 param\"\n> +\n> +\tlocal dirs=\"$1\" &&\n> +\tlocal files=\"$2\" &&\n> +\tlocal basedir=\"$3\" &&\n> +\tlocal counter=0 &&\n\nThe cover letter mentioned that in this round, \"local var=val\" is\navoided?\n\nI've seen instances of local being \"curious\", e.g.\n  https://lore.kernel.org/git/20220125092419.cgtfw32nk2niazfk@carbon/\nand the discussion indicates that it may still be relevant.\n\n> +\tlocal i &&\n> +\tlocal j &&\n> +\ttest_tick &&\n> +\tlocal basedata=$basedir$test_tick &&\n> +\trm -rf \"$basedir\" &&\n> +\tfor i in $(test_seq $dirs)\n> +\tdo\n> +\t\tlocal dir=$basedir/dir$i &&\n\nThis, too.\n\nTo summarize the findings from the thread is:\n\n - very old releases of /bin/dash that predates Git, like 0.3.8,\n   would not correctly handle assignment on \"local\" at all.  It may\n   not matter to us.\n\n - semi-old /bin/dash 0.5.10, which is still in use, mishandled\n   'local var=$val', but 'local var=\"$val\"' is an acceptable\n   workaround for the bug.  \"git grep\" tells us that we use this\n   form in our shell scripts quite a lot, so we may be OK.\n\n - /bin/dash 0.5.11, which was tagged mid 2020, and newer would glok\n   'local var=$val' just fine even $val has $IFS whitespace in it.\n\nSo, I'd say the safe practice we should adopt is to use \"local\" one\nper variable, assignment to the variable on the same line of \"local\"\nis OK, but unlike the regular assignment, older dash may mishandle\nthe right-hand-side unless it is quoted, i.e.\n\n    local var='string literal'\n    local var=\"$variable interpolation\"\n\n"},{"id":"452688","messageId":"CANQDOdfWTufEn0NRSAOG991JcS4x8GsCC62UCLUTEc3gD6tfGA@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqpmm39xhx.fsf@gitster.g","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-30T18:34:58Z","receivedAt":"2022-03-30T18:37:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:11 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > Batched fsync will fit into bulk-checkin by taking advantage of the\n> > plug/unplug functionality to determine the appropriate time to fsync\n> > and make newly-added objects available in the primary object database.\n> >\n> > * Rename 'state' variable to 'bulk_checkin_state', since we will later\n> >   be adding 'bulk_fsync_objdir'.  This also makes the variable easier to\n> >   find in the debugger, since the name is more unique.\n> >\n> > * Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n> >   static variable. Doing this avoids resetting the variable in\n> >   finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n> >   seem to unintentionally disable the plugging functionality the first\n> >   time a new packfile must be created due to packfile size limits. While\n> >   disabling the plugging state only results in suboptimal behavior for\n> >   the current code, it would be fatal for the bulk-fsync functionality\n> >   later in this patch series.\n>\n> Paraphrasing to make sure I understand your reasoning here...\n>\n> In the \"plug and then perform as many changes to the repository and\n> finally unplug\" flow, before or after this series, the \"perform\"\n> step in the middle is unaware of which \"bulk_checkin_state\" instance\n> is being used to keep track of what is done to optimize by deferring\n> some operations until the \"unplug\" time.  So bulk_checkin_state is\n> not there to allow us to create multiple instances of it, pass them\n> around to different sequences of \"plug, perform, unplug\".  Each of\n> its members is inherently a singleton, so in the extreme, we could\n> turn these members into separate file-scope global variables if we\n> wanted to.  The \"plugged\" bit happens to be the only one getting\n> ejected by this patch, because it is inconvenient to \"clear\" other\n> members otherwise.\n>\n> Is that what is going on?\n>\n\nMore or less.  The current state is all about creating a single\npackfile for multiple large objects.  That packfile is a singleton\ntoday (we could have an alternate implementation where there's a\nseparate packfile per thread in the future, so it's not inherent to\nthe API).  We want to do this if the top-level caller is okay with the\nstate being invisible until the \"finish\" call, and that is conveyed by\nthe \"plugged\" flag.\n\n> If it is, I am mildly opposed to the flow of thought, from at least\n> two reasons.  It makes it hard for the next developer to decide if\n> the new members they are adding should be in or out of the struct.\n>\n> More importantly, I think the call of finish_bulk_checkin() we make\n> in deflate_to_pack() you found (and there may possibly other places\n> that we do so; I didn't check) may not appear to be a bug in the\n> original context, but it already is a bug.  And when we change the\n> semantics of plug-unplug to be more \"transaction-like\", it becomes a\n> more serious bug, as you said.\n>\n> There is NO reason to end the ongoing transaction there inside the\n> while() loop that tries to limit the size of the packfile being\n> used.  We may want to flush the \"packfile part\", which may have been\n> almost synonymous to the entirety of bulk_checkin_state, but as you\n> found out, the \"plugged\" bit is *outside* the \"packfile part\", and\n> that makes it a bug to call finish_bulk_checkin() from there.\n>\n> We should add a new function, flush_bulk_checking_packfile(), to\n> flush only the packfile part of the bulk_checkin_state without\n> affecting other things---the \"plugged\" bit is the only one in the\n> current code before this series, but it does not have to stay to be\n> so\n\nI'm happy to rename the packfile related stuff to end with _packfile\nto make it clear that all of that state and functionality is related\nto batching of packfile additions.\nSo from this patch: s/bulk_checkin_state/bulk_checkin_packfile and\ns/finish_bulk_checkin/finish_bulk_checkin_packfile.\n\nMy new state will be bulk_fsync_* (as it is already).  Any future\nODB-related state can go here too (I'm imagining a future\nlog-structured 'new objects pack' that we can append to for adding\nsmall objects, similar to the bulk_checkin_packfile but allowing\nappends from multiple git invocations).\n\n> When you start plugging the loose ref transactions, you may\n> find it handy (this is me handwaving) to have a list of refs that\n> you may have to do something at \"unplug\" time kept in the struct,\n> and you do not want deflate_to_pack() affecting the ongoing\n> \"plugged\" ref operations by calling finish_bulk_checkin() and\n> reinitializing that list, for example.\n>\n\nI don't believe ref transactions will go through this part of the\ninfrastructure.  Refs already have a good transaction system (that\npartly inspired this rebranding, after I saw how Peter implemented\nbatch ref fsync).  I expect this area will remain all about the ODB as\na subsystem that can enlist in a larger repo->transaction.  So a\ntop-level Git command might initiate a repo transaction, which would\ninternally initiate an ODB transaction, index transaction, and ref\ntransaction. The repo transaction would support flushing each of the\nsubtransactions with an optimal number of fsyncs.\n\n> And then we should examine existing calls to finish_bulk_checkin()\n> and replace the ones that should not be finishing, i.e. the ones\n> that wanted \"flush\" but called \"finish\".\n\nSure.  I can fix this, which will only change this file.  The only\ncase of \"finishing\" would be in unplug_bulk_checkin /\nend_odb_transaction.\n\nThanks,\nNeeraj\n"},{"id":"452690","messageId":"CANQDOddUQRwq73e9pxxmp8c2JRvLy4YcbDHEQ+h-6uEoauyb2g@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqy20r8h9o.fsf@gitster.g","subject":"Re: [PATCH v5 05/14] cache-tree: use ODB transaction around writing a tree","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-30T19:04:46Z","receivedAt":"2022-03-30T19:05:09Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:47 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Take advantage of the odb transaction infrastructure around writing the\n> > cached tree to the object database.\n> >\n> > Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> > ---\n> >  cache-tree.c | 3 +++\n> >  1 file changed, 3 insertions(+)\n> >\n> > diff --git a/cache-tree.c b/cache-tree.c\n> > index 6752f69d515..8c5e8822716 100644\n> > --- a/cache-tree.c\n> > +++ b/cache-tree.c\n> > @@ -3,6 +3,7 @@\n> >  #include \"tree.h\"\n> >  #include \"tree-walk.h\"\n> >  #include \"cache-tree.h\"\n> > +#include \"bulk-checkin.h\"\n> >  #include \"object-store.h\"\n> >  #include \"replace-object.h\"\n> >  #include \"promisor-remote.h\"\n> > @@ -474,8 +475,10 @@ int cache_tree_update(struct index_state *istate, int flags)\n> >\n> >       trace_performance_enter();\n> >       trace2_region_enter(\"cache_tree\", \"update\", the_repository);\n>\n> There is no I/O in update_one() when the WRITE_TREE_DRY_RUN bit is\n> set, so we _could_ optimize the begin/end away with\n>\n>         if (!(flags & WRITE_TREE_DRY_RUN))\n>                 begin_odb_transaction()\n>\n> > +     begin_odb_transaction();\n> >       i = update_one(istate->cache_tree, istate->cache, istate->cache_nr,\n> >                      \"\", 0, &skip, flags);\n> > +     end_odb_transaction();\n> >       trace2_region_leave(\"cache_tree\", \"update\", the_repository);\n> >       trace_performance_leave(\"cache_tree_update\");\n> >       if (i < 0)\n>\n> I do not know if that is worth it.  If we do not do any object\n> creation inside begin/end, we don't even create the temporary object\n> directory and there is nothing we need to do when we \"unplug\".  So\n> this would be fine as-is, but I may be overlooking something, so I\n> thought I'd mention it for completeness.\n>\n\nYes, with the current series, beginning and ending a transaction will\njust manipulate a few global variables in bulk-checkin.c unless there\nis something real to flush.\n\nThanks,\nNeeraj\n"},{"id":"452691","messageId":"CANQDOdf5T9iexGryVfZ1epgEdkiBU5WBCX9rBtpr_JuSvvxjmA@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqpmm38h01.fsf@gitster.g","subject":"Re: [PATCH v5 07/14] update-index: use the bulk-checkin infrastructure","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-30T19:09:04Z","receivedAt":"2022-03-30T19:09:47Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:52 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > +     /*\n> > +      * Allow the object layer to optimize adding multiple objects in\n> > +      * a batch.\n> > +      */\n> > +     begin_odb_transaction();\n> >       while (ctx.argc) {\n> >               if (parseopt_state != PARSE_OPT_DONE)\n> >                       parseopt_state = parse_options_step(&ctx, options,\n> > @@ -1167,6 +1174,17 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >               the_index.version = preferred_index_format;\n> >       }\n> >\n> > +     /*\n> > +      * It is possible, though unlikely, that a caller could use the verbose\n> > +      * output to synchronize with addition of objects to the object\n> > +      * database. The current implementation of ODB transactions leaves\n> > +      * objects invisible while a transaction is active, so end the\n> > +      * transaction here if verbose output is enabled.\n> > +      */\n> > +\n> > +     if (verbose)\n> > +             end_odb_transaction();\n> > +\n> >       if (read_from_stdin) {\n> >               struct strbuf buf = STRBUF_INIT;\n> >               struct strbuf unquoted = STRBUF_INIT;\n> > @@ -1190,6 +1208,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n> >               strbuf_release(&buf);\n> >       }\n> >\n> > +     /*\n> > +      * By now we have added all of the new objects\n> > +      */\n> > +     if (!verbose)\n> > +             end_odb_transaction();\n>\n> If we had \"flush\" in addition to \"begin\" and \"end\", then we could,\n> instead of the above\n>\n>     begin_transaction\n>         do things\n>     if condition:\n>         end_transaction\n>     loop:\n>         do thing\n>     if !condition:\n>         end_transaction\n>\n>\n> which is somewhat hard to follow and maintain, consider using a\n> different flow, which is\n>\n>     begin_transaction\n>         do things\n>     loop:\n>         do thing\n>         if condition:\n>             flush\n>     end_transaction\n>\n> and that might make it easier to follow and maintain.  I am not 100%\n> sure if it is worth it, but I am leaning to say it would be.\n>\n> Thanks.\n\nI thought about this, but I was somewhat worried about the extra cost\nof \"flushing\" versus just making things visible immediately if the\ntop-level update-index command knows it's going to need objects to be\nvisible.  I'll go ahead and implement flush in the next iteration.\n"},{"id":"452701","messageId":"xmqqr16j6ve3.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdfWTufEn0NRSAOG991JcS4x8GsCC62UCLUTEc3gD6tfGA@mail.gmail.com","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-30T20:24:52Z","receivedAt":"2022-03-30T20:24:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> We should add a new function, flush_bulk_checking_packfile(), to\n>> flush only the packfile part of the bulk_checkin_state without\n>> affecting other things---the \"plugged\" bit is the only one in the\n>> current code before this series, but it does not have to stay to be\n>> so\n>\n> I'm happy to rename the packfile related stuff to end with _packfile\n> to make it clear that all of that state and functionality is related\n> to batching of packfile additions.\n\nI do not care about names, though.  If you took that I hinted any\nsuch change, sorry about that.  _state is fine as-is.\n\nI do care about not ejecting plugged out of the structure and\ninstead keeping them together, with proper way to flush the part\nthat deflate_to_pack() wants to flush, instead of abusing the\n\"finish\".\n\nThanks.\n"},{"id":"452768","messageId":"CANQDOdfoDSzpWbmnyaPyi9PoftbS7P7=8Y+RS_Bm3t0ggXpLEg@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqbkxn8g0r.fsf@gitster.g","subject":"Re: [PATCH v5 11/14] core.fsyncmethod: tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T03:55:51Z","receivedAt":"2022-03-31T04:08:36Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 11:13 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > diff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\n> > new file mode 100644\n> > index 00000000000..34c01a65256\n> > --- /dev/null\n> > +++ b/t/lib-unique-files.sh\n> > @@ -0,0 +1,34 @@\n> > +# Helper to create files with unique contents\n> > +\n> > +# Create multiple files with unique contents within this test run. Takes the\n> > +# number of directories, the number of files in each directory, and the base\n> > +# directory.\n> > +#\n> > +# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n> > +#                                     each in my_dir, all with contents\n> > +#                                     different from previous invocations\n> > +#                                     of this command in this run.\n> > +\n> > +test_create_unique_files () {\n> > +     test \"$#\" -ne 3 && BUG \"3 param\"\n> > +\n> > +     local dirs=\"$1\" &&\n> > +     local files=\"$2\" &&\n> > +     local basedir=\"$3\" &&\n> > +     local counter=0 &&\n>\n> The cover letter mentioned that in this round, \"local var=val\" is\n> avoided?\n>\n\nWhoops.. I guess I wasn't looking widely enough.  I noticed failures\non my Ubuntu system that uses dash.\n\n> I've seen instances of local being \"curious\", e.g.\n>   https://lore.kernel.org/git/20220125092419.cgtfw32nk2niazfk@carbon/\n> and the discussion indicates that it may still be relevant.\n>\n> > +     local i &&\n> > +     local j &&\n> > +     test_tick &&\n> > +     local basedata=$basedir$test_tick &&\n> > +     rm -rf \"$basedir\" &&\n> > +     for i in $(test_seq $dirs)\n> > +     do\n> > +             local dir=$basedir/dir$i &&\n>\n> This, too.\n>\n> To summarize the findings from the thread is:\n>\n>  - very old releases of /bin/dash that predates Git, like 0.3.8,\n>    would not correctly handle assignment on \"local\" at all.  It may\n>    not matter to us.\n>\n>  - semi-old /bin/dash 0.5.10, which is still in use, mishandled\n>    'local var=$val', but 'local var=\"$val\"' is an acceptable\n>    workaround for the bug.  \"git grep\" tells us that we use this\n>    form in our shell scripts quite a lot, so we may be OK.\n>\n>  - /bin/dash 0.5.11, which was tagged mid 2020, and newer would glok\n>    'local var=$val' just fine even $val has $IFS whitespace in it.\n>\n> So, I'd say the safe practice we should adopt is to use \"local\" one\n> per variable, assignment to the variable on the same line of \"local\"\n> is OK, but unlike the regular assignment, older dash may mishandle\n> the right-hand-side unless it is quoted, i.e.\n>\n>     local var='string literal'\n>     local var=\"$variable interpolation\"\n>\n\nThanks for updating me on the documentation.   I think I'll go back to\none-line assignments with the proper quotations and single variables,\njust to keep things short and consistent. I've been looking up\neverything as I go with these shell scripts, but there's a lot of\nsubtlety.  CMD batch files are even uglier, but I've had many years to\nlearn the common practices there.\n\nI'll ask a question on the perf test patch on a related shell topic.\n\nThanks,\nNeeraj\n"},{"id":"452769","messageId":"CANQDOdea6Ejnsj1r6GhSvR+6ZMBGkRo8yr-GpOv-Gcsc8hEf6w@mail.gmail.com","threadId":"57568","inReplyTo":"26be6ecb28bc1f76fba380fdd10acf59820df997.1648616734.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 13/14] core.fsyncmethod: performance tests for batch mode","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T04:09:28Z","receivedAt":"2022-03-31T04:20:15Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 29, 2022 at 10:05 PM Neeraj Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> Add basic performance tests for git commands that can add data to the\n> object database. We cover:\n> * git add\n> * git stash\n> * git update-index (via git stash)\n> * git unpack-objects\n> * git commit --all\n>\n> We cover all currently available fsync methods as well.\n>\n> Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n> ---\n>  t/perf/p0008-odb-fsync.sh | 81 +++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 81 insertions(+)\n>  create mode 100755 t/perf/p0008-odb-fsync.sh\n>\n> diff --git a/t/perf/p0008-odb-fsync.sh b/t/perf/p0008-odb-fsync.sh\n> new file mode 100755\n> index 00000000000..87092c2627e\n> --- /dev/null\n> +++ b/t/perf/p0008-odb-fsync.sh\n> @@ -0,0 +1,81 @@\n> +#!/bin/sh\n> +#\n> +# This test measures the performance of adding new files to the object\n> +# database. The test was originally added to measure the effect of the\n> +# core.fsyncMethod=batch mode, which is why we are testing different values of\n> +# that setting explicitly and creating a lot of unique objects.\n> +\n> +test_description=\"Tests performance of adding things to the object database\"\n> +\n> +. ./perf-lib.sh\n> +\n> +. $TEST_DIRECTORY/lib-unique-files.sh\n> +\n> +test_perf_fresh_repo\n> +test_checkout_worktree\n> +\n> +dir_count=10\n> +files_per_dir=50\n> +total_files=$((dir_count * files_per_dir))\n> +\n> +populate_files () {\n> +       test_create_unique_files $dir_count $files_per_dir files\n> +}\n> +\n> +setup_repo () {\n> +       (rm -rf .git || 1) &&\n> +       git init &&\n> +       test_commit first &&\n> +       populate_files\n> +}\n> +\n> +test_perf_fsync_cfgs () {\n> +       local method cfg &&\n> +       for method in none fsync batch writeout-only\n> +       do\n> +               case $method in\n> +               none)\n> +                       cfg=\"-c core.fsync=none\"\n> +                       ;;\n> +               *)\n> +                       cfg=\"-c core.fsync=loose-object -c core.fsyncMethod=$method\"\n> +               esac &&\n> +\n\nIn last round, I said I'd go with Ævar's scheme for iterating over\nconfigs.  But when looking at the test output I decided that I wanted\na shorter label for each config rather than the actual command line to\nmake hte output more readable.\n\n> +               # Set GIT_TEST_FSYNC=1 explicitly since fsync is normally\n> +               # disabled by t/test-lib.sh.\n> +               if ! test_perf \"$1 (fsyncMethod=$method)\" \\\n> +                                               --setup \"$2\" \\\n> +                                               \"GIT_TEST_FSYNC=1 git $cfg $3\"\n> +               then\n> +                       break\n> +               fi\n> +       done\n> +}\n\nSo here I split the 'git $cfg' invocation off of the actual command\nbeing executed, since it wasn't clear to me the best way to structure\nthis shell script.\n\nThe overall effect I want to achieve is to be able to iterate over\nevery config for each test case so that the different configs of the\nsame test appear next to each other in the output.\n\n> +\n> +test_perf_fsync_cfgs \"add $total_files files\" \\\n> +       \"setup_repo\" \\\n> +       \"add -- files\"\n> +\n\nI initially tried not substituting the $cfg variable in a test like this:\n'git $cfg add -- files'\n\nAnd then using eval in test_perf_fsync_cfgs to get the variable\nsubstitution to happen later.\n\nIs there a better way to write this?\n\nThanks,\nNeeraj\n"},{"id":"452770","messageId":"CANQDOdcDY-8TZzCHx+tWZJoD0rsULnfaWRhAOox3drSgxW_+ow@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqr16j6ve3.fsf@gitster.g","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T04:17:14Z","receivedAt":"2022-03-31T04:29:33Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 1:24 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> >> We should add a new function, flush_bulk_checking_packfile(), to\n> >> flush only the packfile part of the bulk_checkin_state without\n> >> affecting other things---the \"plugged\" bit is the only one in the\n> >> current code before this series, but it does not have to stay to be\n> >> so\n> >\n> > I'm happy to rename the packfile related stuff to end with _packfile\n> > to make it clear that all of that state and functionality is related\n> > to batching of packfile additions.\n>\n> I do not care about names, though.  If you took that I hinted any\n> such change, sorry about that.  _state is fine as-is.\n>\n> I do care about not ejecting plugged out of the structure and\n> instead keeping them together, with proper way to flush the part\n> that deflate_to_pack() wants to flush, instead of abusing the\n> \"finish\".\n>\n> Thanks.\n\nJust to understand your feedback better, is it a problem to separate\nthe state of each separate \"thing\" under ODB transactions into\nseparate file-scope global(s)?  In this series I declared the fsync\nstate as completely separate from the packfile state.  That's why I\nwas thinking of it as more of a naming problem, since the remaining\nstate aside from the plugged boolean is entirely packfile related.\n\nMy argument in favor of having separate file-scoped variables for each\n'pluggable thing' would be that future implementations can evolve\nseparately without authors first having to disentangle a single\nstruct.\n\nThanks,\nNeeraj\n"},{"id":"452772","messageId":"CANQDOdex_84nhQu=86jrkyVfnRenRtxQ_8B-mmnXE-a1DfuUJg@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqk0cb9x78.fsf@gitster.g","subject":"Re: [PATCH v5 02/14] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T05:51:46Z","receivedAt":"2022-03-31T05:52:08Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:17 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Neeraj Singh <neerajsi@microsoft.com>\n> >\n> > Make it clearer in the naming and documentation of the plug_bulk_checkin\n> > and unplug_bulk_checkin APIs that they can be thought of as\n> > a \"transaction\" to optimize operations on the object database. These\n> > transactions may be nested so that subsystems like the cache-tree\n> > writing code can optimize their operations without caring whether the\n> > top-level code has a transaction active.\n>\n> I can see that \"checkin\" part of the name is too limiting (you may\n> want to do more than optimize checkin, e.g. fsync), and that you may\n> prefer \"begin/end\" over \"plug/unplug\", but I am not sure if we want\n> to limit ourselves to \"odb\".  If we find our code doing things on\n> many instances of something that are not objects (e.g. refs, config\n> variables), don't we want to give them the same chance to be optimized\n> by batching them?\n>\n> {begin,end}_bulk_transaction perhaps?  I dunno.\n\nAt least in the current code where the implementation of each\n'database table' (odb, refs-db, config, index) is pretty separate, it\nseems better to keep bulk-checkin.c and its plugging scoped to the\nODB.  Patrick's (who I previously misnamed as 'Peter') older patch at\nhttps://lore.kernel.org/git/d9aa96913b1730f1d0c238d7d52e27c20bc55390.1636544377.git.ps@pks.im/\nshowed a pretty nice and concise implementation for refs tied to the\nref-transaction infrastructure.\n\nThanks,\nNeeraj\n"},{"id":"452773","messageId":"CANQDOdd-G0VHOKWjWQL75jAJ7Az4izB33HKzayqnk-F-nkHj_A@mail.gmail.com","threadId":"57568","inReplyTo":"xmqq4k3f9w9s.fsf@gitster.g","subject":"Re: [PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T06:28:56Z","receivedAt":"2022-03-31T06:29:20Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Wed, Mar 30, 2022 at 10:37 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Neeraj Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> > index 9da3e5d88f6..3c90ba0b395 100644\n> > --- a/Documentation/config/core.txt\n> > +++ b/Documentation/config/core.txt\n> > @@ -596,6 +596,14 @@ core.fsyncMethod::\n> >  * `writeout-only` issues pagecache writeback requests, but depending on the\n> >    filesystem and storage hardware, data added to the repository may not be\n> >    durable in the event of a system crash. This is the default mode on macOS.\n> > +* `batch` enables a mode that uses writeout-only flushes to stage multiple\n> > +  updates in the disk writeback cache and then does a single full fsync of\n> > +  a dummy file to trigger the disk cache flush at the end of the operation.\n> > ++\n> > +  Currently `batch` mode only applies to loose-object files. Other repository\n> > +  data is made durable as if `fsync` was specified. This mode is expected to\n> > +  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n> > +  and on Windows for repos stored on NTFS or ReFS filesystems.\n>\n> Does this format correctly?  I had an impression that the second and\n> subsequent paragraphs, connected with a line with a single \"+\" on\n> it, has to be flushed left without indentation.\n>\n\nHere's the man format (Not sure how this will render going through gmail):\n           •   batch enables a mode that uses writeout-only flushes to\nstage multiple updates in the disk writeback cache and then does a\nsingle full fsync of a dummy file to trigger the disk cache flush at\nthe end of the\n               operation.\n\n                   Currently `batch` mode only applies to loose-object\nfiles. Other repository\n                   data is made durable as if `fsync` was specified.\nThis mode is expected to\n                   be as safe as `fsync` on macOS for repos stored on\nHFS+ or APFS filesystems\n                   and on Windows for repos stored on NTFS or ReFS filesystems.\n\nTo describe the above if it doesn't render correctly, we have a\nbulleted list where the batch after the * is bolded.  Other instances\nof single backtick quoted text just appears as plaintext. The\ndescriptive \"Currently `batch` mode...\" paragraph is under the bullet\npoint and well-indented.\n\nIn HTML the output looks good as well, except that the descriptive\nparagraph is in monospace for some reason.\n\n> > diff --git a/bulk-checkin.c b/bulk-checkin.c\n> > index 8b0fd5c7723..9799d247cad 100644\n> > --- a/bulk-checkin.c\n> > +++ b/bulk-checkin.c\n> > @@ -3,15 +3,20 @@\n> >   */\n> >  #include \"cache.h\"\n> >  #include \"bulk-checkin.h\"\n> > +#include \"lockfile.h\"\n> >  #include \"repository.h\"\n> >  #include \"csum-file.h\"\n> >  #include \"pack.h\"\n> >  #include \"strbuf.h\"\n> > +#include \"string-list.h\"\n> > +#include \"tmp-objdir.h\"\n> >  #include \"packfile.h\"\n> >  #include \"object-store.h\"\n> >\n> >  static int odb_transaction_nesting;\n> >\n> > +static struct tmp_objdir *bulk_fsync_objdir;\n>\n> I wonder if this should be added to the bulk_checkin_state structure\n> as a new member, especially if we fix the erroneous call to\n> finish_bulk_checkin() as a preliminary fix-up of a bug that existed\n> even before this series.\n>\n\nIt seems like the only thing tying this to the bulk_checkin_state\n(which I've renamed in my local changes to bulk_checkin_packfile) is\nthat they're both generally written when a transaction is active.\nKeeping fsync separate from packfile should help the reader see that\nthe two sets of functions access only their respective global state.\nIf we add another optimization strategy (e.g. appendable pack files),\nit would get its own separate state and functions that are independent\nof the large-blob packfile and loose-object fsync optimizations.\n\n> > +/*\n> > + * Cleanup after batch-mode fsync_object_files.\n> > + */\n> > +static void do_batch_fsync(void)\n> > +{\n> > +     struct strbuf temp_path = STRBUF_INIT;\n> > +     struct tempfile *temp;\n> > +\n> > +     if (!bulk_fsync_objdir)\n> > +             return;\n> > +\n> > +     /*\n> > +      * Issue a full hardware flush against a temporary file to ensure\n> > +      * that all objects are durable before any renames occur. The code in\n> > +      * fsync_loose_object_bulk_checkin has already issued a writeout\n> > +      * request, but it has not flushed any writeback cache in the storage\n> > +      * hardware or any filesystem logs. This fsync call acts as a barrier\n> > +      * to ensure that the data in each new object file is durable before\n> > +      * the final name is visible.\n> > +      */\n> > +     strbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n> > +     temp = xmks_tempfile(temp_path.buf);\n> > +     fsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n> > +     delete_tempfile(&temp);\n> > +     strbuf_release(&temp_path);\n> > +\n> > +     /*\n> > +      * Make the object files visible in the primary ODB after their data is\n> > +      * fully durable.\n> > +      */\n> > +     tmp_objdir_migrate(bulk_fsync_objdir);\n> > +     bulk_fsync_objdir = NULL;\n> > +}\n>\n> OK.\n>\n> > +void prepare_loose_object_bulk_checkin(void)\n> > +{\n> > +     /*\n> > +      * We lazily create the temporary object directory\n> > +      * the first time an object might be added, since\n> > +      * callers may not know whether any objects will be\n> > +      * added at the time they call begin_odb_transaction.\n> > +      */\n> > +     if (!odb_transaction_nesting || bulk_fsync_objdir)\n> > +             return;\n> > +\n> > +     bulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n> > +     if (bulk_fsync_objdir)\n> > +             tmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n> > +}\n>\n> OK.  If we got a failure from tmp_objdir_create(), then we don't\n> swap and end up creating a new loose object file in the primary\n> object store.  I wonder if we at least want to note that fact for\n> later use at \"unplug\" time.  We may create a few loose objects in\n> the primary object store without any fsync, then a later call may\n> successfully create a temporary object directory and we'd create\n> more loose objects in the temporary one, which are flushed with the\n> \"create a dummy and fsync\" trick and migrated, but do we need to do\n> something to the ones we created in the primary object store before\n> all that happens?\n>\n> > +void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n> > +{\n> > +     /*\n> > +      * If we have an active ODB transaction, we issue a call that\n> > +      * cleans the filesystem page cache but avoids a hardware flush\n> > +      * command. Later on we will issue a single hardware flush\n> > +      * before as part of do_batch_fsync.\n> > +      */\n> > +     if (!bulk_fsync_objdir ||\n> > +         git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n> > +             fsync_or_die(fd, filename);\n> > +     }\n> > +}\n>\n> Ah, if we have successfully created the temporary directory, we\n> don't do full fsync but just writeout-only one, so there is no need\n> for the worry I mentioned in the previous paragraph.  OK.\n>\n\nThere is the possibility that we might create the objdir when calling\nprepare_loose_object_bulk_checkin but somehow fail to write the object\nand yet still make it to end_odb_transaction.  In that case, we'd\nissue an extra dummy fsync.  I don't think this case is worth extra\ncode to track, since it's a single fsync in a weird failure case.\n\n> > @@ -301,4 +370,6 @@ void end_odb_transaction(void)\n> >\n> >       if (bulk_checkin_state.f)\n> >               finish_bulk_checkin(&bulk_checkin_state);\n> > +\n> > +     do_batch_fsync();\n> >  }\n>\n> OK.\n>\n> > @@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       static struct strbuf tmp_file = STRBUF_INIT;\n> >       static struct strbuf filename = STRBUF_INIT;\n> >\n> > +     if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> > +             prepare_loose_object_bulk_checkin();\n> > +\n> >       loose_object_path(the_repository, &filename, oid);\n> >\n> >       fd = create_tmpfile(&tmp_file, filename.buf);\n>\n> The necessary change to the \"workhorse\" code path is surprisingly\n> small, which is pleasing to see.\n>\n> Thanks.\n\nThanks for looking at this.\n"},{"id":"452811","messageId":"xmqqy20q2eqn.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdcDY-8TZzCHx+tWZJoD0rsULnfaWRhAOox3drSgxW_+ow@mail.gmail.com","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-31T17:50:24Z","receivedAt":"2022-03-31T17:50:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> Just to understand your feedback better, is it a problem to separate\n> the state of each separate \"thing\" under ODB transactions into\n> separate file-scope global(s)?  In this series I declared the fsync\n> state as completely separate from the packfile state.  That's why I\n> was thinking of it as more of a naming problem, since the remaining\n> state aside from the plugged boolean is entirely packfile related.\n\nAhh, sorry, my mistake.\n\nI somehow thought that you would be making the existing \"struct\nbulk_checkin_state\" infrastructure to cover not just the object\nstore but much more, perhaps because I partly mistook the motivation\nto rename the structure (thinking again about it, since \"checkin\" is\nthe act of adding new objects to the object database from outside\nsources (either from the working tree using \"git add\" command, or\nfrom other sources using unpack-objects), the original name was\nalready fine to signal that it was about the object database, and\nthe need to rename it sounded like we were going to do much more\nthan the object database behind my head).\n\n> My argument in favor of having separate file-scoped variables for each\n> 'pluggable thing' would be that future implementations can evolve\n> separately without authors first having to disentangle a single\n> struct.\n\nThat is fine.  Would the trigger to \"plug\" and \"unplug\" also be\nindependent?  As you said elsewhere, the part to harden refs can\npiggyback on the existing ref-transaction infrastructure.  I do not\nknow offhand what things other than loose objects that want \"plug\"\nand \"unplug\" semantics, but if there are, are we going to have type\nspecific begin- and end-transaction?\n\nThanks.\n"},{"id":"452813","messageId":"xmqqtube2e1d.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdd-G0VHOKWjWQL75jAJ7Az4izB33HKzayqnk-F-nkHj_A@mail.gmail.com","subject":"Re: [PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-31T18:05:34Z","receivedAt":"2022-03-31T18:05:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> To describe the above if it doesn't render correctly, we have a\n> bulleted list where the batch after the * is bolded.  Other instances\n> of single backtick quoted text just appears as plaintext. The\n> descriptive \"Currently `batch` mode...\" paragraph is under the bullet\n> point and well-indented.\n>\n> In HTML the output looks good as well, except that the descriptive\n> paragraph is in monospace for some reason.\n\nThe \"except\" part admits that it does not render well, no?\n\nWhat happens if you modify the second and subsequent paragraph after\nthe \"+\" continuation in the way suggested?  Does it make it worse,\nor does it make it worse?\n\n> Keeping fsync separate from packfile should help the reader see that\n> the two sets of functions access only their respective global state.\n> If we add another optimization strategy (e.g. appendable pack files),\n> it would get its own separate state and functions that are independent\n> of the large-blob packfile and loose-object fsync optimizations.\n\nOK.\n\n>> > +void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n>> > +{\n>> > +     /*\n>> > +      * If we have an active ODB transaction, we issue a call that\n>> > +      * cleans the filesystem page cache but avoids a hardware flush\n>> > +      * command. Later on we will issue a single hardware flush\n>> > +      * before as part of do_batch_fsync.\n>> > +      */\n>> > +     if (!bulk_fsync_objdir ||\n>> > +         git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n>> > +             fsync_or_die(fd, filename);\n>> > +     }\n>> > +}\n>>\n>> Ah, if we have successfully created the temporary directory, we\n>> don't do full fsync but just writeout-only one, so there is no need\n>> for the worry I mentioned in the previous paragraph.  OK.\n>\n> There is the possibility that we might create the objdir when calling\n> prepare_loose_object_bulk_checkin but somehow fail to write the object\n> and yet still make it to end_odb_transaction.  In that case, we'd\n> issue an extra dummy fsync.  I don't think this case is worth extra\n> code to track, since it's a single fsync in a weird failure case.\n\nYup.\n"},{"id":"452819","messageId":"CANQDOde9+x087efFad+y9ir_amv-Km61-+7BnBn2CShw22rkcg@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqy20q2eqn.fsf@gitster.g","subject":"Re: [PATCH v5 01/14] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T19:08:57Z","receivedAt":"2022-03-31T19:09:16Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 31, 2022 at 10:50 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > Just to understand your feedback better, is it a problem to separate\n> > the state of each separate \"thing\" under ODB transactions into\n> > separate file-scope global(s)?  In this series I declared the fsync\n> > state as completely separate from the packfile state.  That's why I\n> > was thinking of it as more of a naming problem, since the remaining\n> > state aside from the plugged boolean is entirely packfile related.\n>\n> Ahh, sorry, my mistake.\n>\n> I somehow thought that you would be making the existing \"struct\n> bulk_checkin_state\" infrastructure to cover not just the object\n> store but much more, perhaps because I partly mistook the motivation\n> to rename the structure (thinking again about it, since \"checkin\" is\n> the act of adding new objects to the object database from outside\n> sources (either from the working tree using \"git add\" command, or\n> from other sources using unpack-objects), the original name was\n> already fine to signal that it was about the object database, and\n> the need to rename it sounded like we were going to do much more\n> than the object database behind my head).\n>\n> > My argument in favor of having separate file-scoped variables for each\n> > 'pluggable thing' would be that future implementations can evolve\n> > separately without authors first having to disentangle a single\n> > struct.\n>\n> That is fine.  Would the trigger to \"plug\" and \"unplug\" also be\n> independent?  As you said elsewhere, the part to harden refs can\n> piggyback on the existing ref-transaction infrastructure.  I do not\n> know offhand what things other than loose objects that want \"plug\"\n> and \"unplug\" semantics, but if there are, are we going to have type\n> specific begin- and end-transaction?\n>\n\nWith regards to bulk-checkin.h, I believe for simplicity of interface\nto the callers, there should be a single pair of APIs for plug or\nunplug of the entire ODB regardless of what optimizations happen under\nthe covers.  For eventual repo-wide transactions, there should be a\nsingle API to initiate a transaction and a single one to commit/abort\nthe transaction at the end.  We may still also want a flush API so\nthat we can make the repo state consistent prior to executing hooks or\ndoing something else where an outside process needs consistent repo\nstate.\n\nThanks,\nNeeraj\n"},{"id":"452824","messageId":"CANQDOdcVnyEo1Om=Odutix1ZT4vbNiZX29_Zo+=PfxnssmUM9g@mail.gmail.com","threadId":"57568","inReplyTo":"xmqqtube2e1d.fsf@gitster.g","subject":"Re: [PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T19:18:57Z","receivedAt":"2022-03-31T19:19:14Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Thu, Mar 31, 2022 at 11:05 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Neeraj Singh <nksingh85@gmail.com> writes:\n>\n> > To describe the above if it doesn't render correctly, we have a\n> > bulleted list where the batch after the * is bolded.  Other instances\n> > of single backtick quoted text just appears as plaintext. The\n> > descriptive \"Currently `batch` mode...\" paragraph is under the bullet\n> > point and well-indented.\n> >\n> > In HTML the output looks good as well, except that the descriptive\n> > paragraph is in monospace for some reason.\n>\n> The \"except\" part admits that it does not render well, no?\n>\n> What happens if you modify the second and subsequent paragraph after\n> the \"+\" continuation in the way suggested?  Does it make it worse,\n> or does it make it worse?\n>\n\nApologies, I misinterpreted your statement that the input \"has to be\nflushed left without indentation\".  Now that I flushed it left I'm\ngetting better output where the follow-on paragraph has a \"normal\"\ntext style and interior backtick quoted things are bolded as expected.\n\nThis will be fixed in the next iteration.\n"},{"id":"452902","messageId":"xmqqy20oyezh.fsf@gitster.g","threadId":"57568","inReplyTo":"CANQDOdcVnyEo1Om=Odutix1ZT4vbNiZX29_Zo+=PfxnssmUM9g@mail.gmail.com","subject":"Re: [PATCH v5 04/14] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-04-01T15:56:18Z","receivedAt":"2022-04-01T16:25:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n> ...  Now that I flushed it left I'm\n> getting better output where the follow-on paragraph has a \"normal\"\n> text style and interior backtick quoted things are bolded as expected.\n\nI was reasonably but not absolutely sure if that works inside\nbulletted list.  Thanks for experimenting it for us ;-)\n\n\n"},{"id":"453085","messageId":"20220405052018.11247-2-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 01/12] bulk-checkin: rename 'state' variable and separate 'plugged' boolean","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:07Z","receivedAt":"2022-04-05T05:20:40Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThis commit prepares for adding batch-fsync to the bulk-checkin\ninfrastructure.\n\nThe bulk-checkin infrastructure is currently used to batch up addition\nof large blobs to a packfile. When a blob is larger than\nbig_file_threshold, we unconditionally add it to a pack. If bulk\ncheckins are 'plugged', we allow multiple large blobs to be added to a\nsingle pack until we reach the packfile size limit; otherwise, we simply\nmake a new packfile for each large blob. The 'unplug' call tells us when\nthe series of blob additions is done so that we can finish the packfiles\nand make their objects available to subsequent operations.\n\nStated another way, bulk-checkin allows callers to define a transaction\nthat adds multiple objects to the object database, where the object\ndatabase can optimize its internal operations within the transaction\nboundary.\n\nBatched fsync will fit into bulk-checkin by taking advantage of the\nplug/unplug functionality to determine the appropriate time to fsync\nand make newly-added objects available in the primary object database.\n\n* Rename 'state' variable to 'bulk_checkin_packfile', since we will\n  later be adding 'bulk_fsync_objdir'. This also makes the variable\n  easier to find in the debugger, since the name is more unique.\n\n* Rename finish_bulk_checkin to flush_bulk_checkin_packfile and call it\n  unconditionally from unplug_bulk_checkin. Internally it will\n  conditionally do a flush if there's any work to do.\n\n* Move the 'plugged' data member of 'bulk_checkin_state' into a separate\n  static variable. Doing this avoids resetting the variable in\n  finish_bulk_checkin when zeroing the 'bulk_checkin_state'. As-is, we\n  seem to unintentionally disable the plugging functionality the first\n  time a new packfile must be created due to packfile size limits. While\n  disabling the plugging state only results in suboptimal behavior for\n  the current code, it would be fatal for the bulk-fsync functionality\n  later in this patch series.\n\nThe net effect of these changes is to make a clear separation between\nthe portion of the bulk-checkin infrastructure that is related to the\npackfile (nearly all of it at present) and the part that is related to\nother future optimizations of the ODB.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n bulk-checkin.c | 33 +++++++++++++++++----------------\n 1 file changed, 17 insertions(+), 16 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 6d6c37171c9..88d72178b2c 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,9 +10,9 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static struct bulk_checkin_state {\n-\tunsigned plugged:1;\n+static int bulk_checkin_plugged;\n \n+static struct bulk_checkin_packfile {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n \toff_t offset;\n@@ -21,7 +21,7 @@ static struct bulk_checkin_state {\n \tstruct pack_idx_entry **written;\n \tuint32_t alloc_written;\n \tuint32_t nr_written;\n-} state;\n+} bulk_checkin_packfile;\n \n static void finish_tmp_packfile(struct strbuf *basename,\n \t\t\t\tconst char *pack_tmp_name,\n@@ -39,7 +39,7 @@ static void finish_tmp_packfile(struct strbuf *basename,\n \tfree(idx_tmp_name);\n }\n \n-static void finish_bulk_checkin(struct bulk_checkin_state *state)\n+static void flush_bulk_checkin_packfile(struct bulk_checkin_packfile *state)\n {\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tstruct strbuf packname = STRBUF_INIT;\n@@ -80,7 +80,7 @@ static void finish_bulk_checkin(struct bulk_checkin_state *state)\n \treprepare_packed_git(the_repository);\n }\n \n-static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n+static int already_written(struct bulk_checkin_packfile *state, struct object_id *oid)\n {\n \tint i;\n \n@@ -112,7 +112,7 @@ static int already_written(struct bulk_checkin_state *state, struct object_id *o\n  * status before calling us just in case we ask it to call us again\n  * with a new pack.\n  */\n-static int stream_to_pack(struct bulk_checkin_state *state,\n+static int stream_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t  git_hash_ctx *ctx, off_t *already_hashed_to,\n \t\t\t  int fd, size_t size, enum object_type type,\n \t\t\t  const char *path, unsigned flags)\n@@ -189,7 +189,7 @@ static int stream_to_pack(struct bulk_checkin_state *state,\n }\n \n /* Lazily create backing packfile for the state */\n-static void prepare_to_stream(struct bulk_checkin_state *state,\n+static void prepare_to_stream(struct bulk_checkin_packfile *state,\n \t\t\t      unsigned flags)\n {\n \tif (!(flags & HASH_WRITE_OBJECT) || state->f)\n@@ -204,7 +204,7 @@ static void prepare_to_stream(struct bulk_checkin_state *state,\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n-static int deflate_to_pack(struct bulk_checkin_state *state,\n+static int deflate_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t   struct object_id *result_oid,\n \t\t\t   int fd, size_t size,\n \t\t\t   enum object_type type, const char *path,\n@@ -251,7 +251,7 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \t\t\tBUG(\"should not happen\");\n \t\thashfile_truncate(state->f, &checkpoint);\n \t\tstate->offset = checkpoint.offset;\n-\t\tfinish_bulk_checkin(state);\n+\t\tflush_bulk_checkin_packfile(state);\n \t\tif (lseek(fd, seekback, SEEK_SET) == (off_t) -1)\n \t\t\treturn error(\"cannot seek back\");\n \t}\n@@ -278,21 +278,22 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&state, oid, fd, size, type,\n+\tint status = deflate_to_pack(&bulk_checkin_packfile, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!state.plugged)\n-\t\tfinish_bulk_checkin(&state);\n+\tif (!bulk_checkin_plugged)\n+\t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n \treturn status;\n }\n \n void plug_bulk_checkin(void)\n {\n-\tstate.plugged = 1;\n+\tassert(!bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 1;\n }\n \n void unplug_bulk_checkin(void)\n {\n-\tstate.plugged = 0;\n-\tif (state.f)\n-\t\tfinish_bulk_checkin(&state);\n+\tassert(bulk_checkin_plugged);\n+\tbulk_checkin_plugged = 0;\n+\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n }\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453086","messageId":"20220405052018.11247-1-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 00/12] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:06Z","receivedAt":"2022-04-05T05:20:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGGG closed this series erroneously, so I'm trying out git-send-email. Apologies for any mistakes.\n\nThis series is also available at https://github.com/neerajsi-msft/git/git.git ns/batched-fsync-v6.\n\nV6 changes:\n\n* Based on master at faa21c1 to pick up ns/fsync-or-die-message-fix. Also resolved a conflict with 8aa0209 in t/perf/p7519-fsmonitor.sh.\n\n* Some independent patches were submitted separately on-list. This series is now dependent on ns/fsync-or-die-message-fix.\n\n* Rename bulk_checkin_state to bulk_checkin_packfile to discourage future authors from adding any non-packfile related stuff to it. Each individual component of bulk_checkin should have its own state variable(s) going forward, and they should only be tied together by odb_transaction_nesting.\n\n* Rename finish_bulk_checkin and do_batch_fsync to flush_bulk_checkin and flush_batch_fsync. The \"finish\" step is going to be the end_odb_transaction. The \"flush\" terminology should be consistently used for making changes visible.\n\n* Add flush_odb_transaction and use it in update-index before printing verbose output to mitigate risk of missing objects for a tricky stdin feeder.\n\n* Re-add shell \"local with assignment\": now these are all on a separate line with quotes around any values, to comply with dash. I'm running on ubuntu 20.04 LTS where I saw some of the dash issues before.\n\nNeeraj Singh (12):\n  bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n  bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n  core.fsyncmethod: batched disk flushes for loose-objects\n  cache-tree: use ODB transaction around writing a tree\n  builtin/add: add ODB transaction around add_files_to_cache\n  update-index: use the bulk-checkin infrastructure\n  unpack-objects: use the bulk-checkin infrastructure\n  core.fsync: use batch mode and sync loose objects by default on\n    Windows\n  test-lib-functions: add parsing helpers for ls-files and ls-tree\n  core.fsyncmethod: tests for batch mode\n  t/perf: add iteration setup mechanism to perf-lib\n  core.fsyncmethod: performance tests for batch mode\n\n Documentation/config/core.txt          |   8 ++\n builtin/add.c                          |  13 ++-\n builtin/unpack-objects.c               |   3 +\n builtin/update-index.c                 |  20 +++++\n bulk-checkin.c                         | 117 +++++++++++++++++++++----\n bulk-checkin.h                         |  27 +++++-\n cache-tree.c                           |   3 +\n cache.h                                |  12 ++-\n compat/mingw.h                         |   3 +\n config.c                               |   4 +-\n git-compat-util.h                      |   2 +\n object-file.c                          |   7 +-\n t/lib-unique-files.sh                  |  34 +++++++\n t/perf/p0008-odb-fsync.sh              |  82 +++++++++++++++++\n t/perf/p4220-log-grep-engines.sh       |   3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |   3 +-\n t/perf/p5302-pack-index.sh             |  15 ++--\n t/perf/p7519-fsmonitor.sh              |  18 +---\n t/perf/p7820-grep-engines.sh           |   6 +-\n t/perf/perf-lib.sh                     |  63 +++++++++++--\n t/t3700-add.sh                         |  28 ++++++\n t/t3903-stash.sh                       |  20 +++++\n t/t5300-pack-object.sh                 |  41 ++++++---\n t/t5317-pack-objects-filter-objects.sh |  91 ++++++++++---------\n t/test-lib-functions.sh                |  10 +++\n 25 files changed, 513 insertions(+), 120 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n create mode 100755 t/perf/p0008-odb-fsync.sh\n\nRange-diff against v5:\n 1:  c7a2a7efe6d <  -:  ----------- bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n -:  ----------- >  1:  adabdaa0290 bulk-checkin: rename 'state' variable and separate 'plugged' boolean\n 2:  d045b13795b !  2:  72a6cd36c9c bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'\n    @@ Commit message\n         writing code can optimize their operations without caring whether the\n         top-level code has a transaction active.\n     \n    +    Add a flush_odb_transaction API that will be used in update-index to\n    +    make objects visible even if a transaction is active. The flush call may\n    +    also be useful in future cases if we hold a transaction active around\n    +    calling hooks.\n    +\n         Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n     \n      ## builtin/add.c ##\n    @@ bulk-checkin.c\n     -static int bulk_checkin_plugged;\n     +static int odb_transaction_nesting;\n      \n    - static struct bulk_checkin_state {\n    + static struct bulk_checkin_packfile {\n      \tchar *pack_tmp_name;\n     @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n      {\n    - \tint status = deflate_to_pack(&bulk_checkin_state, oid, fd, size, type,\n    + \tint status = deflate_to_pack(&bulk_checkin_packfile, oid, fd, size, type,\n      \t\t\t\t     path, flags);\n     -\tif (!bulk_checkin_plugged)\n     +\tif (!odb_transaction_nesting)\n    - \t\tfinish_bulk_checkin(&bulk_checkin_state);\n    + \t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n      \treturn status;\n      }\n      \n    @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n      }\n      \n     -void unplug_bulk_checkin(void)\n    -+void end_odb_transaction(void)\n    ++void flush_odb_transaction(void)\n      {\n     -\tassert(bulk_checkin_plugged);\n     -\tbulk_checkin_plugged = 0;\n    + \tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n    + }\n    ++\n    ++void end_odb_transaction(void)\n    ++{\n     +\todb_transaction_nesting -= 1;\n     +\tif (odb_transaction_nesting < 0)\n     +\t\tBUG(\"Unbalanced ODB transaction nesting\");\n    @@ bulk-checkin.c: int index_bulk_checkin(struct object_id *oid,\n     +\tif (odb_transaction_nesting)\n     +\t\treturn;\n     +\n    - \tif (bulk_checkin_state.f)\n    - \t\tfinish_bulk_checkin(&bulk_checkin_state);\n    - }\n    ++\tflush_odb_transaction();\n    ++}\n     \n      ## bulk-checkin.h ##\n     @@ bulk-checkin.h: int index_bulk_checkin(struct object_id *oid,\n    @@ bulk-checkin.h: int index_bulk_checkin(struct object_id *oid,\n     +/*\n     + * Tell the object database to optimize for adding\n     + * multiple objects. end_odb_transaction must be called\n    -+ * to make new objects visible.\n    ++ * to make new objects visible. Transactions can be nested,\n    ++ * and objects are only visible after the outermost transaction\n    ++ * is complete or the transaction is flushed.\n     + */\n     +void begin_odb_transaction(void);\n     +\n     +/*\n    ++ * Make any objects that are currently part of a pending object\n    ++ * database transaction visible. It is valid to call this function\n    ++ * even if no transaction is active.\n    ++ */\n    ++void flush_odb_transaction(void);\n    ++\n    ++/*\n     + * Tell the object database to make any objects from the\n    -+ * current transaction visible.\n    ++ * current transaction visible if this is the final nested\n    ++ * transaction.\n     + */\n     +void end_odb_transaction(void);\n      \n 3:  2d1bc4568ac <  -:  ----------- object-file: pass filename to fsync_or_die\n 4:  9e7ae22fa4a !  3:  57539f104ef core.fsyncmethod: batched disk flushes for loose-objects\n    @@ Documentation/config/core.txt: core.fsyncMethod::\n     +  updates in the disk writeback cache and then does a single full fsync of\n     +  a dummy file to trigger the disk cache flush at the end of the operation.\n     ++\n    -+  Currently `batch` mode only applies to loose-object files. Other repository\n    -+  data is made durable as if `fsync` was specified. This mode is expected to\n    -+  be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n    -+  and on Windows for repos stored on NTFS or ReFS filesystems.\n    ++Currently `batch` mode only applies to loose-object files. Other repository\n    ++data is made durable as if `fsync` was specified. This mode is expected to\n    ++be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n    ++and on Windows for repos stored on NTFS or ReFS filesystems.\n      \n      core.fsyncObjectFiles::\n      \tThis boolean will enable 'fsync()' when writing object files.\n    @@ bulk-checkin.c\n      \n     +static struct tmp_objdir *bulk_fsync_objdir;\n     +\n    - static struct bulk_checkin_state {\n    + static struct bulk_checkin_packfile {\n      \tchar *pack_tmp_name;\n      \tstruct hashfile *f;\n    -@@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state)\n    +@@ bulk-checkin.c: static void flush_bulk_checkin_packfile(struct bulk_checkin_packfile *state)\n      \treprepare_packed_git(the_repository);\n      }\n      \n     +/*\n     + * Cleanup after batch-mode fsync_object_files.\n     + */\n    -+static void do_batch_fsync(void)\n    ++static void flush_batch_fsync(void)\n     +{\n     +\tstruct strbuf temp_path = STRBUF_INIT;\n     +\tstruct tempfile *temp;\n    @@ bulk-checkin.c: static void finish_bulk_checkin(struct bulk_checkin_state *state\n     +\tbulk_fsync_objdir = NULL;\n     +}\n     +\n    - static int already_written(struct bulk_checkin_state *state, struct object_id *oid)\n    + static int already_written(struct bulk_checkin_packfile *state, struct object_id *oid)\n      {\n      \tint i;\n    -@@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n    +@@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_packfile *state,\n      \treturn 0;\n      }\n      \n    @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n     +\t * If we have an active ODB transaction, we issue a call that\n     +\t * cleans the filesystem page cache but avoids a hardware flush\n     +\t * command. Later on we will issue a single hardware flush\n    -+\t * before as part of do_batch_fsync.\n    ++\t * before renaming the objects to their final names as part of\n    ++\t * flush_batch_fsync.\n     +\t */\n     +\tif (!bulk_fsync_objdir ||\n     +\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n    @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n      int index_bulk_checkin(struct object_id *oid,\n      \t\t       int fd, size_t size, enum object_type type,\n      \t\t       const char *path, unsigned flags)\n    -@@ bulk-checkin.c: void end_odb_transaction(void)\n    +@@ bulk-checkin.c: void begin_odb_transaction(void)\n      \n    - \tif (bulk_checkin_state.f)\n    - \t\tfinish_bulk_checkin(&bulk_checkin_state);\n    -+\n    -+\tdo_batch_fsync();\n    + void flush_odb_transaction(void)\n    + {\n    ++\tflush_batch_fsync();\n    + \tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n      }\n    + \n     \n      ## bulk-checkin.h ##\n     @@\n 5:  83fa4a5f3a5 =  4:  f47631e6a28 cache-tree: use ODB transaction around writing a tree\n 6:  d514842ad49 =  5:  08c9b234942 builtin/add: add ODB transaction around add_files_to_cache\n 7:  8cac94598a5 !  6:  bc37cdbd226 update-index: use the bulk-checkin infrastructure\n    @@ Commit message\n         There is some risk with this change, since under batch fsync, the object\n         files will be in a tmp-objdir until update-index is complete, so callers\n         using the --stdin option will not see them until update-index is done.\n    -    This risk is mitigated by not keeping an ODB transaction open around\n    -    --stdin processing if in --verbose mode. Without --verbose mode,\n    -    a caller feeding update-index via --stdin wouldn't know when\n    -    update-index adds an object, event without an ODB transaction.\n    +    This risk is mitigated by flushing the ODB transaction prior to\n    +    reporting any verbose output so that objects will be visible to callers\n    +    that are synchronizing with update-index by snooping its output.\n     \n         Signed-off-by: Neeraj Singh <neerajsi@microsoft.com>\n     \n    @@ builtin/update-index.c\n      #include \"config.h\"\n      #include \"lockfile.h\"\n      #include \"quote.h\"\n    +@@ builtin/update-index.c: static void report(const char *fmt, ...)\n    + \tif (!verbose)\n    + \t\treturn;\n    + \n    ++\t/*\n    ++\t * It is possible, though unlikely, that a caller could use the verbose\n    ++\t * output to synchronize with addition of objects to the object\n    ++\t * database. The current implementation of ODB transactions leaves\n    ++\t * objects invisible while a transaction is active, so flush the\n    ++\t * transaction here before reporting a change made by update-index.\n    ++\t */\n    ++\tflush_odb_transaction();\n    + \tva_start(vp, fmt);\n    + \tvprintf(fmt, vp);\n    + \tputchar('\\n');\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n      \t */\n      \tparse_options_start(&ctx, argc, argv, prefix,\n    @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n      \twhile (ctx.argc) {\n      \t\tif (parseopt_state != PARSE_OPT_DONE)\n      \t\t\tparseopt_state = parse_options_step(&ctx, options,\n    -@@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n    - \t\tthe_index.version = preferred_index_format;\n    - \t}\n    - \n    -+\t/*\n    -+\t * It is possible, though unlikely, that a caller could use the verbose\n    -+\t * output to synchronize with addition of objects to the object\n    -+\t * database. The current implementation of ODB transactions leaves\n    -+\t * objects invisible while a transaction is active, so end the\n    -+\t * transaction here if verbose output is enabled.\n    -+\t */\n    -+\n    -+\tif (verbose)\n    -+\t\tend_odb_transaction();\n    -+\n    - \tif (read_from_stdin) {\n    - \t\tstruct strbuf buf = STRBUF_INIT;\n    - \t\tstruct strbuf unquoted = STRBUF_INIT;\n     @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const char *prefix)\n      \t\tstrbuf_release(&buf);\n      \t}\n    @@ builtin/update-index.c: int cmd_update_index(int argc, const char **argv, const\n     +\t/*\n     +\t * By now we have added all of the new objects\n     +\t */\n    -+\tif (!verbose)\n    -+\t\tend_odb_transaction();\n    ++\tend_odb_transaction();\n     +\n      \tif (split_index > 0) {\n      \t\tif (git_config_get_split_index() == 0)\n 8:  523e5fbd63e =  7:  9cf584cdb67 unpack-objects: use the bulk-checkin infrastructure\n 9:  faacc19aab2 =  8:  5039b596064 core.fsync: use batch mode and sync loose objects by default on Windows\n10:  4de7300a7b0 =  9:  67205b3ac25 test-lib-functions: add parsing helpers for ls-files and ls-tree\n11:  1a4aff8c350 ! 10:  148e562ddb8 core.fsyncmethod: tests for batch mode\n    @@ t/lib-unique-files.sh (new)\n     +\tlocal dirs=\"$1\" &&\n     +\tlocal files=\"$2\" &&\n     +\tlocal basedir=\"$3\" &&\n    -+\tlocal counter=0 &&\n    ++\tlocal counter=\"0\" &&\n     +\tlocal i &&\n     +\tlocal j &&\n     +\ttest_tick &&\n    -+\tlocal basedata=$basedir$test_tick &&\n    ++\tlocal basedata=\"$basedir$test_tick\" &&\n     +\trm -rf \"$basedir\" &&\n     +\tfor i in $(test_seq $dirs)\n     +\tdo\n    -+\t\tlocal dir=$basedir/dir$i &&\n    ++\t\tlocal dir=\"$basedir/dir$i\" &&\n     +\t\tmkdir -p \"$dir\" &&\n     +\t\tfor j in $(test_seq $files)\n     +\t\tdo\n12:  47cc63e1dda ! 11:  1a8320828f7 t/perf: add iteration setup mechanism to perf-lib\n    @@ t/perf/p7519-fsmonitor.sh: then\n     -\tfi\n     -fi\n     -\n    - trace_start() {\n    + trace_start () {\n      \tif test -n \"$GIT_PERF_7519_TRACE\"\n      \tthen\n    -@@ t/perf/p7519-fsmonitor.sh: setup_for_fsmonitor() {\n    +@@ t/perf/p7519-fsmonitor.sh: setup_for_fsmonitor_hook () {\n      \n      test_perf_w_drop_caches () {\n      \tif test -n \"$GIT_PERF_7519_DROP_CACHE\"; then\n    @@ t/perf/p7519-fsmonitor.sh: setup_for_fsmonitor() {\n     -\ttest_perf \"$@\"\n      }\n      \n    - test_fsmonitor_suite() {\n    + test_fsmonitor_suite () {\n     \n      ## t/perf/p7820-grep-engines.sh ##\n     @@ t/perf/p7820-grep-engines.sh: do\n    @@ t/perf/perf-lib.sh: exit $ret' >&3 2>&4\n      }\n      \n      test_wrapper_ () {\n    -+\tlocal test_wrapper_func_ test_title_\n    - \ttest_wrapper_func_=$1; shift\n    -+\ttest_title_=$1; shift\n    +-\ttest_wrapper_func_=$1; shift\n    ++\tlocal test_wrapper_func_=\"$1\"; shift\n    ++\tlocal test_title_=\"$1\"; shift\n      \ttest_start_\n     -\ttest \"$#\" = 3 && { test_prereq=$1; shift; } || test_prereq=\n     -\ttest \"$#\" = 2 ||\n13:  26be6ecb28b ! 12:  3eb5a6720cb core.fsyncmethod: performance tests for batch mode\n    @@ t/perf/p0008-odb-fsync.sh (new)\n     +}\n     +\n     +test_perf_fsync_cfgs () {\n    -+\tlocal method cfg &&\n    ++\tlocal method &&\n    ++\tlocal cfg &&\n     +\tfor method in none fsync batch writeout-only\n     +\tdo\n     +\t\tcase $method in\n14:  88c1f84d4c3 <  -:  ----------- core.fsyncmethod: correctly camel-case warning message\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453087","messageId":"20220405052018.11247-5-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 04/12] cache-tree: use ODB transaction around writing a tree","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:10Z","receivedAt":"2022-04-05T05:20:46Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTake advantage of the odb transaction infrastructure around writing the\ncached tree to the object database.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache-tree.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6752f69d515..8c5e8822716 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -3,6 +3,7 @@\n #include \"tree.h\"\n #include \"tree-walk.h\"\n #include \"cache-tree.h\"\n+#include \"bulk-checkin.h\"\n #include \"object-store.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n@@ -474,8 +475,10 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \ttrace_performance_enter();\n \ttrace2_region_enter(\"cache_tree\", \"update\", the_repository);\n+\tbegin_odb_transaction();\n \ti = update_one(istate->cache_tree, istate->cache, istate->cache_nr,\n \t\t       \"\", 0, &skip, flags);\n+\tend_odb_transaction();\n \ttrace2_region_leave(\"cache_tree\", \"update\", the_repository);\n \ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453088","messageId":"20220405052018.11247-11-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 10/12] core.fsyncmethod: tests for batch mode","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:16Z","receivedAt":"2022-04-05T05:20:50Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd test cases to exercise batch mode for:\n * 'git add'\n * 'git stash'\n * 'git update-index'\n * 'git unpack-objects'\n\nThese tests ensure that the added data winds up in the object database.\n\nIn this change we introduce a new test helper lib-unique-files.sh. The\ngoal of this library is to create a tree of files that have different\noids from any other files that may have been created in the current test\nrepo. This helps us avoid missing validation of an object being added\ndue to it already being in the repo.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/lib-unique-files.sh  | 34 ++++++++++++++++++++++++++++++++++\n t/t3700-add.sh         | 28 ++++++++++++++++++++++++++++\n t/t3903-stash.sh       | 20 ++++++++++++++++++++\n t/t5300-pack-object.sh | 41 +++++++++++++++++++++++++++--------------\n 4 files changed, 109 insertions(+), 14 deletions(-)\n create mode 100644 t/lib-unique-files.sh\n\ndiff --git a/t/lib-unique-files.sh b/t/lib-unique-files.sh\nnew file mode 100644\nindex 00000000000..a14080fe79b\n--- /dev/null\n+++ b/t/lib-unique-files.sh\n@@ -0,0 +1,34 @@\n+# Helper to create files with unique contents\n+\n+# Create multiple files with unique contents within this test run. Takes the\n+# number of directories, the number of files in each directory, and the base\n+# directory.\n+#\n+# test_create_unique_files 2 3 my_dir -- Creates 2 directories with 3 files\n+#\t\t\t\t\t each in my_dir, all with contents\n+#\t\t\t\t\t different from previous invocations\n+#\t\t\t\t\t of this command in this run.\n+\n+test_create_unique_files () {\n+\ttest \"$#\" -ne 3 && BUG \"3 param\"\n+\n+\tlocal dirs=\"$1\" &&\n+\tlocal files=\"$2\" &&\n+\tlocal basedir=\"$3\" &&\n+\tlocal counter=\"0\" &&\n+\tlocal i &&\n+\tlocal j &&\n+\ttest_tick &&\n+\tlocal basedata=\"$basedir$test_tick\" &&\n+\trm -rf \"$basedir\" &&\n+\tfor i in $(test_seq $dirs)\n+\tdo\n+\t\tlocal dir=\"$basedir/dir$i\" &&\n+\t\tmkdir -p \"$dir\" &&\n+\t\tfor j in $(test_seq $files)\n+\t\tdo\n+\t\t\tcounter=$((counter + 1)) &&\n+\t\t\techo \"$basedata.$counter\">\"$dir/file$j.txt\"\n+\t\tdone\n+\tdone\n+}\ndiff --git a/t/t3700-add.sh b/t/t3700-add.sh\nindex b1f90ba3250..8979c8a5f03 100755\n--- a/t/t3700-add.sh\n+++ b/t/t3700-add.sh\n@@ -8,6 +8,8 @@ test_description='Test of git add, including the -- option.'\n TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n # Test the file mode \"$1\" of the file \"$2\" in the index.\n test_mode_in_index () {\n \tcase \"$(git ls-files -s \"$2\")\" in\n@@ -34,6 +36,32 @@ test_expect_success \\\n     'Test that \"git add -- -q\" works' \\\n     'touch -- -q && git add -- -q'\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'git add: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir1 &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION add -- ./files_base_dir1/ &&\n+\tgit ls-files --stage files_base_dir1/ |\n+\ttest_parse_ls_files_stage_oids >added_files_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files_oids >added_files_actual &&\n+\ttest_cmp added_files_oids added_files_actual\n+\"\n+\n+test_expect_success 'git update-index: core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir2 &&\n+\tfind files_base_dir2 ! -type d -print | xargs git $BATCH_CONFIGURATION update-index --add -- &&\n+\tgit ls-files --stage files_base_dir2 |\n+\ttest_parse_ls_files_stage_oids >added_files2_oids &&\n+\n+\t# We created 2 subdirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 added_files2_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <added_files2_oids >added_files2_actual &&\n+\ttest_cmp added_files2_oids added_files2_actual\n+\"\n+\n test_expect_success \\\n \t'git add: Test that executable bit is not used if core.filemode=0' \\\n \t'git config core.filemode 0 &&\ndiff --git a/t/t3903-stash.sh b/t/t3903-stash.sh\nindex 4abbc8fccae..20e94881964 100755\n--- a/t/t3903-stash.sh\n+++ b/t/t3903-stash.sh\n@@ -9,6 +9,7 @@ GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n \n . ./test-lib.sh\n+. $TEST_DIRECTORY/lib-unique-files.sh\n \n test_expect_success 'usage on cmd and subcommand invalid option' '\n \ttest_expect_code 129 git stash --invalid-option 2>usage &&\n@@ -1410,6 +1411,25 @@ test_expect_success 'stash handles skip-worktree entries nicely' '\n \tgit rev-parse --verify refs/stash:A.t\n '\n \n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'stash with core.fsyncmethod=batch' \"\n+\ttest_create_unique_files 2 4 files_base_dir &&\n+\tGIT_TEST_FSYNC=1 git $BATCH_CONFIGURATION stash push -u -- ./files_base_dir/ &&\n+\n+\t# The files were untracked, so use the third parent,\n+\t# which contains the untracked files\n+\tgit ls-tree -r stash^3 -- ./files_base_dir/ |\n+\ttest_parse_ls_tree_oids >stashed_files_oids &&\n+\n+\t# We created 2 dirs with 4 files each (8 files total) above\n+\ttest_line_count = 8 stashed_files_oids &&\n+\tgit cat-file --batch-check='%(objectname)' <stashed_files_oids >stashed_files_actual &&\n+\ttest_cmp stashed_files_oids stashed_files_actual\n+\"\n+\n+\n test_expect_success 'git stash succeeds despite directory/file change' '\n \ttest_create_repo directory_file_switch_v1 &&\n \t(\ndiff --git a/t/t5300-pack-object.sh b/t/t5300-pack-object.sh\nindex a11d61206ad..f8a0f309e2d 100755\n--- a/t/t5300-pack-object.sh\n+++ b/t/t5300-pack-object.sh\n@@ -161,22 +161,27 @@ test_expect_success 'pack-objects with bogus arguments' '\n '\n \n check_unpack () {\n+\tlocal packname=\"$1\" &&\n+\tlocal object_list=\"$2\" &&\n+\tlocal git_config=\"$3\" &&\n \ttest_when_finished \"rm -rf git2\" &&\n-\tgit init --bare git2 &&\n-\tgit -C git2 unpack-objects -n <\"$1\".pack &&\n-\tgit -C git2 unpack-objects <\"$1\".pack &&\n-\t(cd .git && find objects -type f -print) |\n-\twhile read path\n-\tdo\n-\t\tcmp git2/$path .git/$path || {\n-\t\t\techo $path differs.\n-\t\t\treturn 1\n-\t\t}\n-\tdone\n+\tgit $git_config init --bare git2 &&\n+\t(\n+\t\tgit $git_config -C git2 unpack-objects -n <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 unpack-objects <\"$packname\".pack &&\n+\t\tgit $git_config -C git2 cat-file --batch-check=\"%(objectname)\"\n+\t) <\"$object_list\" >current &&\n+\tcmp \"$object_list\" current\n }\n \n test_expect_success 'unpack without delta' '\n-\tcheck_unpack test-1-${packname_1}\n+\tcheck_unpack test-1-${packname_1} obj-list\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack without delta (core.fsyncmethod=batch)' '\n+\tcheck_unpack test-1-${packname_1} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with REF_DELTA' '\n@@ -185,7 +190,11 @@ test_expect_success 'pack with REF_DELTA' '\n '\n \n test_expect_success 'unpack with REF_DELTA' '\n-\tcheck_unpack test-2-${packname_2}\n+\tcheck_unpack test-2-${packname_2} obj-list\n+'\n+\n+test_expect_success 'unpack with REF_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-2-${packname_2} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'pack with OFS_DELTA' '\n@@ -195,7 +204,11 @@ test_expect_success 'pack with OFS_DELTA' '\n '\n \n test_expect_success 'unpack with OFS_DELTA' '\n-\tcheck_unpack test-3-${packname_3}\n+\tcheck_unpack test-3-${packname_3} obj-list\n+'\n+\n+test_expect_success 'unpack with OFS_DELTA (core.fsyncmethod=batch)' '\n+       check_unpack test-3-${packname_3} obj-list \"$BATCH_CONFIGURATION\"\n '\n \n test_expect_success 'compare delta flavors' '\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453089","messageId":"20220405052018.11247-13-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 12/12] core.fsyncmethod: performance tests for batch mode","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:18Z","receivedAt":"2022-04-05T05:20:52Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nAdd basic performance tests for git commands that can add data to the\nobject database. We cover:\n* git add\n* git stash\n* git update-index (via git stash)\n* git unpack-objects\n* git commit --all\n\nWe cover all currently available fsync methods as well.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p0008-odb-fsync.sh | 82 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 82 insertions(+)\n create mode 100755 t/perf/p0008-odb-fsync.sh\n\ndiff --git a/t/perf/p0008-odb-fsync.sh b/t/perf/p0008-odb-fsync.sh\nnew file mode 100755\nindex 00000000000..b3a90f30eba\n--- /dev/null\n+++ b/t/perf/p0008-odb-fsync.sh\n@@ -0,0 +1,82 @@\n+#!/bin/sh\n+#\n+# This test measures the performance of adding new files to the object\n+# database. The test was originally added to measure the effect of the\n+# core.fsyncMethod=batch mode, which is why we are testing different values of\n+# that setting explicitly and creating a lot of unique objects.\n+\n+test_description=\"Tests performance of adding things to the object database\"\n+\n+. ./perf-lib.sh\n+\n+. $TEST_DIRECTORY/lib-unique-files.sh\n+\n+test_perf_fresh_repo\n+test_checkout_worktree\n+\n+dir_count=10\n+files_per_dir=50\n+total_files=$((dir_count * files_per_dir))\n+\n+populate_files () {\n+\ttest_create_unique_files $dir_count $files_per_dir files\n+}\n+\n+setup_repo () {\n+\t(rm -rf .git || 1) &&\n+\tgit init &&\n+\ttest_commit first &&\n+\tpopulate_files\n+}\n+\n+test_perf_fsync_cfgs () {\n+\tlocal method &&\n+\tlocal cfg &&\n+\tfor method in none fsync batch writeout-only\n+\tdo\n+\t\tcase $method in\n+\t\tnone)\n+\t\t\tcfg=\"-c core.fsync=none\"\n+\t\t\t;;\n+\t\t*)\n+\t\t\tcfg=\"-c core.fsync=loose-object -c core.fsyncMethod=$method\"\n+\t\tesac &&\n+\n+\t\t# Set GIT_TEST_FSYNC=1 explicitly since fsync is normally\n+\t\t# disabled by t/test-lib.sh.\n+\t\tif ! test_perf \"$1 (fsyncMethod=$method)\" \\\n+\t\t\t\t\t\t--setup \"$2\" \\\n+\t\t\t\t\t\t\"GIT_TEST_FSYNC=1 git $cfg $3\"\n+\t\tthen\n+\t\t\tbreak\n+\t\tfi\n+\tdone\n+}\n+\n+test_perf_fsync_cfgs \"add $total_files files\" \\\n+\t\"setup_repo\" \\\n+\t\"add -- files\"\n+\n+test_perf_fsync_cfgs \"stash $total_files files\" \\\n+\t\"setup_repo\" \\\n+\t\"stash push -u -- files\"\n+\n+test_perf_fsync_cfgs \"unpack $total_files files\" \\\n+\t\"\n+\tsetup_repo &&\n+\tgit -c core.fsync=none add -- files &&\n+\tgit -c core.fsync=none commit -q -m second &&\n+\techo HEAD | git pack-objects -q --stdout --revs >test_pack.pack &&\n+\tsetup_repo\n+\t\" \\\n+\t\"unpack-objects -q <test_pack.pack\"\n+\n+test_perf_fsync_cfgs \"commit $total_files files\" \\\n+\t\"\n+\tsetup_repo &&\n+\tgit -c core.fsync=none add -- files &&\n+\tpopulate_files\n+\t\" \\\n+\t\"commit -q -a -m test\"\n+\n+test_done\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453090","messageId":"20220405052018.11247-12-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 11/12] t/perf: add iteration setup mechanism to perf-lib","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:17Z","receivedAt":"2022-04-05T05:20:54Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nTests that affect the repo in stateful ways are easier to write if we\ncan run setup steps outside of the measured portion of perf iteration.\n\nThis change adds a \"--setup 'setup-script'\" parameter to test_perf. To\nmake invocations easier to understand, I also moved the prerequisites to\na new --prereq parameter.\n\nThe setup facility will be used in the upcoming perf tests for batch\nmode, but it already helps in some existing tests, like t5302 and t7820.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/perf/p4220-log-grep-engines.sh       |  3 +-\n t/perf/p4221-log-grep-engines-fixed.sh |  3 +-\n t/perf/p5302-pack-index.sh             | 15 +++---\n t/perf/p7519-fsmonitor.sh              | 18 ++------\n t/perf/p7820-grep-engines.sh           |  6 ++-\n t/perf/perf-lib.sh                     | 63 +++++++++++++++++++++++---\n 6 files changed, 74 insertions(+), 34 deletions(-)\n\ndiff --git a/t/perf/p4220-log-grep-engines.sh b/t/perf/p4220-log-grep-engines.sh\nindex 2bc47ded4d1..03fbfbb85d3 100755\n--- a/t/perf/p4220-log-grep-engines.sh\n+++ b/t/perf/p4220-log-grep-engines.sh\n@@ -36,7 +36,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4220_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4220_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p4221-log-grep-engines-fixed.sh b/t/perf/p4221-log-grep-engines-fixed.sh\nindex 060971265a9..0a6d6dfc219 100755\n--- a/t/perf/p4221-log-grep-engines-fixed.sh\n+++ b/t/perf/p4221-log-grep-engines-fixed.sh\n@@ -26,7 +26,8 @@ do\n \t\telse\n \t\t\tprereq=\"\"\n \t\tfi\n-\t\ttest_perf $prereq \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \"\n+\t\ttest_perf \"$engine log$GIT_PERF_4221_LOG_OPTS --grep='$pattern'\" \\\n+\t\t\t--prereq \"$prereq\" \"\n \t\t\tgit -c grep.patternType=$engine log --pretty=format:%h$GIT_PERF_4221_LOG_OPTS --grep='$pattern' >'out.$engine' || :\n \t\t\"\n \tdone\ndiff --git a/t/perf/p5302-pack-index.sh b/t/perf/p5302-pack-index.sh\nindex c16f6a3ff69..14c601bbf86 100755\n--- a/t/perf/p5302-pack-index.sh\n+++ b/t/perf/p5302-pack-index.sh\n@@ -26,9 +26,8 @@ test_expect_success 'set up thread-counting tests' '\n \tdone\n '\n \n-test_perf PERF_EXTRA 'index-pack 0 threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack 0 threads' --prereq PERF_EXTRA \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --threads=1 --stdin < $PACK\n '\n \n@@ -36,17 +35,15 @@ for t in $threads\n do\n \tTHREADS=$t\n \texport THREADS\n-\ttest_perf PERF_EXTRA \"index-pack $t threads\" '\n-\t\trm -rf repo.git &&\n-\t\tgit init --bare repo.git &&\n+\ttest_perf \"index-pack $t threads\" --prereq PERF_EXTRA \\\n+\t\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \t\tGIT_DIR=repo.git GIT_FORCE_THREADS=1 \\\n \t\tgit index-pack --threads=$THREADS --stdin <$PACK\n \t'\n done\n \n-test_perf 'index-pack default number of threads' '\n-\trm -rf repo.git &&\n-\tgit init --bare repo.git &&\n+test_perf 'index-pack default number of threads' \\\n+\t--setup 'rm -rf repo.git && git init --bare repo.git' '\n \tGIT_DIR=repo.git git index-pack --stdin < $PACK\n '\n \ndiff --git a/t/perf/p7519-fsmonitor.sh b/t/perf/p7519-fsmonitor.sh\nindex 0b9129ca7bc..b1cb23880fb 100755\n--- a/t/perf/p7519-fsmonitor.sh\n+++ b/t/perf/p7519-fsmonitor.sh\n@@ -60,18 +60,6 @@ then\n \tesac\n fi\n \n-if test -n \"$GIT_PERF_7519_DROP_CACHE\"\n-then\n-\t# When using GIT_PERF_7519_DROP_CACHE, GIT_PERF_REPEAT_COUNT must be 1 to\n-\t# generate valid results. Otherwise the caching that happens for the nth\n-\t# run will negate the validity of the comparisons.\n-\tif test \"$GIT_PERF_REPEAT_COUNT\" -ne 1\n-\tthen\n-\t\techo \"warning: Setting GIT_PERF_REPEAT_COUNT=1\" >&2\n-\t\tGIT_PERF_REPEAT_COUNT=1\n-\tfi\n-fi\n-\n trace_start () {\n \tif test -n \"$GIT_PERF_7519_TRACE\"\n \tthen\n@@ -175,10 +163,10 @@ setup_for_fsmonitor_hook () {\n \n test_perf_w_drop_caches () {\n \tif test -n \"$GIT_PERF_7519_DROP_CACHE\"; then\n-\t\ttest-tool drop-caches\n+\t\ttest_perf \"$1\" --setup \"test-tool drop-caches\" \"$2\"\n+\telse\n+\t\ttest_perf \"$@\"\n \tfi\n-\n-\ttest_perf \"$@\"\n }\n \n test_fsmonitor_suite () {\ndiff --git a/t/perf/p7820-grep-engines.sh b/t/perf/p7820-grep-engines.sh\nindex 8b09c5bf328..9bfb86842a9 100755\n--- a/t/perf/p7820-grep-engines.sh\n+++ b/t/perf/p7820-grep-engines.sh\n@@ -49,13 +49,15 @@ do\n \t\tfi\n \t\tif ! test_have_prereq PERF_GREP_ENGINES_THREADS\n \t\tthen\n-\t\t\ttest_perf $prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \"\n+\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern'\" \\\n+\t\t\t\t--prereq \"$prereq\" \"\n \t\t\t\tgit -c grep.patternType=$engine grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine' || :\n \t\t\t\"\n \t\telse\n \t\t\tfor threads in $GIT_PERF_GREP_THREADS\n \t\t\tdo\n-\t\t\t\ttest_perf PTHREADS,$prereq \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\" \"\n+\t\t\t\ttest_perf \"$engine grep$GIT_PERF_7820_GREP_OPTS '$pattern' with $threads threads\"\n+\t\t\t\t\t--prereq PTHREADS,$prereq \"\n \t\t\t\t\tgit -c grep.patternType=$engine -c grep.threads=$threads grep$GIT_PERF_7820_GREP_OPTS -- '$pattern' >'out.$engine.$threads' || :\n \t\t\t\t\"\n \t\t\tdone\ndiff --git a/t/perf/perf-lib.sh b/t/perf/perf-lib.sh\nindex 932105cd12c..ab3687c28d4 100644\n--- a/t/perf/perf-lib.sh\n+++ b/t/perf/perf-lib.sh\n@@ -189,19 +189,39 @@ exit $ret' >&3 2>&4\n }\n \n test_wrapper_ () {\n-\ttest_wrapper_func_=$1; shift\n+\tlocal test_wrapper_func_=\"$1\"; shift\n+\tlocal test_title_=\"$1\"; shift\n \ttest_start_\n-\ttest \"$#\" = 3 && { test_prereq=$1; shift; } || test_prereq=\n-\ttest \"$#\" = 2 ||\n-\tBUG \"not 2 or 3 parameters to test-expect-success\"\n+\ttest_prereq=\n+\ttest_perf_setup_=\n+\twhile test $# != 0\n+\tdo\n+\t\tcase $1 in\n+\t\t--prereq)\n+\t\t\ttest_prereq=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t--setup)\n+\t\t\ttest_perf_setup_=$2\n+\t\t\tshift\n+\t\t\t;;\n+\t\t*)\n+\t\t\tbreak\n+\t\t\t;;\n+\t\tesac\n+\t\tshift\n+\tdone\n+\ttest \"$#\" = 1 || BUG \"test_wrapper_ needs 2 positional parameters\"\n \texport test_prereq\n-\tif ! test_skip \"$@\"\n+\texport test_perf_setup_\n+\n+\tif ! test_skip \"$test_title_\" \"$@\"\n \tthen\n \t\tbase=$(basename \"$0\" .sh)\n \t\techo \"$test_count\" >>\"$perf_results_dir\"/$base.subtests\n \t\techo \"$1\" >\"$perf_results_dir\"/$base.$test_count.descr\n \t\tbase=\"$perf_results_dir\"/\"$PERF_RESULTS_PREFIX$(basename \"$0\" .sh)\".\"$test_count\"\n-\t\t\"$test_wrapper_func_\" \"$@\"\n+\t\t\"$test_wrapper_func_\" \"$test_title_\" \"$@\"\n \tfi\n \n \ttest_finish_\n@@ -214,6 +234,16 @@ test_perf_ () {\n \t\techo \"perf $test_count - $1:\"\n \tfi\n \tfor i in $(test_seq 1 $GIT_PERF_REPEAT_COUNT); do\n+\t\tif test -n \"$test_perf_setup_\"\n+\t\tthen\n+\t\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\t\tif ! test_eval_ $test_perf_setup_\n+\t\t\tthen\n+\t\t\t\ttest_failure_ \"$test_perf_setup_\"\n+\t\t\t\tbreak\n+\t\t\tfi\n+\n+\t\tfi\n \t\tsay >&3 \"running: $2\"\n \t\tif test_run_perf_ \"$2\"\n \t\tthen\n@@ -237,11 +267,24 @@ test_perf_ () {\n \trm test_time.*\n }\n \n+# Usage: test_perf 'title' [options] 'perf-test'\n+#\tRun the performance test script specified in perf-test with\n+#\toptional prerequisite and setup steps.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to each measured iteration\n+#\n test_perf () {\n \ttest_wrapper_ test_perf_ \"$@\"\n }\n \n test_size_ () {\n+\tif test -n \"$test_perf_setup_\"\n+\tthen\n+\t\tsay >&3 \"setup: $test_perf_setup_\"\n+\t\ttest_eval_ $test_perf_setup_\n+\tfi\n+\n \tsay >&3 \"running: $2\"\n \tif test_eval_ \"$2\" 3>\"$base\".result; then\n \t\ttest_ok_ \"$1\"\n@@ -250,6 +293,14 @@ test_size_ () {\n \tfi\n }\n \n+# Usage: test_size 'title' [options] 'size-test'\n+#\tRun the size test script specified in size-test with optional\n+#\tprerequisites and setup steps. Returns the numeric value\n+#\treturned by size-test.\n+# Options:\n+#\t--prereq prerequisites: Skip the test if prequisites aren't met\n+#\t--setup \"setup-steps\": Run setup steps prior to the size measurement\n+\n test_size () {\n \ttest_wrapper_ test_size_ \"$@\"\n }\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453091","messageId":"20220405052018.11247-10-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 09/12] test-lib-functions: add parsing helpers for ls-files and ls-tree","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:15Z","receivedAt":"2022-04-05T05:20:56Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nSeveral tests use awk to parse OIDs from the output of 'git ls-files\n--stage' and 'git ls-tree'. Introduce helpers to centralize these uses\nof awk.\n\nUpdate t5317-pack-objects-filter-objects.sh to use the new ls-files\nhelper so that it has some usages to review. Other updates are left for\nthe future.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n t/t5317-pack-objects-filter-objects.sh | 91 +++++++++++++-------------\n t/test-lib-functions.sh                | 10 +++\n 2 files changed, 54 insertions(+), 47 deletions(-)\n\ndiff --git a/t/t5317-pack-objects-filter-objects.sh b/t/t5317-pack-objects-filter-objects.sh\nindex 33b740ce628..bb633c9b099 100755\n--- a/t/t5317-pack-objects-filter-objects.sh\n+++ b/t/t5317-pack-objects-filter-objects.sh\n@@ -10,9 +10,6 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n # Test blob:none filter.\n \n test_expect_success 'setup r1' '\n-\techo \"{print \\$1}\" >print_1.awk &&\n-\techo \"{print \\$2}\" >print_2.awk &&\n-\n \tgit init r1 &&\n \tfor n in 1 2 3 4 5\n \tdo\n@@ -22,10 +19,13 @@ test_expect_success 'setup r1' '\n \tdone\n '\n \n+parse_verify_pack_blob_oid () {\n+\tawk '{print $1}' -\n+}\n+\n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r1 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -35,7 +35,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -54,12 +54,12 @@ test_expect_success 'verify blob:none packfile has no blobs' '\n test_expect_success 'verify normal and blob:none packfiles have same commits/trees' '\n \tgit -C r1 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r1 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -123,8 +123,8 @@ test_expect_success 'setup r2' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -134,7 +134,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -161,8 +161,8 @@ test_expect_success 'verify blob:limit=1000' '\n '\n \n test_expect_success 'verify blob:limit=1001' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1001 >filter.pack <<-EOF &&\n@@ -172,15 +172,15 @@ test_expect_success 'verify blob:limit=1001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=10001' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=10001 >filter.pack <<-EOF &&\n@@ -190,15 +190,15 @@ test_expect_success 'verify blob:limit=10001' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1k' '\n-\tgit -C r2 ls-files -s large.1000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1k >filter.pack <<-EOF &&\n@@ -208,15 +208,15 @@ test_expect_success 'verify blob:limit=1k' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify explicitly specifying oversized blob in input' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \techo HEAD >objects &&\n@@ -226,15 +226,15 @@ test_expect_success 'verify explicitly specifying oversized blob in input' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify blob:limit=1m' '\n-\tgit -C r2 ls-files -s large.1000 large.10000 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r2 ls-files -s large.1000 large.10000 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r2 pack-objects --revs --stdout --filter=blob:limit=1m >filter.pack <<-EOF &&\n@@ -244,7 +244,7 @@ test_expect_success 'verify blob:limit=1m' '\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -253,12 +253,12 @@ test_expect_success 'verify blob:limit=1m' '\n test_expect_success 'verify normal and blob:limit packfiles have same commits/trees' '\n \tgit -C r2 verify-pack -v ../all.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >expected &&\n \n \tgit -C r2 verify-pack -v ../filter.pack >verify_result &&\n \tgrep -E \"commit|tree\" verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -289,9 +289,8 @@ test_expect_success 'setup r3' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r3 ls-files -s sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r3 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -301,7 +300,7 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r3 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -342,9 +341,8 @@ test_expect_success 'setup r4' '\n '\n \n test_expect_success 'verify blob count in normal packfile' '\n-\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s pattern sparse1 sparse2 dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout >all.pack <<-EOF &&\n@@ -354,19 +352,19 @@ test_expect_success 'verify blob count in normal packfile' '\n \n \tgit -C r4 verify-pack -v ../all.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=OID' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 ls-files -s pattern >staged &&\n-\toid=$(awk -f print_2.awk staged) &&\n+\toid=$(test_parse_ls_files_stage_oids <staged) &&\n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=$oid >filter.pack <<-EOF &&\n \tHEAD\n \tEOF\n@@ -374,15 +372,15 @@ test_expect_success 'verify sparse:oid=OID' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n '\n \n test_expect_success 'verify sparse:oid=oid-ish' '\n-\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 >ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r4 ls-files -s dir1/sparse1 dir1/sparse2 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tgit -C r4 pack-objects --revs --stdout --filter=sparse:oid=main:pattern >filter.pack <<-EOF &&\n@@ -392,7 +390,7 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n \n \tgit -C r4 verify-pack -v ../filter.pack >verify_result &&\n \tgrep blob verify_result |\n-\tawk -f print_1.awk |\n+\tparse_verify_pack_blob_oid |\n \tsort >observed &&\n \n \ttest_cmp expected observed\n@@ -402,9 +400,8 @@ test_expect_success 'verify sparse:oid=oid-ish' '\n # This models previously omitted objects that we did not receive.\n \n test_expect_success 'setup r1 - delete loose blobs' '\n-\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 \\\n-\t\t>ls_files_result &&\n-\tawk -f print_2.awk ls_files_result |\n+\tgit -C r1 ls-files -s file.1 file.2 file.3 file.4 file.5 |\n+\ttest_parse_ls_files_stage_oids |\n \tsort >expected &&\n \n \tfor id in `cat expected | sed \"s|..|&/|\"`\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 93c03380d44..50970d3e03e 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1782,6 +1782,16 @@ test_oid_to_path () {\n \techo \"${1%$basename}/$basename\"\n }\n \n+# Parse oids from git ls-files --staged output\n+test_parse_ls_files_stage_oids () {\n+\tawk '{print $2}' -\n+}\n+\n+# Parse oids from git ls-tree output\n+test_parse_ls_tree_oids () {\n+\tawk '{print $3}' -\n+}\n+\n # Choose a port number based on the test script's number and store it in\n # the given variable name, unless that variable already contains a number.\n test_set_port () {\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453092","messageId":"20220405052018.11247-8-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 07/12] unpack-objects: use the bulk-checkin infrastructure","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:13Z","receivedAt":"2022-04-05T05:20:57Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe unpack-objects functionality is used by fetch, push, and fast-import\nto turn the transfered data into object database entries when there are\nfewer objects than the 'unpacklimit' setting.\n\nBy enabling an odb-transaction when unpacking objects, we can take advantage\nof batched fsyncs.\n\nHere are some performance numbers to justify batch mode for\nunpack-objects, collected on a WSL2 Ubuntu VM.\n\nFsync Mode | Time for 90 objects (ms)\n-------------------------------------\n       Off | 170\n  On,fsync | 760\n  On,batch | 230\n\nNote that the default unpackLimit is 100 objects, so there's a 3x\nbenefit in the worst case. The non-batch mode fsync scales linearly\nwith the number of objects, so there are significant benefits even with\nsmaller numbers of objects.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/unpack-objects.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..56d05e2725d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -1,5 +1,6 @@\n #include \"builtin.h\"\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"object-store.h\"\n #include \"object.h\"\n@@ -503,10 +504,12 @@ static void unpack_all(void)\n \tif (!quiet)\n \t\tprogress = start_progress(_(\"Unpacking objects\"), nr_objects);\n \tCALLOC_ARRAY(obj_list, nr_objects);\n+\tbegin_odb_transaction();\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tunpack_one(i);\n \t\tdisplay_progress(progress, i + 1);\n \t}\n+\tend_odb_transaction();\n \tstop_progress(&progress);\n \n \tif (delta_list)\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453093","messageId":"20220405052018.11247-6-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 05/12] builtin/add: add ODB transaction around add_files_to_cache","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:11Z","receivedAt":"2022-04-05T05:21:00Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe add_files_to_cache function is invoked internally by\nbuiltin/commit.c and builtin/checkout.c for their flags that stage\nmodified files before doing the larger operation. These commands\ncan benefit from batched fsyncing.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c | 9 +++++++++\n 1 file changed, 9 insertions(+)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 9bf37ceae8e..e39770e4746 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -141,7 +141,16 @@ int add_files_to_cache(const char *prefix,\n \trev.diffopt.format_callback_data = &data;\n \trev.diffopt.flags.override_submodule_config = 1;\n \trev.max_count = 0; /* do not compare unmerged paths with stage #2 */\n+\n+\t/*\n+\t * Use an ODB transaction to optimize adding multiple objects.\n+\t * This function is invoked from commands other than 'add', which\n+\t * may not have their own transaction active.\n+\t */\n+\tbegin_odb_transaction();\n \trun_diff_files(&rev, DIFF_RACY_IS_MODIFIED);\n+\tend_odb_transaction();\n+\n \tclear_pathspec(&rev.prune_data);\n \treturn !!data.add_errors;\n }\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453094","messageId":"20220405052018.11247-9-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 08/12] core.fsync: use batch mode and sync loose objects by default on Windows","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:14Z","receivedAt":"2022-04-05T05:21:01Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nGit for Windows has defaulted to core.fsyncObjectFiles=true since\nSeptember 2017. We turn on syncing of loose object files with batch mode\nin upstream Git so that we can get broad coverage of the new code\nupstream.\n\nWe don't actually do fsyncs in the most of the test suite, since\nGIT_TEST_FSYNC is set to 0. However, we do exercise all of the\nsurrounding batch mode code since GIT_TEST_FSYNC merely makes the\nmaybe_fsync wrapper always appear to succeed.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n cache.h           | 4 ++++\n compat/mingw.h    | 3 +++\n config.c          | 2 +-\n git-compat-util.h | 2 ++\n 4 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex ea1466340d7..600375f786a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1031,6 +1031,10 @@ enum fsync_component {\n \t\t\t      FSYNC_COMPONENT_INDEX | \\\n \t\t\t      FSYNC_COMPONENT_REFERENCE)\n \n+#ifndef FSYNC_COMPONENTS_PLATFORM_DEFAULT\n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT FSYNC_COMPONENTS_DEFAULT\n+#endif\n+\n /*\n  * A bitmask indicating which components of the repo should be fsynced.\n  */\ndiff --git a/compat/mingw.h b/compat/mingw.h\nindex 6074a3d3ced..afe30868c04 100644\n--- a/compat/mingw.h\n+++ b/compat/mingw.h\n@@ -332,6 +332,9 @@ int mingw_getpagesize(void);\n int win32_fsync_no_flush(int fd);\n #define fsync_no_flush win32_fsync_no_flush\n \n+#define FSYNC_COMPONENTS_PLATFORM_DEFAULT (FSYNC_COMPONENTS_DEFAULT | FSYNC_COMPONENT_LOOSE_OBJECT)\n+#define FSYNC_METHOD_DEFAULT (FSYNC_METHOD_BATCH)\n+\n struct rlimit {\n \tunsigned int rlim_cur;\n };\ndiff --git a/config.c b/config.c\nindex 8ff25642906..027210a739c 100644\n--- a/config.c\n+++ b/config.c\n@@ -1342,7 +1342,7 @@ static const struct fsync_component_name {\n \n static enum fsync_component parse_fsync_components(const char *var, const char *string)\n {\n-\tenum fsync_component current = FSYNC_COMPONENTS_DEFAULT;\n+\tenum fsync_component current = FSYNC_COMPONENTS_PLATFORM_DEFAULT;\n \tenum fsync_component positive = 0, negative = 0;\n \n \twhile (string) {\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 4d444dca274..aaefd5b60c3 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -1257,11 +1257,13 @@ __attribute__((format (printf, 3, 4))) NORETURN\n void BUG_fl(const char *file, int line, const char *fmt, ...);\n #define BUG(...) BUG_fl(__FILE__, __LINE__, __VA_ARGS__)\n \n+#ifndef FSYNC_METHOD_DEFAULT\n #ifdef __APPLE__\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_WRITEOUT_ONLY\n #else\n #define FSYNC_METHOD_DEFAULT FSYNC_METHOD_FSYNC\n #endif\n+#endif\n \n enum fsync_action {\n \tFSYNC_WRITEOUT_ONLY,\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453095","messageId":"20220405052018.11247-3-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 02/12] bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:08Z","receivedAt":"2022-04-05T05:21:03Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nMake it clearer in the naming and documentation of the plug_bulk_checkin\nand unplug_bulk_checkin APIs that they can be thought of as\na \"transaction\" to optimize operations on the object database. These\ntransactions may be nested so that subsystems like the cache-tree\nwriting code can optimize their operations without caring whether the\ntop-level code has a transaction active.\n\nAdd a flush_odb_transaction API that will be used in update-index to\nmake objects visible even if a transaction is active. The flush call may\nalso be useful in future cases if we hold a transaction active around\ncalling hooks.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/add.c  |  4 ++--\n bulk-checkin.c | 25 +++++++++++++++++--------\n bulk-checkin.h | 24 ++++++++++++++++++++++--\n 3 files changed, 41 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin/add.c b/builtin/add.c\nindex 3ffb86a4338..9bf37ceae8e 100644\n--- a/builtin/add.c\n+++ b/builtin/add.c\n@@ -670,7 +670,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\tstring_list_clear(&only_match_skip_worktree, 0);\n \t}\n \n-\tplug_bulk_checkin();\n+\tbegin_odb_transaction();\n \n \tif (add_renormalize)\n \t\texit_status |= renormalize_tracked_files(&pathspec, flags);\n@@ -682,7 +682,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n \tif (chmod_arg && pathspec.nr)\n \t\texit_status |= chmod_pathspec(&pathspec, chmod_arg[0], show_only);\n-\tunplug_bulk_checkin();\n+\tend_odb_transaction();\n \n finish:\n \tif (write_locked_index(&the_index, &lock_file,\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 88d72178b2c..0fb032c7b69 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -10,7 +10,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n-static int bulk_checkin_plugged;\n+static int odb_transaction_nesting;\n \n static struct bulk_checkin_packfile {\n \tchar *pack_tmp_name;\n@@ -280,20 +280,29 @@ int index_bulk_checkin(struct object_id *oid,\n {\n \tint status = deflate_to_pack(&bulk_checkin_packfile, oid, fd, size, type,\n \t\t\t\t     path, flags);\n-\tif (!bulk_checkin_plugged)\n+\tif (!odb_transaction_nesting)\n \t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n \treturn status;\n }\n \n-void plug_bulk_checkin(void)\n+void begin_odb_transaction(void)\n {\n-\tassert(!bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 1;\n+\todb_transaction_nesting += 1;\n }\n \n-void unplug_bulk_checkin(void)\n+void flush_odb_transaction(void)\n {\n-\tassert(bulk_checkin_plugged);\n-\tbulk_checkin_plugged = 0;\n \tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n }\n+\n+void end_odb_transaction(void)\n+{\n+\todb_transaction_nesting -= 1;\n+\tif (odb_transaction_nesting < 0)\n+\t\tBUG(\"Unbalanced ODB transaction nesting\");\n+\n+\tif (odb_transaction_nesting)\n+\t\treturn;\n+\n+\tflush_odb_transaction();\n+}\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex b26f3dc3b74..ee0832989a8 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -10,7 +10,27 @@ int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\n \n-void plug_bulk_checkin(void);\n-void unplug_bulk_checkin(void);\n+/*\n+ * Tell the object database to optimize for adding\n+ * multiple objects. end_odb_transaction must be called\n+ * to make new objects visible. Transactions can be nested,\n+ * and objects are only visible after the outermost transaction\n+ * is complete or the transaction is flushed.\n+ */\n+void begin_odb_transaction(void);\n+\n+/*\n+ * Make any objects that are currently part of a pending object\n+ * database transaction visible. It is valid to call this function\n+ * even if no transaction is active.\n+ */\n+void flush_odb_transaction(void);\n+\n+/*\n+ * Tell the object database to make any objects from the\n+ * current transaction visible if this is the final nested\n+ * transaction.\n+ */\n+void end_odb_transaction(void);\n \n #endif\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453096","messageId":"20220405052018.11247-4-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 03/12] core.fsyncmethod: batched disk flushes for loose-objects","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:09Z","receivedAt":"2022-04-05T05:21:05Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nWhen adding many objects to a repo with `core.fsync=loose-object`,\nthe cost of fsync'ing each object file can become prohibitive.\n\nOne major source of the cost of fsync is the implied flush of the\nhardware writeback cache within the disk drive. This commit introduces\na new `core.fsyncMethod=batch` option that batches up hardware flushes.\nIt hooks into the bulk-checkin odb-transaction functionality, takes\nadvantage of tmp-objdir, and uses the writeout-only support code.\n\nWhen the new mode is enabled, we do the following for each new object:\n1a. Create the object in a tmp-objdir.\n2a. Issue a pagecache writeback request and wait for it to complete.\n\nAt the end of the entire transaction when unplugging bulk checkin:\n1b. Issue an fsync against a dummy file to flush the log and hardware\n   writeback cache, which should by now have seen the tmp-objdir writes.\n2b. Rename all of the tmp-objdir files to their final names.\n3b. When updating the index and/or refs, we assume that Git will issue\n   another fsync internal to that operation. This is not the default\n   today, but the user now has the option of syncing the index and there\n   is a separate patch series to implement syncing of refs.\n\nOn a filesystem with a singular journal that is updated during name\noperations (e.g. create, link, rename, etc), such as NTFS, HFS+, or XFS\nwe would expect the fsync to trigger a journal writeout so that this\nsequence is enough to ensure that the user's data is durable by the time\nthe git command returns. This sequence also ensures that no object files\nappear in the main object store unless they are fsync-durable.\n\nBatch mode is only enabled if core.fsync includes loose-objects. If\nthe legacy core.fsyncObjectFiles setting is enabled, but core.fsync does\nnot include loose-objects, we will use file-by-file fsyncing.\n\nIn step (1a) of the sequence, the tmp-objdir is created lazily to avoid\nwork if no loose objects are ever added to the ODB. We use a tmp-objdir\nto maintain the invariant that no loose-objects are visible in the main\nODB unless they are properly fsync-durable. This is important since\nfuture ODB operations that try to create an object with specific\ncontents will silently drop the new data if an object with the target\nhash exists without checking that the loose-object contents match the\nhash. Only a full git-fsck would restore the ODB to a functional state\nwhere dataloss doesn't occur.\n\nIn step (1b) of the sequence, we issue a fsync against a dummy file\ncreated specifically for the purpose. This method has a little higher\ncost than using one of the input object files, but makes adding new\ncallers of this mechanism easier, since we don't need to figure out\nwhich object file is \"last\" or risk sharing violations by caching the fd\nof the last object file.\n\n_Performance numbers_:\n\nLinux - Hyper-V VM running Kernel 5.11 (Ubuntu 20.04) on a fast SSD.\nMac - macOS 11.5.1 running on a Mac mini on a 1TB Apple SSD.\nWindows - Same host as Linux, a preview version of Windows 11.\n\nAdding 500 files to the repo with 'git add' Times reported in seconds.\n\nobject file syncing | Linux | Mac   | Windows\n--------------------|-------|-------|--------\n           disabled | 0.06  |  0.35 | 0.61\n              fsync | 1.88  | 11.18 | 2.47\n              batch | 0.15  |  0.41 | 1.53\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n Documentation/config/core.txt |  8 ++++\n bulk-checkin.c                | 71 +++++++++++++++++++++++++++++++++++\n bulk-checkin.h                |  3 ++\n cache.h                       |  8 +++-\n config.c                      |  2 +\n object-file.c                 |  7 +++-\n 6 files changed, 97 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 889522956e4..d543bf12824 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -628,6 +628,14 @@ core.fsyncMethod::\n * `writeout-only` issues pagecache writeback requests, but depending on the\n   filesystem and storage hardware, data added to the repository may not be\n   durable in the event of a system crash. This is the default mode on macOS.\n+* `batch` enables a mode that uses writeout-only flushes to stage multiple\n+  updates in the disk writeback cache and then does a single full fsync of\n+  a dummy file to trigger the disk cache flush at the end of the operation.\n++\n+Currently `batch` mode only applies to loose-object files. Other repository\n+data is made durable as if `fsync` was specified. This mode is expected to\n+be as safe as `fsync` on macOS for repos stored on HFS+ or APFS filesystems\n+and on Windows for repos stored on NTFS or ReFS filesystems.\n \n core.fsyncObjectFiles::\n \tThis boolean will enable 'fsync()' when writing object files.\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 0fb032c7b69..bcf878460b2 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -3,15 +3,20 @@\n  */\n #include \"cache.h\"\n #include \"bulk-checkin.h\"\n+#include \"lockfile.h\"\n #include \"repository.h\"\n #include \"csum-file.h\"\n #include \"pack.h\"\n #include \"strbuf.h\"\n+#include \"string-list.h\"\n+#include \"tmp-objdir.h\"\n #include \"packfile.h\"\n #include \"object-store.h\"\n \n static int odb_transaction_nesting;\n \n+static struct tmp_objdir *bulk_fsync_objdir;\n+\n static struct bulk_checkin_packfile {\n \tchar *pack_tmp_name;\n \tstruct hashfile *f;\n@@ -80,6 +85,40 @@ static void flush_bulk_checkin_packfile(struct bulk_checkin_packfile *state)\n \treprepare_packed_git(the_repository);\n }\n \n+/*\n+ * Cleanup after batch-mode fsync_object_files.\n+ */\n+static void flush_batch_fsync(void)\n+{\n+\tstruct strbuf temp_path = STRBUF_INIT;\n+\tstruct tempfile *temp;\n+\n+\tif (!bulk_fsync_objdir)\n+\t\treturn;\n+\n+\t/*\n+\t * Issue a full hardware flush against a temporary file to ensure\n+\t * that all objects are durable before any renames occur. The code in\n+\t * fsync_loose_object_bulk_checkin has already issued a writeout\n+\t * request, but it has not flushed any writeback cache in the storage\n+\t * hardware or any filesystem logs. This fsync call acts as a barrier\n+\t * to ensure that the data in each new object file is durable before\n+\t * the final name is visible.\n+\t */\n+\tstrbuf_addf(&temp_path, \"%s/bulk_fsync_XXXXXX\", get_object_directory());\n+\ttemp = xmks_tempfile(temp_path.buf);\n+\tfsync_or_die(get_tempfile_fd(temp), get_tempfile_path(temp));\n+\tdelete_tempfile(&temp);\n+\tstrbuf_release(&temp_path);\n+\n+\t/*\n+\t * Make the object files visible in the primary ODB after their data is\n+\t * fully durable.\n+\t */\n+\ttmp_objdir_migrate(bulk_fsync_objdir);\n+\tbulk_fsync_objdir = NULL;\n+}\n+\n static int already_written(struct bulk_checkin_packfile *state, struct object_id *oid)\n {\n \tint i;\n@@ -274,6 +313,37 @@ static int deflate_to_pack(struct bulk_checkin_packfile *state,\n \treturn 0;\n }\n \n+void prepare_loose_object_bulk_checkin(void)\n+{\n+\t/*\n+\t * We lazily create the temporary object directory\n+\t * the first time an object might be added, since\n+\t * callers may not know whether any objects will be\n+\t * added at the time they call begin_odb_transaction.\n+\t */\n+\tif (!odb_transaction_nesting || bulk_fsync_objdir)\n+\t\treturn;\n+\n+\tbulk_fsync_objdir = tmp_objdir_create(\"bulk-fsync\");\n+\tif (bulk_fsync_objdir)\n+\t\ttmp_objdir_replace_primary_odb(bulk_fsync_objdir, 0);\n+}\n+\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n+{\n+\t/*\n+\t * If we have an active ODB transaction, we issue a call that\n+\t * cleans the filesystem page cache but avoids a hardware flush\n+\t * command. Later on we will issue a single hardware flush\n+\t * before renaming the objects to their final names as part of\n+\t * flush_batch_fsync.\n+\t */\n+\tif (!bulk_fsync_objdir ||\n+\t    git_fsync(fd, FSYNC_WRITEOUT_ONLY) < 0) {\n+\t\tfsync_or_die(fd, filename);\n+\t}\n+}\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags)\n@@ -292,6 +362,7 @@ void begin_odb_transaction(void)\n \n void flush_odb_transaction(void)\n {\n+\tflush_batch_fsync();\n \tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n }\n \ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex ee0832989a8..8281b9cb159 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -6,6 +6,9 @@\n \n #include \"cache.h\"\n \n+void prepare_loose_object_bulk_checkin(void);\n+void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n+\n int index_bulk_checkin(struct object_id *oid,\n \t\t       int fd, size_t size, enum object_type type,\n \t\t       const char *path, unsigned flags);\ndiff --git a/cache.h b/cache.h\nindex 6226f6a8a53..ea1466340d7 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1040,7 +1040,8 @@ extern int use_fsync;\n \n enum fsync_method {\n \tFSYNC_METHOD_FSYNC,\n-\tFSYNC_METHOD_WRITEOUT_ONLY\n+\tFSYNC_METHOD_WRITEOUT_ONLY,\n+\tFSYNC_METHOD_BATCH,\n };\n \n extern enum fsync_method fsync_method;\n@@ -1766,6 +1767,11 @@ void fsync_or_die(int fd, const char *);\n int fsync_component(enum fsync_component component, int fd);\n void fsync_component_or_die(enum fsync_component component, int fd, const char *msg);\n \n+static inline int batch_fsync_enabled(enum fsync_component component)\n+{\n+\treturn (fsync_components & component) && (fsync_method == FSYNC_METHOD_BATCH);\n+}\n+\n ssize_t read_in_full(int fd, void *buf, size_t count);\n ssize_t write_in_full(int fd, const void *buf, size_t count);\n ssize_t pread_in_full(int fd, void *buf, size_t count, off_t offset);\ndiff --git a/config.c b/config.c\nindex a5e11aad7fe..8ff25642906 100644\n--- a/config.c\n+++ b/config.c\n@@ -1688,6 +1688,8 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\t\tfsync_method = FSYNC_METHOD_FSYNC;\n \t\telse if (!strcmp(value, \"writeout-only\"))\n \t\t\tfsync_method = FSYNC_METHOD_WRITEOUT_ONLY;\n+\t\telse if (!strcmp(value, \"batch\"))\n+\t\t\tfsync_method = FSYNC_METHOD_BATCH;\n \t\telse\n \t\t\twarning(_(\"ignoring unknown core.fsyncMethod value '%s'\"), value);\n \ndiff --git a/object-file.c b/object-file.c\nindex 5ffbf3d4fd4..d2e0c13198f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1893,7 +1893,9 @@ static void close_loose_object(int fd, const char *filename)\n \tif (the_repository->objects->odb->will_destroy)\n \t\tgoto out;\n \n-\tif (fsync_object_files > 0)\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tfsync_loose_object_bulk_checkin(fd, filename);\n+\telse if (fsync_object_files > 0)\n \t\tfsync_or_die(fd, filename);\n \telse\n \t\tfsync_component_or_die(FSYNC_COMPONENT_LOOSE_OBJECT, fd,\n@@ -1961,6 +1963,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n \tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453097","messageId":"20220405052018.11247-7-neerajsi@microsoft.com","threadId":"57568","inReplyTo":"pull.1134.v5.git.1648616734.gitgitgadget@gmail.com","subject":"[PATCH v6 06/12] update-index: use the bulk-checkin infrastructure","fromName":"","fromEmail":"nksingh85@gmail.com","sentAt":"2022-04-05T05:20:12Z","receivedAt":"2022-04-05T05:21:07Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"From: Neeraj Singh <neerajsi@microsoft.com>\n\nThe update-index functionality is used internally by 'git stash push' to\nsetup the internal stashed commit.\n\nThis change enables odb-transactions for update-index infrastructure to\nspeed up adding new objects to the object database by leveraging the\nbatch fsync functionality.\n\nThere is some risk with this change, since under batch fsync, the object\nfiles will be in a tmp-objdir until update-index is complete, so callers\nusing the --stdin option will not see them until update-index is done.\nThis risk is mitigated by flushing the ODB transaction prior to\nreporting any verbose output so that objects will be visible to callers\nthat are synchronizing with update-index by snooping its output.\n\nSigned-off-by: Neeraj Singh <neerajsi@microsoft.com>\n---\n builtin/update-index.c | 20 ++++++++++++++++++++\n 1 file changed, 20 insertions(+)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 876112abb21..43c713c0873 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -5,6 +5,7 @@\n  */\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"cache.h\"\n+#include \"bulk-checkin.h\"\n #include \"config.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n@@ -57,6 +58,14 @@ static void report(const char *fmt, ...)\n \tif (!verbose)\n \t\treturn;\n \n+\t/*\n+\t * It is possible, though unlikely, that a caller could use the verbose\n+\t * output to synchronize with addition of objects to the object\n+\t * database. The current implementation of ODB transactions leaves\n+\t * objects invisible while a transaction is active, so flush the\n+\t * transaction here before reporting a change made by update-index.\n+\t */\n+\tflush_odb_transaction();\n \tva_start(vp, fmt);\n \tvprintf(fmt, vp);\n \tputchar('\\n');\n@@ -1116,6 +1125,12 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t */\n \tparse_options_start(&ctx, argc, argv, prefix,\n \t\t\t    options, PARSE_OPT_STOP_AT_NON_OPTION);\n+\n+\t/*\n+\t * Allow the object layer to optimize adding multiple objects in\n+\t * a batch.\n+\t */\n+\tbegin_odb_transaction();\n \twhile (ctx.argc) {\n \t\tif (parseopt_state != PARSE_OPT_DONE)\n \t\t\tparseopt_state = parse_options_step(&ctx, options,\n@@ -1190,6 +1205,11 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n+\t/*\n+\t * By now we have added all of the new objects\n+\t */\n+\tend_odb_transaction();\n+\n \tif (split_index > 0) {\n \t\tif (git_config_get_split_index() == 0)\n \t\t\twarning(_(\"core.splitIndex is set to false; \"\n-- \n2.34.1.78.g86e39b8f8d\n\n"},{"id":"453227","messageId":"xmqq7d82x8a9.fsf@gitster.g","threadId":"57568","inReplyTo":"20220405052018.11247-1-neerajsi@microsoft.com","subject":"Re: [PATCH v6 00/12] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-04-06T20:32:14Z","receivedAt":"2022-04-06T21:28:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"nksingh85@gmail.com writes:\n\n> From: Neeraj Singh <neerajsi@microsoft.com>\n>\n> GGG closed this series erroneously, so I'm trying out git-send-email. Apologies for any mistakes.\n>\n> This series is also available at https://github.com/neerajsi-msft/git/git.git ns/batched-fsync-v6.\n>\n> V6 changes:\n>\n> * Based on master at faa21c1 to pick up ns/fsync-or-die-message-fix. Also resolved a conflict with 8aa0209 in t/perf/p7519-fsmonitor.sh.\n>\n> * Some independent patches were submitted separately on-list. This series is now dependent on ns/fsync-or-die-message-fix.\n>\n> * Rename bulk_checkin_state to bulk_checkin_packfile to discourage future authors from adding any non-packfile related stuff to it. Each individual component of bulk_checkin should have its own state variable(s) going forward, and they should only be tied together by odb_transaction_nesting.\n>\n> * Rename finish_bulk_checkin and do_batch_fsync to flush_bulk_checkin and flush_batch_fsync. The \"finish\" step is going to be the end_odb_transaction. The \"flush\" terminology should be consistently used for making changes visible.\n>\n> * Add flush_odb_transaction and use it in update-index before printing verbose output to mitigate risk of missing objects for a tricky stdin feeder.\n>\n> * Re-add shell \"local with assignment\": now these are all on a separate line with quotes around any values, to comply with dash. I'm running on ubuntu 20.04 LTS where I saw some of the dash issues before.\n\nThanks.  These looked all sensible.\n\nWil queue.\n"},{"id":"455540","messageId":"xmqqtu9lkxe5.fsf@gitster.g","threadId":"57568","inReplyTo":"xmqq7d82x8a9.fsf@gitster.g","subject":"Re: [PATCH v6 00/12] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T21:47:30Z","receivedAt":"2022-05-19T21:47:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> nksingh85@gmail.com writes:\n>\n>> From: Neeraj Singh <neerajsi@microsoft.com>\n>>\n>> GGG closed this series erroneously, so I'm trying out\n>> git-send-email. Apologies for any mistakes.\n>>\n>> This series is also available at\n>> https://github.com/neerajsi-msft/git/git.git ns/batched-fsync-v6.\n>>\n>> V6 changes:\n\nWe haven't heard anything on this topic from anybody for this round.\nI am planning to merge it to 'next' soonish.\n\nPlease speak up if anybody has concerns.\n\nThanks.\n\n"},{"id":"455542","messageId":"fb1f2c8c-b749-85d9-b472-587f51dea8a8@gmail.com","threadId":"57568","inReplyTo":"xmqqtu9lkxe5.fsf@gitster.g","subject":"Re: [PATCH v6 00/12] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-05-19T21:54:43Z","receivedAt":"2022-05-19T21:54:59Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On 5/19/2022 2:47 PM, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n>> nksingh85@gmail.com writes:\n>>\n>>> From: Neeraj Singh <neerajsi@microsoft.com>\n>>>\n>>> GGG closed this series erroneously, so I'm trying out\n>>> git-send-email. Apologies for any mistakes.\n>>>\n>>> This series is also available at\n>>> https://github.com/neerajsi-msft/git/git.git ns/batched-fsync-v6.\n>>>\n>>> V6 changes:\n> \n> We haven't heard anything on this topic from anybody for this round.\n> I am planning to merge it to 'next' soonish.\n> \n> Please speak up if anybody has concerns.\n> \n> Thanks.\n> \n\nNo updates on my end. I'll keep my eyes out for any reports of regression.\n\nThanks,\nNeeraj\n"},{"id":"455895","messageId":"nycvar.QRO.7.76.6.2205241425570.352@tvgsbejvaqbjf.bet","threadId":"57568","inReplyTo":"fb1f2c8c-b749-85d9-b472-587f51dea8a8@gmail.com","subject":"Re: [PATCH v6 00/12] core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-24T12:31:06Z","receivedAt":"2022-05-24T12:31:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Neeraj,\n\nOn Thu, 19 May 2022, Neeraj Singh wrote:\n\n> On 5/19/2022 2:47 PM, Junio C Hamano wrote:\n> > Junio C Hamano <gitster@pobox.com> writes:\n> >\n> > > nksingh85@gmail.com writes:\n> > >\n> > > > From: Neeraj Singh <neerajsi@microsoft.com>\n> > > >\n> > > > GGG closed this series erroneously, so I'm trying out\n> > > > git-send-email. Apologies for any mistakes.\n> > > >\n> > > > This series is also available at\n> > > > https://github.com/neerajsi-msft/git/git.git ns/batched-fsync-v6.\n> > > >\n> > > > V6 changes:\n> >\n> > We haven't heard anything on this topic from anybody for this round.\n> > I am planning to merge it to 'next' soonish.\n> >\n> > Please speak up if anybody has concerns.\n> >\n> > Thanks.\n> >\n>\n> No updates on my end. I'll keep my eyes out for any reports of regression.\n\nI asked a colleague to have a go with these patches and the only concern I\nheard back was that with ext4's new `fast_commit` feature, the `fsync`\nseems not to actually flush all metadata. They indicated that they'd be\nhappy with merely documenting this issue, and also pointed out that the\n`fast_commit` feature seems still not to be considered ready for\nproduction workloads.\n\nSo: 👍 from my side.\n\nCiao,\nDscho\n"}]}