{"thread":{"id":"60580","subject":"[RFC PATCH 0/4] add parallel unlink","startedAt":"2023-12-03T13:39:24Z","lastAt":"2023-12-03T13:39:33Z","messageCount":5,"participants":["Han Young"],"isPatch":true,"patchVersion":1,"patchTotal":4},"messages":[{"id":"485339","messageId":"20231203133911.41594-1-hanyoung@protonmail.com","threadId":"60580","inReplyTo":null,"subject":"[RFC PATCH 0/4] add parallel unlink","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2023-12-03T13:39:07Z","receivedAt":"2023-12-03T13:39:24Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"We have had parallel_checkout option since 04155bdad, but the unlink is still performed single threaded.\nWith a very large repository, directory rename or reorganization can lead to a large amount of unlinked entries.\nIn some instance, the unlink process can be slower than the parallel checkout.\n\nThis series of patches introduces basic support for parallel unlink. The removal of individual files\ncan be easily multithreaded, but removing empty directories is a little tricky.\nIf one thread decides to remove the directory, it may still have files that need to be deleted by\nanother thread. I had to use a mutex-guarded hashset to collect these 'race' directories,\nand remove them after all threads have been joined. Maybe there are ways to do this\nwithout mutex and hashmap?\n\nThe speed of unlinking files seems to vary from system to system. I did some tests with a private repo.\nWhen I checkout a commit with 15000 moved files on a Linux machine with btrfs, parallel_unlink yields\n10% speed up. But on a Intel MacBook Pro with APFS, the speed up is over 100%. I find it difficult to\nchoose the default threshold of parallel_unlink.\n\nThis series is by no means complete. Many functions contains duplicated code, and there are some\nmemory leaks. I want to know the community opinion before proceed, if it's worth doing or a waste of time.\n\nHan Young (4):\n  symlinks: add and export threaded rmdir variants\n  entry: add threaded_unlink_entry function\n  parallel-checkout: add parallel_unlink\n  unpack-trees: introduce parallel_unlink\n\n entry.c             |  16 ++++++\n entry.h             |   3 ++\n parallel-checkout.c |  80 +++++++++++++++++++++++++++++\n parallel-checkout.h |  25 +++++++++\n symlinks.c          | 120 ++++++++++++++++++++++++++++++++++++++++++--\n symlinks.h          |   6 +++\n unpack-trees.c      |  15 +-----\n 7 files changed, 249 insertions(+), 16 deletions(-)\n\n-- \n2.43.0\n\n"},{"id":"485340","messageId":"20231203133911.41594-2-hanyoung@protonmail.com","threadId":"60580","inReplyTo":"20231203133911.41594-1-hanyoung@protonmail.com","subject":"[RFC PATCH 1/4] symlinks: add and export threaded rmdir variants","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2023-12-03T13:39:08Z","receivedAt":"2023-12-03T13:39:27Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"From: Han Young <hanyang.tony@bytedance.com>\n\nAdd and export threaded variants of remove dir related functions, these functions will be used by parallel unlink\n---\nMost of the code of threaded_schedule_dir_for_removal and threaded_do_remove_scheduled_dirs is duplicated.\nWe can remove the duplication either via breaking the function into smaller functions, or pass the cache as parameters.\nIf we choose to pass the cache explicitly, default cache in both entry.c and symlinks.c probably need to be moved to\nunpack-trees.c. I'm not satisfied with using mutex guarded hashset to ensure every dir is removed. But I can't come\nup with a better way.\n\n symlinks.c | 120 +++++++++++++++++++++++++++++++++++++++++++++++++++--\n symlinks.h |   6 +++\n 2 files changed, 123 insertions(+), 3 deletions(-)\n\ndiff --git a/symlinks.c b/symlinks.c\nindex b29e340c2d..c8cb0a7eb7 100644\n--- a/symlinks.c\n+++ b/symlinks.c\n@@ -2,9 +2,9 @@\n #include \"gettext.h\"\n #include \"setup.h\"\n #include \"symlinks.h\"\n+#include \"hashmap.h\"\n+#include \"pthread.h\"\n \n-static int threaded_check_leading_path(struct cache_def *cache, const char *name,\n-\t\t\t\t       int len, int warn_on_lstat_err);\n static int threaded_has_dirs_only_path(struct cache_def *cache, const char *name, int len, int prefix_len);\n \n /*\n@@ -229,7 +229,7 @@ int check_leading_path(const char *name, int len, int warn_on_lstat_err)\n  * directory, or if we were unable to lstat() it. If warn_on_lstat_err is true,\n  * also emit a warning for this error.\n  */\n-static int threaded_check_leading_path(struct cache_def *cache, const char *name,\n+int threaded_check_leading_path(struct cache_def *cache, const char *name,\n \t\t\t\t       int len, int warn_on_lstat_err)\n {\n \tint flags;\n@@ -277,6 +277,51 @@ static int threaded_has_dirs_only_path(struct cache_def *cache, const char *name\n }\n \n static struct strbuf removal = STRBUF_INIT;\n+static struct hashmap dir_set;\n+pthread_mutex_t dir_set_mutex = PTHREAD_MUTEX_INITIALIZER;\n+struct rmdir_hash_entry {\n+      struct hashmap_entry hash;\n+      char *dir;\n+      size_t dirlen;\n+};\n+\n+/* rmdir_hashmap comparison function */\n+static int rmdir_hash_entry_cmp(const void *cmp_data UNUSED,\n+\t\t\t       const struct hashmap_entry *eptr,\n+\t\t\t       const struct hashmap_entry *entry_or_key UNUSED,\n+\t\t\t       const void *keydata)\n+{\n+\tconst struct rmdir_hash_entry *a, *b;\n+\n+\ta = container_of(eptr, const struct rmdir_hash_entry, hash);\n+\treturn strcmp(a->dir, (char *)keydata);\n+}\n+\n+void threaded_init_remove_scheduled_dirs(void)\n+{\n+\tunsigned flags = 0;\n+\thashmap_init(&dir_set, rmdir_hash_entry_cmp, &flags, 0);\n+}\n+\n+static void add_dir_to_rmdir_hash(char *dir, size_t dirlen)\n+{\n+\tstruct rmdir_hash_entry *e;\n+\tstruct hashmap_entry *ent;\n+\tint hash = strhash(dir);\n+\tpthread_mutex_lock(&dir_set_mutex);\n+\tent = hashmap_get_from_hash(&dir_set, hash, dir);\n+\n+\tif (!ent) {\n+\t\te = xmalloc(sizeof(struct rmdir_hash_entry));\n+\t\thashmap_entry_init(&e->hash, hash);\n+\t\tchar *_dir= xmallocz(dirlen);\n+\t\tmemcpy(_dir, dir, dirlen+1);\n+\t\te->dir = _dir;\n+\t\te->dirlen = dirlen;\n+\t\thashmap_put_entry(&dir_set, e, hash);\n+\t}\n+\tpthread_mutex_unlock(&dir_set_mutex);\n+}\n \n static void do_remove_scheduled_dirs(int new_len)\n {\n@@ -294,6 +339,26 @@ static void do_remove_scheduled_dirs(int new_len)\n \tremoval.len = new_len;\n }\n \n+\n+static void threaded_do_remove_scheduled_dirs(int new_len, struct strbuf *removal)\n+{\n+\twhile (removal->len > new_len) {\n+\t\tremoval->buf[removal->len] = '\\0';\n+\t\tif (startup_info->original_cwd &&\n+\t\t     !strcmp(removal->buf, startup_info->original_cwd))\n+\t\t\t break;\n+\t\tif (rmdir(removal->buf)) {\n+\t\t\tadd_dir_to_rmdir_hash(removal->buf, removal->len);\n+\t\t\tbreak;\n+\t\t}\n+\t\tdo {\n+\t\t\tremoval->len--;\n+\t\t} while (removal->len > new_len &&\n+\t\t\t removal->buf[removal->len] != '/');\n+\t}\n+\tremoval->len = new_len;\n+}\n+\n void schedule_dir_for_removal(const char *name, int len)\n {\n \tint match_len, last_slash, i, previous_slash;\n@@ -327,11 +392,60 @@ void schedule_dir_for_removal(const char *name, int len)\n \t\tstrbuf_add(&removal, &name[match_len], last_slash - match_len);\n }\n \n+void threaded_schedule_dir_for_removal(const char *name, int len, struct strbuf *removal_cache)\n+{\n+\tint match_len, last_slash, i, previous_slash;\n+\n+\tif (startup_info->original_cwd &&\n+\t    !strcmp(name, startup_info->original_cwd))\n+\t\treturn;\t/* Do not remove the current working directory */\n+\n+\tmatch_len = last_slash = i =\n+\t\tlongest_path_match(name, len, removal_cache->buf, removal_cache->len,\n+\t\t\t\t   &previous_slash);\n+\t/* Find last slash inside 'name' */\n+\twhile (i < len) {\n+\t\tif (name[i] == '/')\n+\t\t\tlast_slash = i;\n+\t\ti++;\n+\t}\n+\n+\t/*\n+\t * If we are about to go down the directory tree, we check if\n+\t * we must first go upwards the tree, such that we then can\n+\t * remove possible empty directories as we go upwards.\n+\t */\n+\tif (match_len < last_slash && match_len < removal_cache->len)\n+\t\tthreaded_do_remove_scheduled_dirs(match_len, removal_cache);\n+\t/*\n+\t * If we go deeper down the directory tree, we only need to\n+\t * save the new path components as we go down.\n+\t */\n+\tif (match_len < last_slash)\n+\t\tstrbuf_add(removal_cache, &name[match_len], last_slash - match_len);\n+}\n+\n void remove_scheduled_dirs(void)\n {\n \tdo_remove_scheduled_dirs(0);\n }\n \n+void threaded_remove_scheduled_dirs_clean_up(void)\n+{\n+\tstruct hashmap_iter iter;\n+\tconst struct rmdir_hash_entry *entry;\n+\n+\thashmap_for_each_entry(&dir_set, &iter, entry, hash /* member name */) {\n+\t\tschedule_dir_for_removal(entry->dir, entry->dirlen);\n+\t}\n+\tremove_scheduled_dirs();\n+}\n+\n+void threaded_remove_scheduled_dirs(struct strbuf *removal_cache)\n+{\n+\tthreaded_do_remove_scheduled_dirs(0, removal_cache);\n+}\n+\n void invalidate_lstat_cache(void)\n {\n \treset_lstat_cache(&default_cache);\ndiff --git a/symlinks.h b/symlinks.h\nindex 7ae3d5b856..7898eae941 100644\n--- a/symlinks.h\n+++ b/symlinks.h\n@@ -20,9 +20,15 @@ static inline void cache_def_clear(struct cache_def *cache)\n int has_symlink_leading_path(const char *name, int len);\n int threaded_has_symlink_leading_path(struct cache_def *, const char *, int);\n int check_leading_path(const char *name, int len, int warn_on_lstat_err);\n+int threaded_check_leading_path(struct cache_def *cache, const char *name,\n+\t\t\t\t       int len, int warn_on_lstat_err);\n int has_dirs_only_path(const char *name, int len, int prefix_len);\n void invalidate_lstat_cache(void);\n void schedule_dir_for_removal(const char *name, int len);\n+void threaded_schedule_dir_for_removal(const char *name, int len, struct strbuf *removal_cache);\n void remove_scheduled_dirs(void);\n+void threaded_remove_scheduled_dirs(struct strbuf *removal_cache);\n+void threaded_init_remove_scheduled_dirs(void);\n+void threaded_remove_scheduled_dirs_clean_up(void);\n \n #endif /* SYMLINKS_H */\n-- \n2.43.0\n\n"},{"id":"485341","messageId":"20231203133911.41594-3-hanyoung@protonmail.com","threadId":"60580","inReplyTo":"20231203133911.41594-1-hanyoung@protonmail.com","subject":"[RFC PATCH 2/4] entry: add threaded_unlink_entry function","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2023-12-03T13:39:09Z","receivedAt":"2023-12-03T13:39:28Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"From: Han Young <hanyang.tony@bytedance.com>\n\nAdd threaded_unlink_entry function, the threaded function uses cache passed by arguments instead of the default cache. It also calls threaded variant of schedule_dir_for_removal to ensure dirs are removed in multithreaded unlink.\n---\nAnother duplicated function. Because default removal cache and default lstat cache live in different source files,\nthreaded variant of check_leading_path and schedule_dir_for_removal must be called here\ninstead of choosing to pass explicit or default cache.\n\n entry.c | 16 ++++++++++++++++\n entry.h |  3 +++\n 2 files changed, 19 insertions(+)\n\ndiff --git a/entry.c b/entry.c\nindex 076e97eb89..04440beb2b 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -567,6 +567,22 @@ int checkout_entry_ca(struct cache_entry *ce, struct conv_attrs *ca,\n \treturn write_entry(ce, path.buf, ca, state, 0, nr_checkouts);\n }\n \n+void threaded_unlink_entry(const struct cache_entry *ce, const char *super_prefix,\n+\t\t\t   struct strbuf *removal, struct cache_def *cache)\n+{\n+\tconst struct submodule *sub = submodule_from_ce(ce);\n+\tif (sub) {\n+\t\t/* state.force is set at the caller. */\n+\t\tsubmodule_move_head(ce->name, super_prefix, \"HEAD\", NULL,\n+\t\t\t\t    SUBMODULE_MOVE_HEAD_FORCE);\n+\t}\n+\tif (threaded_check_leading_path(cache, ce->name, ce_namelen(ce), 1) >= 0)\n+\t\treturn;\n+\tif (remove_or_warn(ce->ce_mode, ce->name))\n+\t\treturn;\n+\tthreaded_schedule_dir_for_removal(ce->name, ce_namelen(ce), removal);\n+}\n+\n void unlink_entry(const struct cache_entry *ce, const char *super_prefix)\n {\n \tconst struct submodule *sub = submodule_from_ce(ce);\ndiff --git a/entry.h b/entry.h\nindex ca3ed35bc0..413ca3822d 100644\n--- a/entry.h\n+++ b/entry.h\n@@ -2,6 +2,7 @@\n #define ENTRY_H\n \n #include \"convert.h\"\n+#include \"symlinks.h\"\n \n struct cache_entry;\n struct index_state;\n@@ -56,6 +57,8 @@ int finish_delayed_checkout(struct checkout *state, int show_progress);\n  * down from \"read-tree\" et al.\n  */\n void unlink_entry(const struct cache_entry *ce, const char *super_prefix);\n+void threaded_unlink_entry(const struct cache_entry *ce, const char *super_prefix,\n+\t\t\t   struct strbuf *removal, struct cache_def *cache);\n \n void *read_blob_entry(const struct cache_entry *ce, size_t *size);\n int fstat_checkout_output(int fd, const struct checkout *state, struct stat *st);\n-- \n2.43.0\n\n"},{"id":"485342","messageId":"20231203133911.41594-4-hanyoung@protonmail.com","threadId":"60580","inReplyTo":"20231203133911.41594-1-hanyoung@protonmail.com","subject":"[RFC PATCH 3/4] parallel-checkout: add parallel_unlink","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2023-12-03T13:39:10Z","receivedAt":"2023-12-03T13:39:30Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"From: Han Young <hanyang.tony@bytedance.com>\n\nAdd parallel_unlink to parallel-checkout, parallel_unlink uses multiple threads to unlink entries. Because the path to be removed is sorted, each thread iterate through the entry list interleaved to distribute the workload as evenly as possible. Due to the multithread nature, it's not possible to remove all the dirs in one pass. The dir one thread is about to remove may have item that are being removed by another thread. Whenever we failed to remove the dir, we save it in a hashset. When every thread has finished its job, we remove all the entries in the hashset.\n---\nNote that we display progress after thread join, the progress count is updated for every thread instead of every path.\nDuring testing, threads almost finished at around the same time. This caused the abrupt progress update.\nWe can use a mutex to display the progress, but that nullified the optimization on environment with fast file deletion time.\n\n parallel-checkout.c | 80 +++++++++++++++++++++++++++++++++++++++++++++\n parallel-checkout.h | 25 ++++++++++++++\n 2 files changed, 105 insertions(+)\n\ndiff --git a/parallel-checkout.c b/parallel-checkout.c\nindex b5a714c711..6e62e044d8 100644\n--- a/parallel-checkout.c\n+++ b/parallel-checkout.c\n@@ -328,6 +328,24 @@ static int close_and_clear(int *fd)\n \treturn ret;\n }\n \n+void *parallel_unlink_proc(void *_data)\n+{\n+\tstruct parallel_unlink_data *data = _data;\n+\tstruct cache_def cache = CACHE_DEF_INIT;\n+\tint i = data->start;\n+\tdata->cnt = 0;\n+\n+\twhile (i < data->len) {\n+\t\tconst struct cache_entry *ce = data->cache[i];\n+\t\tif (ce->ce_flags & CE_WT_REMOVE) {\n+\t\t\t++data->cnt;\n+\t\t\tthreaded_unlink_entry(ce, data->super_prefix, data->removal_cache, &cache);\n+\t\t}\n+\t\ti += data->step;\n+\t}\n+\treturn &data->cnt;\n+}\n+\n void write_pc_item(struct parallel_checkout_item *pc_item,\n \t\t   struct checkout *state)\n {\n@@ -678,3 +696,65 @@ int run_parallel_checkout(struct checkout *state, int num_workers, int threshold\n \tfinish_parallel_checkout();\n \treturn ret;\n }\n+\n+unsigned run_parallel_unlink(struct index_state *index,\n+\t\t\t  struct progress *progress,\n+\t\t\t  const char *super_prefix, int num_workers, int threshold,\n+\t\t\t  unsigned cnt)\n+{\n+\tint i, use_parallel = 0, errs = 0;\n+\tif (num_workers > 1 && index->cache_nr >= threshold) {\n+\t\tint unlink_cnt = 0;\n+\t\tfor (i = 0; i < index->cache_nr; i++) {\n+\t\t\tconst struct cache_entry *ce = index->cache[i];\n+\t\t\tif (ce->ce_flags & CE_WT_REMOVE) {\n+\t\t\t\tunlink_cnt++;\n+\t\t\t}\n+\t\t}\n+\t\tif (unlink_cnt >= threshold) {\n+\t\t\tuse_parallel = 1;\n+\t\t}\n+\t}\n+\tif (use_parallel) {\n+\t\tstruct parallel_unlink_data *unlink_data;\n+\t\tCALLOC_ARRAY(unlink_data, num_workers);\n+\t\tthreaded_init_remove_scheduled_dirs();\n+\t\tstruct strbuf removal_caches[num_workers];\n+\t\tfor (i = 0; i < num_workers; i++) {\n+\t\t\tstruct parallel_unlink_data *data = &unlink_data[i];\n+\t\t\tstrbuf_init(&removal_caches[i], 50);\n+\t\t\tdata->start = i;\n+\t\t\tdata->cache = index->cache;\n+\t\t\tdata->len = index->cache_nr;\n+\t\t\tdata->step = num_workers;\n+\t\t\tdata->super_prefix = super_prefix;\n+\t\t\tdata->removal_cache = &removal_caches[i];\n+\t\t\terrs = pthread_create(&data->pthread, NULL, parallel_unlink_proc, data);\n+\t\t\tif (errs)\n+\t\t\t\tdie(_(\"unable to create parallel_checkout thread: %s\"), strerror(errs));\n+\t\t}\n+\t\tfor (i = 0; i < num_workers; i++) {\n+\t\t\tvoid *t_cnt;\n+\t\t\tif (pthread_join(unlink_data[i].pthread, &t_cnt))\n+\t\t\t\tdie(\"unable to join parallel_unlink_thread\");\n+\t\t\tcnt += *((unsigned *)t_cnt);\n+\t\t\tdisplay_progress(progress, cnt);\n+\t\t}\n+\t\tthreaded_remove_scheduled_dirs_clean_up();\n+\t\tfor (i = 0; i < num_workers; i++) {\n+\t\t\tthreaded_remove_scheduled_dirs(&removal_caches[i]);\n+\t\t}\n+\t\tremove_marked_cache_entries(index, 0);\n+\t} else {\n+\t\tfor (i = 0; i < index->cache_nr; i++) {\n+\t\t\tconst struct cache_entry *ce = index->cache[i];\n+\t\t\tif (ce->ce_flags & CE_WT_REMOVE) {\n+\t\t\t\tdisplay_progress(progress, ++cnt);\n+\t\t\t\tunlink_entry(ce, super_prefix);\n+\t\t\t}\n+\t\t}\n+\t\tremove_marked_cache_entries(index, 0);\n+\t    remove_scheduled_dirs();\n+\t}\n+\treturn cnt;\n+}\ndiff --git a/parallel-checkout.h b/parallel-checkout.h\nindex c575284005..e851b773d9 100644\n--- a/parallel-checkout.h\n+++ b/parallel-checkout.h\n@@ -43,6 +43,18 @@ size_t pc_queue_size(void);\n int run_parallel_checkout(struct checkout *state, int num_workers, int threshold,\n \t\t\t  struct progress *progress, unsigned int *progress_cnt);\n \n+/*\n+ * Unlink all the unlink entries in the index, returning the number of entries\n+ * unlinked plus the origin value of cnt. If the number of entries\n+ * to be removed is smaller than the specified threshold, the operation\n+ * is performed sequentially.\n+ */\n+unsigned run_parallel_unlink(struct index_state *index,\n+\t\t\t  struct progress *progress,\n+\t\t\t  const char *super_prefix,\n+\t\t\t  int num_workers, int threshold,\n+\t\t\t  unsigned cnt);\n+\n /****************************************************************\n  * Interface with checkout--worker\n  ****************************************************************/\n@@ -76,6 +88,19 @@ struct parallel_checkout_item {\n \tstruct stat st;\n };\n \n+struct parallel_unlink_data {\n+\tpthread_t pthread;\n+\tstruct cache_entry **cache;\n+\tstruct strbuf *removal_cache;\n+\tsize_t len;\n+\tint start;\n+\tsize_t step;\n+\tunsigned cnt;\n+\tconst char *super_prefix;\n+};\n+\n+void *parallel_unlink_proc(void *_data);\n+\n /*\n  * The fixed-size portion of `struct parallel_checkout_item` that is sent to the\n  * workers. Following this will be 2 strings: ca.working_tree_encoding and\n-- \n2.43.0\n\n"},{"id":"485343","messageId":"20231203133911.41594-5-hanyoung@protonmail.com","threadId":"60580","inReplyTo":"20231203133911.41594-1-hanyoung@protonmail.com","subject":"[RFC PATCH 4/4] unpack-trees: introduce parallel_unlink","fromName":"Han Young","fromEmail":"hanyang.tony@bytedance.com","sentAt":"2023-12-03T13:39:11Z","receivedAt":"2023-12-03T13:39:33Z","isPatch":true,"sender":{"key":"hanyang.tony@bytedance.com","avatar":"https://avatars.githubusercontent.com/u/108711387?v=4"},"body":"From: Han Young <hanyang.tony@bytedance.com>\n\nWe have parallel_checkout option since 04155bdad, but the unlink is still executed single threaded. On very large repo, checkout across directory rename or restructure commit can lead to large amount of unlinked entries. In some instance, the unlink operation can be slower than the parallel checkout. This commit add parallel unlink support, parallel unlink uses multithreaded removal of entries.\n---\nUnlink operation by itself is way faster than checkout, the default threshold should be way higher\nthan parallel_checkout. I hardcoded the threshold to be 100 times higher, probably need to introduce\na new config option with sensible default.\nTo discover how many entries to remove require us to iterate index->cache, this is fast even for large\nnumber of entries compare to filesystem operation.\nI think we can reuse checkout.workers as the main switch for parallel_unlink, since it's also part of\ncheckout process.\n\n unpack-trees.c | 15 ++-------------\n 1 file changed, 2 insertions(+), 13 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex c2b20b80d5..53589cde8a 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -452,17 +452,8 @@ static int check_updates(struct unpack_trees_options *o,\n \tif (should_update_submodules())\n \t\tload_gitmodules_file(index, NULL);\n \n-\tfor (i = 0; i < index->cache_nr; i++) {\n-\t\tconst struct cache_entry *ce = index->cache[i];\n-\n-\t\tif (ce->ce_flags & CE_WT_REMOVE) {\n-\t\t\tdisplay_progress(progress, ++cnt);\n-\t\t\tunlink_entry(ce, o->super_prefix);\n-\t\t}\n-\t}\n-\n-\tremove_marked_cache_entries(index, 0);\n-\tremove_scheduled_dirs();\n+\tget_parallel_checkout_configs(&pc_workers, &pc_threshold);\n+\tcnt = run_parallel_unlink(index, progress, o->super_prefix, pc_workers, pc_threshold * 100, cnt);\n \n \tif (should_update_submodules())\n \t\tload_gitmodules_file(index, &state);\n@@ -474,8 +465,6 @@ static int check_updates(struct unpack_trees_options *o,\n \t\t */\n \t\tprefetch_cache_entries(index, must_checkout);\n \n-\tget_parallel_checkout_configs(&pc_workers, &pc_threshold);\n-\n \tenable_delayed_checkout(&state);\n \tif (pc_workers > 1)\n \t\tinit_parallel_checkout();\n-- \n2.43.0\n\n"}]}