{"thread":{"id":"61624","subject":"[PATCH 00/16] mktree: support more flexible usage","startedAt":"2024-06-11T18:24:52Z","lastAt":"2024-07-10T21:40:49Z","messageCount":65,"participants":["Victoria Dye via GitGitGadget","Eric Sunshine","Junio C Hamano","Patrick Steinhardt","Victoria Dye"],"isPatch":true,"patchVersion":1,"patchTotal":16},"messages":[{"id":"496938","messageId":"pull.1746.git.1718130288.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":null,"subject":"[PATCH 00/16] mktree: support more flexible usage","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:32Z","receivedAt":"2024-06-11T18:24:52Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"The goal of this series is to make 'git mktree' a much more flexible and\npowerful tool for constructing arbitrary trees in memory without the use of\nan index or worktree. The main additions are:\n\n * Using an optional \"base tree\" to add or replace entries in an existing\n   tree rather than creating a new one from scratch.\n   * Building off of this, having entries with mode \"0\" indicate \"remove\n     this entry, if it exists, from the tree\"\n * Handling tree entries inside of subtrees (e.g., folder1/my-file.txt)\n\nIt also introduces some quality-of-life updates:\n\n * Using the same input parsing as 'update-index' to allow a wider variety\n   of tree entry formats.\n * Adding deduplication of input entries & more thorough validation of\n   inputs (with an option to disable both - plus input sorting - if desired\n   with '--literally').\n\nThe implementation change underpinning the new features is completely\nrevamping how the tree is constructed in memory. Instead of writing a single\ntree object into a strbuf and hashing it into the object database, we\nconstruct an in-core sparse index and write out the root tree, as well as\nany new subtrees, using the cache tree infrastructure.\n\nThe series is organized as follows:\n\n * Commits 1-3 contain miscellaneous small renames/refactors to make the\n   code more readable & prepare for larger refactoring later.\n * Commits 4-7 generalize the input parsing performed by 'read_index_info()'\n   in 'update-index' and update 'mktree' to use it.\n * Commit 8 adds the '--literally' option to 'mktree'. Practically, this\n   option allows tests that currently use 'mktree' to generate corrupt trees\n   to continue functioning after we strengthen input validations.\n * Commits 9 & 10 add input path validation & entry deduplication,\n   respectively.\n * Commit 11 replaces the strbuf-to-object tree creation with construction\n   of an in-core index & writing out the cache tree.\n * Commits 12-14 add the ability to add tree entries to an existing \"base\"\n   tree. Takes 3 commits to do it because it requires a bit of finesse\n   around directory/file deduplication and iterating over a tree with\n   'read_tree()' with a parallel iteration over the input tree entries.\n * Commit 15 allows for deeper paths in the input.\n * Commit 16 adds handling for mode '0' as \"removal\" entries.\n\nI also plan to add a '--strict' option that runs 'fsck' checks on the new\ntree(s) before writing to the object database (similar to 'mkttag\n--strict'), but this series is pretty long as it is and that part can easily\nbe separated out into its own series.\n\nThanks!\n\n * Victoria\n\nVictoria Dye (16):\n  mktree: use OPT_BOOL\n  mktree: rename treeent to tree_entry\n  mktree: use non-static tree_entry array\n  update-index: generalize 'read_index_info'\n  index-info.c: identify empty input lines in read_index_info\n  index-info.c: parse object type in provided in read_index_info\n  mktree: use read_index_info to read stdin lines\n  mktree: add a --literally option\n  mktree: validate paths more carefully\n  mktree: overwrite duplicate entries\n  mktree: create tree using an in-core index\n  mktree: use iterator struct to add tree entries to index\n  mktree: add directory-file conflict hashmap\n  mktree: optionally add to an existing tree\n  mktree: allow deeper paths in input\n  mktree: remove entries when mode is 0\n\n Documentation/git-mktree.txt       |  42 +-\n Makefile                           |   1 +\n builtin/mktree.c                   | 595 +++++++++++++++++++++++------\n builtin/update-index.c             | 119 ++----\n index-info.c                       | 104 +++++\n index-info.h                       |  14 +\n t/t1010-mktree.sh                  | 354 ++++++++++++++++-\n t/t1014-read-tree-confusing.sh     |   6 +-\n t/t1450-fsck.sh                    |   4 +-\n t/t1601-index-bogus.sh             |   2 +-\n t/t1700-split-index.sh             |   6 +-\n t/t2107-update-index-basic.sh      |  32 ++\n t/t7008-filter-branch-null-sha1.sh |   6 +-\n t/t7417-submodule-path-url.sh      |   2 +-\n t/t7450-bad-git-dotfiles.sh        |   8 +-\n 15 files changed, 1055 insertions(+), 240 deletions(-)\n create mode 100644 index-info.c\n create mode 100644 index-info.h\n\n\nbase-commit: 8d94cfb54504f2ec9edc7ca3eb5c29a3dd3675ae\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1746%2Fvdye%2Fvdye%2Fmktree-recursive-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1746/vdye/vdye/mktree-recursive-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1746\n-- \ngitgitgadget\n"},{"id":"496939","messageId":"074dc98acc79e08d07cf4f5c8105b872ec57980c.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 01/16] mktree: use OPT_BOOL","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:33Z","receivedAt":"2024-06-11T18:24:52Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace 'OPT_SET_INT' with 'OPT_BOOL' for the options '--missing' and\n'--batch'. The use of 'OPT_SET_INT' in these options is identical to\n'OPT_BOOL', but 'OPT_BOOL' provides slightly simpler syntax.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 9a22d4e2773..8b19d440747 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -162,8 +162,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n-\t\tOPT_SET_INT( 0 , \"missing\", &allow_missing, N_(\"allow missing objects\"), 1),\n-\t\tOPT_SET_INT( 0 , \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\"), 1),\n+\t\tOPT_BOOL(0, \"missing\", &allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n \n-- \ngitgitgadget\n\n"},{"id":"496940","messageId":"4558f35e7bf9a1594510951ee54252069bdcfc5b.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 02/16] mktree: rename treeent to tree_entry","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:34Z","receivedAt":"2024-06-11T18:24:54Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nRename the type for better readability, clearly specifying \"entry\" (instead\nof the \"ent\" abbreviation) and separating \"tree\" from \"entry\".\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 10 +++++-----\n 1 file changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 8b19d440747..c02feb06aff 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,7 +12,7 @@\n #include \"parse-options.h\"\n #include \"object-store-ll.h\"\n \n-static struct treeent {\n+static struct tree_entry {\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n@@ -22,7 +22,7 @@ static int alloc, used;\n \n static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n {\n-\tstruct treeent *ent;\n+\tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n \tif (strchr(path, '/'))\n \t\tdie(\"path %s contains slash\", path);\n@@ -38,8 +38,8 @@ static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n \n static int ent_compare(const void *a_, const void *b_)\n {\n-\tstruct treeent *a = *(struct treeent **)a_;\n-\tstruct treeent *b = *(struct treeent **)b_;\n+\tstruct tree_entry *a = *(struct tree_entry **)a_;\n+\tstruct tree_entry *b = *(struct tree_entry **)b_;\n \treturn base_name_compare(a->name, a->len, a->mode,\n \t\t\t\t b->name, b->len, b->mode);\n }\n@@ -56,7 +56,7 @@ static void write_tree(struct object_id *oid)\n \n \tstrbuf_init(&buf, size);\n \tfor (i = 0; i < used; i++) {\n-\t\tstruct treeent *ent = entries[i];\n+\t\tstruct tree_entry *ent = entries[i];\n \t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n \t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"496941","messageId":"5ade145352f44b431c16a2ec29cd87de489e8032.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 03/16] mktree: use non-static tree_entry array","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:35Z","receivedAt":"2024-06-11T18:24:54Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace the static 'struct tree_entry **entries' with a non-static 'struct\ntree_entry_array' instance. In later commits, we'll want to be able to\ncreate additional 'struct tree_entry_array' instances utilizing common\nfunctionality (create, push, clear, free). To avoid code duplication, create\nthe 'struct tree_entry_array' type and add functions that perform those\nbasic operations.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 67 +++++++++++++++++++++++++++++++++---------------\n 1 file changed, 47 insertions(+), 20 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex c02feb06aff..15bd908702a 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,15 +12,39 @@\n #include \"parse-options.h\"\n #include \"object-store-ll.h\"\n \n-static struct tree_entry {\n+struct tree_entry {\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n \tchar name[FLEX_ARRAY];\n-} **entries;\n-static int alloc, used;\n+};\n+\n+struct tree_entry_array {\n+\tsize_t nr, alloc;\n+\tstruct tree_entry **entries;\n+};\n \n-static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n+static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n+{\n+\tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n+\tarr->entries[arr->nr++] = ent;\n+}\n+\n+static void clear_tree_entry_array(struct tree_entry_array *arr)\n+{\n+\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\tFREE_AND_NULL(arr->entries[i]);\n+\tarr->nr = 0;\n+}\n+\n+static void release_tree_entry_array(struct tree_entry_array *arr)\n+{\n+\tFREE_AND_NULL(arr->entries);\n+\tarr->nr = arr->alloc = 0;\n+}\n+\n+static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n+\t\t\t   struct tree_entry_array *arr)\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n@@ -32,8 +56,8 @@ static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n \tent->len = len;\n \toidcpy(&ent->oid, oid);\n \n-\tALLOC_GROW(entries, used + 1, alloc);\n-\tentries[used++] = ent;\n+\t/* Append the update */\n+\ttree_entry_array_push(arr, ent);\n }\n \n static int ent_compare(const void *a_, const void *b_)\n@@ -44,19 +68,18 @@ static int ent_compare(const void *a_, const void *b_)\n \t\t\t\t b->name, b->len, b->mode);\n }\n \n-static void write_tree(struct object_id *oid)\n+static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n \tstruct strbuf buf;\n-\tsize_t size;\n-\tint i;\n+\tsize_t size = 0;\n \n-\tQSORT(entries, used, ent_compare);\n-\tfor (size = i = 0; i < used; i++)\n-\t\tsize += 32 + entries[i]->len;\n+\tQSORT(arr->entries, arr->nr, ent_compare);\n+\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\tsize += 32 + arr->entries[i]->len;\n \n \tstrbuf_init(&buf, size);\n-\tfor (i = 0; i < used; i++) {\n-\t\tstruct tree_entry *ent = entries[i];\n+\tfor (size_t i = 0; i < arr->nr; i++) {\n+\t\tstruct tree_entry *ent = arr->entries[i];\n \t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n \t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n \t}\n@@ -70,7 +93,8 @@ static const char *mktree_usage[] = {\n \tNULL\n };\n \n-static void mktree_line(char *buf, int nul_term_line, int allow_missing)\n+static void mktree_line(char *buf, int nul_term_line, int allow_missing,\n+\t\t\tstruct tree_entry_array *arr)\n {\n \tchar *ptr, *ntr;\n \tconst char *p;\n@@ -146,7 +170,7 @@ static void mktree_line(char *buf, int nul_term_line, int allow_missing)\n \t\t}\n \t}\n \n-\tappend_to_tree(mode, &oid, path);\n+\tappend_to_tree(mode, &oid, path, arr);\n \tfree(to_free);\n }\n \n@@ -158,6 +182,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tint allow_missing = 0;\n \tint is_batch_mode = 0;\n \tint got_eof = 0;\n+\tstruct tree_entry_array arr = { 0 };\n \tstrbuf_getline_fn getline_fn;\n \n \tconst struct option option[] = {\n@@ -182,9 +207,9 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t\t\tbreak;\n \t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n \t\t\t}\n-\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing);\n+\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n \t\t}\n-\t\tif (is_batch_mode && got_eof && used < 1) {\n+\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n \t\t\t/*\n \t\t\t * Execution gets here if the last tree entry is terminated with a\n \t\t\t * new-line.  The final new-line has been made optional to be\n@@ -192,12 +217,14 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t */\n \t\t\t; /* skip creating an empty tree */\n \t\t} else {\n-\t\t\twrite_tree(&oid);\n+\t\t\twrite_tree(&arr, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\n-\t\tused=0; /* reset tree entry buffer for re-use in batch mode */\n+\t\tclear_tree_entry_array(&arr); /* reset tree entry buffer for re-use in batch mode */\n \t}\n+\n+\trelease_tree_entry_array(&arr);\n \tstrbuf_release(&sb);\n \treturn 0;\n }\n-- \ngitgitgadget\n\n"},{"id":"496942","messageId":"9d0689e9c285b375b0067760929011038c085d65.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 04/16] update-index: generalize 'read_index_info'","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:36Z","receivedAt":"2024-06-11T18:24:55Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nMove 'read_index_info()' into a new header 'index-info.h' and generalize the\nfunction to call a provided callback for each parsed line. Update\n'update-index.c' to use this generalized 'read_index_info()', adding the\ncallback 'apply_index_info()' to verify the parsed line and update the index\naccording to its contents.\n\nThe input parsing done by 'read_index_info()' is similar to, but more\nflexible than, the parsing done in 'mktree' by 'mktree_line()' (handling not\nonly 'git ls-tree' output but also the outputs of 'git apply --index-info'\nand 'git ls-files --stage' outputs). To make 'mktree' more flexible, a later\npatch will replace mktree's custom parsing with 'read_index_info()'.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Makefile                      |   1 +\n builtin/update-index.c        | 116 ++++++++--------------------------\n index-info.c                  |  91 ++++++++++++++++++++++++++\n index-info.h                  |  11 ++++\n t/t2107-update-index-basic.sh |  27 ++++++++\n 5 files changed, 155 insertions(+), 91 deletions(-)\n create mode 100644 index-info.c\n create mode 100644 index-info.h\n\ndiff --git a/Makefile b/Makefile\nindex 2f5f16847ae..db9604e59c3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1037,6 +1037,7 @@ LIB_OBJS += hex.o\n LIB_OBJS += hex-ll.o\n LIB_OBJS += hook.o\n LIB_OBJS += ident.o\n+LIB_OBJS += index-info.o\n LIB_OBJS += json-writer.o\n LIB_OBJS += kwset.o\n LIB_OBJS += levenshtein.o\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex d343416ae26..77df380cb54 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -11,6 +11,7 @@\n #include \"gettext.h\"\n #include \"hash.h\"\n #include \"hex.h\"\n+#include \"index-info.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n #include \"cache-tree.h\"\n@@ -509,100 +510,29 @@ static void update_one(const char *path)\n \treport(\"add '%s'\", path);\n }\n \n-static void read_index_info(int nul_term_line)\n+static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n+\t\t\t    const char *path_name, void *cbdata UNUSED)\n {\n-\tconst int hexsz = the_hash_algo->hexsz;\n-\tstruct strbuf buf = STRBUF_INIT;\n-\tstruct strbuf uq = STRBUF_INIT;\n-\tstrbuf_getline_fn getline_fn;\n+\tif (!verify_path(path_name, mode)) {\n+\t\tfprintf(stderr, \"Ignoring path %s\\n\", path_name);\n+\t\treturn 0;\n+\t}\n \n-\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n-\twhile (getline_fn(&buf, stdin) != EOF) {\n-\t\tchar *ptr, *tab;\n-\t\tchar *path_name;\n-\t\tstruct object_id oid;\n-\t\tunsigned int mode;\n-\t\tunsigned long ul;\n-\t\tint stage;\n-\n-\t\t/* This reads lines formatted in one of three formats:\n-\t\t *\n-\t\t * (1) mode         SP sha1          TAB path\n-\t\t * The first format is what \"git apply --index-info\"\n-\t\t * reports, and used to reconstruct a partial tree\n-\t\t * that is used for phony merge base tree when falling\n-\t\t * back on 3-way merge.\n-\t\t *\n-\t\t * (2) mode SP type SP sha1          TAB path\n-\t\t * The second format is to stuff \"git ls-tree\" output\n-\t\t * into the index file.\n-\t\t *\n-\t\t * (3) mode         SP sha1 SP stage TAB path\n-\t\t * This format is to put higher order stages into the\n-\t\t * index file and matches \"git ls-files --stage\" output.\n+\tif (!mode) {\n+\t\t/* mode == 0 means there is no such path -- remove */\n+\t\tif (remove_file_from_index(the_repository->index, path_name))\n+\t\t\tdie(\"git update-index: unable to remove %s\", path_name);\n+\t}\n+\telse {\n+\t\t/* mode ' ' sha1 '\\t' name\n+\t\t * ptr[-1] points at tab,\n+\t\t * ptr[-41] is at the beginning of sha1\n \t\t */\n-\t\terrno = 0;\n-\t\tul = strtoul(buf.buf, &ptr, 8);\n-\t\tif (ptr == buf.buf || *ptr != ' '\n-\t\t    || errno || (unsigned int) ul != ul)\n-\t\t\tgoto bad_line;\n-\t\tmode = ul;\n-\n-\t\ttab = strchr(ptr, '\\t');\n-\t\tif (!tab || tab - ptr < hexsz + 1)\n-\t\t\tgoto bad_line;\n-\n-\t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n-\t\t\tstage = tab[-1] - '0';\n-\t\t\tptr = tab + 1; /* point at the head of path */\n-\t\t\ttab = tab - 2; /* point at tail of sha1 */\n-\t\t}\n-\t\telse {\n-\t\t\tstage = 0;\n-\t\t\tptr = tab + 1; /* point at the head of path */\n-\t\t}\n-\n-\t\tif (get_oid_hex(tab - hexsz, &oid) ||\n-\t\t\ttab[-(hexsz + 1)] != ' ')\n-\t\t\tgoto bad_line;\n-\n-\t\tpath_name = ptr;\n-\t\tif (!nul_term_line && path_name[0] == '\"') {\n-\t\t\tstrbuf_reset(&uq);\n-\t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n-\t\t\t\tdie(\"git update-index: bad quoting of path name\");\n-\t\t\t}\n-\t\t\tpath_name = uq.buf;\n-\t\t}\n-\n-\t\tif (!verify_path(path_name, mode)) {\n-\t\t\tfprintf(stderr, \"Ignoring path %s\\n\", path_name);\n-\t\t\tcontinue;\n-\t\t}\n-\n-\t\tif (!mode) {\n-\t\t\t/* mode == 0 means there is no such path -- remove */\n-\t\t\tif (remove_file_from_index(the_repository->index, path_name))\n-\t\t\t\tdie(\"git update-index: unable to remove %s\",\n-\t\t\t\t    ptr);\n-\t\t}\n-\t\telse {\n-\t\t\t/* mode ' ' sha1 '\\t' name\n-\t\t\t * ptr[-1] points at tab,\n-\t\t\t * ptr[-41] is at the beginning of sha1\n-\t\t\t */\n-\t\t\tptr[-(hexsz + 2)] = ptr[-1] = 0;\n-\t\t\tif (add_cacheinfo(mode, &oid, path_name, stage))\n-\t\t\t\tdie(\"git update-index: unable to update %s\",\n-\t\t\t\t    path_name);\n-\t\t}\n-\t\tcontinue;\n-\n-\tbad_line:\n-\t\tdie(\"malformed index info %s\", buf.buf);\n+\t\tif (add_cacheinfo(mode, oid, path_name, stage))\n+\t\t\tdie(\"git update-index: unable to update %s\", path_name);\n \t}\n-\tstrbuf_release(&buf);\n-\tstrbuf_release(&uq);\n+\n+\treturn 0;\n }\n \n static const char * const update_index_usage[] = {\n@@ -849,6 +779,7 @@ static enum parse_opt_result stdin_cacheinfo_callback(\n \tconst char *arg, int unset)\n {\n \tint *nul_term_line = opt->value;\n+\tint ret;\n \n \tBUG_ON_OPT_NEG(unset);\n \tBUG_ON_OPT_ARG(arg);\n@@ -856,7 +787,10 @@ static enum parse_opt_result stdin_cacheinfo_callback(\n \tif (ctx->argc != 1)\n \t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n \tallow_add = allow_replace = allow_remove = 1;\n-\tread_index_info(*nul_term_line);\n+\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n+\tif (ret)\n+\t\treturn -1;\n+\n \treturn 0;\n }\n \ndiff --git a/index-info.c b/index-info.c\nnew file mode 100644\nindex 00000000000..0b68e34c361\n--- /dev/null\n+++ b/index-info.c\n@@ -0,0 +1,91 @@\n+#include \"git-compat-util.h\"\n+#include \"index-info.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"strbuf.h\"\n+#include \"quote.h\"\n+\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n+{\n+\tconst int hexsz = the_hash_algo->hexsz;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf uq = STRBUF_INIT;\n+\tstrbuf_getline_fn getline_fn;\n+\tint ret = 0;\n+\n+\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n+\twhile (getline_fn(&buf, stdin) != EOF) {\n+\t\tchar *ptr, *tab;\n+\t\tchar *path_name;\n+\t\tstruct object_id oid;\n+\t\tunsigned int mode;\n+\t\tunsigned long ul;\n+\t\tint stage;\n+\n+\t\t/* This reads lines formatted in one of three formats:\n+\t\t *\n+\t\t * (1) mode         SP sha1          TAB path\n+\t\t * The first format is what \"git apply --index-info\"\n+\t\t * reports, and used to reconstruct a partial tree\n+\t\t * that is used for phony merge base tree when falling\n+\t\t * back on 3-way merge.\n+\t\t *\n+\t\t * (2) mode SP type SP sha1          TAB path\n+\t\t * The second format is to stuff \"git ls-tree\" output\n+\t\t * into the index file.\n+\t\t *\n+\t\t * (3) mode         SP sha1 SP stage TAB path\n+\t\t * This format is to put higher order stages into the\n+\t\t * index file and matches \"git ls-files --stage\" output.\n+\t\t */\n+\t\terrno = 0;\n+\t\tul = strtoul(buf.buf, &ptr, 8);\n+\t\tif (ptr == buf.buf || *ptr != ' '\n+\t\t    || errno || (unsigned int) ul != ul)\n+\t\t\tgoto bad_line;\n+\t\tmode = ul;\n+\n+\t\ttab = strchr(ptr, '\\t');\n+\t\tif (!tab || tab - ptr < hexsz + 1)\n+\t\t\tgoto bad_line;\n+\n+\t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n+\t\t\tstage = tab[-1] - '0';\n+\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\ttab = tab - 2; /* point at tail of sha1 */\n+\t\t} else {\n+\t\t\tstage = 0;\n+\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t}\n+\n+\t\tif (get_oid_hex(tab - hexsz, &oid) ||\n+\t\t\ttab[-(hexsz + 1)] != ' ')\n+\t\t\tgoto bad_line;\n+\n+\t\tpath_name = ptr;\n+\t\tif (!nul_term_line && path_name[0] == '\"') {\n+\t\t\tstrbuf_reset(&uq);\n+\t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n+\t\t\t\tret = error(\"bad quoting of path name\");\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tpath_name = uq.buf;\n+\t\t}\n+\n+\t\tret = fn(mode, &oid, stage, path_name, cbdata);\n+\t\tif (ret) {\n+\t\t\tret = -1;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tcontinue;\n+\n+\tbad_line:\n+\t\tret = error(\"malformed input line '%s'\", buf.buf);\n+\t\tbreak;\n+\t}\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&uq);\n+\n+\treturn ret;\n+}\ndiff --git a/index-info.h b/index-info.h\nnew file mode 100644\nindex 00000000000..d650498325a\n--- /dev/null\n+++ b/index-info.h\n@@ -0,0 +1,11 @@\n+#ifndef INDEX_INFO_H\n+#define INDEX_INFO_H\n+\n+#include \"hash.h\"\n+\n+typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n+\n+/* Iterate over parsed index info from stdin */\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n+\n+#endif /* INDEX_INFO_H */\ndiff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\nindex cc72ead79f3..29696ade0d0 100755\n--- a/t/t2107-update-index-basic.sh\n+++ b/t/t2107-update-index-basic.sh\n@@ -142,4 +142,31 @@ test_expect_success '--index-version' '\n \ttest_must_be_empty actual\n '\n \n+test_expect_success '--index-info fails on malformed input' '\n+\t# empty line\n+\techo \"\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# bad whitespace\n+\tprintf \"100644 $EMPTY_BLOB A\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# invalid stage value\n+\tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# invalid OID length\n+\tprintf \"100755 abc123\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# bad quoting\n+\tprintf \"100644 $EMPTY_BLOB\\t\\\"A\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"bad quoting of path name\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496943","messageId":"f56eee0b48da907a27edc99ca135cf8f6c19af35.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 06/16] index-info.c: parse object type in provided in read_index_info","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:38Z","receivedAt":"2024-06-11T18:24:56Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf the object type (e.g. \"blob\", \"tree\") is identified on a stdin line read\nby 'read_index_info()' (i.e. on lines formatted like the output of 'git\nls-tree'), parse it into an 'enum object_type' and provide it to the\n'read_index_info()' callback as an argument. If the type is not provided,\npass 'OBJ_NONE' instead. If the object type is invalid, return an error.\n\nThe goal of this change is to allow for more thorough validation of the\nprovided object type (e.g. against the provided mode) in 'mktree' once\n'mktree_line' is replaced with 'read_index_info()'. Note, though, that this\nchange also strengthens the validation done by 'update-index', since invalid\ntype names now trigger an error.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/update-index.c        |  3 ++-\n index-info.c                  | 16 ++++++++++++----\n index-info.h                  |  3 ++-\n t/t2107-update-index-basic.sh |  5 +++++\n 4 files changed, 21 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex b1b334807f8..8882433b644 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -510,7 +510,8 @@ static void update_one(const char *path)\n \treport(\"add '%s'\", path);\n }\n \n-static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n+static int apply_index_info(unsigned int mode, struct object_id *oid,\n+\t\t\t    enum object_type obj_type UNUSED, int stage,\n \t\t\t    const char *path_name, void *cbdata UNUSED)\n {\n \tif (!verify_path(path_name, mode)) {\ndiff --git a/index-info.c b/index-info.c\nindex 735cbf1f476..5d61e61e28f 100644\n--- a/index-info.c\n+++ b/index-info.c\n@@ -18,6 +18,7 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \t\tchar *ptr, *tab;\n \t\tchar *path_name;\n \t\tstruct object_id oid;\n+\t\tenum object_type obj_type = OBJ_NONE;\n \t\tunsigned int mode;\n \t\tunsigned long ul;\n \t\tint stage;\n@@ -56,18 +57,17 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \n \t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n \t\t\tstage = tab[-1] - '0';\n-\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\tpath_name = tab + 1; /* point at the head of path */\n \t\t\ttab = tab - 2; /* point at tail of sha1 */\n \t\t} else {\n \t\t\tstage = 0;\n-\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\tpath_name = tab + 1; /* point at the head of path */\n \t\t}\n \n \t\tif (get_oid_hex(tab - hexsz, &oid) ||\n \t\t\ttab[-(hexsz + 1)] != ' ')\n \t\t\tgoto bad_line;\n \n-\t\tpath_name = ptr;\n \t\tif (!nul_term_line && path_name[0] == '\"') {\n \t\t\tstrbuf_reset(&uq);\n \t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n@@ -77,7 +77,15 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \t\t\tpath_name = uq.buf;\n \t\t}\n \n-\t\tret = fn(mode, &oid, stage, path_name, cbdata);\n+\t\t/* Get the type, if provided */\n+\t\tif (tab - hexsz - 1 > ptr + 1) {\n+\t\t\tif (*(tab - hexsz - 1) != ' ')\n+\t\t\t\tgoto bad_line;\n+\t\t\t*(tab - hexsz - 1) = '\\0';\n+\t\t\tobj_type = type_from_string(ptr + 1);\n+\t\t}\n+\n+\t\tret = fn(mode, &oid, obj_type, stage, path_name, cbdata);\n \t\tif (ret) {\n \t\t\tret = -1;\n \t\t\tbreak;\ndiff --git a/index-info.h b/index-info.h\nindex 1884972021d..767cf304213 100644\n--- a/index-info.h\n+++ b/index-info.h\n@@ -2,8 +2,9 @@\n #define INDEX_INFO_H\n \n #include \"hash.h\"\n+#include \"object.h\"\n \n-typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n+typedef int (*each_index_info_fn)(unsigned int, struct object_id *, enum object_type, int, const char *, void *);\n \n #define INDEX_INFO_EMPTY_LINE 1\n \ndiff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\nindex 29696ade0d0..9c19d24cd4a 100755\n--- a/t/t2107-update-index-basic.sh\n+++ b/t/t2107-update-index-basic.sh\n@@ -153,6 +153,11 @@ test_expect_success '--index-info fails on malformed input' '\n \ttest_must_fail git update-index --index-info 2>err &&\n \tgrep \"malformed input line\" err &&\n \n+\t# invalid type\n+\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\tgrep \"invalid object type\" err &&\n+\n \t# invalid stage value\n \tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n \ttest_must_fail git update-index --index-info 2>err &&\n-- \ngitgitgadget\n\n"},{"id":"496944","messageId":"7e3bcc16e23c97d8a4efbb9e14b230ef9f44a1a7.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 05/16] index-info.c: identify empty input lines in read_index_info","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:37Z","receivedAt":"2024-06-11T18:24:57Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUpdate 'read_index_info()' to return INDEX_INFO_EMPTY_LINE (value 1), rather\nthan the default error code (value -1) when the function encounters an empty\nline in stdin. This grants the caller the flexibility to handle such\nscenarios differently than a typical error. In the case of 'update-index',\nwe'll still exit with a \"malformed input line\" error. However, when\n'read_index_info()' is used to process the input to 'mktree' in a later\npatch, the empty line return value will signal a new tree in --batch mode.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/update-index.c | 4 +++-\n index-info.c           | 5 +++++\n index-info.h           | 2 ++\n 3 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 77df380cb54..b1b334807f8 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -788,7 +788,9 @@ static enum parse_opt_result stdin_cacheinfo_callback(\n \t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n \tallow_add = allow_replace = allow_remove = 1;\n \tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n-\tif (ret)\n+\tif (ret == INDEX_INFO_EMPTY_LINE)\n+\t\treturn error(\"malformed input line ''\");\n+\telse if (ret < 0)\n \t\treturn -1;\n \n \treturn 0;\ndiff --git a/index-info.c b/index-info.c\nindex 0b68e34c361..735cbf1f476 100644\n--- a/index-info.c\n+++ b/index-info.c\n@@ -22,6 +22,11 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \t\tunsigned long ul;\n \t\tint stage;\n \n+\t\tif (!buf.len) {\n+\t\t\tret = INDEX_INFO_EMPTY_LINE;\n+\t\t\tbreak;\n+\t\t}\n+\n \t\t/* This reads lines formatted in one of three formats:\n \t\t *\n \t\t * (1) mode         SP sha1          TAB path\ndiff --git a/index-info.h b/index-info.h\nindex d650498325a..1884972021d 100644\n--- a/index-info.h\n+++ b/index-info.h\n@@ -5,6 +5,8 @@\n \n typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n \n+#define INDEX_INFO_EMPTY_LINE 1\n+\n /* Iterate over parsed index info from stdin */\n int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n \n-- \ngitgitgadget\n\n"},{"id":"496946","messageId":"8d1e1eaa70b96779416f2f48a862d31a730c4521.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 07/16] mktree: use read_index_info to read stdin lines","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:39Z","receivedAt":"2024-06-11T18:24:57Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace the custom input parsing of 'mktree' with 'read_index_info()', which\nhandles not only the 'ls-tree' output format it already handles but also the\nother formats compatible with 'update-index'. This lends some consistency\nacross the commands (avoiding the need for two similar implementations for\ninput parsing) and adds flexibility to mktree.\n\nUpdate 'Documentation/git-mktree.txt' to reflect the more permissive input\nformat.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  17 +++--\n builtin/mktree.c             | 139 ++++++++++++-----------------------\n t/t1010-mktree.sh            |  66 +++++++++++++++++\n 3 files changed, 125 insertions(+), 97 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 383f09dd333..507682ed23e 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -3,7 +3,7 @@ git-mktree(1)\n \n NAME\n ----\n-git-mktree - Build a tree-object from ls-tree formatted text\n+git-mktree - Build a tree-object from formatted tree entries\n \n \n SYNOPSIS\n@@ -13,15 +13,13 @@ SYNOPSIS\n \n DESCRIPTION\n -----------\n-Reads standard input in non-recursive `ls-tree` output format, and creates\n-a tree object.  The order of the tree entries is normalized by mktree so\n-pre-sorting the input is not required.  The object name of the tree object\n-built is written to the standard output.\n+Reads entry information from stdin and creates a tree object from those entries.\n+The object name of the tree object built is written to the standard output.\n \n OPTIONS\n -------\n -z::\n-\tRead the NUL-terminated `ls-tree -z` output instead.\n+\tInput lines are separated with NUL rather than LF.\n \n --missing::\n \tAllow missing objects.  The default behaviour (without this option)\n@@ -35,6 +33,13 @@ OPTIONS\n \toptional.  Note - if the `-z` option is used, lines are terminated\n \twith NUL.\n \n+INPUT FORMAT\n+------------\n+Tree entries may be specified in any of the formats compatible with the\n+`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n+entries is normalized by `mktree` so pre-sorting the input by path is not\n+required.\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 15bd908702a..5530257252d 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -6,6 +6,7 @@\n #include \"builtin.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n+#include \"index-info.h\"\n #include \"quote.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n@@ -93,123 +94,80 @@ static const char *mktree_usage[] = {\n \tNULL\n };\n \n-static void mktree_line(char *buf, int nul_term_line, int allow_missing,\n-\t\t\tstruct tree_entry_array *arr)\n+struct mktree_line_data {\n+\tstruct tree_entry_array *arr;\n+\tint allow_missing;\n+};\n+\n+static int mktree_line(unsigned int mode, struct object_id *oid,\n+\t\t       enum object_type obj_type, int stage UNUSED,\n+\t\t       const char *path, void *cbdata)\n {\n-\tchar *ptr, *ntr;\n-\tconst char *p;\n-\tunsigned mode;\n-\tenum object_type mode_type; /* object type derived from mode */\n-\tenum object_type obj_type; /* object type derived from sha */\n+\tstruct mktree_line_data *data = cbdata;\n+\tenum object_type mode_type = object_type(mode);\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tchar *path, *to_free = NULL;\n-\tstruct object_id oid;\n+\tenum object_type parsed_obj_type;\n \n-\tptr = buf;\n-\t/*\n-\t * Read non-recursive ls-tree output format:\n-\t *     mode SP type SP sha1 TAB name\n-\t */\n-\tmode = strtoul(ptr, &ntr, 8);\n-\tif (ptr == ntr || !ntr || *ntr != ' ')\n-\t\tdie(\"input format error: %s\", buf);\n-\tptr = ntr + 1; /* type */\n-\tntr = strchr(ptr, ' ');\n-\tif (!ntr || parse_oid_hex(ntr + 1, &oid, &p) ||\n-\t    *p != '\\t')\n-\t\tdie(\"input format error: %s\", buf);\n-\n-\t/* It is perfectly normal if we do not have a commit from a submodule */\n-\tif (S_ISGITLINK(mode))\n-\t\tallow_missing = 1;\n-\n-\n-\t*ntr++ = 0; /* now at the beginning of SHA1 */\n-\n-\tpath = (char *)p + 1;  /* at the beginning of name */\n-\tif (!nul_term_line && path[0] == '\"') {\n-\t\tstruct strbuf p_uq = STRBUF_INIT;\n-\t\tif (unquote_c_style(&p_uq, path, NULL))\n-\t\t\tdie(\"invalid quoting\");\n-\t\tpath = to_free = strbuf_detach(&p_uq, NULL);\n-\t}\n+\tif (obj_type && mode_type != obj_type)\n+\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n+\t\t    type_name(obj_type), type_name(mode_type));\n \n-\t/*\n-\t * Object type is redundantly derivable three ways.\n-\t * These should all agree.\n-\t */\n-\tmode_type = object_type(mode);\n-\tif (mode_type != type_from_string(ptr)) {\n-\t\tdie(\"entry '%s' object type (%s) doesn't match mode type (%s)\",\n-\t\t\tpath, ptr, type_name(mode_type));\n-\t}\n+\toi.typep = &parsed_obj_type;\n \n-\t/* Check the type of object identified by oid without fetching objects */\n-\toi.typep = &obj_type;\n-\tif (oid_object_info_extended(the_repository, &oid, &oi,\n+\tif (oid_object_info_extended(the_repository, oid, &oi,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n \t\t\t\t     OBJECT_INFO_QUICK |\n \t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n-\t\tobj_type = -1;\n+\t\tparsed_obj_type = -1;\n \n-\tif (obj_type < 0) {\n-\t\tif (allow_missing) {\n-\t\t\t; /* no problem - missing objects are presumed to be of the right type */\n+\tif (parsed_obj_type < 0) {\n+\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n+\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n \t\t} else {\n-\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(&oid));\n-\t\t}\n-\t} else {\n-\t\tif (obj_type != mode_type) {\n-\t\t\t/*\n-\t\t\t * The object exists but is of the wrong type.\n-\t\t\t * This is a problem regardless of allow_missing\n-\t\t\t * because the new tree entry will never be correct.\n-\t\t\t */\n-\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n-\t\t\t\tpath, oid_to_hex(&oid), type_name(obj_type), type_name(mode_type));\n+\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n \t\t}\n+\t} else if (parsed_obj_type != mode_type) {\n+\t\t/*\n+\t\t * The object exists but is of the wrong type.\n+\t\t * This is a problem regardless of allow_missing\n+\t\t * because the new tree entry will never be correct.\n+\t\t */\n+\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n+\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n \t}\n \n-\tappend_to_tree(mode, &oid, path, arr);\n-\tfree(to_free);\n+\tappend_to_tree(mode, oid, path, data->arr);\n+\treturn 0;\n }\n \n int cmd_mktree(int ac, const char **av, const char *prefix)\n {\n-\tstruct strbuf sb = STRBUF_INIT;\n \tstruct object_id oid;\n \tint nul_term_line = 0;\n-\tint allow_missing = 0;\n \tint is_batch_mode = 0;\n-\tint got_eof = 0;\n \tstruct tree_entry_array arr = { 0 };\n-\tstrbuf_getline_fn getline_fn;\n+\tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n+\tint ret;\n \n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n-\t\tOPT_BOOL(0, \"missing\", &allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"missing\", &mktree_line_data.allow_missing, N_(\"allow missing objects\")),\n \t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n-\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n-\n-\twhile (!got_eof) {\n-\t\twhile (1) {\n-\t\t\tif (getline_fn(&sb, stdin) == EOF) {\n-\t\t\t\tgot_eof = 1;\n-\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tif (sb.buf[0] == '\\0') {\n-\t\t\t\t/* empty lines denote tree boundaries in batch mode */\n-\t\t\t\tif (is_batch_mode)\n-\t\t\t\t\tbreak;\n-\t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n-\t\t\t}\n-\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n-\t\t}\n-\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n+\n+\tdo {\n+\t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data);\n+\t\tif (ret < 0)\n+\t\t\tbreak;\n+\n+\t\t/* empty lines denote tree boundaries in batch mode */\n+\t\tif (ret > 0 && !is_batch_mode)\n+\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n+\n+\t\tif (is_batch_mode && !ret && arr.nr < 1) {\n \t\t\t/*\n \t\t\t * Execution gets here if the last tree entry is terminated with a\n \t\t\t * new-line.  The final new-line has been made optional to be\n@@ -222,9 +180,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\tfflush(stdout);\n \t\t}\n \t\tclear_tree_entry_array(&arr); /* reset tree entry buffer for re-use in batch mode */\n-\t}\n+\t} while (ret > 0);\n \n \trelease_tree_entry_array(&arr);\n-\tstrbuf_release(&sb);\n-\treturn 0;\n+\treturn !!ret;\n }\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 22875ba598c..9b2ab0c97ad 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -54,11 +54,36 @@ test_expect_success 'ls-tree output in wrong order given to mktree (2)' '\n \ttest_cmp tree.withsub actual\n '\n \n+test_expect_success '--batch creates multiple trees' '\n+\tcat top >multi-tree &&\n+\techo \"\" >>multi-tree &&\n+\tcat top.withsub >>multi-tree &&\n+\n+\tcat tree >expect &&\n+\tcat tree.withsub >>expect &&\n+\tgit mktree --batch <multi-tree >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'allow missing object with --missing' '\n \tgit mktree --missing <top.missing >actual &&\n \ttest_cmp tree.missing actual\n '\n \n+test_expect_success 'mktree with invalid submodule OIDs' '\n+\t# non-existent OID - ok\n+\tprintf \"160000 commit $(test_oid numeric)\\tA\\n\" >in &&\n+\tgit mktree <in >tree.actual &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\ttest_cmp in actual &&\n+\n+\t# existing OID, wrong type - error\n+\ttree_oid=\"$(cat tree)\" &&\n+\tprintf \"160000 commit $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"object $tree_oid is a tree but specified type was (commit)\" err\n+'\n+\n test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n \ttest_must_fail git mktree <all\n '\n@@ -67,4 +92,45 @@ test_expect_success 'mktree refuses to read ls-tree -r output (2)' '\n \ttest_must_fail git mktree <all.withsub\n '\n \n+test_expect_success 'mktree fails on malformed input' '\n+\t# empty line without --batch\n+\techo \"\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"blank line only valid in batch mode\" err &&\n+\n+\t# bad whitespace\n+\tprintf \"100644 blob $EMPTY_BLOB A\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# invalid type\n+\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"invalid object type\" err &&\n+\n+\t# invalid OID length\n+\tprintf \"100755 blob abc123\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"malformed input line\" err &&\n+\n+\t# bad quoting\n+\tprintf \"100644 blob $EMPTY_BLOB\\t\\\"A\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"bad quoting of path name\" err\n+'\n+\n+test_expect_success 'mktree fails on mode mismatch' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\n+\t# mode-type mismatch\n+\tprintf \"100644 tree $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"object type (tree) doesn${SQ}t match mode type (blob)\" err &&\n+\n+\t# mode-object mismatch (no --missing)\n+\tprintf \"100644 $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"object $tree_oid is a tree but specified type was (blob)\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496945","messageId":"b497dc90687a7c77a4d21c3a12fe5fa3bfdabc16.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 08/16] mktree: add a --literally option","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:40Z","receivedAt":"2024-06-11T18:24:58Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nAdd the '--literally' option to 'git mktree' to allow constructing a tree\nwith invalid contents. For now, the only change this represents compared to\nthe normal 'git mktree' behavior is no longer sorting the inputs; in later\ncommits, deduplicaton and path validation will be added to the command and\n'--literally' will skip those as well.\n\nCertain tests use 'git mktree' to intentionally generate corrupt trees.\nUpdate these tests to use '--literally' so that they continue functioning\nproperly when additional input cleanup & validation is added to the base\ncommand. Note that, because 'mktree --literally' does not sort entries, some\nof the tests are updated to provide their inputs in tree order; otherwise,\nthe test would fail with an \"incorrect order\" error instead of the error the\ntest expects.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt       |  9 ++++++-\n builtin/mktree.c                   | 36 +++++++++++++++++++++++----\n t/t1010-mktree.sh                  | 40 ++++++++++++++++++++++++++++++\n t/t1014-read-tree-confusing.sh     |  6 ++---\n t/t1450-fsck.sh                    |  4 +--\n t/t1601-index-bogus.sh             |  2 +-\n t/t1700-split-index.sh             |  6 ++---\n t/t7008-filter-branch-null-sha1.sh |  6 ++---\n t/t7417-submodule-path-url.sh      |  2 +-\n t/t7450-bad-git-dotfiles.sh        |  8 +++---\n 10 files changed, 96 insertions(+), 23 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 507682ed23e..fb07e40cef0 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -9,7 +9,7 @@ git-mktree - Build a tree-object from formatted tree entries\n SYNOPSIS\n --------\n [verse]\n-'git mktree' [-z] [--missing] [--batch]\n+'git mktree' [-z] [--missing] [--literally] [--batch]\n \n DESCRIPTION\n -----------\n@@ -27,6 +27,13 @@ OPTIONS\n \tobject.  This option has no effect on the treatment of gitlink entries\n \t(aka \"submodules\") which are always allowed to be missing.\n \n+--literally::\n+\tCreate the tree from the tree entries provided to stdin in the order\n+\tthey are provided without performing additional sorting, deduplication,\n+\tor path validation on them. This option is primarily useful for creating\n+\tinvalid tree objects to use in tests of how Git deals with various forms\n+\tof tree corruption.\n+\n --batch::\n \tAllow building of more than one tree object before exiting.  Each\n \ttree is separated by a single blank line. The final newline is\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 5530257252d..48019448c1f 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -45,11 +45,11 @@ static void release_tree_entry_array(struct tree_entry_array *arr)\n }\n \n static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n-\t\t\t   struct tree_entry_array *arr)\n+\t\t\t   struct tree_entry_array *arr, int literally)\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n-\tif (strchr(path, '/'))\n+\tif (!literally && strchr(path, '/'))\n \t\tdie(\"path %s contains slash\", path);\n \n \tFLEX_ALLOC_MEM(ent, name, path, len);\n@@ -89,14 +89,35 @@ static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n \tstrbuf_release(&buf);\n }\n \n+static void write_tree_literally(struct tree_entry_array *arr,\n+\t\t\t\t struct object_id *oid)\n+{\n+\tstruct strbuf buf;\n+\tsize_t size = 0;\n+\n+\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\tsize += 32 + arr->entries[i]->len;\n+\n+\tstrbuf_init(&buf, size);\n+\tfor (size_t i = 0; i < arr->nr; i++) {\n+\t\tstruct tree_entry *ent = arr->entries[i];\n+\t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n+\t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n+\t}\n+\n+\twrite_object_file(buf.buf, buf.len, OBJ_TREE, oid);\n+\tstrbuf_release(&buf);\n+}\n+\n static const char *mktree_usage[] = {\n-\t\"git mktree [-z] [--missing] [--batch]\",\n+\t\"git mktree [-z] [--missing] [--literally] [--batch]\",\n \tNULL\n };\n \n struct mktree_line_data {\n \tstruct tree_entry_array *arr;\n \tint allow_missing;\n+\tint literally;\n };\n \n static int mktree_line(unsigned int mode, struct object_id *oid,\n@@ -136,7 +157,7 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n \t}\n \n-\tappend_to_tree(mode, oid, path, data->arr);\n+\tappend_to_tree(mode, oid, path, data->arr, data->literally);\n \treturn 0;\n }\n \n@@ -152,6 +173,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n \t\tOPT_BOOL(0, \"missing\", &mktree_line_data.allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"literally\", &mktree_line_data.literally,\n+\t\t\t N_(\"do not sort, deduplicate, or validate paths of tree entries\")),\n \t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n@@ -175,7 +198,10 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t */\n \t\t\t; /* skip creating an empty tree */\n \t\t} else {\n-\t\t\twrite_tree(&arr, &oid);\n+\t\t\tif (mktree_line_data.literally)\n+\t\t\t\twrite_tree_literally(&arr, &oid);\n+\t\t\telse\n+\t\t\t\twrite_tree(&arr, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 9b2ab0c97ad..e0687cb529f 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -133,4 +133,44 @@ test_expect_success 'mktree fails on mode mismatch' '\n \tgrep \"object $tree_oid is a tree but specified type was (blob)\" err\n '\n \n+test_expect_success '--literally can create invalid trees' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse ${tree_oid}:one)\" &&\n+\n+\t# duplicate entries\n+\t{\n+\t\tprintf \"040000 tree $tree_oid\\tmy-tree\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest-file\\n\" &&\n+\t\tprintf \"100755 blob $blob_oid\\ttest-file\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\tgrep \"contains duplicate file entries\" err &&\n+\n+\t# disallowed path\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\t.git\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\tgrep \"contains ${SQ}.git${SQ}\" err &&\n+\n+\t# nested entry\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tdeeper/my-file\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\tgrep \"contains full pathnames\" err &&\n+\n+\t# bad entry ordering\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tB\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tA\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\tgrep \"not properly sorted\" err\n+'\n+\n test_done\ndiff --git a/t/t1014-read-tree-confusing.sh b/t/t1014-read-tree-confusing.sh\nindex 8ea8d36818b..762eb789704 100755\n--- a/t/t1014-read-tree-confusing.sh\n+++ b/t/t1014-read-tree-confusing.sh\n@@ -30,13 +30,13 @@ while read path pretty; do\n \tesac\n \ttest_expect_success \"reject $pretty at end of path\" '\n \t\tprintf \"100644 blob %s\\t%s\" \"$blob\" \"$path\" >tree &&\n-\t\tbogus=$(git mktree <tree) &&\n+\t\tbogus=$(git mktree --literally <tree) &&\n \t\ttest_must_fail git read-tree $bogus\n \t'\n \n \ttest_expect_success \"reject $pretty as subtree\" '\n \t\tprintf \"040000 tree %s\\t%s\" \"$tree\" \"$path\" >tree &&\n-\t\tbogus=$(git mktree <tree) &&\n+\t\tbogus=$(git mktree --literally <tree) &&\n \t\ttest_must_fail git read-tree $bogus\n \t'\n done <<-EOF\n@@ -58,7 +58,7 @@ test_expect_success 'utf-8 paths allowed with core.protectHFS off' '\n \ttest_when_finished \"git read-tree HEAD\" &&\n \ttest_config core.protectHFS false &&\n \tprintf \"100644 blob %s\\t%s\" \"$blob\" \".gi${u200c}t\" >tree &&\n-\tok=$(git mktree <tree) &&\n+\tok=$(git mktree --literally <tree) &&\n \tgit read-tree $ok\n '\n \ndiff --git a/t/t1450-fsck.sh b/t/t1450-fsck.sh\nindex 8a456b1142d..532d2770e88 100755\n--- a/t/t1450-fsck.sh\n+++ b/t/t1450-fsck.sh\n@@ -316,7 +316,7 @@ check_duplicate_names () {\n \t\t\t*)  printf \"100644 blob %s\\t%s\\n\" $blob \"$name\" ;;\n \t\t\tesac\n \t\tdone >badtree &&\n-\t\tbadtree=$(git mktree <badtree) &&\n+\t\tbadtree=$(git mktree --literally <badtree) &&\n \t\ttest_must_fail git fsck 2>out &&\n \t\ttest_grep \"$badtree\" out &&\n \t\ttest_grep \"error in tree .*contains duplicate file entries\" out\n@@ -614,7 +614,7 @@ while read name path pretty; do\n \t\t\ttree=$(git rev-parse HEAD^{tree}) &&\n \t\t\tvalue=$(eval \"echo \\$$type\") &&\n \t\t\tprintf \"$mode $type %s\\t%s\" \"$value\" \"$path\" >bad &&\n-\t\t\tbad_tree=$(git mktree <bad) &&\n+\t\t\tbad_tree=$(git mktree --literally <bad) &&\n \t\t\tgit fsck 2>out &&\n \t\t\ttest_grep \"warning.*tree $bad_tree\" out\n \t\t)'\ndiff --git a/t/t1601-index-bogus.sh b/t/t1601-index-bogus.sh\nindex 4171f1e1410..54e8ae038b7 100755\n--- a/t/t1601-index-bogus.sh\n+++ b/t/t1601-index-bogus.sh\n@@ -4,7 +4,7 @@ test_description='test handling of bogus index entries'\n . ./test-lib.sh\n \n test_expect_success 'create tree with null sha1' '\n-\ttree=$(printf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" | git mktree)\n+\ttree=$(printf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" | git mktree --literally)\n '\n \n test_expect_success 'read-tree refuses to read null sha1' '\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex ac4a5b2734c..97b58aa3cca 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -478,12 +478,12 @@ test_expect_success 'writing split index with null sha1 does not write cache tre\n \tgit config splitIndex.maxPercentChange 0 &&\n \tgit commit -m \"commit\" &&\n \t{\n-\t\tgit ls-tree HEAD &&\n-\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\"\n+\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" &&\n+\t\tgit ls-tree HEAD\n \t} >broken-tree &&\n \techo \"add broken entry\" >msg &&\n \n-\ttree=$(git mktree <broken-tree) &&\n+\ttree=$(git mktree --literally <broken-tree) &&\n \ttest_tick &&\n \tcommit=$(git commit-tree $tree -p HEAD <msg) &&\n \tgit update-ref HEAD \"$commit\" &&\ndiff --git a/t/t7008-filter-branch-null-sha1.sh b/t/t7008-filter-branch-null-sha1.sh\nindex 93fbc92b8db..a1b4c295c01 100755\n--- a/t/t7008-filter-branch-null-sha1.sh\n+++ b/t/t7008-filter-branch-null-sha1.sh\n@@ -12,12 +12,12 @@ test_expect_success 'setup: base commits' '\n \n test_expect_success 'setup: a commit with a bogus null sha1 in the tree' '\n \t{\n-\t\tgit ls-tree HEAD &&\n-\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\"\n+\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" &&\n+\t\tgit ls-tree HEAD\n \t} >broken-tree &&\n \techo \"add broken entry\" >msg &&\n \n-\ttree=$(git mktree <broken-tree) &&\n+\ttree=$(git mktree --literally <broken-tree) &&\n \ttest_tick &&\n \tcommit=$(git commit-tree $tree -p HEAD <msg) &&\n \tgit update-ref HEAD \"$commit\"\ndiff --git a/t/t7417-submodule-path-url.sh b/t/t7417-submodule-path-url.sh\nindex dbbb3853dc0..5d3c98e99a7 100755\n--- a/t/t7417-submodule-path-url.sh\n+++ b/t/t7417-submodule-path-url.sh\n@@ -42,7 +42,7 @@ test_expect_success MINGW 'submodule paths disallows trailing spaces' '\n \ttree=$(git -C super write-tree) &&\n \tgit -C super ls-tree $tree >tree &&\n \tsed \"s/sub/sub /\" <tree >tree.new &&\n-\ttree=$(git -C super mktree <tree.new) &&\n+\ttree=$(git -C super mktree --literally <tree.new) &&\n \tcommit=$(echo with space | git -C super commit-tree $tree) &&\n \tgit -C super update-ref refs/heads/main $commit &&\n \ndiff --git a/t/t7450-bad-git-dotfiles.sh b/t/t7450-bad-git-dotfiles.sh\nindex 4a9c22c9e2b..de2d45d2244 100755\n--- a/t/t7450-bad-git-dotfiles.sh\n+++ b/t/t7450-bad-git-dotfiles.sh\n@@ -203,11 +203,11 @@ check_dotx_symlink () {\n \t\t\tcontent=$(git hash-object -w ../.gitmodules) &&\n \t\t\ttarget=$(printf \"$tricky\" | git hash-object -w --stdin) &&\n \t\t\t{\n-\t\t\t\tprintf \"100644 blob $content\\t$tricky\\n\" &&\n-\t\t\t\tprintf \"120000 blob $target\\t$path\\n\"\n+\t\t\t\tprintf \"120000 blob $target\\t$path\\n\" &&\n+\t\t\t\tprintf \"100644 blob $content\\t$tricky\\n\"\n \t\t\t} >bad-tree\n \t\t) &&\n-\t\ttree=$(git -C $dir mktree <$dir/bad-tree)\n+\t\ttree=$(git -C $dir mktree --literally <$dir/bad-tree)\n \t'\n \n \ttest_expect_success \"fsck detects symlinked $name ($type)\" '\n@@ -261,7 +261,7 @@ test_expect_success 'fsck detects non-blob .gitmodules' '\n \t\tcp ../.gitmodules subdir/file &&\n \t\tgit add subdir/file &&\n \t\tgit commit -m ok &&\n-\t\tgit ls-tree HEAD | sed s/subdir/.gitmodules/ | git mktree &&\n+\t\tgit ls-tree HEAD | sed s/subdir/.gitmodules/ | git mktree --literally &&\n \n \t\ttest_must_fail git fsck 2>output &&\n \t\ttest_grep gitmodulesBlob output\n-- \ngitgitgadget\n\n"},{"id":"496947","messageId":"4f9f77e693cfc4fbe72a2ae739bc7e236a3b82d3.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 09/16] mktree: validate paths more carefully","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:41Z","receivedAt":"2024-06-11T18:25:00Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUse 'verify_path' to validate the paths provided as tree entries, ensuring\nwe do not create entries with paths not allowed in trees (e.g., .git). Also,\nremove trailing slashes on directories before validating, allowing users to\nprovide 'folder-name/' as the path for a tree object entry.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c  | 20 +++++++++++++++++---\n t/t1010-mktree.sh | 33 +++++++++++++++++++++++++++++++++\n 2 files changed, 50 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 48019448c1f..29e9dc6ce69 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -8,6 +8,7 @@\n #include \"hex.h\"\n #include \"index-info.h\"\n #include \"quote.h\"\n+#include \"read-cache-ll.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n #include \"parse-options.h\"\n@@ -49,10 +50,23 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n-\tif (!literally && strchr(path, '/'))\n-\t\tdie(\"path %s contains slash\", path);\n \n-\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\tif (literally) {\n+\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\t} else {\n+\t\t/* Normalize and validate entry path */\n+\t\tif (S_ISDIR(mode)) {\n+\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n+\t\t\t\tlen--;\n+\t\t}\n+\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\n+\t\tif (!verify_path(ent->name, mode))\n+\t\t\tdie(_(\"invalid path '%s'\"), path);\n+\t\tif (strchr(ent->name, '/'))\n+\t\t\tdie(\"path %s contains slash\", path);\n+\t}\n+\n \tent->mode = mode;\n \tent->len = len;\n \toidcpy(&ent->oid, oid);\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex e0687cb529f..e0263cb2bf8 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -173,4 +173,37 @@ test_expect_success '--literally can create invalid trees' '\n \tgrep \"not properly sorted\" err\n '\n \n+test_expect_success 'mktree validates path' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:a/one)\" &&\n+\thead_oid=\"$(git rev-parse HEAD)\" &&\n+\n+\t# Valid: tree with or without trailing slash, blob without trailing slash\n+\t{\n+\t\tprintf \"040000 tree $tree_oid\\tfolder1/\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tfolder2\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfile.txt\\n\"\n+\t} | git mktree >actual &&\n+\n+\t# Invalid: blob with trailing slash\n+\tprintf \"100644 blob $blob_oid\\ttest/\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"invalid path ${SQ}test/${SQ}\" err &&\n+\n+\t# Invalid: dotdot\n+\tprintf \"040000 tree $tree_oid\\t../\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"invalid path ${SQ}../${SQ}\" err &&\n+\n+\t# Invalid: dot\n+\tprintf \"040000 tree $tree_oid\\t.\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"invalid path ${SQ}.${SQ}\" err &&\n+\n+\t# Invalid: .git\n+\tprintf \"040000 tree $tree_oid\\t.git/\" |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"invalid path ${SQ}.git/${SQ}\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496948","messageId":"b59a4ad8ab4b0e47373f811700eba59141fdc6c6.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 10/16] mktree: overwrite duplicate entries","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:42Z","receivedAt":"2024-06-11T18:25:00Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf multiple tree entries with the same name are provided as input to\n'mktree', only write the last one to the tree. Entries are considered\nduplicates if they have identical names (*not* considering mode); if a blob\nand a tree with the same name are provided, only the last one will be\nwritten to the tree. A tree with duplicate entries is invalid (per 'git\nfsck'), so that condition should be avoided wherever possible.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  8 ++++---\n builtin/mktree.c             | 45 ++++++++++++++++++++++++++++++++----\n t/t1010-mktree.sh            | 36 +++++++++++++++++++++++++++--\n 3 files changed, 80 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex fb07e40cef0..afbc846d077 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -43,9 +43,11 @@ OPTIONS\n INPUT FORMAT\n ------------\n Tree entries may be specified in any of the formats compatible with the\n-`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n-entries is normalized by `mktree` so pre-sorting the input by path is not\n-required.\n+`--index-info` option to linkgit:git-update-index[1].\n+\n+The order of the tree entries is normalized by `mktree` so pre-sorting the input\n+by path is not required. Multiple entries provided with the same path are\n+deduplicated, with only the last one specified added to the tree.\n \n GIT\n ---\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 29e9dc6ce69..e9e2134136f 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -15,6 +15,9 @@\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n+\t/* Internal */\n+\tsize_t order;\n+\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n@@ -72,15 +75,49 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \toidcpy(&ent->oid, oid);\n \n \t/* Append the update */\n+\tent->order = arr->nr;\n \ttree_entry_array_push(arr, ent);\n }\n \n-static int ent_compare(const void *a_, const void *b_)\n+static int ent_compare(const void *a_, const void *b_, void *ctx)\n {\n+\tint cmp;\n \tstruct tree_entry *a = *(struct tree_entry **)a_;\n \tstruct tree_entry *b = *(struct tree_entry **)b_;\n-\treturn base_name_compare(a->name, a->len, a->mode,\n-\t\t\t\t b->name, b->len, b->mode);\n+\tint ignore_mode = *((int *)ctx);\n+\n+\tif (ignore_mode)\n+\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n+\telse\n+\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n+\t\t\t\t\tb->name, b->len, b->mode);\n+\treturn cmp ? cmp : b->order - a->order;\n+}\n+\n+static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n+{\n+\tsize_t count = arr->nr;\n+\tstruct tree_entry *prev = NULL;\n+\n+\tint ignore_mode = 1;\n+\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n+\n+\tarr->nr = 0;\n+\tfor (size_t i = 0; i < count; i++) {\n+\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\tif (prev &&\n+\t\t    !name_compare(prev->name, prev->len,\n+\t\t\t\t  curr->name, curr->len)) {\n+\t\t\tFREE_AND_NULL(curr);\n+\t\t} else {\n+\t\t\tarr->entries[arr->nr++] = curr;\n+\t\t\tprev = curr;\n+\t\t}\n+\t}\n+\n+\t/* Sort again to order the entries for tree insertion */\n+\tignore_mode = 0;\n+\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n@@ -88,7 +125,7 @@ static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n \tstruct strbuf buf;\n \tsize_t size = 0;\n \n-\tQSORT(arr->entries, arr->nr, ent_compare);\n+\tsort_and_dedup_tree_entry_array(arr);\n \tfor (size_t i = 0; i < arr->nr; i++)\n \t\tsize += 32 + arr->entries[i]->len;\n \ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex e0263cb2bf8..956692347f0 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -6,11 +6,16 @@ TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n test_expect_success setup '\n-\tfor d in a a- a0\n+\tfor d in folder folder- folder0\n \tdo\n \t\tmkdir \"$d\" && echo \"$d/one\" >\"$d/one\" &&\n \t\tgit add \"$d\" || return 1\n \tdone &&\n+\tfor f in before folder.txt later\n+\tdo\n+\t\techo \"$f\" >\"$f\" &&\n+\t\tgit add \"$f\" || return 1\n+\tdone &&\n \techo zero >one &&\n \tgit update-index --add --info-only one &&\n \tgit write-tree --missing-ok >tree.missing &&\n@@ -175,7 +180,7 @@ test_expect_success '--literally can create invalid trees' '\n \n test_expect_success 'mktree validates path' '\n \ttree_oid=\"$(cat tree)\" &&\n-\tblob_oid=\"$(git rev-parse $tree_oid:a/one)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n \thead_oid=\"$(git rev-parse HEAD)\" &&\n \n \t# Valid: tree with or without trailing slash, blob without trailing slash\n@@ -206,4 +211,31 @@ test_expect_success 'mktree validates path' '\n \tgrep \"invalid path ${SQ}.git/${SQ}\" err\n '\n \n+test_expect_success 'mktree with duplicate entries' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\ttest\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest0\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest-\\n\"\n+\t} >top.dup &&\n+\tgit mktree <top.dup >tree.actual &&\n+\n+\t{\n+\t\tprintf \"160000 commit $head_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest0\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496949","messageId":"130413f2404bb27a2ede4fb00041227c90587e8e.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 11/16] mktree: create tree using an in-core index","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:43Z","receivedAt":"2024-06-11T18:25:01Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nRather than manually write out the contents of a tree object file, construct\nan in-memory sparse index from the provided tree entries and create the tree\nby writing out its corresponding cache tree.\n\nThis patch does not change the behavior of the 'mktree' command. However,\nconstructing the tree this way will substantially simplify future extensions\nto the command's functionality, including handling deeper-than-toplevel tree\nentries and applying the provided entries to an existing tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 74 +++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 55 insertions(+), 19 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex e9e2134136f..12f68187221 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -4,6 +4,7 @@\n  * Copyright (c) Junio C Hamano, 2006, 2009\n  */\n #include \"builtin.h\"\n+#include \"cache-tree.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n #include \"index-info.h\"\n@@ -24,6 +25,11 @@ struct tree_entry {\n \tchar name[FLEX_ARRAY];\n };\n \n+static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n+{\n+\treturn S_ISDIR(mode) ? pathlen - 1 : pathlen;\n+}\n+\n struct tree_entry_array {\n \tsize_t nr, alloc;\n \tstruct tree_entry **entries;\n@@ -57,17 +63,25 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \tif (literally) {\n \t\tFLEX_ALLOC_MEM(ent, name, path, len);\n \t} else {\n+\t\tsize_t len_to_copy = len;\n+\n \t\t/* Normalize and validate entry path */\n \t\tif (S_ISDIR(mode)) {\n-\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n-\t\t\t\tlen--;\n+\t\t\twhile(len_to_copy > 0 && is_dir_sep(path[len_to_copy - 1]))\n+\t\t\t\tlen_to_copy--;\n+\t\t\tlen = len_to_copy + 1; /* add space for trailing slash */\n \t\t}\n-\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\t\tent = xcalloc(1, st_add3(sizeof(struct tree_entry), len, 1));\n+\t\tmemcpy(ent->name, path, len_to_copy);\n \n \t\tif (!verify_path(ent->name, mode))\n \t\t\tdie(_(\"invalid path '%s'\"), path);\n \t\tif (strchr(ent->name, '/'))\n \t\t\tdie(\"path %s contains slash\", path);\n+\n+\t\t/* Add trailing slash to dir */\n+\t\tif (S_ISDIR(mode))\n+\t\t\tent->name[len - 1] = '/';\n \t}\n \n \tent->mode = mode;\n@@ -86,11 +100,14 @@ static int ent_compare(const void *a_, const void *b_, void *ctx)\n \tstruct tree_entry *b = *(struct tree_entry **)b_;\n \tint ignore_mode = *((int *)ctx);\n \n-\tif (ignore_mode)\n-\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n-\telse\n-\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n-\t\t\t\t\tb->name, b->len, b->mode);\n+\tsize_t a_len = a->len, b_len = b->len;\n+\n+\tif (ignore_mode) {\n+\t\ta_len = df_path_len(a_len, a->mode);\n+\t\tb_len = df_path_len(b_len, b->mode);\n+\t}\n+\n+\tcmp = name_compare(a->name, a_len, b->name, b_len);\n \treturn cmp ? cmp : b->order - a->order;\n }\n \n@@ -106,8 +123,8 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tfor (size_t i = 0; i < count; i++) {\n \t\tstruct tree_entry *curr = arr->entries[i];\n \t\tif (prev &&\n-\t\t    !name_compare(prev->name, prev->len,\n-\t\t\t\t  curr->name, curr->len)) {\n+\t\t    !name_compare(prev->name, df_path_len(prev->len, prev->mode),\n+\t\t\t\t  curr->name, df_path_len(curr->len, curr->mode))) {\n \t\t\tFREE_AND_NULL(curr);\n \t\t} else {\n \t\t\tarr->entries[arr->nr++] = curr;\n@@ -120,24 +137,43 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n+static int add_tree_entry_to_index(struct index_state *istate,\n+\t\t\t\t   struct tree_entry *ent)\n+{\n+\tstruct cache_entry *ce;\n+\tstruct strbuf ce_name = STRBUF_INIT;\n+\tstrbuf_add(&ce_name, ent->name, ent->len);\n+\n+\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n+\tif (!ce)\n+\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\n+\tadd_index_entry(istate, ce, ADD_CACHE_JUST_APPEND);\n+\tstrbuf_release(&ce_name);\n+\treturn 0;\n+}\n+\n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n-\tstruct strbuf buf;\n-\tsize_t size = 0;\n+\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n+\tistate.sparse_index = 1;\n \n \tsort_and_dedup_tree_entry_array(arr);\n-\tfor (size_t i = 0; i < arr->nr; i++)\n-\t\tsize += 32 + arr->entries[i]->len;\n \n-\tstrbuf_init(&buf, size);\n+\t/* Construct an in-memory index from the provided entries */\n \tfor (size_t i = 0; i < arr->nr; i++) {\n \t\tstruct tree_entry *ent = arr->entries[i];\n-\t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n-\t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n+\n+\t\tif (add_tree_entry_to_index(&istate, ent))\n+\t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n \t}\n \n-\twrite_object_file(buf.buf, buf.len, OBJ_TREE, oid);\n-\tstrbuf_release(&buf);\n+\t/* Write out new tree */\n+\tif (cache_tree_update(&istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n+\t\tdie(_(\"failed to write tree\"));\n+\toidcpy(oid, &istate.cache_tree->oid);\n+\n+\trelease_index(&istate);\n }\n \n static void write_tree_literally(struct tree_entry_array *arr,\n-- \ngitgitgadget\n\n"},{"id":"496950","messageId":"94d6615d634c4f78c88d3e01abbb27f13f85828c.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 12/16] mktree: use iterator struct to add tree entries to index","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:44Z","receivedAt":"2024-06-11T18:25:02Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nCreate 'struct tree_entry_iterator' to manage iteration through a 'struct\ntree_entry_array'. Using an iterator allows for conditional iteration; this\nfunctionality will be necessary in later commits when performing parallel\niteration through multiple sets of tree entries.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 40 +++++++++++++++++++++++++++++++++++++---\n 1 file changed, 37 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 12f68187221..bee359e9978 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -137,6 +137,38 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n+struct tree_entry_iterator {\n+\tstruct tree_entry *current;\n+\n+\t/* private */\n+\tstruct {\n+\t\tstruct tree_entry_array *arr;\n+\t\tsize_t idx;\n+\t} priv;\n+};\n+\n+static void init_tree_entry_iterator(struct tree_entry_iterator *iter,\n+\t\t\t\t     struct tree_entry_array *arr)\n+{\n+\titer->priv.arr = arr;\n+\titer->priv.idx = 0;\n+\titer->current = 0 < arr->nr ? arr->entries[0] : NULL;\n+}\n+\n+/*\n+ * Advance the tree entry iterator to the next entry in the array. If no entries\n+ * remain, 'current' is set to NULL. Returns the previous 'current' value of the\n+ * iterator.\n+ */\n+static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator *iter)\n+{\n+\tstruct tree_entry *prev = iter->current;\n+\titer->current = (iter->priv.idx + 1) < iter->priv.arr->nr\n+\t\t\t? iter->priv.arr->entries[++iter->priv.idx]\n+\t\t\t: NULL;\n+\treturn prev;\n+}\n+\n static int add_tree_entry_to_index(struct index_state *istate,\n \t\t\t\t   struct tree_entry *ent)\n {\n@@ -155,15 +187,17 @@ static int add_tree_entry_to_index(struct index_state *istate,\n \n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n+\tstruct tree_entry_iterator iter = { NULL };\n+\tstruct tree_entry *ent;\n \tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n \tistate.sparse_index = 1;\n \n \tsort_and_dedup_tree_entry_array(arr);\n \n-\t/* Construct an in-memory index from the provided entries */\n-\tfor (size_t i = 0; i < arr->nr; i++) {\n-\t\tstruct tree_entry *ent = arr->entries[i];\n+\tinit_tree_entry_iterator(&iter, arr);\n \n+\t/* Construct an in-memory index from the provided entries & base tree */\n+\twhile ((ent = advance_tree_entry_iterator(&iter))) {\n \t\tif (add_tree_entry_to_index(&istate, ent))\n \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"496951","messageId":"68acdd3c5ee266fd0c1d3ae45f69af4ef9012e07.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 13/16] mktree: add directory-file conflict hashmap","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:45Z","receivedAt":"2024-06-11T18:25:03Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nCreate a hashmap member of a 'struct tree_entry_array' that contains all of\nthe (de-duplicated) provided tree entries, indexed by the hash of their path\nwith *no* trailing slash. This hashmap will be used in a later commit to\navoid adding a file to an existing tree that has the same path as a\ndirectory, or vice versa.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 39 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 39 insertions(+)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex bee359e9978..09b3c5c6244 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -16,6 +16,8 @@\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n+\tstruct hashmap_entry ent;\n+\n \t/* Internal */\n \tsize_t order;\n \n@@ -33,8 +35,33 @@ static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n struct tree_entry_array {\n \tsize_t nr, alloc;\n \tstruct tree_entry **entries;\n+\n+\tstruct hashmap df_name_hash;\n };\n \n+static int df_name_hash_cmp(const void *cmp_data UNUSED,\n+\t\t\t    const struct hashmap_entry *eptr,\n+\t\t\t    const struct hashmap_entry *entry_or_key,\n+\t\t\t    const void *keydata UNUSED)\n+{\n+\tconst struct tree_entry *e1, *e2;\n+\tsize_t e1_len, e2_len;\n+\n+\te1 = container_of(eptr, const struct tree_entry, ent);\n+\te2 = container_of(entry_or_key, const struct tree_entry, ent);\n+\n+\te1_len = df_path_len(e1->len, e1->mode);\n+\te2_len = df_path_len(e2->len, e2->mode);\n+\n+\treturn e1_len != e2_len ||\n+\t       name_compare(e1->name, e1_len, e2->name, e2_len);\n+}\n+\n+static void init_tree_entry_array(struct tree_entry_array *arr)\n+{\n+\thashmap_init(&arr->df_name_hash, df_name_hash_cmp, NULL, 0);\n+}\n+\n static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n {\n \tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n@@ -43,6 +70,7 @@ static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entr\n \n static void clear_tree_entry_array(struct tree_entry_array *arr)\n {\n+\thashmap_clear(&arr->df_name_hash);\n \tfor (size_t i = 0; i < arr->nr; i++)\n \t\tFREE_AND_NULL(arr->entries[i]);\n \tarr->nr = 0;\n@@ -50,6 +78,7 @@ static void clear_tree_entry_array(struct tree_entry_array *arr)\n \n static void release_tree_entry_array(struct tree_entry_array *arr)\n {\n+\thashmap_clear(&arr->df_name_hash);\n \tFREE_AND_NULL(arr->entries);\n \tarr->nr = arr->alloc = 0;\n }\n@@ -135,6 +164,14 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \t/* Sort again to order the entries for tree insertion */\n \tignore_mode = 0;\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n+\n+\t/* Finally, initialize the directory-file conflict hash map */\n+\tfor (size_t i = 0; i < count; i++) {\n+\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\thashmap_entry_init(&curr->ent,\n+\t\t\t\t   memhash(curr->name, df_path_len(curr->len, curr->mode)));\n+\t\thashmap_put(&arr->df_name_hash, &curr->ent);\n+\t}\n }\n \n struct tree_entry_iterator {\n@@ -302,6 +339,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n \n+\tinit_tree_entry_array(&arr);\n+\n \tdo {\n \t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data);\n \t\tif (ret < 0)\n-- \ngitgitgadget\n\n"},{"id":"496952","messageId":"df0c50dfea3cb77e0070246efdf7a3f070b2ad97.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 14/16] mktree: optionally add to an existing tree","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:46Z","receivedAt":"2024-06-11T18:25:04Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nAllow users to specify a single \"tree-ish\" value as a positional argument.\nIf provided, the contents of the given tree serve as the basis for the new\ntree (or trees, in --batch mode) created by 'mktree', on top of which all of\nthe stdin-provided tree entries are applied.\n\nAt a high level, the entries are \"applied\" to a base tree by iterating\nthrough the base tree using 'read_tree' in parallel with iterating through\nthe sorted & deduplicated stdin entries via their iterator. That is, for\neach call to the 'build_index_from_tree callback of 'read_tree':\n\n* If the iterator entry precedes the base tree entry, add it to the in-core\n  index, increment the iterator, and repeat.\n* If the iterator entry has the same name as the base tree entry, add the\n  iterator entry to the index, increment the iterator, and return from the\n  callback to continue the 'read_tree' iteration.\n* If the iterator entry follows the base tree entry, first check\n  'df_name_hash' to ensure we won't be adding an entry with the same name\n  later (with a different mode). If there's no directory/file conflict, add\n  the base tree entry to the index. In either case, return from the callback\n  to continue the 'read_tree' iteration.\n\nFinally, once 'read_tree' is complete, add the remaining entries in the\niterator to the index and write out the index as a tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |   7 +-\n builtin/mktree.c             | 134 ++++++++++++++++++++++++++++++-----\n t/t1010-mktree.sh            |  36 ++++++++++\n 3 files changed, 157 insertions(+), 20 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex afbc846d077..99abd3c31a6 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -9,7 +9,7 @@ git-mktree - Build a tree-object from formatted tree entries\n SYNOPSIS\n --------\n [verse]\n-'git mktree' [-z] [--missing] [--literally] [--batch]\n+'git mktree' [-z] [--missing] [--literally] [--batch] [--] [<tree-ish>]\n \n DESCRIPTION\n -----------\n@@ -40,6 +40,11 @@ OPTIONS\n \toptional.  Note - if the `-z` option is used, lines are terminated\n \twith NUL.\n \n+<tree-ish>::\n+\tIf provided, the tree entries provided in stdin are added to this tree\n+\trather than a new empty one, replacing existing entries with identical\n+\tnames. Not compatible with `--literally`.\n+\n INPUT FORMAT\n ------------\n Tree entries may be specified in any of the formats compatible with the\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 09b3c5c6244..9e9d2554cad 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,7 +12,9 @@\n #include \"read-cache-ll.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n+#include \"object-name.h\"\n #include \"parse-options.h\"\n+#include \"pathspec.h\"\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n@@ -206,45 +208,122 @@ static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator\n \treturn prev;\n }\n \n-static int add_tree_entry_to_index(struct index_state *istate,\n+struct build_index_data {\n+\tstruct tree_entry_iterator iter;\n+\tstruct hashmap *df_name_hash;\n+\tstruct index_state istate;\n+};\n+\n+static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n \tstruct cache_entry *ce;\n-\tstruct strbuf ce_name = STRBUF_INIT;\n-\tstrbuf_add(&ce_name, ent->name, ent->len);\n-\n-\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n+\tce = make_cache_entry(&data->istate, ent->mode, &ent->oid, ent->name, 0, 0);\n \tif (!ce)\n \t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n \n-\tadd_index_entry(istate, ce, ADD_CACHE_JUST_APPEND);\n-\tstrbuf_release(&ce_name);\n+\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n \treturn 0;\n }\n \n-static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n+static int build_index_from_tree(const struct object_id *oid,\n+\t\t\t\t struct strbuf *base, const char *filename,\n+\t\t\t\t unsigned mode, void *context)\n {\n-\tstruct tree_entry_iterator iter = { NULL };\n+\tint result;\n+\tstruct tree_entry *base_tree_ent;\n+\tstruct build_index_data *cbdata = context;\n+\tsize_t filename_len = strlen(filename);\n+\tsize_t path_len = S_ISDIR(mode) ? st_add3(filename_len, base->len, 1)\n+\t\t\t\t\t: st_add(filename_len, base->len);\n+\n+\t/* Create a tree entry from the current entry in read_tree iteration */\n+\tbase_tree_ent = xcalloc(1, st_add3(sizeof(struct tree_entry), path_len, 1));\n+\tbase_tree_ent->len = path_len;\n+\tbase_tree_ent->mode = mode;\n+\toidcpy(&base_tree_ent->oid, oid);\n+\n+\tmemcpy(base_tree_ent->name, base->buf, base->len);\n+\tmemcpy(base_tree_ent->name + base->len, filename, filename_len);\n+\tif (S_ISDIR(mode))\n+\t\tbase_tree_ent->name[base_tree_ent->len - 1] = '/';\n+\n+\twhile (cbdata->iter.current) {\n+\t\tstruct tree_entry *ent = cbdata->iter.current;\n+\n+\t\tint cmp = name_compare(ent->name, ent->len,\n+\t\t\t\t       base_tree_ent->name, base_tree_ent->len);\n+\t\tif (!cmp || cmp < 0) {\n+\t\t\tadvance_tree_entry_iterator(&cbdata->iter);\n+\n+\t\t\tif (add_tree_entry_to_index(cbdata, ent) < 0) {\n+\t\t\t\tresult = error(_(\"failed to add tree entry '%s'\"), ent->name);\n+\t\t\t\tgoto cleanup_and_return;\n+\t\t\t}\n+\n+\t\t\tif (!cmp) {\n+\t\t\t\tresult = 0;\n+\t\t\t\tgoto cleanup_and_return;\n+\t\t\t} else\n+\t\t\t\tcontinue;\n+\t\t}\n+\n+\t\tbreak;\n+\t}\n+\n+\t/*\n+\t * If the tree entry should be replaced with an entry with the same name\n+\t * (but different mode), skip it.\n+\t */\n+\thashmap_entry_init(&base_tree_ent->ent,\n+\t\t\t   memhash(base_tree_ent->name, df_path_len(base_tree_ent->len, base_tree_ent->mode)));\n+\tif (hashmap_get_entry(cbdata->df_name_hash, base_tree_ent, ent, NULL)) {\n+\t\tresult = 0;\n+\t\tgoto cleanup_and_return;\n+\t}\n+\n+\tif (add_tree_entry_to_index(cbdata, base_tree_ent)) {\n+\t\tresult = -1;\n+\t\tgoto cleanup_and_return;\n+\t}\n+\n+\tresult = 0;\n+\n+cleanup_and_return:\n+\tFREE_AND_NULL(base_tree_ent);\n+\treturn result;\n+}\n+\n+static void write_tree(struct tree_entry_array *arr, struct tree *base_tree,\n+\t\t       struct object_id *oid)\n+{\n+\tstruct build_index_data cbdata = { 0 };\n \tstruct tree_entry *ent;\n-\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n-\tistate.sparse_index = 1;\n+\tstruct pathspec ps = { 0 };\n \n \tsort_and_dedup_tree_entry_array(arr);\n \n-\tinit_tree_entry_iterator(&iter, arr);\n+\tindex_state_init(&cbdata.istate, the_repository);\n+\tcbdata.istate.sparse_index = 1;\n+\tinit_tree_entry_iterator(&cbdata.iter, arr);\n+\tcbdata.df_name_hash = &arr->df_name_hash;\n \n \t/* Construct an in-memory index from the provided entries & base tree */\n-\twhile ((ent = advance_tree_entry_iterator(&iter))) {\n-\t\tif (add_tree_entry_to_index(&istate, ent))\n+\tif (base_tree &&\n+\t    read_tree(the_repository, base_tree, &ps, build_index_from_tree, &cbdata) < 0)\n+\t\tdie(_(\"failed to create tree\"));\n+\n+\twhile ((ent = advance_tree_entry_iterator(&cbdata.iter))) {\n+\t\tif (add_tree_entry_to_index(&cbdata, ent))\n \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n \t}\n \n \t/* Write out new tree */\n-\tif (cache_tree_update(&istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n+\tif (cache_tree_update(&cbdata.istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n \t\tdie(_(\"failed to write tree\"));\n-\toidcpy(oid, &istate.cache_tree->oid);\n+\toidcpy(oid, &cbdata.istate.cache_tree->oid);\n \n-\trelease_index(&istate);\n+\trelease_index(&cbdata.istate);\n }\n \n static void write_tree_literally(struct tree_entry_array *arr,\n@@ -268,7 +347,7 @@ static void write_tree_literally(struct tree_entry_array *arr,\n }\n \n static const char *mktree_usage[] = {\n-\t\"git mktree [-z] [--missing] [--literally] [--batch]\",\n+\t\"git mktree [-z] [--missing] [--literally] [--batch] [--] [<tree-ish>]\",\n \tNULL\n };\n \n@@ -326,6 +405,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tint is_batch_mode = 0;\n \tstruct tree_entry_array arr = { 0 };\n \tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n+\tstruct tree *base_tree = NULL;\n \tint ret;\n \n \tconst struct option option[] = {\n@@ -338,6 +418,22 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t};\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n+\tif (ac > 1)\n+\t\tusage_with_options(mktree_usage, option);\n+\n+\tif (ac) {\n+\t\tstruct object_id base_tree_oid;\n+\n+\t\tif (mktree_line_data.literally)\n+\t\t\tdie(_(\"option '%s' and tree-ish cannot be used together\"), \"--literally\");\n+\n+\t\tif (repo_get_oid(the_repository, av[0], &base_tree_oid))\n+\t\t\tdie(_(\"not a valid object name %s\"), av[0]);\n+\n+\t\tbase_tree = parse_tree_indirect(&base_tree_oid);\n+\t\tif (!base_tree)\n+\t\t\tdie(_(\"not a tree object: %s\"), oid_to_hex(&base_tree_oid));\n+\t}\n \n \tinit_tree_entry_array(&arr);\n \n@@ -361,7 +457,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\tif (mktree_line_data.literally)\n \t\t\t\twrite_tree_literally(&arr, &oid);\n \t\t\telse\n-\t\t\t\twrite_tree(&arr, &oid);\n+\t\t\t\twrite_tree(&arr, base_tree, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 956692347f0..ea5a011405e 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -238,4 +238,40 @@ test_expect_success 'mktree with duplicate entries' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree with base tree' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest0\\n\"\n+\t} >top.base &&\n+\tgit mktree <top.base >tree.base &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tz\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.xyz\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ta\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest\\n\"\n+\t} >top.append &&\n+\tgit mktree $(cat tree.base) <top.append >tree.actual &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\ta\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.xyz\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest0\\n\" &&\n+\t\tprintf \"100755 blob $before_oid\\tz\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496953","messageId":"058354f45f7b837ebeb08337a8dfd6e0ec1e9d1b.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 15/16] mktree: allow deeper paths in input","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:47Z","receivedAt":"2024-06-11T18:25:05Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUpdate 'git mktree' to handle entries nested inside of directories (e.g.\n'path/to/a/file.txt'). This functionality requires a series of changes:\n\n* In 'sort_and_dedup_tree_entry_array()', remove entries inside of\n  directories that come after them in input order.\n* Also in 'sort_and_dedup_tree_entry_array()', mark directories that contain\n  entries that come after them in input order (e.g., 'folder/' followed by\n  'folder/file.txt') as \"need to expand\".\n* In 'add_tree_entry_to_index()', if a tree entry is marked as \"need to\n  expand\", recurse into it with 'read_tree_at()' & 'build_index_from_tree'.\n* In 'build_index_from_tree()', if a user-specified tree entry is contained\n  within the current iterated entry, return 'READ_TREE_RECURSIVE' to recurse\n  into the iterated tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |   5 ++\n builtin/mktree.c             | 101 ++++++++++++++++++++++++++++++---\n t/t1010-mktree.sh            | 107 +++++++++++++++++++++++++++++++++--\n 3 files changed, 200 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 99abd3c31a6..db90fdcdc8f 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -50,6 +50,11 @@ INPUT FORMAT\n Tree entries may be specified in any of the formats compatible with the\n `--index-info` option to linkgit:git-update-index[1].\n \n+Entries may use full pathnames containing directory separators to specify\n+entries nested within one or more directories. These entries are inserted into\n+the appropriate tree in the base tree-ish if one exists. Otherwise, empty parent\n+trees are created to contain the entries.\n+\n The order of the tree entries is normalized by `mktree` so pre-sorting the input\n by path is not required. Multiple entries provided with the same path are\n deduplicated, with only the last one specified added to the tree.\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 9e9d2554cad..00b77869a56 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -22,6 +22,7 @@ struct tree_entry {\n \n \t/* Internal */\n \tsize_t order;\n+\tint expand_dir;\n \n \tunsigned mode;\n \tstruct object_id oid;\n@@ -39,6 +40,7 @@ struct tree_entry_array {\n \tstruct tree_entry **entries;\n \n \tstruct hashmap df_name_hash;\n+\tint has_nested_entries;\n };\n \n static int df_name_hash_cmp(const void *cmp_data UNUSED,\n@@ -70,6 +72,13 @@ static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entr\n \tarr->entries[arr->nr++] = ent;\n }\n \n+static struct tree_entry *tree_entry_array_pop(struct tree_entry_array *arr)\n+{\n+\tif (!arr->nr)\n+\t\treturn NULL;\n+\treturn arr->entries[--arr->nr];\n+}\n+\n static void clear_tree_entry_array(struct tree_entry_array *arr)\n {\n \thashmap_clear(&arr->df_name_hash);\n@@ -107,8 +116,10 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \n \t\tif (!verify_path(ent->name, mode))\n \t\t\tdie(_(\"invalid path '%s'\"), path);\n-\t\tif (strchr(ent->name, '/'))\n-\t\t\tdie(\"path %s contains slash\", path);\n+\n+\t\t/* mark has_nested_entries if needed */\n+\t\tif (!arr->has_nested_entries && strchr(ent->name, '/'))\n+\t\t\tarr->has_nested_entries = 1;\n \n \t\t/* Add trailing slash to dir */\n \t\tif (S_ISDIR(mode))\n@@ -167,6 +178,46 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tignore_mode = 0;\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n \n+\tif (arr->has_nested_entries) {\n+\t\tstruct tree_entry_array parent_dir_ents = { 0 };\n+\n+\t\tcount = arr->nr;\n+\t\tarr->nr = 0;\n+\n+\t\t/* Remove any entries where one of its parent dirs has a higher 'order' */\n+\t\tfor (size_t i = 0; i < count; i++) {\n+\t\t\tconst char *skipped_prefix;\n+\t\t\tstruct tree_entry *parent;\n+\t\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\t\tint skip_entry = 0;\n+\n+\t\t\twhile ((parent = tree_entry_array_pop(&parent_dir_ents))) {\n+\t\t\t\tif (!skip_prefix(curr->name, parent->name, &skipped_prefix))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\t/* entry in dir, so we push the parent back onto the stack */\n+\t\t\t\ttree_entry_array_push(&parent_dir_ents, parent);\n+\n+\t\t\t\tif (parent->order > curr->order)\n+\t\t\t\t\tskip_entry = 1;\n+\t\t\t\telse\n+\t\t\t\t\tparent->expand_dir = 1;\n+\n+\t\t\t\tbreak;\n+\t\t\t}\n+\n+\t\t\tif (!skip_entry) {\n+\t\t\t\tarr->entries[arr->nr++] = curr;\n+\t\t\t\tif (S_ISDIR(curr->mode))\n+\t\t\t\t\ttree_entry_array_push(&parent_dir_ents, curr);\n+\t\t\t} else {\n+\t\t\t\tFREE_AND_NULL(curr);\n+\t\t\t}\n+\t\t}\n+\n+\t\trelease_tree_entry_array(&parent_dir_ents);\n+\t}\n+\n \t/* Finally, initialize the directory-file conflict hash map */\n \tfor (size_t i = 0; i < count; i++) {\n \t\tstruct tree_entry *curr = arr->entries[i];\n@@ -214,15 +265,40 @@ struct build_index_data {\n \tstruct index_state istate;\n };\n \n+static int build_index_from_tree(const struct object_id *oid,\n+\t\t\t\t struct strbuf *base, const char *filename,\n+\t\t\t\t unsigned mode, void *context);\n+\n static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n-\tstruct cache_entry *ce;\n-\tce = make_cache_entry(&data->istate, ent->mode, &ent->oid, ent->name, 0, 0);\n-\tif (!ce)\n-\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\tif (ent->expand_dir) {\n+\t\tint ret = 0;\n+\t\tstruct pathspec ps = { 0 };\n+\t\tstruct tree *subtree = parse_tree_indirect(&ent->oid);\n+\t\tstruct strbuf base_path = STRBUF_INIT;\n+\t\tstrbuf_add(&base_path, ent->name, ent->len);\n+\n+\t\tif (!subtree)\n+\t\t\tret = error(_(\"not a tree object: %s\"), oid_to_hex(&ent->oid));\n+\t\telse if (read_tree_at(the_repository, subtree, &base_path, 0, &ps,\n+\t\t\t\t build_index_from_tree, data) < 0)\n+\t\t\tret = -1;\n+\n+\t\tstrbuf_release(&base_path);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\n+\t} else {\n+\t\tstruct cache_entry *ce = make_cache_entry(&data->istate,\n+\t\t\t\t\t\t\t  ent->mode, &ent->oid,\n+\t\t\t\t\t\t\t  ent->name, 0, 0);\n+\t\tif (!ce)\n+\t\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\n+\t\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n+\t}\n \n-\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n \treturn 0;\n }\n \n@@ -249,10 +325,12 @@ static int build_index_from_tree(const struct object_id *oid,\n \t\tbase_tree_ent->name[base_tree_ent->len - 1] = '/';\n \n \twhile (cbdata->iter.current) {\n+\t\tconst char *skipped_prefix;\n \t\tstruct tree_entry *ent = cbdata->iter.current;\n+\t\tint cmp;\n \n-\t\tint cmp = name_compare(ent->name, ent->len,\n-\t\t\t\t       base_tree_ent->name, base_tree_ent->len);\n+\t\tcmp = name_compare(ent->name, ent->len,\n+\t\t\t\t   base_tree_ent->name, base_tree_ent->len);\n \t\tif (!cmp || cmp < 0) {\n \t\t\tadvance_tree_entry_iterator(&cbdata->iter);\n \n@@ -266,6 +344,11 @@ static int build_index_from_tree(const struct object_id *oid,\n \t\t\t\tgoto cleanup_and_return;\n \t\t\t} else\n \t\t\t\tcontinue;\n+\t\t} else if (skip_prefix(ent->name, base_tree_ent->name, &skipped_prefix) &&\n+\t\t\t   S_ISDIR(base_tree_ent->mode)) {\n+\t\t\t/* The entry is in the current traversed tree entry, so we recurse */\n+\t\t\tresult = READ_TREE_RECURSIVE;\n+\t\t\tgoto cleanup_and_return;\n \t\t}\n \n \t\tbreak;\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex ea5a011405e..1d6365141fc 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -89,12 +89,21 @@ test_expect_success 'mktree with invalid submodule OIDs' '\n \tgrep \"object $tree_oid is a tree but specified type was (commit)\" err\n '\n \n-test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n-\ttest_must_fail git mktree <all\n+test_expect_success 'mktree reads ls-tree -r output (1)' '\n+\tgit mktree <all >actual &&\n+\ttest_cmp tree actual\n '\n \n-test_expect_success 'mktree refuses to read ls-tree -r output (2)' '\n-\ttest_must_fail git mktree <all.withsub\n+test_expect_success 'mktree reads ls-tree -r output (2)' '\n+\tgit mktree <all.withsub >actual &&\n+\ttest_cmp tree.withsub actual\n+'\n+\n+test_expect_success 'mktree de-duplicates files inside directories' '\n+\tgit ls-tree $(cat tree) >everything &&\n+\tcat <all >top_and_all &&\n+\tgit mktree <top_and_all >actual &&\n+\ttest_cmp tree actual\n '\n \n test_expect_success 'mktree fails on malformed input' '\n@@ -238,6 +247,50 @@ test_expect_success 'mktree with duplicate entries' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree adds entry after nested entry' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tone_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tearly\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tearly/one\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tlater\\n\" &&\n+\t\tprintf \"040000 tree $EMPTY_TREE\\tnew-tree\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tnew-tree/one\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tzzz\\n\"\n+\t} >top.rec &&\n+\tgit mktree <top.rec >tree.actual &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tearly\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tlater\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tnew-tree\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tzzz\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'mktree inserts entries into directories' '\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tone_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\tblob_oid=$(git rev-parse ${tree_oid}:before) &&\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tfolder\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/two\\n\"\n+\t} | git mktree >actual &&\n+\n+\t{\n+\t\tprintf \"100644 blob $one_oid\\tfolder/one\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/two\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'mktree with base tree' '\n \ttree_oid=$(cat tree) &&\n \tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n@@ -274,4 +327,50 @@ test_expect_success 'mktree with base tree' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree with base tree (deep)' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\tfolder_one_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tfolder/before\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\tfolder/one.txt\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\tfolder/sub\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tfolder/one\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tfolder/one/deeper\\n\"\n+\t} >top.append &&\n+\tgit mktree <top.append $(cat tree) >tree.actual &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tfolder/before\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\tfolder/one.txt\\n\" &&\n+\t\tprintf \"100644 blob $folder_one_oid\\tfolder/one/deeper/one\\n\" &&\n+\t\tprintf \"100644 blob $folder_one_oid\\tfolder/one/one\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\tfolder/sub\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat tree.actual) -- folder/ >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'mktree fails on directory-file conflict' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest/deeper\\n\"\n+\t} |\n+\ttest_must_fail git mktree 2>err &&\n+\tgrep \"You have both test and test/deeper\" err &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/one/deeper/deep\\n\"\n+\t} |\n+\ttest_must_fail git mktree $tree_oid 2>err &&\n+\tgrep \"You have both folder/one and folder/one/deeper/deep\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"496954","messageId":"a90d6d0c943283e9e7bd181cd6e9bb6d4572aaeb.1718130288.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH 16/16] mktree: remove entries when mode is 0","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-11T18:24:48Z","receivedAt":"2024-06-11T18:25:05Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf tree entries are specified with a mode with value '0', remove them from\nthe tree instead of adding/updating them. If the mode is '0', both the\nprovided type string (if specified) and the object ID of the entry are\nignored.\n\nNote that entries with mode '0' are added to the 'struct tree_ent_array'\nwith a trailing slash so that it's always treated like a directory. This is\na bit of a hack to ensure that the removal supercedes any preceding entries\nwith matching names, as well as any nested inside a directory matching its\nname.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  4 +++\n builtin/mktree.c             | 64 ++++++++++++++++++++----------------\n t/t1010-mktree.sh            | 38 +++++++++++++++++++++\n 3 files changed, 77 insertions(+), 29 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex db90fdcdc8f..a660438c67f 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -55,6 +55,10 @@ entries nested within one or more directories. These entries are inserted into\n the appropriate tree in the base tree-ish if one exists. Otherwise, empty parent\n trees are created to contain the entries.\n \n+An entry with a mode of \"0\" will remove an entry of the same name from the base\n+tree-ish. If no tree-ish argument is given, or the entry does not exist in that\n+tree, the entry is ignored.\n+\n The order of the tree entries is normalized by `mktree` so pre-sorting the input\n by path is not required. Multiple entries provided with the same path are\n deduplicated, with only the last one specified added to the tree.\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 00b77869a56..e94c9ca7e87 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -32,7 +32,7 @@ struct tree_entry {\n \n static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n {\n-\treturn S_ISDIR(mode) ? pathlen - 1 : pathlen;\n+\treturn (S_ISDIR(mode) || !mode) ? pathlen - 1 : pathlen;\n }\n \n struct tree_entry_array {\n@@ -106,7 +106,7 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \t\tsize_t len_to_copy = len;\n \n \t\t/* Normalize and validate entry path */\n-\t\tif (S_ISDIR(mode)) {\n+\t\tif (S_ISDIR(mode) || !mode) {\n \t\t\twhile(len_to_copy > 0 && is_dir_sep(path[len_to_copy - 1]))\n \t\t\t\tlen_to_copy--;\n \t\t\tlen = len_to_copy + 1; /* add space for trailing slash */\n@@ -122,7 +122,7 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \t\t\tarr->has_nested_entries = 1;\n \n \t\t/* Add trailing slash to dir */\n-\t\tif (S_ISDIR(mode))\n+\t\tif (S_ISDIR(mode) || !mode)\n \t\t\tent->name[len - 1] = '/';\n \t}\n \n@@ -208,7 +208,7 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \n \t\t\tif (!skip_entry) {\n \t\t\t\tarr->entries[arr->nr++] = curr;\n-\t\t\t\tif (S_ISDIR(curr->mode))\n+\t\t\t\tif (S_ISDIR(curr->mode) || !curr->mode)\n \t\t\t\t\ttree_entry_array_push(&parent_dir_ents, curr);\n \t\t\t} else {\n \t\t\t\tFREE_AND_NULL(curr);\n@@ -272,6 +272,9 @@ static int build_index_from_tree(const struct object_id *oid,\n static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n+\tif (!ent->mode)\n+\t\treturn 0;\n+\n \tif (ent->expand_dir) {\n \t\tint ret = 0;\n \t\tstruct pathspec ps = { 0 };\n@@ -445,36 +448,39 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \t\t       const char *path, void *cbdata)\n {\n \tstruct mktree_line_data *data = cbdata;\n-\tenum object_type mode_type = object_type(mode);\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tenum object_type parsed_obj_type;\n \n-\tif (obj_type && mode_type != obj_type)\n-\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n-\t\t    type_name(obj_type), type_name(mode_type));\n+\tif (mode) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tenum object_type parsed_obj_type;\n+\t\tenum object_type mode_type = object_type(mode);\n \n-\toi.typep = &parsed_obj_type;\n+\t\tif (obj_type && mode_type != obj_type)\n+\t\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n+\t\t\t    type_name(obj_type), type_name(mode_type));\n \n-\tif (oid_object_info_extended(the_repository, oid, &oi,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n-\t\t\t\t     OBJECT_INFO_QUICK |\n-\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n-\t\tparsed_obj_type = -1;\n+\t\toi.typep = &parsed_obj_type;\n \n-\tif (parsed_obj_type < 0) {\n-\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n-\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n-\t\t} else {\n-\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n+\t\tif (oid_object_info_extended(the_repository, oid, &oi,\n+\t\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n+\t\t\t\t\t     OBJECT_INFO_QUICK |\n+\t\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n+\t\t\tparsed_obj_type = -1;\n+\n+\t\tif (parsed_obj_type < 0) {\n+\t\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n+\t\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n+\t\t\t} else {\n+\t\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n+\t\t\t}\n+\t\t} else if (parsed_obj_type != mode_type) {\n+\t\t\t/*\n+\t\t\t* The object exists but is of the wrong type.\n+\t\t\t* This is a problem regardless of allow_missing\n+\t\t\t* because the new tree entry will never be correct.\n+\t\t\t*/\n+\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n+\t\t\tpath, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n \t\t}\n-\t} else if (parsed_obj_type != mode_type) {\n-\t\t/*\n-\t\t * The object exists but is of the wrong type.\n-\t\t * This is a problem regardless of allow_missing\n-\t\t * because the new tree entry will never be correct.\n-\t\t */\n-\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n-\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n \t}\n \n \tappend_to_tree(mode, oid, path, data->arr, data->literally);\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 1d6365141fc..7cb88e32d4f 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -373,4 +373,42 @@ test_expect_success 'mktree fails on directory-file conflict' '\n \tgrep \"You have both folder/one and folder/one/deeper/deep\" err\n '\n \n+test_expect_success 'mktree with remove entries' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\ttest/deeper/deep.txt\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.a/file\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.txt\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tfolder\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\tfolder\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\tmissing\\n\"\n+\t} | git mktree >tree.base &&\n+\n+\t{\n+\t\tprintf \"0 $ZERO_OID\\texample.txt\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\ttest/deeper\\n\"\n+\t} | git mktree $(cat tree.base) >tree.actual &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\texample\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.a/file\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest.txt\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'type and oid not checked if entry mode is 0' '\n+\t# type and oid do not match\n+\tprintf \"0 commit $EMPTY_TREE\\tfolder.txt\\n\" |\n+\tgit mktree >tree.actual &&\n+\n+\ttest \"$(cat tree.actual)\" = $EMPTY_TREE\n+'\n+\n test_done\n-- \ngitgitgadget\n"},{"id":"496955","messageId":"CAPig+cSg_SgU_ETktkz7KapLmzdDjzvjEx-_jS8L-eguBGk4oQ@mail.gmail.com","threadId":"61624","inReplyTo":"5ade145352f44b431c16a2ec29cd87de489e8032.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 03/16] mktree: use non-static tree_entry array","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2024-06-11T18:45:08Z","receivedAt":"2024-06-11T18:45:20Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Tue, Jun 11, 2024 at 2:25 PM Victoria Dye via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n> Replace the static 'struct tree_entry **entries' with a non-static 'struct\n> tree_entry_array' instance. In later commits, we'll want to be able to\n> create additional 'struct tree_entry_array' instances utilizing common\n> functionality (create, push, clear, free). To avoid code duplication, create\n> the 'struct tree_entry_array' type and add functions that perform those\n> basic operations.\n>\n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> @@ -12,15 +12,39 @@\n> +struct tree_entry_array {\n> +       size_t nr, alloc;\n> +       struct tree_entry **entries;\n> +};\n>\n> +static void clear_tree_entry_array(struct tree_entry_array *arr)\n> +{\n> +       for (size_t i = 0; i < arr->nr; i++)\n> +               FREE_AND_NULL(arr->entries[i]);\n> +       arr->nr = 0;\n> +}\n> +\n> +static void release_tree_entry_array(struct tree_entry_array *arr)\n> +{\n> +       FREE_AND_NULL(arr->entries);\n> +       arr->nr = arr->alloc = 0;\n> +}\n\nFor robustness, to make it less likely for future code to leak the\nitems pointed to by `arr->entries`, it might make sense for\nrelease_tree_entry_array() to call clear_tree_entry_array() before\ncalling FREE_AND_NULL().\n\n> -       ALLOC_GROW(entries, used + 1, alloc);\n> -       entries[used++] = ent;\n> +       /* Append the update */\n> +       tree_entry_array_push(arr, ent);\n\nNit: the new comment seems superfluous\n"},{"id":"496962","messageId":"xmqqa5jrt7x4.fsf@gitster.g","threadId":"61624","inReplyTo":"9d0689e9c285b375b0067760929011038c085d65.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 04/16] update-index: generalize 'read_index_info'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-11T22:45:59Z","receivedAt":"2024-06-11T22:46:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Move 'read_index_info()' into a new header 'index-info.h' and generalize the\n> function to call a provided callback for each parsed line. Update\n> 'update-index.c' to use this generalized 'read_index_info()', adding the\n> callback 'apply_index_info()' to verify the parsed line and update the index\n> according to its contents.\n>\n> The input parsing done by 'read_index_info()' is similar to, but more\n> flexible than, the parsing done in 'mktree' by 'mktree_line()' (handling not\n> only 'git ls-tree' output but also the outputs of 'git apply --index-info'\n> and 'git ls-files --stage' outputs). To make 'mktree' more flexible, a later\n> patch will replace mktree's custom parsing with 'read_index_info()'.\n\n\"git apply --index-info\"?  \n\nThat is a blast from the past.  It no longer exists since 7a988699\n(apply: get rid of --index-info in favor of --build-fake-ancestor,\n2007-09-17).\n\nAs to the scriptability, supporting \"ls-files -s\" and \"ls-tree -r\"\noutput as our input do help, but the third one is not natively\nemitted and it is very unlikely that there are third-party tools\nthat give output in that format.  After all these years, I suspect\nthat it is sufficient to say\n\n    \"update-index --index-info\" and \"mktree\" both read information\n    necessary to eventually build trees, but having two separate\n    parsers is a maintenance burden, so we are massaging the code\n    from the former to be reusable.\n\nwithout mentioning where the old third format comes from.\n\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index d343416ae26..77df380cb54 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -11,6 +11,7 @@\n>  #include \"gettext.h\"\n>  #include \"hash.h\"\n>  #include \"hex.h\"\n> +#include \"index-info.h\"\n>  #include \"lockfile.h\"\n>  #include \"quote.h\"\n>  #include \"cache-tree.h\"\n> @@ -509,100 +510,29 @@ static void update_one(const char *path)\n>  \treport(\"add '%s'\", path);\n>  }\n>  \n> +static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n> +\t\t\t    const char *path_name, void *cbdata UNUSED)\n>  {\n> +\tif (!verify_path(path_name, mode)) {\n> +\t\tfprintf(stderr, \"Ignoring path %s\\n\", path_name);\n> +\t\treturn 0;\n> +\t}\n>  \n> +\tif (!mode) {\n> +\t\t/* mode == 0 means there is no such path -- remove */\n> +\t\tif (remove_file_from_index(the_repository->index, path_name))\n> +\t\t\tdie(\"git update-index: unable to remove %s\", path_name);\n\nThis changes the error message.  We used to feed \"ptr\" (no longer\nvisible to this function, as the caller unquotes before calling us)\nthat pointed at the original the user gave to the program; now we\nreport the path_name which is the result of the unquoting.\n\n> +\t}\n> +\telse {\n> +\t\t/* mode ' ' sha1 '\\t' name\n> +\t\t * ptr[-1] points at tab,\n> +\t\t * ptr[-41] is at the beginning of sha1\n>  \t\t */\n> +\t\tif (add_cacheinfo(mode, oid, path_name, stage))\n> +\t\t\tdie(\"git update-index: unable to update %s\", path_name);\n\nBut this side used to report the path_name as the result of\nunquoting in the original.  So the above change would probably be OK\nin the name of consistency?\n\n973d6a20 (update-index --index-info: adjust for funny-path quoting.,\n2005-10-16) was the origin of the unquoting, and looking at that\ncommit, I have a feeling that the \"ptr\" thing above (i.e., the one I\npointed out as changing the behaviour) was simply forgotten (as\nopposed to deliberately made to report the original) while updating\nthe code to deal with quoted original into unquoted paths.\n\nSo I think the change is more than OK.  It is a very welcome (belated)\nbugfix for 973d6a20 ;-).\n\n>  \t}\n> +\n> +\treturn 0;\n>  }\n\nIt looks a bit disappointing that we die in the callback like above,\nwhen the main parser loop that moved to the other file to be more\nreusable is now capable of returning to the caller with an error,\nbut at this step, it is a good place to stop.  A refactor that does\nnot change the behaviour.\n\nNicely done.\n\n> diff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\n> index cc72ead79f3..29696ade0d0 100755\n> --- a/t/t2107-update-index-basic.sh\n> +++ b/t/t2107-update-index-basic.sh\n> @@ -142,4 +142,31 @@ test_expect_success '--index-version' '\n>  \ttest_must_be_empty actual\n>  '\n>  \n> +test_expect_success '--index-info fails on malformed input' '\n> +\t# empty line\n> +\techo \"\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"malformed input line\" err &&\n\nUsing \"test_grep\" would make it easier to diagnose when test breaks.\nA failing \"grep\" will be silent.  A failing \"test_grep\" will tell us\n\"I was told to find THIS, but didn't find any in THAT\".\n\n> +\t# bad whitespace\n> +\tprintf \"100644 $EMPTY_BLOB A\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"malformed input line\" err &&\n> +\n> +\t# invalid stage value\n> +\tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"malformed input line\" err &&\n> +\n> +\t# invalid OID length\n> +\tprintf \"100755 abc123\\tA\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"malformed input line\" err &&\n> +\n> +\t# bad quoting\n> +\tprintf \"100644 $EMPTY_BLOB\\t\\\"A\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"bad quoting of path name\" err\n> +'\n> +\n>  test_done\n"},{"id":"496963","messageId":"xmqq34pjt7m2.fsf@gitster.g","threadId":"61624","inReplyTo":"7e3bcc16e23c97d8a4efbb9e14b230ef9f44a1a7.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 05/16] index-info.c: identify empty input lines in read_index_info","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-11T22:52:37Z","receivedAt":"2024-06-11T22:52:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Update 'read_index_info()' to return INDEX_INFO_EMPTY_LINE (value 1), rather\n> than the default error code (value -1) when the function encounters an empty\n> line in stdin. This grants the caller the flexibility to handle such\n> scenarios differently than a typical error. In the case of 'update-index',\n> we'll still exit with a \"malformed input line\" error. However, when\n> 'read_index_info()' is used to process the input to 'mktree' in a later\n> patch, the empty line return value will signal a new tree in --batch mode.\n\nInteresting.  We could even introduce \"# commented input\" but that\nis a different story ;-).\n\nI also wonder if we can flip it around and teach read_index_info()\nto (1) silently accept and do a callback when it recognises the\ninput line is one of the supported formats, and (2) send any\nunrecognised line, not just an empty one, with \"unrecognised\" status\ncode.  That way, the caller can handle more than single kind of\n\"special input line\" more easily, perhaps?\n\nThanks.\n"},{"id":"496970","messageId":"xmqqcyonrkms.fsf@gitster.g","threadId":"61624","inReplyTo":"f56eee0b48da907a27edc99ca135cf8f6c19af35.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 06/16] index-info.c: parse object type in provided in read_index_info","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T01:54:19Z","receivedAt":"2024-06-12T01:54:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> If the object type (e.g. \"blob\", \"tree\") is identified on a stdin line read\n> by 'read_index_info()' (i.e. on lines formatted like the output of 'git\n> ls-tree'), parse it into an 'enum object_type' and provide it to the\n> 'read_index_info()' callback as an argument. If the type is not provided,\n> pass 'OBJ_NONE' instead. If the object type is invalid, return an error.\n\nMy recollection is, when we do not know what to expect, we tend to\nuse OBJ_ANY rather than OBJ_NONE as convention to signal that fact\n(e.g., object-name.c:peel_to_type()).\n\nAs long as the code path this series touches is internally\nconsistent, using OBJ_NONE may not hurt but once they need to start\ninteracting with existing code paths that use OBJ_ANY for that\npurpose, we may need to adjust one to match the other.\n\n> The goal of this change is to allow for more thorough validation of the\n> provided object type (e.g. against the provided mode) in 'mktree' once\n> 'mktree_line' is replaced with 'read_index_info()'. Note, though, that this\n> change also strengthens the validation done by 'update-index', since invalid\n> type names now trigger an error.\n\nNice.\n\n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n>  builtin/update-index.c        |  3 ++-\n>  index-info.c                  | 16 ++++++++++++----\n>  index-info.h                  |  3 ++-\n>  t/t2107-update-index-basic.sh |  5 +++++\n>  4 files changed, 21 insertions(+), 6 deletions(-)\n>\n> diff --git a/builtin/update-index.c b/builtin/update-index.c\n> index b1b334807f8..8882433b644 100644\n> --- a/builtin/update-index.c\n> +++ b/builtin/update-index.c\n> @@ -510,7 +510,8 @@ static void update_one(const char *path)\n>  \treport(\"add '%s'\", path);\n>  }\n>  \n> -static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n> +static int apply_index_info(unsigned int mode, struct object_id *oid,\n> +\t\t\t    enum object_type obj_type UNUSED, int stage,\n>  \t\t\t    const char *path_name, void *cbdata UNUSED)\n>  {\n>  \tif (!verify_path(path_name, mode)) {\n> diff --git a/index-info.c b/index-info.c\n> index 735cbf1f476..5d61e61e28f 100644\n> --- a/index-info.c\n> +++ b/index-info.c\n> @@ -18,6 +18,7 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n>  \t\tchar *ptr, *tab;\n>  \t\tchar *path_name;\n>  \t\tstruct object_id oid;\n> +\t\tenum object_type obj_type = OBJ_NONE;\n>  \t\tunsigned int mode;\n>  \t\tunsigned long ul;\n>  \t\tint stage;\n> @@ -56,18 +57,17 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n>  \n>  \t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n>  \t\t\tstage = tab[-1] - '0';\n> -\t\t\tptr = tab + 1; /* point at the head of path */\n> +\t\t\tpath_name = tab + 1; /* point at the head of path */\n>  \t\t\ttab = tab - 2; /* point at tail of sha1 */\n>  \t\t} else {\n>  \t\t\tstage = 0;\n> -\t\t\tptr = tab + 1; /* point at the head of path */\n> +\t\t\tpath_name = tab + 1; /* point at the head of path */\n>  \t\t}\n>  \n>  \t\tif (get_oid_hex(tab - hexsz, &oid) ||\n>  \t\t\ttab[-(hexsz + 1)] != ' ')\n>  \t\t\tgoto bad_line;\n>  \n> -\t\tpath_name = ptr;\n>  \t\tif (!nul_term_line && path_name[0] == '\"') {\n>  \t\t\tstrbuf_reset(&uq);\n>  \t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n> @@ -77,7 +77,15 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n>  \t\t\tpath_name = uq.buf;\n>  \t\t}\n>  \n> -\t\tret = fn(mode, &oid, stage, path_name, cbdata);\n> +\t\t/* Get the type, if provided */\n> +\t\tif (tab - hexsz - 1 > ptr + 1) {\n> +\t\t\tif (*(tab - hexsz - 1) != ' ')\n> +\t\t\t\tgoto bad_line;\n> +\t\t\t*(tab - hexsz - 1) = '\\0';\n> +\t\t\tobj_type = type_from_string(ptr + 1);\n> +\t\t}\n> +\n> +\t\tret = fn(mode, &oid, obj_type, stage, path_name, cbdata);\n>  \t\tif (ret) {\n>  \t\t\tret = -1;\n>  \t\t\tbreak;\n> diff --git a/index-info.h b/index-info.h\n> index 1884972021d..767cf304213 100644\n> --- a/index-info.h\n> +++ b/index-info.h\n> @@ -2,8 +2,9 @@\n>  #define INDEX_INFO_H\n>  \n>  #include \"hash.h\"\n> +#include \"object.h\"\n>  \n> -typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n> +typedef int (*each_index_info_fn)(unsigned int, struct object_id *, enum object_type, int, const char *, void *);\n>  \n>  #define INDEX_INFO_EMPTY_LINE 1\n>  \n> diff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\n> index 29696ade0d0..9c19d24cd4a 100755\n> --- a/t/t2107-update-index-basic.sh\n> +++ b/t/t2107-update-index-basic.sh\n> @@ -153,6 +153,11 @@ test_expect_success '--index-info fails on malformed input' '\n>  \ttest_must_fail git update-index --index-info 2>err &&\n>  \tgrep \"malformed input line\" err &&\n>  \n> +\t# invalid type\n> +\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n> +\ttest_must_fail git update-index --index-info 2>err &&\n> +\tgrep \"invalid object type\" err &&\n> +\n>  \t# invalid stage value\n>  \tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n>  \ttest_must_fail git update-index --index-info 2>err &&\n"},{"id":"496971","messageId":"xmqqplsmrjtc.fsf@gitster.g","threadId":"61624","inReplyTo":"8d1e1eaa70b96779416f2f48a862d31a730c4521.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 07/16] mktree: use read_index_info to read stdin lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T02:11:59Z","receivedAt":"2024-06-12T02:12:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Replace the custom input parsing of 'mktree' with 'read_index_info()', which\n> handles not only the 'ls-tree' output format it already handles but also the\n> other formats compatible with 'update-index'.\n\nYay.\n\n> This lends some consistency\n> across the commands (avoiding the need for two similar implementations for\n> input parsing) and adds flexibility to mktree.\n>\n> Update 'Documentation/git-mktree.txt' to reflect the more permissive input\n> format.\n\nNice.\n>  DESCRIPTION\n>  -----------\n> -Reads standard input in non-recursive `ls-tree` output format, and creates\n> -a tree object.  The order of the tree entries is normalized by mktree so\n> -pre-sorting the input is not required.  The object name of the tree object\n> -built is written to the standard output.\n> +Reads entry information from stdin and creates a tree object from those entries.\n> +The object name of the tree object built is written to the standard output.\n\npre-sorting is now required?  Ah, such details are left to the\nsection dedicated for the input format.  Makes sense.\n\nThe line is getting overly long (the first line now is exactly\n80-columns); wrapping them to leave a bit of room to grow, like\nat around 72-76 columns, would be appreciated.\n\n> +INPUT FORMAT\n> +------------\n> +Tree entries may be specified in any of the formats compatible with the\n> +`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n> +entries is normalized by `mktree` so pre-sorting the input by path is not\n> +required.\n\nOK.  We might want to split the description of the three-formats\ninto a separate file and include it in here and in the original (I'd\ncertainly insist doing so if we had three places that want to refer\nto it), but we have only two so let's just remember to do so when we\nmay want to add the third place in the future.\n\n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index 15bd908702a..5530257252d 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -6,6 +6,7 @@\n>  #include \"builtin.h\"\n>  #include \"gettext.h\"\n>  #include \"hex.h\"\n> +#include \"index-info.h\"\n>  #include \"quote.h\"\n>  #include \"strbuf.h\"\n>  #include \"tree.h\"\n> @@ -93,123 +94,80 @@ static const char *mktree_usage[] = {\n>  \tNULL\n>  };\n>  \n> -static void mktree_line(char *buf, int nul_term_line, int allow_missing,\n> -\t\t\tstruct tree_entry_array *arr)\n> +struct mktree_line_data {\n> +\tstruct tree_entry_array *arr;\n> +\tint allow_missing;\n> +};\n> +\n> +static int mktree_line(unsigned int mode, struct object_id *oid,\n> +\t\t       enum object_type obj_type, int stage UNUSED,\n> +\t\t       const char *path, void *cbdata)\n>  {\n> +\tstruct mktree_line_data *data = cbdata;\n> +\tenum object_type mode_type = object_type(mode);\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n> +\tenum object_type parsed_obj_type;\n>  \n> +\tif (obj_type && mode_type != obj_type)\n> +\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n> +\t\t    type_name(obj_type), type_name(mode_type));\n>  \n> +\toi.typep = &parsed_obj_type;\n>  \n> +\tif (oid_object_info_extended(the_repository, oid, &oi,\n>  \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n>  \t\t\t\t     OBJECT_INFO_QUICK |\n>  \t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n> +\t\tparsed_obj_type = -1;\n>  \n> +\tif (parsed_obj_type < 0) {\n> +\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n> +\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n\nOverlong line?\n\n>  \t\t} else {\n> +\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n>  \t\t}\n\nEach side of if/else has only a single statement block that does not\nwant {braces} around it.  I wonder if flipping the polarity makes it\neasier to follow the logic flow:\n\n\t\tif (!data->allow_missing && !S_ISGITLINK(mode))\n\t\t\tdie(\"...\");\n\nI wonder if we even want to do the oid_object_info_extended() when\nwe are expecting to see a gitlink.  We do not expect to have the\ncommit in our history (as it is part of the history of a submodule,\nwhich is from a separate project), so even if we found such an\nobject in our object database, we do not want to do anything with\nthe information we learn about the object.\n\nSo I am wondering if the whole cascade should read more like\n\n\tif (S_ISGITILNK(mode)) {\n\t\t... anything goes ...\n\t} else if (oid_object_info_extended(...) < 0 &&\n\t\t   !data->allow_missing) {\n        \t... not found ...\n\t} else if (parsed_obj_type != mode_type) {\n        \t... found something different from what we expected ...\n\t}\n\nThe main loop, thanks to read_index_info() refactoring, got really\neasier to read, i.e. compact and clear.\n"},{"id":"496972","messageId":"xmqqh6dyrjia.fsf@gitster.g","threadId":"61624","inReplyTo":"b497dc90687a7c77a4d21c3a12fe5fa3bfdabc16.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 08/16] mktree: add a --literally option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T02:18:37Z","receivedAt":"2024-06-12T02:18:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Add the '--literally' option to 'git mktree' to allow constructing a tree\n> with invalid contents. For now, the only change this represents compared to\n> the normal 'git mktree' behavior is no longer sorting the inputs; in later\n> commits, deduplicaton and path validation will be added to the command and\n> '--literally' will skip those as well.\n\nHmph, the end state of the the series as a whole may be good, but\nthe above makes me wonder if we broke bisectability with the\nprevious step 07/16 where we introduced type checks without touching\nany existing tests?\n\n> Certain tests use 'git mktree' to intentionally generate corrupt trees.\n> Update these tests to use '--literally' so that they continue functioning\n> properly when additional input cleanup & validation is added to the base\n> command. Note that, because 'mktree --literally' does not sort entries, some\n> of the tests are updated to provide their inputs in tree order; otherwise,\n> the test would fail with an \"incorrect order\" error instead of the error the\n> test expects.\n\nMakes sense.\n\n> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n> index 507682ed23e..fb07e40cef0 100644\n> --- a/Documentation/git-mktree.txt\n> +++ b/Documentation/git-mktree.txt\n> @@ -9,7 +9,7 @@ git-mktree - Build a tree-object from formatted tree entries\n>  SYNOPSIS\n>  --------\n>  [verse]\n> -'git mktree' [-z] [--missing] [--batch]\n> +'git mktree' [-z] [--missing] [--literally] [--batch]\n>  \n>  DESCRIPTION\n>  -----------\n> @@ -27,6 +27,13 @@ OPTIONS\n>  \tobject.  This option has no effect on the treatment of gitlink entries\n>  \t(aka \"submodules\") which are always allowed to be missing.\n>  \n> +--literally::\n> +\tCreate the tree from the tree entries provided to stdin in the order\n> +\tthey are provided without performing additional sorting, deduplication,\n> +\tor path validation on them. This option is primarily useful for creating\n> +\tinvalid tree objects to use in tests of how Git deals with various forms\n> +\tof tree corruption.\n> +\n\nOK.\n\n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index 5530257252d..48019448c1f 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -45,11 +45,11 @@ static void release_tree_entry_array(struct tree_entry_array *arr)\n>  }\n>  \n>  static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n> -\t\t\t   struct tree_entry_array *arr)\n> +\t\t\t   struct tree_entry_array *arr, int literally)\n>  {\n>  \tstruct tree_entry *ent;\n>  \tsize_t len = strlen(path);\n> -\tif (strchr(path, '/'))\n> +\tif (!literally && strchr(path, '/'))\n>  \t\tdie(\"path %s contains slash\", path);\n\n;-).\n\nA tree_entry with a slash in it.  Our fsck should be catching them\nalready, but this will allow constructing a test case more easily.\n\n"},{"id":"496973","messageId":"xmqq34pirj51.fsf@gitster.g","threadId":"61624","inReplyTo":"4f9f77e693cfc4fbe72a2ae739bc7e236a3b82d3.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 09/16] mktree: validate paths more carefully","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T02:26:34Z","receivedAt":"2024-06-12T02:26:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Use 'verify_path' to validate the paths provided as tree entries, ensuring\n> we do not create entries with paths not allowed in trees (e.g.,\n> .git).\n\nSensible.\n\n> Also,\n> remove trailing slashes on directories before validating, allowing users to\n> provide 'folder-name/' as the path for a tree object entry.\n\nIs that a good idea for a plumbing like this command?  We would\nsilently accept these after silently stripping the trailing slash?\n\n040000 tree 82a33d5150d9316378ef1955a49f2a5bf21aaeb2    templates/\n100644 blob 1f89ffab4c32bc02b5d955851401628a5b9a540e    thread-utils.c/\n\nThe former _might_ count as \"usability improvement\", but if we are\ndoing the same for the latter we might be going a bit too lenient.\n\nLet's see what really happens in the code.\n\n> @@ -49,10 +50,23 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n>  {\n>  \tstruct tree_entry *ent;\n>  \tsize_t len = strlen(path);\n> -\tif (!literally && strchr(path, '/'))\n> -\t\tdie(\"path %s contains slash\", path);\n>  \n> -\tFLEX_ALLOC_MEM(ent, name, path, len);\n> +\tif (literally) {\n> +\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n> +\t} else {\n> +\t\t/* Normalize and validate entry path */\n> +\t\tif (S_ISDIR(mode)) {\n> +\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n> +\t\t\t\tlen--;\n> +\t\t}\n\nLeave a single SP after \"while\", please.\n\nWe do this only to subtree entries, and all trailing slashes, not\njust a single one.  OK, but I am not sure if the extra leniency is a\ngood idea to begin with.  \"ls-tree\" output does not have such a\ntrailing slashes, so it is unclear whom we are trying to be extra\nnice with this.\n\n> +\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n> +\n> +\t\tif (!verify_path(ent->name, mode))\n> +\t\t\tdie(_(\"invalid path '%s'\"), path);\n\nThis is the crux of the change.  And it is so simple.  Very nice.\n\n> +\t\tif (strchr(ent->name, '/'))\n> +\t\t\tdie(\"path %s contains slash\", path);\n> +\t}\n\n> diff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\n> index e0687cb529f..e0263cb2bf8 100755\n> --- a/t/t1010-mktree.sh\n> +++ b/t/t1010-mktree.sh\n> @@ -173,4 +173,37 @@ test_expect_success '--literally can create invalid trees' '\n>  \tgrep \"not properly sorted\" err\n>  '\n>  \n> +test_expect_success 'mktree validates path' '\n> +\ttree_oid=\"$(cat tree)\" &&\n> +\tblob_oid=\"$(git rev-parse $tree_oid:a/one)\" &&\n> +\thead_oid=\"$(git rev-parse HEAD)\" &&\n> +\n> +\t# Valid: tree with or without trailing slash, blob without trailing slash\n> +\t{\n> +\t\tprintf \"040000 tree $tree_oid\\tfolder1/\\n\" &&\n> +\t\tprintf \"040000 tree $tree_oid\\tfolder2\\n\" &&\n> +\t\tprintf \"100644 blob $blob_oid\\tfile.txt\\n\"\n> +\t} | git mktree >actual &&\n> +\n> +\t# Invalid: blob with trailing slash\n> +\tprintf \"100644 blob $blob_oid\\ttest/\" |\n> +\ttest_must_fail git mktree 2>err &&\n> +\tgrep \"invalid path ${SQ}test/${SQ}\" err &&\n> +\n> +\t# Invalid: dotdot\n> +\tprintf \"040000 tree $tree_oid\\t../\" |\n> +\ttest_must_fail git mktree 2>err &&\n> +\tgrep \"invalid path ${SQ}../${SQ}\" err &&\n> +\n> +\t# Invalid: dot\n> +\tprintf \"040000 tree $tree_oid\\t.\" |\n> +\ttest_must_fail git mktree 2>err &&\n> +\tgrep \"invalid path ${SQ}.${SQ}\" err &&\n> +\n> +\t# Invalid: .git\n> +\tprintf \"040000 tree $tree_oid\\t.git/\" |\n> +\ttest_must_fail git mktree 2>err &&\n> +\tgrep \"invalid path ${SQ}.git/${SQ}\" err\n> +'\n> +\n>  test_done\n"},{"id":"497013","messageId":"ZmltCIoeYgRdj-CN@tanuki","threadId":"61624","inReplyTo":"4558f35e7bf9a1594510951ee54252069bdcfc5b.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 02/16] mktree: rename treeent to tree_entry","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:24Z","receivedAt":"2024-06-12T09:40:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:34PM +0000, Victoria Dye via GitGitGadget wrote:\n> From: Victoria Dye <vdye@github.com>\n> \n> Rename the type for better readability, clearly specifying \"entry\" (instead\n> of the \"ent\" abbreviation) and separating \"tree\" from \"entry\".\n> \n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n>  builtin/mktree.c | 10 +++++-----\n>  1 file changed, 5 insertions(+), 5 deletions(-)\n> \n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index 8b19d440747..c02feb06aff 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -12,7 +12,7 @@\n>  #include \"parse-options.h\"\n>  #include \"object-store-ll.h\"\n>  \n> -static struct treeent {\n> +static struct tree_entry {\n>  \tunsigned mode;\n>  \tstruct object_id oid;\n>  \tint len;\n\nThis reads a ton better compared to `treeent`, thanks! I've never been a\nfan of abbreviations like this in code.\n\nPatrick\n"},{"id":"497014","messageId":"ZmltDQ5SlVvrEDGP@tanuki","threadId":"61624","inReplyTo":"5ade145352f44b431c16a2ec29cd87de489e8032.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 03/16] mktree: use non-static tree_entry array","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:29Z","receivedAt":"2024-06-12T09:40:33Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:35PM +0000, Victoria Dye via GitGitGadget wrote:\n> From: Victoria Dye <vdye@github.com>\n> \n> Replace the static 'struct tree_entry **entries' with a non-static 'struct\n> tree_entry_array' instance. In later commits, we'll want to be able to\n> create additional 'struct tree_entry_array' instances utilizing common\n> functionality (create, push, clear, free). To avoid code duplication, create\n> the 'struct tree_entry_array' type and add functions that perform those\n> basic operations.\n\nThanks for getting rid of more global state, I really appreciate this.\n\n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n>  builtin/mktree.c | 67 +++++++++++++++++++++++++++++++++---------------\n>  1 file changed, 47 insertions(+), 20 deletions(-)\n> \n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index c02feb06aff..15bd908702a 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -12,15 +12,39 @@\n>  #include \"parse-options.h\"\n>  #include \"object-store-ll.h\"\n>  \n> -static struct tree_entry {\n> +struct tree_entry {\n>  \tunsigned mode;\n>  \tstruct object_id oid;\n>  \tint len;\n>  \tchar name[FLEX_ARRAY];\n> -} **entries;\n> -static int alloc, used;\n> +};\n> +\n> +struct tree_entry_array {\n> +\tsize_t nr, alloc;\n> +\tstruct tree_entry **entries;\n> +};\n>  \n> -static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n> +static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n> +{\n> +\tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n> +\tarr->entries[arr->nr++] = ent;\n> +}\n> +\n> +static void clear_tree_entry_array(struct tree_entry_array *arr)\n> +{\n> +\tfor (size_t i = 0; i < arr->nr; i++)\n> +\t\tFREE_AND_NULL(arr->entries[i]);\n> +\tarr->nr = 0;\n> +}\n> +\n> +static void release_tree_entry_array(struct tree_entry_array *arr)\n> +{\n> +\tFREE_AND_NULL(arr->entries);\n> +\tarr->nr = arr->alloc = 0;\n> +}\n\nNit: should these be called `tree_entry_array_clear()` and\n`tree_entry_array_release()`? This is one of the areas where our coding\nguidelines aren't sufficiently clear in my opinion. I personally\nstrongly prefer `<noun>_<action>` syntax because it groups together\nrelated functionality much better. And while I personally do not use\ncode completion, it does help others that do because they can simply\ntype in the noun as prefix and then, via completion, learn about all\nrelated functions.\n\nMost of our general-purpose interfaces follow this naming schema\n(strbuf, string_list, strvec, oidset, ...). I think we should document\nthis accordingly.\n\nPatrick\n"},{"id":"497015","messageId":"ZmltEti7TRpaiCD-@tanuki","threadId":"61624","inReplyTo":"8d1e1eaa70b96779416f2f48a862d31a730c4521.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 07/16] mktree: use read_index_info to read stdin lines","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:34Z","receivedAt":"2024-06-12T09:40:39Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:39PM +0000, Victoria Dye via GitGitGadget wrote:\n> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n> index 383f09dd333..507682ed23e 100644\n> --- a/Documentation/git-mktree.txt\n> +++ b/Documentation/git-mktree.txt\n> @@ -13,15 +13,13 @@ SYNOPSIS\n>  \n>  DESCRIPTION\n>  -----------\n> -Reads standard input in non-recursive `ls-tree` output format, and creates\n> -a tree object.  The order of the tree entries is normalized by mktree so\n> -pre-sorting the input is not required.  The object name of the tree object\n> -built is written to the standard output.\n> +Reads entry information from stdin and creates a tree object from those entries.\n> +The object name of the tree object built is written to the standard output.\n\nIt makes perfect sense to not single out git-ls-tree(1) anymore. But I\nthink we should help the reader a bit by continuing to point out which\ncommands can be used as input here. That can be either here in the\ndescription, further down in the new \"INPUT FORMAT\" section, or in both\nplaces.\n\nPatrick\n"},{"id":"497016","messageId":"ZmltGAPQ2dAfW0kG@tanuki","threadId":"61624","inReplyTo":"b59a4ad8ab4b0e47373f811700eba59141fdc6c6.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 10/16] mktree: overwrite duplicate entries","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:40Z","receivedAt":"2024-06-12T09:40:44Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:42PM +0000, Victoria Dye via GitGitGadget wrote:\n> From: Victoria Dye <vdye@github.com>\n> \n> If multiple tree entries with the same name are provided as input to\n> 'mktree', only write the last one to the tree. Entries are considered\n> duplicates if they have identical names (*not* considering mode); if a blob\n> and a tree with the same name are provided, only the last one will be\n> written to the tree. A tree with duplicate entries is invalid (per 'git\n> fsck'), so that condition should be avoided wherever possible.\n> \n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n>  Documentation/git-mktree.txt |  8 ++++---\n>  builtin/mktree.c             | 45 ++++++++++++++++++++++++++++++++----\n>  t/t1010-mktree.sh            | 36 +++++++++++++++++++++++++++--\n>  3 files changed, 80 insertions(+), 9 deletions(-)\n> \n> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n> index fb07e40cef0..afbc846d077 100644\n> --- a/Documentation/git-mktree.txt\n> +++ b/Documentation/git-mktree.txt\n> @@ -43,9 +43,11 @@ OPTIONS\n>  INPUT FORMAT\n>  ------------\n>  Tree entries may be specified in any of the formats compatible with the\n> -`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n> -entries is normalized by `mktree` so pre-sorting the input by path is not\n> -required.\n> +`--index-info` option to linkgit:git-update-index[1].\n> +\n> +The order of the tree entries is normalized by `mktree` so pre-sorting the input\n> +by path is not required. Multiple entries provided with the same path are\n> +deduplicated, with only the last one specified added to the tree.\n\nHm. I'm not sure whether this is a good idea. With git-mktree(1) being\npart of our plumbing layer, you can expect that it's mostly going to be\nfed input from scripts. And any script that generates duplicate tree\nentries is broken, but we now start to paper over such brokenness\nwithout giving the user any indicator of this. As user of git-mktree(1)\nin Gitaly I can certainly say that I'd rather want to see it die instead\nof silently fixing my inputs so that I start to notice my own bugs.\n\nSo without seeing a strong motivating usecase for this feature I'd think\nthat git-mktree(1) should reject such inputs and return an error such\nthat the user can fix their tooling.\n\nPatrick\n"},{"id":"497017","messageId":"ZmltHUeuCum4daB2@tanuki","threadId":"61624","inReplyTo":"130413f2404bb27a2ede4fb00041227c90587e8e.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 11/16] mktree: create tree using an in-core index","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:45Z","receivedAt":"2024-06-12T09:40:49Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:43PM +0000, Victoria Dye via GitGitGadget wrote:\n> From: Victoria Dye <vdye@github.com>\n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index e9e2134136f..12f68187221 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -24,6 +25,11 @@ struct tree_entry {\n>  \tchar name[FLEX_ARRAY];\n>  };\n>  \n> +static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n> +{\n> +\treturn S_ISDIR(mode) ? pathlen - 1 : pathlen;\n\nI wonder whether we want to have a sanity check that ensures that the\npath really does have a trailing slash.\n\n> @@ -120,24 +137,43 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>  \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n>  }\n>  \n> +static int add_tree_entry_to_index(struct index_state *istate,\n> +\t\t\t\t   struct tree_entry *ent)\n> +{\n> +\tstruct cache_entry *ce;\n> +\tstruct strbuf ce_name = STRBUF_INIT;\n> +\tstrbuf_add(&ce_name, ent->name, ent->len);\n> +\n> +\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n> +\tif (!ce)\n> +\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n\nI noticed that `make_cache_entry()` will skip over index entries which\nare up-to-date, and it will replace entries which are part of the index\nbut with different information. Is this the motivator for the preceding\ncommit where we start to overwrite duplicate entries?\n\nPatrick\n"},{"id":"497018","messageId":"ZmltI7HA7O4w2E-6@tanuki","threadId":"61624","inReplyTo":"94d6615d634c4f78c88d3e01abbb27f13f85828c.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 12/16] mktree: use iterator struct to add tree entries to index","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:51Z","receivedAt":"2024-06-12T09:40:55Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:44PM +0000, Victoria Dye via GitGitGadget wrote:\n> From: Victoria Dye <vdye@github.com>\n> \n> Create 'struct tree_entry_iterator' to manage iteration through a 'struct\n> tree_entry_array'. Using an iterator allows for conditional iteration; this\n> functionality will be necessary in later commits when performing parallel\n> iteration through multiple sets of tree entries.\n> \n> Signed-off-by: Victoria Dye <vdye@github.com>\n> ---\n>  builtin/mktree.c | 40 +++++++++++++++++++++++++++++++++++++---\n>  1 file changed, 37 insertions(+), 3 deletions(-)\n> \n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index 12f68187221..bee359e9978 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> @@ -137,6 +137,38 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>  \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n>  }\n>  \n> +struct tree_entry_iterator {\n> +\tstruct tree_entry *current;\n> +\n> +\t/* private */\n> +\tstruct {\n> +\t\tstruct tree_entry_array *arr;\n> +\t\tsize_t idx;\n> +\t} priv;\n> +};\n> +\n> +static void init_tree_entry_iterator(struct tree_entry_iterator *iter,\n> +\t\t\t\t     struct tree_entry_array *arr)\n> +{\n> +\titer->priv.arr = arr;\n> +\titer->priv.idx = 0;\n> +\titer->current = 0 < arr->nr ? arr->entries[0] : NULL;\n> +}\n\nNit: Same comment as before, I think these should rather be named\n`tree_entry_iterator_init()` and `tree_entry_iterator_advance()`.\n\n> +/*\n> + * Advance the tree entry iterator to the next entry in the array. If no entries\n> + * remain, 'current' is set to NULL. Returns the previous 'current' value of the\n> + * iterator.\n> + */\n> +static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator *iter)\n> +{\n> +\tstruct tree_entry *prev = iter->current;\n> +\titer->current = (iter->priv.idx + 1) < iter->priv.arr->nr\n> +\t\t\t? iter->priv.arr->entries[++iter->priv.idx]\n> +\t\t\t: NULL;\n> +\treturn prev;\n> +}\n\nI think it's somewhat confusing to have this return a different value\nthan `current`. When I call `next()`, then I expect the iterator to\nreturn the next item. And after having called `next()`, I expect that\nthe current value is the one that the previous call to `next()` has\nreturned.\n\nTo avoid confusion, I'd propose to get rid of the `current` member\naltogether. It's not needed as we already save the current index and\navoids the confusion.\n\nPatrick\n"},{"id":"497019","messageId":"ZmltKHI-Vz1L44r8@tanuki","threadId":"61624","inReplyTo":"df0c50dfea3cb77e0070246efdf7a3f070b2ad97.1718130288.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 14/16] mktree: optionally add to an existing tree","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-12T09:40:56Z","receivedAt":"2024-06-12T09:41:00Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Jun 11, 2024 at 06:24:46PM +0000, Victoria Dye via GitGitGadget wrote:\n> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n> index afbc846d077..99abd3c31a6 100644\n> --- a/Documentation/git-mktree.txt\n> +++ b/Documentation/git-mktree.txt\n> @@ -40,6 +40,11 @@ OPTIONS\n>  \toptional.  Note - if the `-z` option is used, lines are terminated\n>  \twith NUL.\n>  \n> +<tree-ish>::\n> +\tIf provided, the tree entries provided in stdin are added to this tree\n> +\trather than a new empty one, replacing existing entries with identical\n> +\tnames. Not compatible with `--literally`.\n\nI think it'd be a bit more intuitive is this was an option, like\n`--base-tree=` or just `--base=`.\n\nOne question that comes up naturally in this context: when I have a base\ntree, how do I remove entries from it?\n\nPatrick\n"},{"id":"497044","messageId":"xmqqle3aovpq.fsf@gitster.g","threadId":"61624","inReplyTo":"ZmltEti7TRpaiCD-@tanuki","subject":"Re: [PATCH 07/16] mktree: use read_index_info to read stdin lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T18:35:29Z","receivedAt":"2024-06-12T18:35:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> It makes perfect sense to not single out git-ls-tree(1) anymore. But I\n> think we should help the reader a bit by continuing to point out which\n> commands can be used as input here. That can be either here in the\n> description, further down in the new \"INPUT FORMAT\" section, or in both\n> places.\n\nHere is a way to do so, which I alluded to earlier.  The original\ntext is too specific to \"update-index\" in that it talked about\n\"stuffing them into the index\", which does not apply in the context\nof \"mktree\".\n\nAnd then it made me realize that \"ls-files -s\" output has the stage\ninformation, which of course is needed for \"update-index\" to be able\nto recreate the index state from a textual dump, but \"mktree\" should\nreject if given a higher stage entry.\n\nIt seems that the code after applying all these 16 patches does not\ndiagnose it as an error if you feed a non-zero stage.  The callback\nstarts like so.\n\n    static int mktree_line(unsigned int mode, struct object_id *oid,\n                           enum object_type obj_type, int stage UNUSED,\n                           const char *path, void *cbdata)\n    {\n    \nI _think_ it should be made an error if the input has non-zero\nstage, which would be a sign that it was taken from \"ls-files -s\"\n(or even \"ls-files -u\"), out of which \"git write-tree\" will REFUSE\nto create a tree object.  \"mktree\" should behave the same way, no?\n\nIn any case, here is the documentation split/refactor.\n\n Documentation/git-mktree.txt         |  4 +++-\n Documentation/git-update-index.txt   | 14 +-------------\n Documentation/index-info-formats.txt | 13 +++++++++++++\n 3 files changed, 17 insertions(+), 14 deletions(-)\n\ndiff --git c/Documentation/git-mktree.txt w/Documentation/git-mktree.txt\nindex a660438c67..fefaa83d29 100644\n--- c/Documentation/git-mktree.txt\n+++ w/Documentation/git-mktree.txt\n@@ -48,7 +48,9 @@ OPTIONS\n INPUT FORMAT\n ------------\n Tree entries may be specified in any of the formats compatible with the\n-`--index-info` option to linkgit:git-update-index[1].\n+`--index-info` option to linkgit:git-update-index[1].  That is:\n+\n+include::index-info-formats.txt[]\n \n Entries may use full pathnames containing directory separators to specify\n entries nested within one or more directories. These entries are inserted into\ndiff --git c/Documentation/git-update-index.txt w/Documentation/git-update-index.txt\nindex 7128aed540..2287a5d4be 100644\n--- c/Documentation/git-update-index.txt\n+++ w/Documentation/git-update-index.txt\n@@ -280,19 +280,7 @@ USING --INDEX-INFO\n multiple entry definitions from the standard input, and designed\n specifically for scripts.  It can take inputs of three formats:\n \n-    . mode SP type SP sha1          TAB path\n-+\n-This format is to stuff `git ls-tree` output into the index.\n-\n-    . mode         SP sha1 SP stage TAB path\n-+\n-This format is to put higher order stages into the\n-index file and matches 'git ls-files --stage' output.\n-\n-    . mode         SP sha1          TAB path\n-+\n-This format is no longer produced by any Git command, but is\n-and will continue to be supported by `update-index --index-info`.\n+include::index-info-formats.txt[]\n \n To place a higher stage entry to the index, the path should\n first be removed by feeding a mode=0 entry for the path, and\ndiff --git c/Documentation/index-info-formats.txt w/Documentation/index-info-formats.txt\nnew file mode 100644\nindex 0000000000..037ebd2432\n--- /dev/null\n+++ w/Documentation/index-info-formats.txt\n@@ -0,0 +1,13 @@\n+    . mode SP type SP sha1          TAB path\n++\n+This format is to use `git ls-tree` output.\n+\n+    . mode         SP sha1 SP stage TAB path\n++\n+This format allows higher order stages to appear and\n+matches 'git ls-files --stage' output.\n+\n+    . mode         SP sha1          TAB path\n++\n+This format is no longer produced by any Git command, but is\n+and will continue to be supported.\n"},{"id":"497046","messageId":"dab4b0e3-8000-465e-8f0a-61df3d9168a3@github.com","threadId":"61624","inReplyTo":"ZmltGAPQ2dAfW0kG@tanuki","subject":"Re: [PATCH 10/16] mktree: overwrite duplicate entries","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2024-06-12T18:48:37Z","receivedAt":"2024-06-12T18:48:39Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Patrick Steinhardt wrote:\n> On Tue, Jun 11, 2024 at 06:24:42PM +0000, Victoria Dye via GitGitGadget wrote:\n>> From: Victoria Dye <vdye@github.com>\n>>\n>> If multiple tree entries with the same name are provided as input to\n>> 'mktree', only write the last one to the tree. Entries are considered\n>> duplicates if they have identical names (*not* considering mode); if a blob\n>> and a tree with the same name are provided, only the last one will be\n>> written to the tree. A tree with duplicate entries is invalid (per 'git\n>> fsck'), so that condition should be avoided wherever possible.\n>>\n>> Signed-off-by: Victoria Dye <vdye@github.com>\n>> ---\n>>  Documentation/git-mktree.txt |  8 ++++---\n>>  builtin/mktree.c             | 45 ++++++++++++++++++++++++++++++++----\n>>  t/t1010-mktree.sh            | 36 +++++++++++++++++++++++++++--\n>>  3 files changed, 80 insertions(+), 9 deletions(-)\n>>\n>> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n>> index fb07e40cef0..afbc846d077 100644\n>> --- a/Documentation/git-mktree.txt\n>> +++ b/Documentation/git-mktree.txt\n>> @@ -43,9 +43,11 @@ OPTIONS\n>>  INPUT FORMAT\n>>  ------------\n>>  Tree entries may be specified in any of the formats compatible with the\n>> -`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n>> -entries is normalized by `mktree` so pre-sorting the input by path is not\n>> -required.\n>> +`--index-info` option to linkgit:git-update-index[1].\n>> +\n>> +The order of the tree entries is normalized by `mktree` so pre-sorting the input\n>> +by path is not required. Multiple entries provided with the same path are\n>> +deduplicated, with only the last one specified added to the tree.\n> \n> Hm. I'm not sure whether this is a good idea. With git-mktree(1) being\n> part of our plumbing layer, you can expect that it's mostly going to be\n> fed input from scripts. And any script that generates duplicate tree\n> entries is broken, but we now start to paper over such brokenness\n> without giving the user any indicator of this. As user of git-mktree(1)\n> in Gitaly I can certainly say that I'd rather want to see it die instead\n> of silently fixing my inputs so that I start to notice my own bugs.\n\n'git mktree' already does some cleaning of the inputs by sorting the\nentries, presumably so that a valid tree is created rather than one with\nordering errors. Deduplication is also a cleanup of user inputs to ensure a\nvalid tree is created, so to me it's a consistent extension to existing\nbehavior. Conversely, rejecting the inputs and failing would be introducing\nan error scenario where none existed previously, which to me would be a\nbigger deviation.\n\nOne potential way to get the kind of functionality you're looking for,\nthough, might be to combine something like '--literally' and a '--strict'\nthat validates the tree before writing. Like I mentioned in the cover letter\n[1], I do plan to submit a follow-up series with '--strict' (it's just that\nthis series is already pretty long and it would add 4-ish more patches). \n\n[1] https://lore.kernel.org/git/pull.1746.git.1718130288.gitgitgadget@gmail.com/\n\n> So without seeing a strong motivating usecase for this feature I'd think\n> that git-mktree(1) should reject such inputs and return an error such\n> that the user can fix their tooling.\n\nPractically, there are a couple of reasons that led me to wanting this\nbehavior. One is that it allows using data structures with more rigid\nintegrity checks (like the index & cache tree). The other is that, once the\nability to add nested entries is introduced, the concept of a \"duplicate\"\ngets fuzzier and blocking them entirely could lead to inconsistencies and/or\nlimited flexibility. If, for example a user wants to create a tree with a\ndirectory 'folder1/' with OID '0123456789012345678901234567890123456789',\nbut update a blob 'folder1/file1' in it to OID\n'0987654321098765432109876543210987654321', the latter is technically a\n\"duplicate\" but rejecting it would avoid being able to create the tree\nwithout first expanding 'folder1/'with something like 'ls-tree', replacing the\nappropriate entry, then calling 'mktree'.\n\n> \n> Patrick\n\n\n"},{"id":"497047","messageId":"63fd367e-4246-46a8-9b95-6353a5a54b36@github.com","threadId":"61624","inReplyTo":"xmqq34pirj51.fsf@gitster.g","subject":"Re: [PATCH 09/16] mktree: validate paths more carefully","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2024-06-12T19:01:24Z","receivedAt":"2024-06-12T19:01:26Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Junio C Hamano wrote:\n> \"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> Also,\n>> remove trailing slashes on directories before validating, allowing users to\n>> provide 'folder-name/' as the path for a tree object entry.\n> \n> Is that a good idea for a plumbing like this command?  We would\n> silently accept these after silently stripping the trailing slash?\n> \n> 040000 tree 82a33d5150d9316378ef1955a49f2a5bf21aaeb2    templates/\n> 100644 blob 1f89ffab4c32bc02b5d955851401628a5b9a540e    thread-utils.c/\n> \n> The former _might_ count as \"usability improvement\", but if we are\n> doing the same for the latter we might be going a bit too lenient.\n\nThe trailing slashes are only ignored on tree entries (with mode 040000), so\nthe latter case would not be allowed (and triggers a 'die()' as it would\ntoday).\n\n> \n> Let's see what really happens in the code.\n> \n>> @@ -49,10 +50,23 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n>>  {\n>>  \tstruct tree_entry *ent;\n>>  \tsize_t len = strlen(path);\n>> -\tif (!literally && strchr(path, '/'))\n>> -\t\tdie(\"path %s contains slash\", path);\n>>  \n>> -\tFLEX_ALLOC_MEM(ent, name, path, len);\n>> +\tif (literally) {\n>> +\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n>> +\t} else {\n>> +\t\t/* Normalize and validate entry path */\n>> +\t\tif (S_ISDIR(mode)) {\n>> +\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n>> +\t\t\t\tlen--;\n>> +\t\t}\n> \n> Leave a single SP after \"while\", please.\n\nAh, sorry about that, thanks for catching it.\n\n> We do this only to subtree entries, and all trailing slashes, not\n> just a single one.  OK, but I am not sure if the extra leniency is a\n> good idea to begin with.  \"ls-tree\" output does not have such a\n> trailing slashes, so it is unclear whom we are trying to be extra\n> nice with this.\n\nIt might be a bit niche, but 'git ls-files -s --sparse' does print\ndirectories with a trailing slash, and in a format that is otherwise\naccepted by the command after switching to 'read_index_info' for input\nparsing. \n\n"},{"id":"497048","messageId":"xmqqh6dyosgn.fsf@gitster.g","threadId":"61624","inReplyTo":"63fd367e-4246-46a8-9b95-6353a5a54b36@github.com","subject":"Re: [PATCH 09/16] mktree: validate paths more carefully","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T19:45:44Z","receivedAt":"2024-06-12T19:45:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Victoria Dye <vdye@github.com> writes:\n\n> It might be a bit niche, but 'git ls-files -s --sparse' does print\n> directories with a trailing slash, ...\n\nOK.\n"},{"id":"497049","messageId":"xmqqcyomos8n.fsf@gitster.g","threadId":"61624","inReplyTo":"ZmltKHI-Vz1L44r8@tanuki","subject":"Re: [PATCH 14/16] mktree: optionally add to an existing tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-12T19:50:32Z","receivedAt":"2024-06-12T19:50:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Tue, Jun 11, 2024 at 06:24:46PM +0000, Victoria Dye via GitGitGadget wrote:\n>> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n>> index afbc846d077..99abd3c31a6 100644\n>> --- a/Documentation/git-mktree.txt\n>> +++ b/Documentation/git-mktree.txt\n>> @@ -40,6 +40,11 @@ OPTIONS\n>>  \toptional.  Note - if the `-z` option is used, lines are terminated\n>>  \twith NUL.\n>>  \n>> +<tree-ish>::\n>> +\tIf provided, the tree entries provided in stdin are added to this tree\n>> +\trather than a new empty one, replacing existing entries with identical\n>> +\tnames. Not compatible with `--literally`.\n>\n> I think it'd be a bit more intuitive is this was an option, like\n> `--base-tree=` or just `--base=`.\n>\n> One question that comes up naturally in this context: when I have a base\n> tree, how do I remove entries from it?\n\nPresumably the same way how you remove entries with \"update-index --index-info\"?\nI.e. mode=0 entry in the input would serve as a signal to remove the\npath?\n"},{"id":"497099","messageId":"da036540-531f-4ad8-8be8-93c104930976@github.com","threadId":"61624","inReplyTo":"ZmltI7HA7O4w2E-6@tanuki","subject":"Re: [PATCH 12/16] mktree: use iterator struct to add tree entries to index","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2024-06-13T18:38:54Z","receivedAt":"2024-06-13T18:38:56Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Patrick Steinhardt wrote:\n> On Tue, Jun 11, 2024 at 06:24:44PM +0000, Victoria Dye via GitGitGadget wrote:\n>> From: Victoria Dye <vdye@github.com>\n>>\n>> Create 'struct tree_entry_iterator' to manage iteration through a 'struct\n>> tree_entry_array'. Using an iterator allows for conditional iteration; this\n>> functionality will be necessary in later commits when performing parallel\n>> iteration through multiple sets of tree entries.\n>>\n>> Signed-off-by: Victoria Dye <vdye@github.com>\n>> ---\n>>  builtin/mktree.c | 40 +++++++++++++++++++++++++++++++++++++---\n>>  1 file changed, 37 insertions(+), 3 deletions(-)\n>>\n>> diff --git a/builtin/mktree.c b/builtin/mktree.c\n>> index 12f68187221..bee359e9978 100644\n>> --- a/builtin/mktree.c\n>> +++ b/builtin/mktree.c\n>> @@ -137,6 +137,38 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>>  \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n>>  }\n>>  \n>> +struct tree_entry_iterator {\n>> +\tstruct tree_entry *current;\n>> +\n>> +\t/* private */\n>> +\tstruct {\n>> +\t\tstruct tree_entry_array *arr;\n>> +\t\tsize_t idx;\n>> +\t} priv;\n>> +};\n>> +\n>> +static void init_tree_entry_iterator(struct tree_entry_iterator *iter,\n>> +\t\t\t\t     struct tree_entry_array *arr)\n>> +{\n>> +\titer->priv.arr = arr;\n>> +\titer->priv.idx = 0;\n>> +\titer->current = 0 < arr->nr ? arr->entries[0] : NULL;\n>> +}\n> \n> Nit: Same comment as before, I think these should rather be named\n> `tree_entry_iterator_init()` and `tree_entry_iterator_advance()`.\n\nThat works for me. I'm not attached to the naming convention I used and your\njustification for changing it in [1] is reasonable.\n\n[1] https://lore.kernel.org/git/ZmltDQ5SlVvrEDGP@tanuki/\n\n>> +/*\n>> + * Advance the tree entry iterator to the next entry in the array. If no entries\n>> + * remain, 'current' is set to NULL. Returns the previous 'current' value of the\n>> + * iterator.\n>> + */\n>> +static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator *iter)\n>> +{\n>> +\tstruct tree_entry *prev = iter->current;\n>> +\titer->current = (iter->priv.idx + 1) < iter->priv.arr->nr\n>> +\t\t\t? iter->priv.arr->entries[++iter->priv.idx]\n>> +\t\t\t: NULL;\n>> +\treturn prev;\n>> +}\n> \n> I think it's somewhat confusing to have this return a different value\n> than `current`. When I call `next()`, then I expect the iterator to\n> return the next item. And after having called `next()`, I expect that\n> the current value is the one that the previous call to `next()` has\n> returned.\n\nI do see how it's confusing. I was attempting to mimic the various\narray/stack \"pop\" methods throughout the codebase (which return the \"popped\"\nvalue while moving the stack pointer), but that doesn't really work here\nwith an iterator. \n\nThe only real benefit of this was that it simplified a loop somewhere later\non, but not by a ton. I'll drop the 'tree_entry *' return value from the\nmethod and access 'iter->current' directly where it's needed.\n\n> To avoid confusion, I'd propose to get rid of the `current` member\n> altogether. It's not needed as we already save the current index and\n> avoids the confusion.\n\nThe idea of the iterator is to have callers only ever reference the\n'current' value to avoid needing to deal with the array & current index\ndirectly; I find that it majorly simplifies the parallel iteration through\nthe base tree and entry array in [2]. IOW, in a language with support for\nit, 'idx' would be private & 'current' would be public. So I would like to\nkeep the 'current' value as the publicly-accessible way of interacting with\nthe iterator (although, as mentioned above, I'm happy to drop it from the\n'advance' method return value).\n\n[2] https://lore.kernel.org/git/df0c50dfea3cb77e0070246efdf7a3f070b2ad97.1718130288.git.gitgitgadget@gmail.com/\n\n> \n> Patrick\n\n"},{"id":"497252","messageId":"55e06e1b-580e-45fe-a914-d1ef252d5981@github.com","threadId":"61624","inReplyTo":"ZmltKHI-Vz1L44r8@tanuki","subject":"Re: [PATCH 14/16] mktree: optionally add to an existing tree","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2024-06-17T19:23:39Z","receivedAt":"2024-06-17T19:23:41Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Patrick Steinhardt wrote:\n> On Tue, Jun 11, 2024 at 06:24:46PM +0000, Victoria Dye via GitGitGadget wrote:\n>> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n>> index afbc846d077..99abd3c31a6 100644\n>> --- a/Documentation/git-mktree.txt\n>> +++ b/Documentation/git-mktree.txt\n>> @@ -40,6 +40,11 @@ OPTIONS\n>>  \toptional.  Note - if the `-z` option is used, lines are terminated\n>>  \twith NUL.\n>>  \n>> +<tree-ish>::\n>> +\tIf provided, the tree entries provided in stdin are added to this tree\n>> +\trather than a new empty one, replacing existing entries with identical\n>> +\tnames. Not compatible with `--literally`.\n> \n> I think it'd be a bit more intuitive is this was an option, like\n> `--base-tree=` or just `--base=`.\n\nTo me, the positional '<tree-ish>' is more intuitive; it's reminiscent of\n'read-tree' (but with '--empty' being the default, since there's no\nequivalent to the existing index to overwrite). I consider 'read-tree'\nrelevant in this case because the updated 'mktree' allows a users to create\ntrees like:\n\n$ git read-tree <tree-ish>\n$ git update-index <entries\n$ git write-tree\n\nwithout the intermediate on-disk index. Conversely, there isn't really an\nequivalent option to base the name on ('--base' is a bit overloaded, as it\ntypically refers to a merge/diff base), and I'd like to avoid adding more\npotentially-confusing names to the overall Git UX if I can help it (even if\nthis is a plumbing command).\n\nHowever, looking at other command documentation, I should at least drop\n'[--]' from the usage string. While that is a separator used to signify \"end\nof options\" using 'parse_options()', it's typically only included in the\nusage string to separate different sets of positional arguments (e.g.\nrevisions from pathspecs). \n\n> \n> One question that comes up naturally in this context: when I have a base\n> tree, how do I remove entries from it?\n\nIn patch 16 [1], entries with mode \"0\" are removed from the tree (similar to\n'update-index').\n\n[1] https://lore.kernel.org/git/a90d6d0c943283e9e7bd181cd6e9bb6d4572aaeb.1718130288.git.gitgitgadget@gmail.com/\n\n> \n> Patrick\n\n"},{"id":"497294","messageId":"96f0ca92-b30f-4a7a-bf67-1967d8399177@github.com","threadId":"61624","inReplyTo":"xmqq34pjt7m2.fsf@gitster.g","subject":"Re: [PATCH 05/16] index-info.c: identify empty input lines in read_index_info","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2024-06-18T17:33:03Z","receivedAt":"2024-06-18T17:33:06Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Junio C Hamano wrote:\n> \"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Victoria Dye <vdye@github.com>\n>>\n>> Update 'read_index_info()' to return INDEX_INFO_EMPTY_LINE (value 1), rather\n>> than the default error code (value -1) when the function encounters an empty\n>> line in stdin. This grants the caller the flexibility to handle such\n>> scenarios differently than a typical error. In the case of 'update-index',\n>> we'll still exit with a \"malformed input line\" error. However, when\n>> 'read_index_info()' is used to process the input to 'mktree' in a later\n>> patch, the empty line return value will signal a new tree in --batch mode.\n> \n> Interesting.  We could even introduce \"# commented input\" but that\n> is a different story ;-).\n> \n> I also wonder if we can flip it around and teach read_index_info()\n> to (1) silently accept and do a callback when it recognises the\n> input line is one of the supported formats, and (2) send any\n> unrecognised line, not just an empty one, with \"unrecognised\" status\n> code.  That way, the caller can handle more than single kind of\n> \"special input line\" more easily, perhaps?\n\nThis is an interesting idea. The simplest way to do this would probably to\nhave 'bad_line' stop printing the \"malformed input line\" error and instead\nreturn INDEX_INFO_UNRECOGNIZED_LINE, and pass in a strbuf so that the\n\"malformed\" line is available to the caller.\n\nThat seems simple enough, so I'll include it in V2; I can always revert back\nto INDEX_INFO_EMPTY_LINE if the generalized approach doesn't work out for\nsome reason.\n\n> \n> Thanks.\n\n"},{"id":"497374","messageId":"074dc98acc79e08d07cf4f5c8105b872ec57980c.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 01/17] mktree: use OPT_BOOL","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:49Z","receivedAt":"2024-06-19T21:58:10Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace 'OPT_SET_INT' with 'OPT_BOOL' for the options '--missing' and\n'--batch'. The use of 'OPT_SET_INT' in these options is identical to\n'OPT_BOOL', but 'OPT_BOOL' provides slightly simpler syntax.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 9a22d4e2773..8b19d440747 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -162,8 +162,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n-\t\tOPT_SET_INT( 0 , \"missing\", &allow_missing, N_(\"allow missing objects\"), 1),\n-\t\tOPT_SET_INT( 0 , \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\"), 1),\n+\t\tOPT_BOOL(0, \"missing\", &allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n \n-- \ngitgitgadget\n\n"},{"id":"497375","messageId":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.git.1718130288.gitgitgadget@gmail.com","subject":"[PATCH v2 00/17] mktree: support more flexible usage","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:48Z","receivedAt":"2024-06-19T21:58:10Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"The goal of this series is to make 'git mktree' a much more flexible and\npowerful tool for constructing arbitrary trees in memory without the use of\nan index or worktree. The main additions are:\n\n * Using an optional \"base tree\" to add or replace entries in an existing\n   tree rather than creating a new one from scratch.\n   * Building off of this, having entries with mode \"0\" indicate \"remove\n     this entry, if it exists, from the tree\"\n * Handling tree entries inside of subtrees (e.g., folder1/my-file.txt)\n\nIt also introduces some quality-of-life updates:\n\n * Using the same input parsing as 'update-index' to allow a wider variety\n   of tree entry formats.\n * Adding deduplication of input entries & more thorough validation of\n   inputs (with an option to disable both - plus input sorting - if desired\n   with '--literally').\n\nThe implementation change underpinning the new features is completely\nrevamping how the tree is constructed in memory. Instead of writing a single\ntree object into a strbuf and hashing it into the object database, we\nconstruct an in-core sparse index and write out the root tree, as well as\nany new subtrees, using the cache tree infrastructure.\n\nThe series is organized as follows:\n\n * Commits 1-3 contain miscellaneous small renames/refactors to make the\n   code more readable & prepare for larger refactoring later.\n * Commits 4-7 generalize the input parsing performed by 'read_index_info()'\n   in 'update-index' and update 'mktree' to use it.\n * Commit 8 removes the check on object existence & type from submodule\n   entries.\n * Commit 9 adds the '--literally' option to 'mktree'. Practically, this\n   option allows tests that currently use 'mktree' to generate corrupt trees\n   to continue functioning after we strengthen input validations.\n * Commits 10 & 11 add input path validation & entry deduplication,\n   respectively.\n * Commit 12 replaces the strbuf-to-object tree creation with construction\n   of an in-core index & writing out the cache tree.\n * Commits 13-15 add the ability to add tree entries to an existing \"base\"\n   tree. Takes 3 commits to do it because it requires a bit of finesse\n   around directory/file deduplication and iterating over a tree with\n   'read_tree()' with a parallel iteration over the input tree entries.\n * Commit 16 allows for deeper paths in the input.\n * Commit 17 adds handling for mode '0' as \"removal\" entries.\n\nI also plan to add a '--strict' option that runs 'fsck' checks on the new\ntree(s) before writing to the object database (similar to 'mkttag\n--strict'), but this series is pretty long as it is and that part can easily\nbe separated out into its own series.\n\n\nChanges since V1\n================\n\n * Renamed 'tree_entry_array' & 'tree_entry_iterator' functions to\n   'tree_entry_array' & 'tree_entry_iterator', respectively.\n * Removed the return value from 'tree_entry_iterator_advance'.\n * Added call to 'tree_entry_array_clear' in 'tree_entry_array_release',\n   updated both methods to optionally free tree entries based on a\n   'free_entries' arg.\n * Updated 'read_index_info()':\n   * Replaced INDEX_INFO_EMPTY_LINE with INDEX_INFO_UNRECOGNIZED_LINE which\n     is returned when any \"malformed\" line is found, setting an input strbuf\n     to the line contents for the caller to deal with.\n   * Updated default object type value (when no type info is specified) to\n     'OBJ_ANY' from 'OBJ_NONE'.\n * Updated 'mktree' to 'die()' on malformed line with \"input format error\",\n   rather than 'error()' with \"malformed input line\", to avoid unnecessarily\n   changing the error message.\n * Removed check for object existence & type when the entry type is a\n   submodule, rearranged checks for readability.\n * Replaced use of 'grep' with 'test_grep' in new/updated tests.\n * Wrapped lines in documentation updates to 76 characters.\n * Applied documentation refactor patch in [1].\n * Dropped '[--]' from the 'mktree' usage string.\n\n[1] https://lore.kernel.org/git/xmqqle3aovpq.fsf@gitster.g/\n\nThanks!\n\n * Victoria\n\nVictoria Dye (17):\n  mktree: use OPT_BOOL\n  mktree: rename treeent to tree_entry\n  mktree: use non-static tree_entry array\n  update-index: generalize 'read_index_info'\n  index-info.c: return unrecognized lines to caller\n  index-info.c: parse object type in provided in read_index_info\n  mktree: use read_index_info to read stdin lines\n  mktree.c: do not fail on mismatched submodule type\n  mktree: add a --literally option\n  mktree: validate paths more carefully\n  mktree: overwrite duplicate entries\n  mktree: create tree using an in-core index\n  mktree: use iterator struct to add tree entries to index\n  mktree: add directory-file conflict hashmap\n  mktree: optionally add to an existing tree\n  mktree: allow deeper paths in input\n  mktree: remove entries when mode is 0\n\n Documentation/git-mktree.txt         |  50 ++-\n Documentation/git-update-index.txt   |  16 +-\n Documentation/index-info-formats.txt |  13 +\n Makefile                             |   1 +\n builtin/mktree.c                     | 592 ++++++++++++++++++++++-----\n builtin/update-index.c               | 135 ++----\n index-info.c                         | 100 +++++\n index-info.h                         |  15 +\n t/t1010-mktree.sh                    | 350 +++++++++++++++-\n t/t1014-read-tree-confusing.sh       |   6 +-\n t/t1450-fsck.sh                      |   4 +-\n t/t1601-index-bogus.sh               |   2 +-\n t/t1700-split-index.sh               |   6 +-\n t/t2107-update-index-basic.sh        |  32 ++\n t/t7008-filter-branch-null-sha1.sh   |   6 +-\n t/t7417-submodule-path-url.sh        |   2 +-\n t/t7450-bad-git-dotfiles.sh          |   8 +-\n 17 files changed, 1088 insertions(+), 250 deletions(-)\n create mode 100644 Documentation/index-info-formats.txt\n create mode 100644 index-info.c\n create mode 100644 index-info.h\n\n\nbase-commit: 8d94cfb54504f2ec9edc7ca3eb5c29a3dd3675ae\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1746%2Fvdye%2Fvdye%2Fmktree-recursive-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1746/vdye/vdye/mktree-recursive-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1746\n\nRange-diff vs v1:\n\n  1:  074dc98acc7 =  1:  074dc98acc7 mktree: use OPT_BOOL\n  2:  4558f35e7bf =  2:  4558f35e7bf mktree: rename treeent to tree_entry\n  3:  5ade145352f !  3:  d0d5523a32b mktree: use non-static tree_entry array\n     @@ builtin/mktree.c\n      +\tsize_t nr, alloc;\n      +\tstruct tree_entry **entries;\n      +};\n     - \n     --static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n     ++\n      +static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n      +{\n      +\tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n      +\tarr->entries[arr->nr++] = ent;\n      +}\n      +\n     -+static void clear_tree_entry_array(struct tree_entry_array *arr)\n     ++static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entries)\n      +{\n     -+\tfor (size_t i = 0; i < arr->nr; i++)\n     -+\t\tFREE_AND_NULL(arr->entries[i]);\n     ++\tif (free_entries) {\n     ++\t\tfor (size_t i = 0; i < arr->nr; i++)\n     ++\t\t\tFREE_AND_NULL(arr->entries[i]);\n     ++\t}\n      +\tarr->nr = 0;\n      +}\n     -+\n     -+static void release_tree_entry_array(struct tree_entry_array *arr)\n     + \n     +-static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n     ++static void tree_entry_array_release(struct tree_entry_array *arr, int free_entries)\n      +{\n     ++\ttree_entry_array_clear(arr, free_entries);\n      +\tFREE_AND_NULL(arr->entries);\n     -+\tarr->nr = arr->alloc = 0;\n     ++\tarr->alloc = 0;\n      +}\n      +\n      +static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n     @@ builtin/mktree.c: static void append_to_tree(unsigned mode, struct object_id *oi\n       \n      -\tALLOC_GROW(entries, used + 1, alloc);\n      -\tentries[used++] = ent;\n     -+\t/* Append the update */\n      +\ttree_entry_array_push(arr, ent);\n       }\n       \n     @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n       \t\t\tfflush(stdout);\n       \t\t}\n      -\t\tused=0; /* reset tree entry buffer for re-use in batch mode */\n     -+\t\tclear_tree_entry_array(&arr); /* reset tree entry buffer for re-use in batch mode */\n     ++\t\ttree_entry_array_clear(&arr, 1); /* reset tree entry buffer for re-use in batch mode */\n       \t}\n      +\n     -+\trelease_tree_entry_array(&arr);\n     ++\ttree_entry_array_release(&arr, 1);\n       \tstrbuf_release(&sb);\n       \treturn 0;\n       }\n  4:  9d0689e9c28 !  4:  f5473764236 update-index: generalize 'read_index_info'\n     @@ Commit message\n          callback 'apply_index_info()' to verify the parsed line and update the index\n          according to its contents.\n      \n     -    The input parsing done by 'read_index_info()' is similar to, but more\n     -    flexible than, the parsing done in 'mktree' by 'mktree_line()' (handling not\n     -    only 'git ls-tree' output but also the outputs of 'git apply --index-info'\n     -    and 'git ls-files --stage' outputs). To make 'mktree' more flexible, a later\n     -    patch will replace mktree's custom parsing with 'read_index_info()'.\n     +    Switching to using a callback to validate the parsed entry in 'update-index'\n     +    results in a slight change to the error message indicating a file could not\n     +    be removed from the index. The original implementation uses the raw, quoted\n     +    pathname in the error message, whereas the callback (without access to the\n     +    raw pathname) uses the unquoted value. However, this change makes the failed\n     +    removal message consistent with all other error messages in the function,\n     +    and that consistency is likely more beneficial than not to a user.\n      \n     +    The motivation for this change is to consolidate the already-similar input\n     +    parsing logic in 'git update-index' and 'git mktree', avoiding code\n     +    duplication and the associated maintenance burden. The input formats\n     +    accepted by 'update-index' are a superset of those accepted by 'mktree', so\n     +    in a later commit we can replace the input parsing of the latter with\n     +    'read_index_info()' without breaking existing usage.\n     +\n     +    Co-authored-by: Junio C Hamano <gitster@pobox.com>\n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n     + ## Documentation/git-update-index.txt ##\n     +@@ Documentation/git-update-index.txt: USING --INDEX-INFO\n     + \n     + `--index-info` is a more powerful mechanism that lets you feed\n     + multiple entry definitions from the standard input, and designed\n     +-specifically for scripts.  It can take inputs of three formats:\n     ++specifically for scripts.  It can take inputs in the following formats:\n     + \n     +-    . mode SP type SP sha1          TAB path\n     +-+\n     +-This format is to stuff `git ls-tree` output into the index.\n     +-\n     +-    . mode         SP sha1 SP stage TAB path\n     +-+\n     +-This format is to put higher order stages into the\n     +-index file and matches 'git ls-files --stage' output.\n     +-\n     +-    . mode         SP sha1          TAB path\n     +-+\n     +-This format is no longer produced by any Git command, but is\n     +-and will continue to be supported by `update-index --index-info`.\n     ++include::index-info-formats.txt[]\n     + \n     + To place a higher stage entry to the index, the path should\n     + first be removed by feeding a mode=0 entry for the path, and\n     +\n     + ## Documentation/index-info-formats.txt (new) ##\n     +@@\n     ++    . mode SP type SP sha1          TAB path\n     +++\n     ++This format is to use `git ls-tree` output.\n     ++\n     ++    . mode         SP sha1 SP stage TAB path\n     +++\n     ++This format allows higher order stages to appear and\n     ++matches 'git ls-files --stage' output.\n     ++\n     ++    . mode         SP sha1          TAB path\n     +++\n     ++This format is no longer produced by any Git command, but is\n     ++and will continue to be supported.\n     +\n       ## Makefile ##\n      @@ Makefile: LIB_OBJS += hex.o\n       LIB_OBJS += hex-ll.o\n     @@ builtin/update-index.c: static void update_one(const char *path)\n       \n       static const char * const update_index_usage[] = {\n      @@ builtin/update-index.c: static enum parse_opt_result stdin_cacheinfo_callback(\n     + \tstruct parse_opt_ctx_t *ctx, const struct option *opt,\n       \tconst char *arg, int unset)\n       {\n     - \tint *nul_term_line = opt->value;\n     -+\tint ret;\n     +-\tint *nul_term_line = opt->value;\n     ++\tint ret = 0;\n       \n       \tBUG_ON_OPT_NEG(unset);\n       \tBUG_ON_OPT_ARG(arg);\n     -@@ builtin/update-index.c: static enum parse_opt_result stdin_cacheinfo_callback(\n     - \tif (ctx->argc != 1)\n     - \t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n     - \tallow_add = allow_replace = allow_remove = 1;\n     + \n     +-\tif (ctx->argc != 1)\n     +-\t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n     +-\tallow_add = allow_replace = allow_remove = 1;\n      -\tread_index_info(*nul_term_line);\n     -+\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n     -+\tif (ret)\n     -+\t\treturn -1;\n     +-\treturn 0;\n     ++\tif (ctx->argc != 1) {\n     ++\t\tret = error(\"option '%s' must be the last argument\", opt->long_name);\n     ++\t} else {\n     ++\t\tint *nul_term_line = opt->value;\n      +\n     - \treturn 0;\n     ++\t\tallow_add = allow_replace = allow_remove = 1;\n     ++\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n     ++\t\tif (ret)\n     ++\t\t\tret = -1;\n     ++\t}\n     ++\n     ++\treturn ret;\n       }\n       \n     + static enum parse_opt_result stdin_callback(\n      \n       ## index-info.c (new) ##\n      @@\n     @@ index-info.c (new)\n      +\t\tcontinue;\n      +\n      +\tbad_line:\n     -+\t\tret = error(\"malformed input line '%s'\", buf.buf);\n     -+\t\tbreak;\n     ++\t\tdie(\"malformed input line '%s'\", buf.buf);\n      +\t}\n      +\tstrbuf_release(&buf);\n      +\tstrbuf_release(&uq);\n     @@ t/t2107-update-index-basic.sh: test_expect_success '--index-version' '\n      +\t# empty line\n      +\techo \"\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"malformed input line\" err &&\n      +\n      +\t# bad whitespace\n      +\tprintf \"100644 $EMPTY_BLOB A\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"malformed input line\" err &&\n      +\n      +\t# invalid stage value\n      +\tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"malformed input line\" err &&\n      +\n      +\t# invalid OID length\n      +\tprintf \"100755 abc123\\tA\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"malformed input line\" err &&\n      +\n      +\t# bad quoting\n      +\tprintf \"100644 $EMPTY_BLOB\\t\\\"A\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"bad quoting of path name\" err\n     ++\ttest_grep \"bad quoting of path name\" err\n      +'\n      +\n       test_done\n  5:  7e3bcc16e23 !  5:  4f4d54c8d07 index-info.c: identify empty input lines in read_index_info\n     @@ Metadata\n      Author: Victoria Dye <vdye@github.com>\n      \n       ## Commit message ##\n     -    index-info.c: identify empty input lines in read_index_info\n     +    index-info.c: return unrecognized lines to caller\n      \n     -    Update 'read_index_info()' to return INDEX_INFO_EMPTY_LINE (value 1), rather\n     -    than the default error code (value -1) when the function encounters an empty\n     -    line in stdin. This grants the caller the flexibility to handle such\n     -    scenarios differently than a typical error. In the case of 'update-index',\n     -    we'll still exit with a \"malformed input line\" error. However, when\n     -    'read_index_info()' is used to process the input to 'mktree' in a later\n     -    patch, the empty line return value will signal a new tree in --batch mode.\n     +    Update 'read_index_info()' to return INDEX_INFO_UNRECOGNIZED_LINE (value 1),\n     +    rather than die()-ing when the function encounters a line that cannot be\n     +    parsed according to one of the accepted formats. This grants the caller the\n     +    flexibility to fall back on custom handling for such lines rather than a\n     +    returning a catch-all error. In the case of 'update-index', we'll still exit\n     +    with a \"malformed input line\" error. However, when 'read_index_info()' is\n     +    used to process the input to 'mktree' in a later patch, an empty line return\n     +    value will signal a new tree in --batch mode.\n      \n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n       ## builtin/update-index.c ##\n      @@ builtin/update-index.c: static enum parse_opt_result stdin_cacheinfo_callback(\n     - \t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n     - \tallow_add = allow_replace = allow_remove = 1;\n     - \tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n     --\tif (ret)\n     -+\tif (ret == INDEX_INFO_EMPTY_LINE)\n     -+\t\treturn error(\"malformed input line ''\");\n     -+\telse if (ret < 0)\n     - \t\treturn -1;\n     - \n     - \treturn 0;\n     + \t\tret = error(\"option '%s' must be the last argument\", opt->long_name);\n     + \t} else {\n     + \t\tint *nul_term_line = opt->value;\n     ++\t\tstruct strbuf line = STRBUF_INIT;\n     + \n     + \t\tallow_add = allow_replace = allow_remove = 1;\n     +-\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n     +-\t\tif (ret)\n     ++\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL, &line);\n     ++\n     ++\t\tif (ret == INDEX_INFO_UNRECOGNIZED_LINE)\n     ++\t\t\tret = error(\"malformed input line '%s'\", line.buf);\n     ++\t\telse if (ret)\n     + \t\t\tret = -1;\n     ++\t\tstrbuf_release(&line);\n     + \t}\n     + \n     + \treturn ret;\n      \n       ## index-info.c ##\n     +@@\n     + #include \"strbuf.h\"\n     + #include \"quote.h\"\n     + \n     +-int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     ++int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n     ++\t\t    struct strbuf *line)\n     + {\n     + \tconst int hexsz = the_hash_algo->hexsz;\n     +-\tstruct strbuf buf = STRBUF_INIT;\n     + \tstruct strbuf uq = STRBUF_INIT;\n     + \tstrbuf_getline_fn getline_fn;\n     + \tint ret = 0;\n     + \n     + \tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n     +-\twhile (getline_fn(&buf, stdin) != EOF) {\n     ++\twhile (getline_fn(line, stdin) != EOF) {\n     + \t\tchar *ptr, *tab;\n     + \t\tchar *path_name;\n     + \t\tstruct object_id oid;\n     +@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     + \t\t * index file and matches \"git ls-files --stage\" output.\n     + \t\t */\n     + \t\terrno = 0;\n     +-\t\tul = strtoul(buf.buf, &ptr, 8);\n     +-\t\tif (ptr == buf.buf || *ptr != ' '\n     ++\t\tul = strtoul(line->buf, &ptr, 8);\n     ++\t\tif (ptr == line->buf || *ptr != ' '\n     + \t\t    || errno || (unsigned int) ul != ul)\n     + \t\t\tgoto bad_line;\n     + \t\tmode = ul;\n      @@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     - \t\tunsigned long ul;\n     - \t\tint stage;\n     + \t\tcontinue;\n       \n     -+\t\tif (!buf.len) {\n     -+\t\t\tret = INDEX_INFO_EMPTY_LINE;\n     -+\t\t\tbreak;\n     -+\t\t}\n     -+\n     - \t\t/* This reads lines formatted in one of three formats:\n     - \t\t *\n     - \t\t * (1) mode         SP sha1          TAB path\n     + \tbad_line:\n     +-\t\tdie(\"malformed input line '%s'\", buf.buf);\n     ++\t\tret = INDEX_INFO_UNRECOGNIZED_LINE;\n     ++\t\tbreak;\n     + \t}\n     +-\tstrbuf_release(&buf);\n     + \tstrbuf_release(&uq);\n     ++\tif (!ret)\n     ++\t\tstrbuf_reset(line);\n     + \n     + \treturn ret;\n     + }\n      \n       ## index-info.h ##\n      @@\n       \n       typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n       \n     -+#define INDEX_INFO_EMPTY_LINE 1\n     ++#define INDEX_INFO_UNRECOGNIZED_LINE 1\n      +\n       /* Iterate over parsed index info from stdin */\n     - int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n     +-int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n     ++int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n     ++\t\t    struct strbuf *line);\n       \n     + #endif /* INDEX_INFO_H */\n  6:  f56eee0b48d !  6:  472efcaf1dd index-info.c: parse object type in provided in read_index_info\n     @@ Commit message\n          by 'read_index_info()' (i.e. on lines formatted like the output of 'git\n          ls-tree'), parse it into an 'enum object_type' and provide it to the\n          'read_index_info()' callback as an argument. If the type is not provided,\n     -    pass 'OBJ_NONE' instead. If the object type is invalid, return an error.\n     +    pass 'OBJ_ANY' instead. If the object type is invalid, return an error.\n      \n          The goal of this change is to allow for more thorough validation of the\n          provided object type (e.g. against the provided mode) in 'mktree' once\n     @@ builtin/update-index.c: static void update_one(const char *path)\n       \tif (!verify_path(path_name, mode)) {\n      \n       ## index-info.c ##\n     -@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     +@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n       \t\tchar *ptr, *tab;\n       \t\tchar *path_name;\n       \t\tstruct object_id oid;\n     -+\t\tenum object_type obj_type = OBJ_NONE;\n     ++\t\tenum object_type obj_type = OBJ_ANY;\n       \t\tunsigned int mode;\n       \t\tunsigned long ul;\n       \t\tint stage;\n     -@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     +@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n       \n       \t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n       \t\t\tstage = tab[-1] - '0';\n     @@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void\n       \t\tif (!nul_term_line && path_name[0] == '\"') {\n       \t\t\tstrbuf_reset(&uq);\n       \t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n     -@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n     +@@ index-info.c: int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n       \t\t\tpath_name = uq.buf;\n       \t\t}\n       \n     @@ index-info.h\n      -typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n      +typedef int (*each_index_info_fn)(unsigned int, struct object_id *, enum object_type, int, const char *, void *);\n       \n     - #define INDEX_INFO_EMPTY_LINE 1\n     + #define INDEX_INFO_UNRECOGNIZED_LINE 1\n       \n      \n       ## t/t2107-update-index-basic.sh ##\n      @@ t/t2107-update-index-basic.sh: test_expect_success '--index-info fails on malformed input' '\n       \ttest_must_fail git update-index --index-info 2>err &&\n     - \tgrep \"malformed input line\" err &&\n     + \ttest_grep \"malformed input line\" err &&\n       \n      +\t# invalid type\n      +\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n      +\ttest_must_fail git update-index --index-info 2>err &&\n     -+\tgrep \"invalid object type\" err &&\n     ++\ttest_grep \"invalid object type\" err &&\n      +\n       \t# invalid stage value\n       \tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n  7:  8d1e1eaa70b !  7:  9dc8e16a7fc mktree: use read_index_info to read stdin lines\n     @@ Commit message\n          across the commands (avoiding the need for two similar implementations for\n          input parsing) and adds flexibility to mktree.\n      \n     +    It should be noted that, while the error messages are largely preserved in\n     +    the refactor, one does change: \"fatal: invalid quoting\" is now \"error: bad\n     +    quoting of path name\".\n     +\n          Update 'Documentation/git-mktree.txt' to reflect the more permissive input\n     -    format.\n     +    format, as well as make a note about rejecting stage values higher than 0.\n      \n     +    Helped-by: Junio C Hamano <gitster@pobox.com>\n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n       ## Documentation/git-mktree.txt ##\n     @@ Documentation/git-mktree.txt: SYNOPSIS\n      -a tree object.  The order of the tree entries is normalized by mktree so\n      -pre-sorting the input is not required.  The object name of the tree object\n      -built is written to the standard output.\n     -+Reads entry information from stdin and creates a tree object from those entries.\n     -+The object name of the tree object built is written to the standard output.\n     ++Reads entry information from stdin and creates a tree object from those\n     ++entries. The object name of the tree object built is written to the standard\n     ++output.\n       \n       OPTIONS\n       -------\n     @@ Documentation/git-mktree.txt: OPTIONS\n      +INPUT FORMAT\n      +------------\n      +Tree entries may be specified in any of the formats compatible with the\n     -+`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n     -+entries is normalized by `mktree` so pre-sorting the input by path is not\n     -+required.\n     ++`--index-info` option to linkgit:git-update-index[1]:\n     ++\n     ++include::index-info-formats.txt[]\n     ++\n     ++Note that if the `stage` of a tree entry is given, the value must be 0.\n     ++Higher stages represent conflicted files in an index; this information\n     ++cannot be represented in a tree object. The command will fail without\n     ++writing the tree if a higher order stage is specified for any entry.\n     ++\n     ++The order of the tree entries is normalized by `mktree` so pre-sorting the\n     ++input by path is not required.\n      +\n       GIT\n       ---\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      +};\n      +\n      +static int mktree_line(unsigned int mode, struct object_id *oid,\n     -+\t\t       enum object_type obj_type, int stage UNUSED,\n     ++\t\t       enum object_type obj_type, int stage,\n      +\t\t       const char *path, void *cbdata)\n       {\n      -\tchar *ptr, *ntr;\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      -\t\t\tdie(\"invalid quoting\");\n      -\t\tpath = to_free = strbuf_detach(&p_uq, NULL);\n      -\t}\n     -+\tif (obj_type && mode_type != obj_type)\n     -+\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n     -+\t\t    type_name(obj_type), type_name(mode_type));\n     ++\tif (stage)\n     ++\t\tdie(_(\"path '%s' is unmerged\"), path);\n       \n      -\t/*\n      -\t * Object type is redundantly derivable three ways.\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      -\t\tdie(\"entry '%s' object type (%s) doesn't match mode type (%s)\",\n      -\t\t\tpath, ptr, type_name(mode_type));\n      -\t}\n     ++\tif (obj_type != OBJ_ANY && mode_type != obj_type)\n     ++\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n     ++\t\t    type_name(obj_type), type_name(mode_type));\n     ++\n      +\toi.typep = &parsed_obj_type;\n       \n      -\t/* Check the type of object identified by oid without fetching objects */\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n       \t\t\t\t     OBJECT_INFO_QUICK |\n       \t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n      -\t\tobj_type = -1;\n     -+\t\tparsed_obj_type = -1;\n     - \n     +-\n      -\tif (obj_type < 0) {\n      -\t\tif (allow_missing) {\n      -\t\t\t; /* no problem - missing objects are presumed to be of the right type */\n     -+\tif (parsed_obj_type < 0) {\n     -+\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n     -+\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n     - \t\t} else {\n     +-\t\t} else {\n      -\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(&oid));\n      -\t\t}\n      -\t} else {\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      -\t\t\t */\n      -\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n      -\t\t\t\tpath, oid_to_hex(&oid), type_name(obj_type), type_name(mode_type));\n     +-\t\t}\n     ++\t\tparsed_obj_type = -1;\n     ++\n     ++\tif (parsed_obj_type < 0) {\n     ++\t\t/*\n     ++\t\t * There are two conditions where the object being missing\n     ++\t\t * is acceptable:\n     ++\t\t *\n     ++\t\t * - We're explicitly allowing it with --missing.\n     ++\t\t * - The object is a submodule, which we wouldn't expect to\n     ++\t\t *   be in this repo anyway.\n     ++\t\t *\n     ++\t\t * If neither condition is met, die().\n     ++\t\t */\n     ++\t\tif (!data->allow_missing && !S_ISGITLINK(mode))\n      +\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n     - \t\t}\n     ++\n      +\t} else if (parsed_obj_type != mode_type) {\n      +\t\t/*\n      +\t\t * The object exists but is of the wrong type.\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n       \tstruct tree_entry_array arr = { 0 };\n      -\tstrbuf_getline_fn getline_fn;\n      +\tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n     ++\tstruct strbuf line = STRBUF_INIT;\n      +\tint ret;\n       \n       \tconst struct option option[] = {\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      -\t\t\t\tbreak;\n      -\t\t\t}\n      -\t\t\tif (sb.buf[0] == '\\0') {\n     --\t\t\t\t/* empty lines denote tree boundaries in batch mode */\n     --\t\t\t\tif (is_batch_mode)\n     --\t\t\t\t\tbreak;\n     --\t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n     --\t\t\t}\n     --\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n     --\t\t}\n     --\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n      +\n      +\tdo {\n     -+\t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data);\n     ++\t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data, &line);\n      +\t\tif (ret < 0)\n      +\t\t\tbreak;\n      +\n     -+\t\t/* empty lines denote tree boundaries in batch mode */\n     -+\t\tif (ret > 0 && !is_batch_mode)\n     -+\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n     ++\t\tif (ret == INDEX_INFO_UNRECOGNIZED_LINE) {\n     ++\t\t\tif (line.len)\n     ++\t\t\t\tdie(\"input format error: %s\", line.buf);\n     ++\t\t\telse if (!is_batch_mode)\n     + \t\t\t\t/* empty lines denote tree boundaries in batch mode */\n     +-\t\t\t\tif (is_batch_mode)\n     +-\t\t\t\t\tbreak;\n     + \t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n     +-\t\t\t}\n     +-\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n     + \t\t}\n     +-\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n      +\n      +\t\tif (is_batch_mode && !ret && arr.nr < 1) {\n       \t\t\t/*\n     @@ builtin/mktree.c: static const char *mktree_usage[] = {\n      @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n       \t\t\tfflush(stdout);\n       \t\t}\n     - \t\tclear_tree_entry_array(&arr); /* reset tree entry buffer for re-use in batch mode */\n     + \t\ttree_entry_array_clear(&arr, 1); /* reset tree entry buffer for re-use in batch mode */\n      -\t}\n      +\t} while (ret > 0);\n       \n     - \trelease_tree_entry_array(&arr);\n     ++\tstrbuf_release(&line);\n     + \ttree_entry_array_release(&arr, 1);\n      -\tstrbuf_release(&sb);\n      -\treturn 0;\n      +\treturn !!ret;\n     @@ t/t1010-mktree.sh: test_expect_success 'ls-tree output in wrong order given to m\n      +\ttree_oid=\"$(cat tree)\" &&\n      +\tprintf \"160000 commit $tree_oid\\tA\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"object $tree_oid is a tree but specified type was (commit)\" err\n     ++\ttest_grep \"object $tree_oid is a tree but specified type was (commit)\" err\n      +'\n      +\n       test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree refuses to read ls-tree -r output\n      +\t# empty line without --batch\n      +\techo \"\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"blank line only valid in batch mode\" err &&\n     ++\ttest_grep \"blank line only valid in batch mode\" err &&\n      +\n      +\t# bad whitespace\n      +\tprintf \"100644 blob $EMPTY_BLOB A\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"input format error\" err &&\n      +\n      +\t# invalid type\n      +\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"invalid object type\" err &&\n     ++\ttest_grep \"invalid object type\" err &&\n      +\n      +\t# invalid OID length\n      +\tprintf \"100755 blob abc123\\tA\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"malformed input line\" err &&\n     ++\ttest_grep \"input format error\" err &&\n      +\n      +\t# bad quoting\n      +\tprintf \"100644 blob $EMPTY_BLOB\\t\\\"A\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"bad quoting of path name\" err\n     ++\ttest_grep \"bad quoting of path name\" err\n      +'\n      +\n      +test_expect_success 'mktree fails on mode mismatch' '\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree refuses to read ls-tree -r output\n      +\t# mode-type mismatch\n      +\tprintf \"100644 tree $tree_oid\\tA\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"object type (tree) doesn${SQ}t match mode type (blob)\" err &&\n     ++\ttest_grep \"object type (tree) doesn${SQ}t match mode type (blob)\" err &&\n      +\n      +\t# mode-object mismatch (no --missing)\n      +\tprintf \"100644 $tree_oid\\tA\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"object $tree_oid is a tree but specified type was (blob)\" err\n     ++\ttest_grep \"object $tree_oid is a tree but specified type was (blob)\" err\n      +'\n      +\n       test_done\n  -:  ----------- >  8:  8a3264afd0c mktree.c: do not fail on mismatched submodule type\n  8:  b497dc90687 !  9:  e640a385b3d mktree: add a --literally option\n     @@ Documentation/git-mktree.txt: OPTIONS\n       \n      +--literally::\n      +\tCreate the tree from the tree entries provided to stdin in the order\n     -+\tthey are provided without performing additional sorting, deduplication,\n     -+\tor path validation on them. This option is primarily useful for creating\n     -+\tinvalid tree objects to use in tests of how Git deals with various forms\n     -+\tof tree corruption.\n     ++\tthey are provided without performing additional sorting,\n     ++\tdeduplication, or path validation on them. This option is primarily\n     ++\tuseful for creating invalid tree objects to use in tests of how Git\n     ++\tdeals with various forms of tree corruption.\n      +\n       --batch::\n       \tAllow building of more than one tree object before exiting.  Each\n       \ttree is separated by a single blank line. The final newline is\n      \n       ## builtin/mktree.c ##\n     -@@ builtin/mktree.c: static void release_tree_entry_array(struct tree_entry_array *arr)\n     +@@ builtin/mktree.c: static void tree_entry_array_release(struct tree_entry_array *arr, int free_entr\n       }\n       \n       static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n     @@ builtin/mktree.c: static void write_tree(struct tree_entry_array *arr, struct ob\n       \n       static int mktree_line(unsigned int mode, struct object_id *oid,\n      @@ builtin/mktree.c: static int mktree_line(unsigned int mode, struct object_id *oid,\n     - \t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n     + \t\t}\n       \t}\n       \n      -\tappend_to_tree(mode, oid, path, data->arr);\n     @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n      \n       ## t/t1010-mktree.sh ##\n      @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on mode mismatch' '\n     - \tgrep \"object $tree_oid is a tree but specified type was (blob)\" err\n     + \ttest_grep \"object $tree_oid is a tree but specified type was (blob)\" err\n       '\n       \n      +test_expect_success '--literally can create invalid trees' '\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on mode mismatch' '\n      +\t} | git mktree --literally >tree.bad &&\n      +\tgit cat-file tree $(cat tree.bad) >top.bad &&\n      +\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n     -+\tgrep \"contains duplicate file entries\" err &&\n     ++\ttest_grep \"contains duplicate file entries\" err &&\n      +\n      +\t# disallowed path\n      +\t{\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on mode mismatch' '\n      +\t} | git mktree --literally >tree.bad &&\n      +\tgit cat-file tree $(cat tree.bad) >top.bad &&\n      +\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n     -+\tgrep \"contains ${SQ}.git${SQ}\" err &&\n     ++\ttest_grep \"contains ${SQ}.git${SQ}\" err &&\n      +\n      +\t# nested entry\n      +\t{\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on mode mismatch' '\n      +\t} | git mktree --literally >tree.bad &&\n      +\tgit cat-file tree $(cat tree.bad) >top.bad &&\n      +\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n     -+\tgrep \"contains full pathnames\" err &&\n     ++\ttest_grep \"contains full pathnames\" err &&\n      +\n      +\t# bad entry ordering\n      +\t{\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on mode mismatch' '\n      +\t} | git mktree --literally >tree.bad &&\n      +\tgit cat-file tree $(cat tree.bad) >top.bad &&\n      +\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n     -+\tgrep \"not properly sorted\" err\n     ++\ttest_grep \"not properly sorted\" err\n      +'\n      +\n       test_done\n  9:  4f9f77e693c ! 10:  2eb207064f8 mktree: validate paths more carefully\n     @@ builtin/mktree.c: static void append_to_tree(unsigned mode, struct object_id *oi\n      \n       ## t/t1010-mktree.sh ##\n      @@ t/t1010-mktree.sh: test_expect_success '--literally can create invalid trees' '\n     - \tgrep \"not properly sorted\" err\n     + \ttest_grep \"not properly sorted\" err\n       '\n       \n      +test_expect_success 'mktree validates path' '\n     @@ t/t1010-mktree.sh: test_expect_success '--literally can create invalid trees' '\n      +\t# Invalid: blob with trailing slash\n      +\tprintf \"100644 blob $blob_oid\\ttest/\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"invalid path ${SQ}test/${SQ}\" err &&\n     ++\ttest_grep \"invalid path ${SQ}test/${SQ}\" err &&\n      +\n      +\t# Invalid: dotdot\n      +\tprintf \"040000 tree $tree_oid\\t../\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"invalid path ${SQ}../${SQ}\" err &&\n     ++\ttest_grep \"invalid path ${SQ}../${SQ}\" err &&\n      +\n      +\t# Invalid: dot\n      +\tprintf \"040000 tree $tree_oid\\t.\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"invalid path ${SQ}.${SQ}\" err &&\n     ++\ttest_grep \"invalid path ${SQ}.${SQ}\" err &&\n      +\n      +\t# Invalid: .git\n      +\tprintf \"040000 tree $tree_oid\\t.git/\" |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"invalid path ${SQ}.git/${SQ}\" err\n     ++\ttest_grep \"invalid path ${SQ}.git/${SQ}\" err\n      +'\n      +\n       test_done\n 10:  b59a4ad8ab4 ! 11:  fb555658057 mktree: overwrite duplicate entries\n     @@ Commit message\n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n       ## Documentation/git-mktree.txt ##\n     -@@ Documentation/git-mktree.txt: OPTIONS\n     - INPUT FORMAT\n     - ------------\n     - Tree entries may be specified in any of the formats compatible with the\n     --`--index-info` option to linkgit:git-update-index[1]. The order of the tree\n     --entries is normalized by `mktree` so pre-sorting the input by path is not\n     --required.\n     -+`--index-info` option to linkgit:git-update-index[1].\n     -+\n     -+The order of the tree entries is normalized by `mktree` so pre-sorting the input\n     -+by path is not required. Multiple entries provided with the same path are\n     -+deduplicated, with only the last one specified added to the tree.\n     +@@ Documentation/git-mktree.txt: cannot be represented in a tree object. The command will fail without\n     + writing the tree if a higher order stage is specified for any entry.\n     + \n     + The order of the tree entries is normalized by `mktree` so pre-sorting the\n     +-input by path is not required.\n     ++input by path is not required. Multiple entries provided with the same path\n     ++are deduplicated, with only the last one specified added to the tree.\n       \n       GIT\n       ---\n     @@ builtin/mktree.c\n       \tstruct object_id oid;\n       \tint len;\n      @@ builtin/mktree.c: static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n     + \tent->len = len;\n       \toidcpy(&ent->oid, oid);\n       \n     - \t/* Append the update */\n      +\tent->order = arr->nr;\n       \ttree_entry_array_push(arr, ent);\n       }\n     @@ t/t1010-mktree.sh: test_expect_success '--literally can create invalid trees' '\n       \n       \t# Valid: tree with or without trailing slash, blob without trailing slash\n      @@ t/t1010-mktree.sh: test_expect_success 'mktree validates path' '\n     - \tgrep \"invalid path ${SQ}.git/${SQ}\" err\n     + \ttest_grep \"invalid path ${SQ}.git/${SQ}\" err\n       '\n       \n      +test_expect_success 'mktree with duplicate entries' '\n 11:  130413f2404 = 12:  2333775ba5b mktree: create tree using an in-core index\n 12:  94d6615d634 ! 13:  56f28efff54 mktree: use iterator struct to add tree entries to index\n     @@ builtin/mktree.c: static void sort_and_dedup_tree_entry_array(struct tree_entry_\n      +\t} priv;\n      +};\n      +\n     -+static void init_tree_entry_iterator(struct tree_entry_iterator *iter,\n     ++static void tree_entry_iterator_init(struct tree_entry_iterator *iter,\n      +\t\t\t\t     struct tree_entry_array *arr)\n      +{\n      +\titer->priv.arr = arr;\n     @@ builtin/mktree.c: static void sort_and_dedup_tree_entry_array(struct tree_entry_\n      +}\n      +\n      +/*\n     -+ * Advance the tree entry iterator to the next entry in the array. If no entries\n     -+ * remain, 'current' is set to NULL. Returns the previous 'current' value of the\n     -+ * iterator.\n     ++ * Advance the tree entry iterator to the next entry in the array. If no\n     ++ * entries remain, 'current' is set to NULL.\n      + */\n     -+static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator *iter)\n     ++static void tree_entry_iterator_advance(struct tree_entry_iterator *iter)\n      +{\n     -+\tstruct tree_entry *prev = iter->current;\n      +\titer->current = (iter->priv.idx + 1) < iter->priv.arr->nr\n      +\t\t\t? iter->priv.arr->entries[++iter->priv.idx]\n      +\t\t\t: NULL;\n     -+\treturn prev;\n      +}\n      +\n       static int add_tree_entry_to_index(struct index_state *istate,\n     @@ builtin/mktree.c: static int add_tree_entry_to_index(struct index_state *istate,\n       static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n       {\n      +\tstruct tree_entry_iterator iter = { NULL };\n     -+\tstruct tree_entry *ent;\n       \tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n       \tistate.sparse_index = 1;\n       \n     @@ builtin/mktree.c: static int add_tree_entry_to_index(struct index_state *istate,\n      -\t/* Construct an in-memory index from the provided entries */\n      -\tfor (size_t i = 0; i < arr->nr; i++) {\n      -\t\tstruct tree_entry *ent = arr->entries[i];\n     -+\tinit_tree_entry_iterator(&iter, arr);\n     - \n     ++\ttree_entry_iterator_init(&iter, arr);\n     ++\n      +\t/* Construct an in-memory index from the provided entries & base tree */\n     -+\twhile ((ent = advance_tree_entry_iterator(&iter))) {\n     ++\twhile (iter.current) {\n     ++\t\tstruct tree_entry *ent = iter.current;\n     ++\t\ttree_entry_iterator_advance(&iter);\n     + \n       \t\tif (add_tree_entry_to_index(&istate, ent))\n       \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n     - \t}\n 13:  68acdd3c5ee ! 14:  6f6d78ae7ac mktree: add directory-file conflict hashmap\n     @@ builtin/mktree.c: static inline size_t df_path_len(size_t pathlen, unsigned int\n      +\t       name_compare(e1->name, e1_len, e2->name, e2_len);\n      +}\n      +\n     -+static void init_tree_entry_array(struct tree_entry_array *arr)\n     ++static void tree_entry_array_init(struct tree_entry_array *arr)\n      +{\n      +\thashmap_init(&arr->df_name_hash, df_name_hash_cmp, NULL, 0);\n      +}\n     @@ builtin/mktree.c: static inline size_t df_path_len(size_t pathlen, unsigned int\n       static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n       {\n       \tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n     -@@ builtin/mktree.c: static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entr\n     - \n     - static void clear_tree_entry_array(struct tree_entry_array *arr)\n     - {\n     -+\thashmap_clear(&arr->df_name_hash);\n     - \tfor (size_t i = 0; i < arr->nr; i++)\n     - \t\tFREE_AND_NULL(arr->entries[i]);\n     +@@ builtin/mktree.c: static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entrie\n     + \t\t\tFREE_AND_NULL(arr->entries[i]);\n     + \t}\n       \tarr->nr = 0;\n     -@@ builtin/mktree.c: static void clear_tree_entry_array(struct tree_entry_array *arr)\n     - \n     - static void release_tree_entry_array(struct tree_entry_array *arr)\n     - {\n      +\thashmap_clear(&arr->df_name_hash);\n     - \tFREE_AND_NULL(arr->entries);\n     - \tarr->nr = arr->alloc = 0;\n       }\n     + \n     + static void tree_entry_array_release(struct tree_entry_array *arr, int free_entries)\n      @@ builtin/mktree.c: static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n       \t/* Sort again to order the entries for tree insertion */\n       \tignore_mode = 0;\n     @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n       \n       \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n       \n     -+\tinit_tree_entry_array(&arr);\n     ++\ttree_entry_array_init(&arr);\n      +\n       \tdo {\n     - \t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data);\n     + \t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data, &line);\n       \t\tif (ret < 0)\n 14:  df0c50dfea3 ! 15:  4b88f84b933 mktree: optionally add to an existing tree\n     @@ Documentation/git-mktree.txt: git-mktree - Build a tree-object from formatted tr\n       --------\n       [verse]\n      -'git mktree' [-z] [--missing] [--literally] [--batch]\n     -+'git mktree' [-z] [--missing] [--literally] [--batch] [--] [<tree-ish>]\n     ++'git mktree' [-z] [--missing] [--literally] [--batch] [<tree-ish>]\n       \n       DESCRIPTION\n       -----------\n     @@ Documentation/git-mktree.txt: OPTIONS\n       \twith NUL.\n       \n      +<tree-ish>::\n     -+\tIf provided, the tree entries provided in stdin are added to this tree\n     -+\trather than a new empty one, replacing existing entries with identical\n     -+\tnames. Not compatible with `--literally`.\n     ++\tIf provided, the tree entries provided in stdin are added to this\n     ++\ttree rather than a new empty one, replacing existing entries with\n     ++\tidentical names. Not compatible with `--literally`.\n      +\n       INPUT FORMAT\n       ------------\n     @@ builtin/mktree.c\n       #include \"object-store-ll.h\"\n       \n       struct tree_entry {\n     -@@ builtin/mktree.c: static struct tree_entry *advance_tree_entry_iterator(struct tree_entry_iterator\n     - \treturn prev;\n     +@@ builtin/mktree.c: static void tree_entry_iterator_advance(struct tree_entry_iterator *iter)\n     + \t\t\t: NULL;\n       }\n       \n      -static int add_tree_entry_to_index(struct index_state *istate,\n     @@ builtin/mktree.c: static struct tree_entry *advance_tree_entry_iterator(struct t\n      +\t\t\t\t unsigned mode, void *context)\n       {\n      -\tstruct tree_entry_iterator iter = { NULL };\n     +-\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n     +-\tistate.sparse_index = 1;\n      +\tint result;\n      +\tstruct tree_entry *base_tree_ent;\n      +\tstruct build_index_data *cbdata = context;\n     @@ builtin/mktree.c: static struct tree_entry *advance_tree_entry_iterator(struct t\n      +\t\tint cmp = name_compare(ent->name, ent->len,\n      +\t\t\t\t       base_tree_ent->name, base_tree_ent->len);\n      +\t\tif (!cmp || cmp < 0) {\n     -+\t\t\tadvance_tree_entry_iterator(&cbdata->iter);\n     ++\t\t\ttree_entry_iterator_advance(&cbdata->iter);\n      +\n      +\t\t\tif (add_tree_entry_to_index(cbdata, ent) < 0) {\n      +\t\t\t\tresult = error(_(\"failed to add tree entry '%s'\"), ent->name);\n     @@ builtin/mktree.c: static struct tree_entry *advance_tree_entry_iterator(struct t\n      +\t\t       struct object_id *oid)\n      +{\n      +\tstruct build_index_data cbdata = { 0 };\n     - \tstruct tree_entry *ent;\n     --\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n     --\tistate.sparse_index = 1;\n      +\tstruct pathspec ps = { 0 };\n       \n       \tsort_and_dedup_tree_entry_array(arr);\n       \n     --\tinit_tree_entry_iterator(&iter, arr);\n     +-\ttree_entry_iterator_init(&iter, arr);\n      +\tindex_state_init(&cbdata.istate, the_repository);\n      +\tcbdata.istate.sparse_index = 1;\n     -+\tinit_tree_entry_iterator(&cbdata.iter, arr);\n     ++\ttree_entry_iterator_init(&cbdata.iter, arr);\n      +\tcbdata.df_name_hash = &arr->df_name_hash;\n       \n       \t/* Construct an in-memory index from the provided entries & base tree */\n     --\twhile ((ent = advance_tree_entry_iterator(&iter))) {\n     --\t\tif (add_tree_entry_to_index(&istate, ent))\n     +-\twhile (iter.current) {\n     +-\t\tstruct tree_entry *ent = iter.current;\n     +-\t\ttree_entry_iterator_advance(&iter);\n      +\tif (base_tree &&\n      +\t    read_tree(the_repository, base_tree, &ps, build_index_from_tree, &cbdata) < 0)\n      +\t\tdie(_(\"failed to create tree\"));\n      +\n     -+\twhile ((ent = advance_tree_entry_iterator(&cbdata.iter))) {\n     ++\twhile (cbdata.iter.current) {\n     ++\t\tstruct tree_entry *ent = cbdata.iter.current;\n     ++\t\ttree_entry_iterator_advance(&cbdata.iter);\n     + \n     +-\t\tif (add_tree_entry_to_index(&istate, ent))\n      +\t\tif (add_tree_entry_to_index(&cbdata, ent))\n       \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n       \t}\n     @@ builtin/mktree.c: static void write_tree_literally(struct tree_entry_array *arr,\n       \n       static const char *mktree_usage[] = {\n      -\t\"git mktree [-z] [--missing] [--literally] [--batch]\",\n     -+\t\"git mktree [-z] [--missing] [--literally] [--batch] [--] [<tree-ish>]\",\n     ++\t\"git mktree [-z] [--missing] [--literally] [--batch] [<tree-ish>]\",\n       \tNULL\n       };\n       \n      @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n     - \tint is_batch_mode = 0;\n       \tstruct tree_entry_array arr = { 0 };\n       \tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n     + \tstruct strbuf line = STRBUF_INIT;\n      +\tstruct tree *base_tree = NULL;\n       \tint ret;\n       \n     @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n      +\t\t\tdie(_(\"not a tree object: %s\"), oid_to_hex(&base_tree_oid));\n      +\t}\n       \n     - \tinit_tree_entry_array(&arr);\n     + \ttree_entry_array_init(&arr);\n       \n      @@ builtin/mktree.c: int cmd_mktree(int ac, const char **av, const char *prefix)\n       \t\t\tif (mktree_line_data.literally)\n 15:  058354f45f7 ! 16:  46756c4e314 mktree: allow deeper paths in input\n     @@ Commit message\n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n       ## Documentation/git-mktree.txt ##\n     -@@ Documentation/git-mktree.txt: INPUT FORMAT\n     - Tree entries may be specified in any of the formats compatible with the\n     - `--index-info` option to linkgit:git-update-index[1].\n     +@@ Documentation/git-mktree.txt: Higher stages represent conflicted files in an index; this information\n     + cannot be represented in a tree object. The command will fail without\n     + writing the tree if a higher order stage is specified for any entry.\n       \n      +Entries may use full pathnames containing directory separators to specify\n     -+entries nested within one or more directories. These entries are inserted into\n     -+the appropriate tree in the base tree-ish if one exists. Otherwise, empty parent\n     -+trees are created to contain the entries.\n     ++entries nested within one or more directories. These entries are inserted\n     ++into the appropriate tree in the base tree-ish if one exists. Otherwise,\n     ++empty parent trees are created to contain the entries.\n      +\n     - The order of the tree entries is normalized by `mktree` so pre-sorting the input\n     - by path is not required. Multiple entries provided with the same path are\n     - deduplicated, with only the last one specified added to the tree.\n     + The order of the tree entries is normalized by `mktree` so pre-sorting the\n     + input by path is not required. Multiple entries provided with the same path\n     + are deduplicated, with only the last one specified added to the tree.\n      \n       ## builtin/mktree.c ##\n      @@ builtin/mktree.c: struct tree_entry {\n     @@ builtin/mktree.c: static void tree_entry_array_push(struct tree_entry_array *arr\n      +\treturn arr->entries[--arr->nr];\n      +}\n      +\n     - static void clear_tree_entry_array(struct tree_entry_array *arr)\n     + static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entries)\n       {\n     - \thashmap_clear(&arr->df_name_hash);\n     + \tif (free_entries) {\n      @@ builtin/mktree.c: static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n       \n       \t\tif (!verify_path(ent->name, mode))\n     @@ builtin/mktree.c: static void sort_and_dedup_tree_entry_array(struct tree_entry_\n      +\t\t\t}\n      +\t\t}\n      +\n     -+\t\trelease_tree_entry_array(&parent_dir_ents);\n     ++\t\ttree_entry_array_release(&parent_dir_ents, 0);\n      +\t}\n      +\n       \t/* Finally, initialize the directory-file conflict hash map */\n     @@ builtin/mktree.c: static int build_index_from_tree(const struct object_id *oid,\n      +\t\tcmp = name_compare(ent->name, ent->len,\n      +\t\t\t\t   base_tree_ent->name, base_tree_ent->len);\n       \t\tif (!cmp || cmp < 0) {\n     - \t\t\tadvance_tree_entry_iterator(&cbdata->iter);\n     + \t\t\ttree_entry_iterator_advance(&cbdata->iter);\n       \n      @@ builtin/mktree.c: static int build_index_from_tree(const struct object_id *oid,\n       \t\t\t\tgoto cleanup_and_return;\n     @@ builtin/mktree.c: static int build_index_from_tree(const struct object_id *oid,\n      \n       ## t/t1010-mktree.sh ##\n      @@ t/t1010-mktree.sh: test_expect_success 'mktree with invalid submodule OIDs' '\n     - \tgrep \"object $tree_oid is a tree but specified type was (commit)\" err\n     + \tdone\n       '\n       \n      -test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n     @@ t/t1010-mktree.sh: test_expect_success 'mktree with base tree' '\n      +\t\tprintf \"100644 blob $blob_oid\\ttest/deeper\\n\"\n      +\t} |\n      +\ttest_must_fail git mktree 2>err &&\n     -+\tgrep \"You have both test and test/deeper\" err &&\n     ++\ttest_grep \"You have both test and test/deeper\" err &&\n      +\n      +\t{\n      +\t\tprintf \"100644 blob $blob_oid\\tfolder/one/deeper/deep\\n\"\n      +\t} |\n      +\ttest_must_fail git mktree $tree_oid 2>err &&\n     -+\tgrep \"You have both folder/one and folder/one/deeper/deep\" err\n     ++\ttest_grep \"You have both folder/one and folder/one/deeper/deep\" err\n      +'\n      +\n       test_done\n 16:  a90d6d0c943 ! 17:  d392c440b8a mktree: remove entries when mode is 0\n     @@ Commit message\n          Signed-off-by: Victoria Dye <vdye@github.com>\n      \n       ## Documentation/git-mktree.txt ##\n     -@@ Documentation/git-mktree.txt: entries nested within one or more directories. These entries are inserted into\n     - the appropriate tree in the base tree-ish if one exists. Otherwise, empty parent\n     - trees are created to contain the entries.\n     +@@ Documentation/git-mktree.txt: entries nested within one or more directories. These entries are inserted\n     + into the appropriate tree in the base tree-ish if one exists. Otherwise,\n     + empty parent trees are created to contain the entries.\n       \n     -+An entry with a mode of \"0\" will remove an entry of the same name from the base\n     -+tree-ish. If no tree-ish argument is given, or the entry does not exist in that\n     -+tree, the entry is ignored.\n     ++An entry with a mode of \"0\" will remove an entry of the same name from the\n     ++base tree-ish. If no tree-ish argument is given, or the entry does not exist\n     ++in that tree, the entry is ignored.\n      +\n     - The order of the tree entries is normalized by `mktree` so pre-sorting the input\n     - by path is not required. Multiple entries provided with the same path are\n     - deduplicated, with only the last one specified added to the tree.\n     + The order of the tree entries is normalized by `mktree` so pre-sorting the\n     + input by path is not required. Multiple entries provided with the same path\n     + are deduplicated, with only the last one specified added to the tree.\n      \n       ## builtin/mktree.c ##\n      @@ builtin/mktree.c: struct tree_entry {\n     @@ builtin/mktree.c: static int build_index_from_tree(const struct object_id *oid,\n       \t\tint ret = 0;\n       \t\tstruct pathspec ps = { 0 };\n      @@ builtin/mktree.c: static int mktree_line(unsigned int mode, struct object_id *oid,\n     - \t\t       const char *path, void *cbdata)\n     - {\n     - \tstruct mktree_line_data *data = cbdata;\n     --\tenum object_type mode_type = object_type(mode);\n     --\tstruct object_info oi = OBJECT_INFO_INIT;\n     --\tenum object_type parsed_obj_type;\n     - \n     --\tif (obj_type && mode_type != obj_type)\n     --\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n     --\t\t    type_name(obj_type), type_name(mode_type));\n     -+\tif (mode) {\n     -+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n     -+\t\tenum object_type parsed_obj_type;\n     -+\t\tenum object_type mode_type = object_type(mode);\n     - \n     --\toi.typep = &parsed_obj_type;\n     -+\t\tif (obj_type && mode_type != obj_type)\n     -+\t\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n     -+\t\t\t    type_name(obj_type), type_name(mode_type));\n     + \tif (stage)\n     + \t\tdie(_(\"path '%s' is unmerged\"), path);\n       \n     --\tif (oid_object_info_extended(the_repository, oid, &oi,\n     --\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n     --\t\t\t\t     OBJECT_INFO_QUICK |\n     --\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n     --\t\tparsed_obj_type = -1;\n     -+\t\toi.typep = &parsed_obj_type;\n     - \n     --\tif (parsed_obj_type < 0) {\n     --\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n     --\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n     --\t\t} else {\n     --\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n     -+\t\tif (oid_object_info_extended(the_repository, oid, &oi,\n     -+\t\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n     -+\t\t\t\t\t     OBJECT_INFO_QUICK |\n     -+\t\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n     -+\t\t\tparsed_obj_type = -1;\n     ++\t/* OID ignored for zero-mode entries; append unconditionally */\n     ++\tif (!mode)\n     ++\t\tgoto append_entry;\n      +\n     -+\t\tif (parsed_obj_type < 0) {\n     -+\t\t\tif (data->allow_missing || S_ISGITLINK(mode)) {\n     -+\t\t\t\t; /* no problem - missing objects & submodules are presumed to be of the right type */\n     -+\t\t\t} else {\n     -+\t\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n     -+\t\t\t}\n     -+\t\t} else if (parsed_obj_type != mode_type) {\n     -+\t\t\t/*\n     -+\t\t\t* The object exists but is of the wrong type.\n     -+\t\t\t* This is a problem regardless of allow_missing\n     -+\t\t\t* because the new tree entry will never be correct.\n     -+\t\t\t*/\n     -+\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n     -+\t\t\tpath, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n     + \tif (obj_type != OBJ_ANY && mode_type != obj_type)\n     + \t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n     + \t\t    type_name(obj_type), type_name(mode_type));\n     +@@ builtin/mktree.c: static int mktree_line(unsigned int mode, struct object_id *oid,\n       \t\t}\n     --\t} else if (parsed_obj_type != mode_type) {\n     --\t\t/*\n     --\t\t * The object exists but is of the wrong type.\n     --\t\t * This is a problem regardless of allow_missing\n     --\t\t * because the new tree entry will never be correct.\n     --\t\t */\n     --\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n     --\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n       \t}\n       \n     ++append_entry:\n       \tappend_to_tree(mode, oid, path, data->arr, data->literally);\n     + \treturn 0;\n     + }\n      \n       ## t/t1010-mktree.sh ##\n      @@ t/t1010-mktree.sh: test_expect_success 'mktree fails on directory-file conflict' '\n     - \tgrep \"You have both folder/one and folder/one/deeper/deep\" err\n     + \ttest_grep \"You have both folder/one and folder/one/deeper/deep\" err\n       '\n       \n      +test_expect_success 'mktree with remove entries' '\n\n-- \ngitgitgadget\n"},{"id":"497376","messageId":"4558f35e7bf9a1594510951ee54252069bdcfc5b.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 02/17] mktree: rename treeent to tree_entry","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:50Z","receivedAt":"2024-06-19T21:58:12Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nRename the type for better readability, clearly specifying \"entry\" (instead\nof the \"ent\" abbreviation) and separating \"tree\" from \"entry\".\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 10 +++++-----\n 1 file changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 8b19d440747..c02feb06aff 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,7 +12,7 @@\n #include \"parse-options.h\"\n #include \"object-store-ll.h\"\n \n-static struct treeent {\n+static struct tree_entry {\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n@@ -22,7 +22,7 @@ static int alloc, used;\n \n static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n {\n-\tstruct treeent *ent;\n+\tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n \tif (strchr(path, '/'))\n \t\tdie(\"path %s contains slash\", path);\n@@ -38,8 +38,8 @@ static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n \n static int ent_compare(const void *a_, const void *b_)\n {\n-\tstruct treeent *a = *(struct treeent **)a_;\n-\tstruct treeent *b = *(struct treeent **)b_;\n+\tstruct tree_entry *a = *(struct tree_entry **)a_;\n+\tstruct tree_entry *b = *(struct tree_entry **)b_;\n \treturn base_name_compare(a->name, a->len, a->mode,\n \t\t\t\t b->name, b->len, b->mode);\n }\n@@ -56,7 +56,7 @@ static void write_tree(struct object_id *oid)\n \n \tstrbuf_init(&buf, size);\n \tfor (i = 0; i < used; i++) {\n-\t\tstruct treeent *ent = entries[i];\n+\t\tstruct tree_entry *ent = entries[i];\n \t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n \t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"497377","messageId":"d0d5523a32b2f56f48772367651eba3d52d16c35.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 03/17] mktree: use non-static tree_entry array","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:51Z","receivedAt":"2024-06-19T21:58:14Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace the static 'struct tree_entry **entries' with a non-static 'struct\ntree_entry_array' instance. In later commits, we'll want to be able to\ncreate additional 'struct tree_entry_array' instances utilizing common\nfunctionality (create, push, clear, free). To avoid code duplication, create\nthe 'struct tree_entry_array' type and add functions that perform those\nbasic operations.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 69 ++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 49 insertions(+), 20 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex c02feb06aff..a96ea10bf95 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,15 +12,42 @@\n #include \"parse-options.h\"\n #include \"object-store-ll.h\"\n \n-static struct tree_entry {\n+struct tree_entry {\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n \tchar name[FLEX_ARRAY];\n-} **entries;\n-static int alloc, used;\n+};\n+\n+struct tree_entry_array {\n+\tsize_t nr, alloc;\n+\tstruct tree_entry **entries;\n+};\n+\n+static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n+{\n+\tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n+\tarr->entries[arr->nr++] = ent;\n+}\n+\n+static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entries)\n+{\n+\tif (free_entries) {\n+\t\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\t\tFREE_AND_NULL(arr->entries[i]);\n+\t}\n+\tarr->nr = 0;\n+}\n \n-static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n+static void tree_entry_array_release(struct tree_entry_array *arr, int free_entries)\n+{\n+\ttree_entry_array_clear(arr, free_entries);\n+\tFREE_AND_NULL(arr->entries);\n+\tarr->alloc = 0;\n+}\n+\n+static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n+\t\t\t   struct tree_entry_array *arr)\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n@@ -32,8 +59,7 @@ static void append_to_tree(unsigned mode, struct object_id *oid, char *path)\n \tent->len = len;\n \toidcpy(&ent->oid, oid);\n \n-\tALLOC_GROW(entries, used + 1, alloc);\n-\tentries[used++] = ent;\n+\ttree_entry_array_push(arr, ent);\n }\n \n static int ent_compare(const void *a_, const void *b_)\n@@ -44,19 +70,18 @@ static int ent_compare(const void *a_, const void *b_)\n \t\t\t\t b->name, b->len, b->mode);\n }\n \n-static void write_tree(struct object_id *oid)\n+static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n \tstruct strbuf buf;\n-\tsize_t size;\n-\tint i;\n+\tsize_t size = 0;\n \n-\tQSORT(entries, used, ent_compare);\n-\tfor (size = i = 0; i < used; i++)\n-\t\tsize += 32 + entries[i]->len;\n+\tQSORT(arr->entries, arr->nr, ent_compare);\n+\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\tsize += 32 + arr->entries[i]->len;\n \n \tstrbuf_init(&buf, size);\n-\tfor (i = 0; i < used; i++) {\n-\t\tstruct tree_entry *ent = entries[i];\n+\tfor (size_t i = 0; i < arr->nr; i++) {\n+\t\tstruct tree_entry *ent = arr->entries[i];\n \t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n \t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n \t}\n@@ -70,7 +95,8 @@ static const char *mktree_usage[] = {\n \tNULL\n };\n \n-static void mktree_line(char *buf, int nul_term_line, int allow_missing)\n+static void mktree_line(char *buf, int nul_term_line, int allow_missing,\n+\t\t\tstruct tree_entry_array *arr)\n {\n \tchar *ptr, *ntr;\n \tconst char *p;\n@@ -146,7 +172,7 @@ static void mktree_line(char *buf, int nul_term_line, int allow_missing)\n \t\t}\n \t}\n \n-\tappend_to_tree(mode, &oid, path);\n+\tappend_to_tree(mode, &oid, path, arr);\n \tfree(to_free);\n }\n \n@@ -158,6 +184,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tint allow_missing = 0;\n \tint is_batch_mode = 0;\n \tint got_eof = 0;\n+\tstruct tree_entry_array arr = { 0 };\n \tstrbuf_getline_fn getline_fn;\n \n \tconst struct option option[] = {\n@@ -182,9 +209,9 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t\t\tbreak;\n \t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n \t\t\t}\n-\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing);\n+\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n \t\t}\n-\t\tif (is_batch_mode && got_eof && used < 1) {\n+\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n \t\t\t/*\n \t\t\t * Execution gets here if the last tree entry is terminated with a\n \t\t\t * new-line.  The final new-line has been made optional to be\n@@ -192,12 +219,14 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t */\n \t\t\t; /* skip creating an empty tree */\n \t\t} else {\n-\t\t\twrite_tree(&oid);\n+\t\t\twrite_tree(&arr, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\n-\t\tused=0; /* reset tree entry buffer for re-use in batch mode */\n+\t\ttree_entry_array_clear(&arr, 1); /* reset tree entry buffer for re-use in batch mode */\n \t}\n+\n+\ttree_entry_array_release(&arr, 1);\n \tstrbuf_release(&sb);\n \treturn 0;\n }\n-- \ngitgitgadget\n\n"},{"id":"497378","messageId":"f5473764236be36c6e23714ce99c533ba83ac18e.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 04/17] update-index: generalize 'read_index_info'","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:52Z","receivedAt":"2024-06-19T21:58:15Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nMove 'read_index_info()' into a new header 'index-info.h' and generalize the\nfunction to call a provided callback for each parsed line. Update\n'update-index.c' to use this generalized 'read_index_info()', adding the\ncallback 'apply_index_info()' to verify the parsed line and update the index\naccording to its contents.\n\nSwitching to using a callback to validate the parsed entry in 'update-index'\nresults in a slight change to the error message indicating a file could not\nbe removed from the index. The original implementation uses the raw, quoted\npathname in the error message, whereas the callback (without access to the\nraw pathname) uses the unquoted value. However, this change makes the failed\nremoval message consistent with all other error messages in the function,\nand that consistency is likely more beneficial than not to a user.\n\nThe motivation for this change is to consolidate the already-similar input\nparsing logic in 'git update-index' and 'git mktree', avoiding code\nduplication and the associated maintenance burden. The input formats\naccepted by 'update-index' are a superset of those accepted by 'mktree', so\nin a later commit we can replace the input parsing of the latter with\n'read_index_info()' without breaking existing usage.\n\nCo-authored-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-update-index.txt   |  16 +---\n Documentation/index-info-formats.txt |  13 +++\n Makefile                             |   1 +\n builtin/update-index.c               | 129 +++++++--------------------\n index-info.c                         |  90 +++++++++++++++++++\n index-info.h                         |  11 +++\n t/t2107-update-index-basic.sh        |  27 ++++++\n 7 files changed, 177 insertions(+), 110 deletions(-)\n create mode 100644 Documentation/index-info-formats.txt\n create mode 100644 index-info.c\n create mode 100644 index-info.h\n\ndiff --git a/Documentation/git-update-index.txt b/Documentation/git-update-index.txt\nindex 7128aed5405..e52aecb845d 100644\n--- a/Documentation/git-update-index.txt\n+++ b/Documentation/git-update-index.txt\n@@ -278,21 +278,9 @@ USING --INDEX-INFO\n \n `--index-info` is a more powerful mechanism that lets you feed\n multiple entry definitions from the standard input, and designed\n-specifically for scripts.  It can take inputs of three formats:\n+specifically for scripts.  It can take inputs in the following formats:\n \n-    . mode SP type SP sha1          TAB path\n-+\n-This format is to stuff `git ls-tree` output into the index.\n-\n-    . mode         SP sha1 SP stage TAB path\n-+\n-This format is to put higher order stages into the\n-index file and matches 'git ls-files --stage' output.\n-\n-    . mode         SP sha1          TAB path\n-+\n-This format is no longer produced by any Git command, but is\n-and will continue to be supported by `update-index --index-info`.\n+include::index-info-formats.txt[]\n \n To place a higher stage entry to the index, the path should\n first be removed by feeding a mode=0 entry for the path, and\ndiff --git a/Documentation/index-info-formats.txt b/Documentation/index-info-formats.txt\nnew file mode 100644\nindex 00000000000..037ebd24321\n--- /dev/null\n+++ b/Documentation/index-info-formats.txt\n@@ -0,0 +1,13 @@\n+    . mode SP type SP sha1          TAB path\n++\n+This format is to use `git ls-tree` output.\n+\n+    . mode         SP sha1 SP stage TAB path\n++\n+This format allows higher order stages to appear and\n+matches 'git ls-files --stage' output.\n+\n+    . mode         SP sha1          TAB path\n++\n+This format is no longer produced by any Git command, but is\n+and will continue to be supported.\ndiff --git a/Makefile b/Makefile\nindex 2f5f16847ae..db9604e59c3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1037,6 +1037,7 @@ LIB_OBJS += hex.o\n LIB_OBJS += hex-ll.o\n LIB_OBJS += hook.o\n LIB_OBJS += ident.o\n+LIB_OBJS += index-info.o\n LIB_OBJS += json-writer.o\n LIB_OBJS += kwset.o\n LIB_OBJS += levenshtein.o\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex d343416ae26..fddf59b54c1 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -11,6 +11,7 @@\n #include \"gettext.h\"\n #include \"hash.h\"\n #include \"hex.h\"\n+#include \"index-info.h\"\n #include \"lockfile.h\"\n #include \"quote.h\"\n #include \"cache-tree.h\"\n@@ -509,100 +510,29 @@ static void update_one(const char *path)\n \treport(\"add '%s'\", path);\n }\n \n-static void read_index_info(int nul_term_line)\n+static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n+\t\t\t    const char *path_name, void *cbdata UNUSED)\n {\n-\tconst int hexsz = the_hash_algo->hexsz;\n-\tstruct strbuf buf = STRBUF_INIT;\n-\tstruct strbuf uq = STRBUF_INIT;\n-\tstrbuf_getline_fn getline_fn;\n+\tif (!verify_path(path_name, mode)) {\n+\t\tfprintf(stderr, \"Ignoring path %s\\n\", path_name);\n+\t\treturn 0;\n+\t}\n \n-\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n-\twhile (getline_fn(&buf, stdin) != EOF) {\n-\t\tchar *ptr, *tab;\n-\t\tchar *path_name;\n-\t\tstruct object_id oid;\n-\t\tunsigned int mode;\n-\t\tunsigned long ul;\n-\t\tint stage;\n-\n-\t\t/* This reads lines formatted in one of three formats:\n-\t\t *\n-\t\t * (1) mode         SP sha1          TAB path\n-\t\t * The first format is what \"git apply --index-info\"\n-\t\t * reports, and used to reconstruct a partial tree\n-\t\t * that is used for phony merge base tree when falling\n-\t\t * back on 3-way merge.\n-\t\t *\n-\t\t * (2) mode SP type SP sha1          TAB path\n-\t\t * The second format is to stuff \"git ls-tree\" output\n-\t\t * into the index file.\n-\t\t *\n-\t\t * (3) mode         SP sha1 SP stage TAB path\n-\t\t * This format is to put higher order stages into the\n-\t\t * index file and matches \"git ls-files --stage\" output.\n+\tif (!mode) {\n+\t\t/* mode == 0 means there is no such path -- remove */\n+\t\tif (remove_file_from_index(the_repository->index, path_name))\n+\t\t\tdie(\"git update-index: unable to remove %s\", path_name);\n+\t}\n+\telse {\n+\t\t/* mode ' ' sha1 '\\t' name\n+\t\t * ptr[-1] points at tab,\n+\t\t * ptr[-41] is at the beginning of sha1\n \t\t */\n-\t\terrno = 0;\n-\t\tul = strtoul(buf.buf, &ptr, 8);\n-\t\tif (ptr == buf.buf || *ptr != ' '\n-\t\t    || errno || (unsigned int) ul != ul)\n-\t\t\tgoto bad_line;\n-\t\tmode = ul;\n-\n-\t\ttab = strchr(ptr, '\\t');\n-\t\tif (!tab || tab - ptr < hexsz + 1)\n-\t\t\tgoto bad_line;\n-\n-\t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n-\t\t\tstage = tab[-1] - '0';\n-\t\t\tptr = tab + 1; /* point at the head of path */\n-\t\t\ttab = tab - 2; /* point at tail of sha1 */\n-\t\t}\n-\t\telse {\n-\t\t\tstage = 0;\n-\t\t\tptr = tab + 1; /* point at the head of path */\n-\t\t}\n-\n-\t\tif (get_oid_hex(tab - hexsz, &oid) ||\n-\t\t\ttab[-(hexsz + 1)] != ' ')\n-\t\t\tgoto bad_line;\n-\n-\t\tpath_name = ptr;\n-\t\tif (!nul_term_line && path_name[0] == '\"') {\n-\t\t\tstrbuf_reset(&uq);\n-\t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n-\t\t\t\tdie(\"git update-index: bad quoting of path name\");\n-\t\t\t}\n-\t\t\tpath_name = uq.buf;\n-\t\t}\n-\n-\t\tif (!verify_path(path_name, mode)) {\n-\t\t\tfprintf(stderr, \"Ignoring path %s\\n\", path_name);\n-\t\t\tcontinue;\n-\t\t}\n-\n-\t\tif (!mode) {\n-\t\t\t/* mode == 0 means there is no such path -- remove */\n-\t\t\tif (remove_file_from_index(the_repository->index, path_name))\n-\t\t\t\tdie(\"git update-index: unable to remove %s\",\n-\t\t\t\t    ptr);\n-\t\t}\n-\t\telse {\n-\t\t\t/* mode ' ' sha1 '\\t' name\n-\t\t\t * ptr[-1] points at tab,\n-\t\t\t * ptr[-41] is at the beginning of sha1\n-\t\t\t */\n-\t\t\tptr[-(hexsz + 2)] = ptr[-1] = 0;\n-\t\t\tif (add_cacheinfo(mode, &oid, path_name, stage))\n-\t\t\t\tdie(\"git update-index: unable to update %s\",\n-\t\t\t\t    path_name);\n-\t\t}\n-\t\tcontinue;\n-\n-\tbad_line:\n-\t\tdie(\"malformed index info %s\", buf.buf);\n+\t\tif (add_cacheinfo(mode, oid, path_name, stage))\n+\t\t\tdie(\"git update-index: unable to update %s\", path_name);\n \t}\n-\tstrbuf_release(&buf);\n-\tstrbuf_release(&uq);\n+\n+\treturn 0;\n }\n \n static const char * const update_index_usage[] = {\n@@ -848,16 +778,23 @@ static enum parse_opt_result stdin_cacheinfo_callback(\n \tstruct parse_opt_ctx_t *ctx, const struct option *opt,\n \tconst char *arg, int unset)\n {\n-\tint *nul_term_line = opt->value;\n+\tint ret = 0;\n \n \tBUG_ON_OPT_NEG(unset);\n \tBUG_ON_OPT_ARG(arg);\n \n-\tif (ctx->argc != 1)\n-\t\treturn error(\"option '%s' must be the last argument\", opt->long_name);\n-\tallow_add = allow_replace = allow_remove = 1;\n-\tread_index_info(*nul_term_line);\n-\treturn 0;\n+\tif (ctx->argc != 1) {\n+\t\tret = error(\"option '%s' must be the last argument\", opt->long_name);\n+\t} else {\n+\t\tint *nul_term_line = opt->value;\n+\n+\t\tallow_add = allow_replace = allow_remove = 1;\n+\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n+\t\tif (ret)\n+\t\t\tret = -1;\n+\t}\n+\n+\treturn ret;\n }\n \n static enum parse_opt_result stdin_callback(\ndiff --git a/index-info.c b/index-info.c\nnew file mode 100644\nindex 00000000000..8ccaac5487b\n--- /dev/null\n+++ b/index-info.c\n@@ -0,0 +1,90 @@\n+#include \"git-compat-util.h\"\n+#include \"index-info.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"strbuf.h\"\n+#include \"quote.h\"\n+\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n+{\n+\tconst int hexsz = the_hash_algo->hexsz;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf uq = STRBUF_INIT;\n+\tstrbuf_getline_fn getline_fn;\n+\tint ret = 0;\n+\n+\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n+\twhile (getline_fn(&buf, stdin) != EOF) {\n+\t\tchar *ptr, *tab;\n+\t\tchar *path_name;\n+\t\tstruct object_id oid;\n+\t\tunsigned int mode;\n+\t\tunsigned long ul;\n+\t\tint stage;\n+\n+\t\t/* This reads lines formatted in one of three formats:\n+\t\t *\n+\t\t * (1) mode         SP sha1          TAB path\n+\t\t * The first format is what \"git apply --index-info\"\n+\t\t * reports, and used to reconstruct a partial tree\n+\t\t * that is used for phony merge base tree when falling\n+\t\t * back on 3-way merge.\n+\t\t *\n+\t\t * (2) mode SP type SP sha1          TAB path\n+\t\t * The second format is to stuff \"git ls-tree\" output\n+\t\t * into the index file.\n+\t\t *\n+\t\t * (3) mode         SP sha1 SP stage TAB path\n+\t\t * This format is to put higher order stages into the\n+\t\t * index file and matches \"git ls-files --stage\" output.\n+\t\t */\n+\t\terrno = 0;\n+\t\tul = strtoul(buf.buf, &ptr, 8);\n+\t\tif (ptr == buf.buf || *ptr != ' '\n+\t\t    || errno || (unsigned int) ul != ul)\n+\t\t\tgoto bad_line;\n+\t\tmode = ul;\n+\n+\t\ttab = strchr(ptr, '\\t');\n+\t\tif (!tab || tab - ptr < hexsz + 1)\n+\t\t\tgoto bad_line;\n+\n+\t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n+\t\t\tstage = tab[-1] - '0';\n+\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\ttab = tab - 2; /* point at tail of sha1 */\n+\t\t} else {\n+\t\t\tstage = 0;\n+\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t}\n+\n+\t\tif (get_oid_hex(tab - hexsz, &oid) ||\n+\t\t\ttab[-(hexsz + 1)] != ' ')\n+\t\t\tgoto bad_line;\n+\n+\t\tpath_name = ptr;\n+\t\tif (!nul_term_line && path_name[0] == '\"') {\n+\t\t\tstrbuf_reset(&uq);\n+\t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n+\t\t\t\tret = error(\"bad quoting of path name\");\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t\tpath_name = uq.buf;\n+\t\t}\n+\n+\t\tret = fn(mode, &oid, stage, path_name, cbdata);\n+\t\tif (ret) {\n+\t\t\tret = -1;\n+\t\t\tbreak;\n+\t\t}\n+\n+\t\tcontinue;\n+\n+\tbad_line:\n+\t\tdie(\"malformed input line '%s'\", buf.buf);\n+\t}\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&uq);\n+\n+\treturn ret;\n+}\ndiff --git a/index-info.h b/index-info.h\nnew file mode 100644\nindex 00000000000..d650498325a\n--- /dev/null\n+++ b/index-info.h\n@@ -0,0 +1,11 @@\n+#ifndef INDEX_INFO_H\n+#define INDEX_INFO_H\n+\n+#include \"hash.h\"\n+\n+typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n+\n+/* Iterate over parsed index info from stdin */\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n+\n+#endif /* INDEX_INFO_H */\ndiff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\nindex cc72ead79f3..794a5b1a184 100755\n--- a/t/t2107-update-index-basic.sh\n+++ b/t/t2107-update-index-basic.sh\n@@ -142,4 +142,31 @@ test_expect_success '--index-version' '\n \ttest_must_be_empty actual\n '\n \n+test_expect_success '--index-info fails on malformed input' '\n+\t# empty line\n+\techo \"\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"malformed input line\" err &&\n+\n+\t# bad whitespace\n+\tprintf \"100644 $EMPTY_BLOB A\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"malformed input line\" err &&\n+\n+\t# invalid stage value\n+\tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"malformed input line\" err &&\n+\n+\t# invalid OID length\n+\tprintf \"100755 abc123\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"malformed input line\" err &&\n+\n+\t# bad quoting\n+\tprintf \"100644 $EMPTY_BLOB\\t\\\"A\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"bad quoting of path name\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497379","messageId":"4f4d54c8d075a43960af36ccb025b2ddb34266f8.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 05/17] index-info.c: return unrecognized lines to caller","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:53Z","receivedAt":"2024-06-19T21:58:17Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUpdate 'read_index_info()' to return INDEX_INFO_UNRECOGNIZED_LINE (value 1),\nrather than die()-ing when the function encounters a line that cannot be\nparsed according to one of the accepted formats. This grants the caller the\nflexibility to fall back on custom handling for such lines rather than a\nreturning a catch-all error. In the case of 'update-index', we'll still exit\nwith a \"malformed input line\" error. However, when 'read_index_info()' is\nused to process the input to 'mktree' in a later patch, an empty line return\nvalue will signal a new tree in --batch mode.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/update-index.c |  9 +++++++--\n index-info.c           | 16 +++++++++-------\n index-info.h           |  5 ++++-\n 3 files changed, 20 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex fddf59b54c1..8d0b40a6fd6 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -787,11 +787,16 @@ static enum parse_opt_result stdin_cacheinfo_callback(\n \t\tret = error(\"option '%s' must be the last argument\", opt->long_name);\n \t} else {\n \t\tint *nul_term_line = opt->value;\n+\t\tstruct strbuf line = STRBUF_INIT;\n \n \t\tallow_add = allow_replace = allow_remove = 1;\n-\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL);\n-\t\tif (ret)\n+\t\tret = read_index_info(*nul_term_line, apply_index_info, NULL, &line);\n+\n+\t\tif (ret == INDEX_INFO_UNRECOGNIZED_LINE)\n+\t\t\tret = error(\"malformed input line '%s'\", line.buf);\n+\t\telse if (ret)\n \t\t\tret = -1;\n+\t\tstrbuf_release(&line);\n \t}\n \n \treturn ret;\ndiff --git a/index-info.c b/index-info.c\nindex 8ccaac5487b..7a02f66426a 100644\n--- a/index-info.c\n+++ b/index-info.c\n@@ -5,16 +5,16 @@\n #include \"strbuf.h\"\n #include \"quote.h\"\n \n-int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n+\t\t    struct strbuf *line)\n {\n \tconst int hexsz = the_hash_algo->hexsz;\n-\tstruct strbuf buf = STRBUF_INIT;\n \tstruct strbuf uq = STRBUF_INIT;\n \tstrbuf_getline_fn getline_fn;\n \tint ret = 0;\n \n \tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n-\twhile (getline_fn(&buf, stdin) != EOF) {\n+\twhile (getline_fn(line, stdin) != EOF) {\n \t\tchar *ptr, *tab;\n \t\tchar *path_name;\n \t\tstruct object_id oid;\n@@ -39,8 +39,8 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \t\t * index file and matches \"git ls-files --stage\" output.\n \t\t */\n \t\terrno = 0;\n-\t\tul = strtoul(buf.buf, &ptr, 8);\n-\t\tif (ptr == buf.buf || *ptr != ' '\n+\t\tul = strtoul(line->buf, &ptr, 8);\n+\t\tif (ptr == line->buf || *ptr != ' '\n \t\t    || errno || (unsigned int) ul != ul)\n \t\t\tgoto bad_line;\n \t\tmode = ul;\n@@ -81,10 +81,12 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata)\n \t\tcontinue;\n \n \tbad_line:\n-\t\tdie(\"malformed input line '%s'\", buf.buf);\n+\t\tret = INDEX_INFO_UNRECOGNIZED_LINE;\n+\t\tbreak;\n \t}\n-\tstrbuf_release(&buf);\n \tstrbuf_release(&uq);\n+\tif (!ret)\n+\t\tstrbuf_reset(line);\n \n \treturn ret;\n }\ndiff --git a/index-info.h b/index-info.h\nindex d650498325a..9258011462d 100644\n--- a/index-info.h\n+++ b/index-info.h\n@@ -5,7 +5,10 @@\n \n typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n \n+#define INDEX_INFO_UNRECOGNIZED_LINE 1\n+\n /* Iterate over parsed index info from stdin */\n-int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata);\n+int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n+\t\t    struct strbuf *line);\n \n #endif /* INDEX_INFO_H */\n-- \ngitgitgadget\n\n"},{"id":"497380","messageId":"472efcaf1dde3dd590f34bde63c5ce6dcf72a531.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 06/17] index-info.c: parse object type in provided in read_index_info","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:54Z","receivedAt":"2024-06-19T21:58:17Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf the object type (e.g. \"blob\", \"tree\") is identified on a stdin line read\nby 'read_index_info()' (i.e. on lines formatted like the output of 'git\nls-tree'), parse it into an 'enum object_type' and provide it to the\n'read_index_info()' callback as an argument. If the type is not provided,\npass 'OBJ_ANY' instead. If the object type is invalid, return an error.\n\nThe goal of this change is to allow for more thorough validation of the\nprovided object type (e.g. against the provided mode) in 'mktree' once\n'mktree_line' is replaced with 'read_index_info()'. Note, though, that this\nchange also strengthens the validation done by 'update-index', since invalid\ntype names now trigger an error.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/update-index.c        |  3 ++-\n index-info.c                  | 16 ++++++++++++----\n index-info.h                  |  3 ++-\n t/t2107-update-index-basic.sh |  5 +++++\n 4 files changed, 21 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/update-index.c b/builtin/update-index.c\nindex 8d0b40a6fd6..42a274f9ce4 100644\n--- a/builtin/update-index.c\n+++ b/builtin/update-index.c\n@@ -510,7 +510,8 @@ static void update_one(const char *path)\n \treport(\"add '%s'\", path);\n }\n \n-static int apply_index_info(unsigned int mode, struct object_id *oid, int stage,\n+static int apply_index_info(unsigned int mode, struct object_id *oid,\n+\t\t\t    enum object_type obj_type UNUSED, int stage,\n \t\t\t    const char *path_name, void *cbdata UNUSED)\n {\n \tif (!verify_path(path_name, mode)) {\ndiff --git a/index-info.c b/index-info.c\nindex 7a02f66426a..9c986cd9093 100644\n--- a/index-info.c\n+++ b/index-info.c\n@@ -18,6 +18,7 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n \t\tchar *ptr, *tab;\n \t\tchar *path_name;\n \t\tstruct object_id oid;\n+\t\tenum object_type obj_type = OBJ_ANY;\n \t\tunsigned int mode;\n \t\tunsigned long ul;\n \t\tint stage;\n@@ -51,18 +52,17 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n \n \t\tif (tab[-2] == ' ' && '0' <= tab[-1] && tab[-1] <= '3') {\n \t\t\tstage = tab[-1] - '0';\n-\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\tpath_name = tab + 1; /* point at the head of path */\n \t\t\ttab = tab - 2; /* point at tail of sha1 */\n \t\t} else {\n \t\t\tstage = 0;\n-\t\t\tptr = tab + 1; /* point at the head of path */\n+\t\t\tpath_name = tab + 1; /* point at the head of path */\n \t\t}\n \n \t\tif (get_oid_hex(tab - hexsz, &oid) ||\n \t\t\ttab[-(hexsz + 1)] != ' ')\n \t\t\tgoto bad_line;\n \n-\t\tpath_name = ptr;\n \t\tif (!nul_term_line && path_name[0] == '\"') {\n \t\t\tstrbuf_reset(&uq);\n \t\t\tif (unquote_c_style(&uq, path_name, NULL)) {\n@@ -72,7 +72,15 @@ int read_index_info(int nul_term_line, each_index_info_fn fn, void *cbdata,\n \t\t\tpath_name = uq.buf;\n \t\t}\n \n-\t\tret = fn(mode, &oid, stage, path_name, cbdata);\n+\t\t/* Get the type, if provided */\n+\t\tif (tab - hexsz - 1 > ptr + 1) {\n+\t\t\tif (*(tab - hexsz - 1) != ' ')\n+\t\t\t\tgoto bad_line;\n+\t\t\t*(tab - hexsz - 1) = '\\0';\n+\t\t\tobj_type = type_from_string(ptr + 1);\n+\t\t}\n+\n+\t\tret = fn(mode, &oid, obj_type, stage, path_name, cbdata);\n \t\tif (ret) {\n \t\t\tret = -1;\n \t\t\tbreak;\ndiff --git a/index-info.h b/index-info.h\nindex 9258011462d..adea453b197 100644\n--- a/index-info.h\n+++ b/index-info.h\n@@ -2,8 +2,9 @@\n #define INDEX_INFO_H\n \n #include \"hash.h\"\n+#include \"object.h\"\n \n-typedef int (*each_index_info_fn)(unsigned int, struct object_id *, int, const char *, void *);\n+typedef int (*each_index_info_fn)(unsigned int, struct object_id *, enum object_type, int, const char *, void *);\n \n #define INDEX_INFO_UNRECOGNIZED_LINE 1\n \ndiff --git a/t/t2107-update-index-basic.sh b/t/t2107-update-index-basic.sh\nindex 794a5b1a184..9e0e77bbf9e 100755\n--- a/t/t2107-update-index-basic.sh\n+++ b/t/t2107-update-index-basic.sh\n@@ -153,6 +153,11 @@ test_expect_success '--index-info fails on malformed input' '\n \ttest_must_fail git update-index --index-info 2>err &&\n \ttest_grep \"malformed input line\" err &&\n \n+\t# invalid type\n+\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n+\ttest_must_fail git update-index --index-info 2>err &&\n+\ttest_grep \"invalid object type\" err &&\n+\n \t# invalid stage value\n \tprintf \"100644 $EMPTY_BLOB 5\\tA\" |\n \ttest_must_fail git update-index --index-info 2>err &&\n-- \ngitgitgadget\n\n"},{"id":"497381","messageId":"8a3264afd0c072d10ec0571e2038f009733c4de5.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 08/17] mktree.c: do not fail on mismatched submodule type","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:56Z","receivedAt":"2024-06-19T21:58:19Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nAdjust the 'git mktree' tree entry intake logic to no longer fail if an OID\nspecified with a S_IFGITLINK mode exists in the current repository's object\ndatabase with a different type.\n\nWhile this scenario likely represents a mistake by the user, submodule OIDs\nare not validated as part of object writes or in 'git fsck'. In other\ncommands, any object info would be ignored if such an OID was found in the\ncurrent repository with a different type.\n\nSince this check is not needed to avoid creation of a corrupt tree, let's\nremove it and make 'git mktree' less opinionated as a result.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c  | 58 ++++++++++++++++++++++-------------------------\n t/t1010-mktree.sh | 18 ++++++---------\n 2 files changed, 34 insertions(+), 42 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 03a9899bc11..f509ed1a81f 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -107,8 +107,6 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n {\n \tstruct mktree_line_data *data = cbdata;\n \tenum object_type mode_type = object_type(mode);\n-\tstruct object_info oi = OBJECT_INFO_INIT;\n-\tenum object_type parsed_obj_type;\n \n \tif (stage)\n \t\tdie(_(\"path '%s' is unmerged\"), path);\n@@ -117,36 +115,34 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n \t\t    type_name(obj_type), type_name(mode_type));\n \n-\toi.typep = &parsed_obj_type;\n-\n-\tif (oid_object_info_extended(the_repository, oid, &oi,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n+\tif (!S_ISGITLINK(mode)) {\n+\t\tstruct object_info oi = OBJECT_INFO_INIT;\n+\t\tenum object_type parsed_obj_type;\n+\t\tunsigned int flags = OBJECT_INFO_LOOKUP_REPLACE |\n \t\t\t\t     OBJECT_INFO_QUICK |\n-\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n-\t\tparsed_obj_type = -1;\n-\n-\tif (parsed_obj_type < 0) {\n-\t\t/*\n-\t\t * There are two conditions where the object being missing\n-\t\t * is acceptable:\n-\t\t *\n-\t\t * - We're explicitly allowing it with --missing.\n-\t\t * - The object is a submodule, which we wouldn't expect to\n-\t\t *   be in this repo anyway.\n-\t\t *\n-\t\t * If neither condition is met, die().\n-\t\t */\n-\t\tif (!data->allow_missing && !S_ISGITLINK(mode))\n-\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n-\n-\t} else if (parsed_obj_type != mode_type) {\n-\t\t/*\n-\t\t * The object exists but is of the wrong type.\n-\t\t * This is a problem regardless of allow_missing\n-\t\t * because the new tree entry will never be correct.\n-\t\t */\n-\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n-\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n+\t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT;\n+\n+\t\toi.typep = &parsed_obj_type;\n+\n+\t\tif (oid_object_info_extended(the_repository, oid, &oi, flags) < 0) {\n+\t\t\t/*\n+\t\t\t * If the object is missing and we aren't explicitly\n+\t\t\t * allowing missing objects, die(). Otherwise, continue\n+\t\t\t * without error.\n+\t\t\t */\n+\t\t\tif (!data->allow_missing)\n+\t\t\t\tdie(\"entry '%s' object %s is unavailable\", path,\n+\t\t\t\t    oid_to_hex(oid));\n+\t\t} else if (parsed_obj_type != mode_type) {\n+\t\t\t/*\n+\t\t\t * The object exists but is of the wrong type.\n+\t\t\t * This is a problem regardless of allow_missing\n+\t\t\t * because the new tree entry will never be correct.\n+\t\t\t */\n+\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n+\t\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type),\n+\t\t\t    type_name(mode_type));\n+\t\t}\n \t}\n \n \tappend_to_tree(mode, oid, path, data->arr);\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 649842fa27c..48fc532e7af 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -71,17 +71,13 @@ test_expect_success 'allow missing object with --missing' '\n '\n \n test_expect_success 'mktree with invalid submodule OIDs' '\n-\t# non-existent OID - ok\n-\tprintf \"160000 commit $(test_oid numeric)\\tA\\n\" >in &&\n-\tgit mktree <in >tree.actual &&\n-\tgit ls-tree $(cat tree.actual) >actual &&\n-\ttest_cmp in actual &&\n-\n-\t# existing OID, wrong type - error\n-\ttree_oid=\"$(cat tree)\" &&\n-\tprintf \"160000 commit $tree_oid\\tA\" |\n-\ttest_must_fail git mktree 2>err &&\n-\ttest_grep \"object $tree_oid is a tree but specified type was (commit)\" err\n+\tfor oid in \"$(test_oid numeric)\" \"$(cat tree)\"\n+\tdo\n+\t\tprintf \"160000 commit $oid\\tA\\n\" >in &&\n+\t\tgit mktree <in >tree.actual &&\n+\t\tgit ls-tree $(cat tree.actual) >actual &&\n+\t\ttest_cmp in actual || return 1\n+\tdone\n '\n \n test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n-- \ngitgitgadget\n\n"},{"id":"497382","messageId":"9dc8e16a7fca886ec378d74a8e2ac61921a7f6ea.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 07/17] mktree: use read_index_info to read stdin lines","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:55Z","receivedAt":"2024-06-19T21:58:19Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nReplace the custom input parsing of 'mktree' with 'read_index_info()', which\nhandles not only the 'ls-tree' output format it already handles but also the\nother formats compatible with 'update-index'. This lends some consistency\nacross the commands (avoiding the need for two similar implementations for\ninput parsing) and adds flexibility to mktree.\n\nIt should be noted that, while the error messages are largely preserved in\nthe refactor, one does change: \"fatal: invalid quoting\" is now \"error: bad\nquoting of path name\".\n\nUpdate 'Documentation/git-mktree.txt' to reflect the more permissive input\nformat, as well as make a note about rejecting stage values higher than 0.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  26 ++++--\n builtin/mktree.c             | 156 +++++++++++++++--------------------\n t/t1010-mktree.sh            |  66 +++++++++++++++\n 3 files changed, 151 insertions(+), 97 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 383f09dd333..c187403c6bd 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -3,7 +3,7 @@ git-mktree(1)\n \n NAME\n ----\n-git-mktree - Build a tree-object from ls-tree formatted text\n+git-mktree - Build a tree-object from formatted tree entries\n \n \n SYNOPSIS\n@@ -13,15 +13,14 @@ SYNOPSIS\n \n DESCRIPTION\n -----------\n-Reads standard input in non-recursive `ls-tree` output format, and creates\n-a tree object.  The order of the tree entries is normalized by mktree so\n-pre-sorting the input is not required.  The object name of the tree object\n-built is written to the standard output.\n+Reads entry information from stdin and creates a tree object from those\n+entries. The object name of the tree object built is written to the standard\n+output.\n \n OPTIONS\n -------\n -z::\n-\tRead the NUL-terminated `ls-tree -z` output instead.\n+\tInput lines are separated with NUL rather than LF.\n \n --missing::\n \tAllow missing objects.  The default behaviour (without this option)\n@@ -35,6 +34,21 @@ OPTIONS\n \toptional.  Note - if the `-z` option is used, lines are terminated\n \twith NUL.\n \n+INPUT FORMAT\n+------------\n+Tree entries may be specified in any of the formats compatible with the\n+`--index-info` option to linkgit:git-update-index[1]:\n+\n+include::index-info-formats.txt[]\n+\n+Note that if the `stage` of a tree entry is given, the value must be 0.\n+Higher stages represent conflicted files in an index; this information\n+cannot be represented in a tree object. The command will fail without\n+writing the tree if a higher order stage is specified for any entry.\n+\n+The order of the tree entries is normalized by `mktree` so pre-sorting the\n+input by path is not required.\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex a96ea10bf95..03a9899bc11 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -6,6 +6,7 @@\n #include \"builtin.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n+#include \"index-info.h\"\n #include \"quote.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n@@ -95,123 +96,96 @@ static const char *mktree_usage[] = {\n \tNULL\n };\n \n-static void mktree_line(char *buf, int nul_term_line, int allow_missing,\n-\t\t\tstruct tree_entry_array *arr)\n+struct mktree_line_data {\n+\tstruct tree_entry_array *arr;\n+\tint allow_missing;\n+};\n+\n+static int mktree_line(unsigned int mode, struct object_id *oid,\n+\t\t       enum object_type obj_type, int stage,\n+\t\t       const char *path, void *cbdata)\n {\n-\tchar *ptr, *ntr;\n-\tconst char *p;\n-\tunsigned mode;\n-\tenum object_type mode_type; /* object type derived from mode */\n-\tenum object_type obj_type; /* object type derived from sha */\n+\tstruct mktree_line_data *data = cbdata;\n+\tenum object_type mode_type = object_type(mode);\n \tstruct object_info oi = OBJECT_INFO_INIT;\n-\tchar *path, *to_free = NULL;\n-\tstruct object_id oid;\n+\tenum object_type parsed_obj_type;\n \n-\tptr = buf;\n-\t/*\n-\t * Read non-recursive ls-tree output format:\n-\t *     mode SP type SP sha1 TAB name\n-\t */\n-\tmode = strtoul(ptr, &ntr, 8);\n-\tif (ptr == ntr || !ntr || *ntr != ' ')\n-\t\tdie(\"input format error: %s\", buf);\n-\tptr = ntr + 1; /* type */\n-\tntr = strchr(ptr, ' ');\n-\tif (!ntr || parse_oid_hex(ntr + 1, &oid, &p) ||\n-\t    *p != '\\t')\n-\t\tdie(\"input format error: %s\", buf);\n-\n-\t/* It is perfectly normal if we do not have a commit from a submodule */\n-\tif (S_ISGITLINK(mode))\n-\t\tallow_missing = 1;\n-\n-\n-\t*ntr++ = 0; /* now at the beginning of SHA1 */\n-\n-\tpath = (char *)p + 1;  /* at the beginning of name */\n-\tif (!nul_term_line && path[0] == '\"') {\n-\t\tstruct strbuf p_uq = STRBUF_INIT;\n-\t\tif (unquote_c_style(&p_uq, path, NULL))\n-\t\t\tdie(\"invalid quoting\");\n-\t\tpath = to_free = strbuf_detach(&p_uq, NULL);\n-\t}\n+\tif (stage)\n+\t\tdie(_(\"path '%s' is unmerged\"), path);\n \n-\t/*\n-\t * Object type is redundantly derivable three ways.\n-\t * These should all agree.\n-\t */\n-\tmode_type = object_type(mode);\n-\tif (mode_type != type_from_string(ptr)) {\n-\t\tdie(\"entry '%s' object type (%s) doesn't match mode type (%s)\",\n-\t\t\tpath, ptr, type_name(mode_type));\n-\t}\n+\tif (obj_type != OBJ_ANY && mode_type != obj_type)\n+\t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n+\t\t    type_name(obj_type), type_name(mode_type));\n+\n+\toi.typep = &parsed_obj_type;\n \n-\t/* Check the type of object identified by oid without fetching objects */\n-\toi.typep = &obj_type;\n-\tif (oid_object_info_extended(the_repository, &oid, &oi,\n+\tif (oid_object_info_extended(the_repository, oid, &oi,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE |\n \t\t\t\t     OBJECT_INFO_QUICK |\n \t\t\t\t     OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)\n-\t\tobj_type = -1;\n-\n-\tif (obj_type < 0) {\n-\t\tif (allow_missing) {\n-\t\t\t; /* no problem - missing objects are presumed to be of the right type */\n-\t\t} else {\n-\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(&oid));\n-\t\t}\n-\t} else {\n-\t\tif (obj_type != mode_type) {\n-\t\t\t/*\n-\t\t\t * The object exists but is of the wrong type.\n-\t\t\t * This is a problem regardless of allow_missing\n-\t\t\t * because the new tree entry will never be correct.\n-\t\t\t */\n-\t\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n-\t\t\t\tpath, oid_to_hex(&oid), type_name(obj_type), type_name(mode_type));\n-\t\t}\n+\t\tparsed_obj_type = -1;\n+\n+\tif (parsed_obj_type < 0) {\n+\t\t/*\n+\t\t * There are two conditions where the object being missing\n+\t\t * is acceptable:\n+\t\t *\n+\t\t * - We're explicitly allowing it with --missing.\n+\t\t * - The object is a submodule, which we wouldn't expect to\n+\t\t *   be in this repo anyway.\n+\t\t *\n+\t\t * If neither condition is met, die().\n+\t\t */\n+\t\tif (!data->allow_missing && !S_ISGITLINK(mode))\n+\t\t\tdie(\"entry '%s' object %s is unavailable\", path, oid_to_hex(oid));\n+\n+\t} else if (parsed_obj_type != mode_type) {\n+\t\t/*\n+\t\t * The object exists but is of the wrong type.\n+\t\t * This is a problem regardless of allow_missing\n+\t\t * because the new tree entry will never be correct.\n+\t\t */\n+\t\tdie(\"entry '%s' object %s is a %s but specified type was (%s)\",\n+\t\t    path, oid_to_hex(oid), type_name(parsed_obj_type), type_name(mode_type));\n \t}\n \n-\tappend_to_tree(mode, &oid, path, arr);\n-\tfree(to_free);\n+\tappend_to_tree(mode, oid, path, data->arr);\n+\treturn 0;\n }\n \n int cmd_mktree(int ac, const char **av, const char *prefix)\n {\n-\tstruct strbuf sb = STRBUF_INIT;\n \tstruct object_id oid;\n \tint nul_term_line = 0;\n-\tint allow_missing = 0;\n \tint is_batch_mode = 0;\n-\tint got_eof = 0;\n \tstruct tree_entry_array arr = { 0 };\n-\tstrbuf_getline_fn getline_fn;\n+\tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n+\tstruct strbuf line = STRBUF_INIT;\n+\tint ret;\n \n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n-\t\tOPT_BOOL(0, \"missing\", &allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"missing\", &mktree_line_data.allow_missing, N_(\"allow missing objects\")),\n \t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n-\tgetline_fn = nul_term_line ? strbuf_getline_nul : strbuf_getline_lf;\n-\n-\twhile (!got_eof) {\n-\t\twhile (1) {\n-\t\t\tif (getline_fn(&sb, stdin) == EOF) {\n-\t\t\t\tgot_eof = 1;\n-\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tif (sb.buf[0] == '\\0') {\n+\n+\tdo {\n+\t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data, &line);\n+\t\tif (ret < 0)\n+\t\t\tbreak;\n+\n+\t\tif (ret == INDEX_INFO_UNRECOGNIZED_LINE) {\n+\t\t\tif (line.len)\n+\t\t\t\tdie(\"input format error: %s\", line.buf);\n+\t\t\telse if (!is_batch_mode)\n \t\t\t\t/* empty lines denote tree boundaries in batch mode */\n-\t\t\t\tif (is_batch_mode)\n-\t\t\t\t\tbreak;\n \t\t\t\tdie(\"input format error: (blank line only valid in batch mode)\");\n-\t\t\t}\n-\t\t\tmktree_line(sb.buf, nul_term_line, allow_missing, &arr);\n \t\t}\n-\t\tif (is_batch_mode && got_eof && arr.nr < 1) {\n+\n+\t\tif (is_batch_mode && !ret && arr.nr < 1) {\n \t\t\t/*\n \t\t\t * Execution gets here if the last tree entry is terminated with a\n \t\t\t * new-line.  The final new-line has been made optional to be\n@@ -224,9 +198,9 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\tfflush(stdout);\n \t\t}\n \t\ttree_entry_array_clear(&arr, 1); /* reset tree entry buffer for re-use in batch mode */\n-\t}\n+\t} while (ret > 0);\n \n+\tstrbuf_release(&line);\n \ttree_entry_array_release(&arr, 1);\n-\tstrbuf_release(&sb);\n-\treturn 0;\n+\treturn !!ret;\n }\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 22875ba598c..649842fa27c 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -54,11 +54,36 @@ test_expect_success 'ls-tree output in wrong order given to mktree (2)' '\n \ttest_cmp tree.withsub actual\n '\n \n+test_expect_success '--batch creates multiple trees' '\n+\tcat top >multi-tree &&\n+\techo \"\" >>multi-tree &&\n+\tcat top.withsub >>multi-tree &&\n+\n+\tcat tree >expect &&\n+\tcat tree.withsub >>expect &&\n+\tgit mktree --batch <multi-tree >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'allow missing object with --missing' '\n \tgit mktree --missing <top.missing >actual &&\n \ttest_cmp tree.missing actual\n '\n \n+test_expect_success 'mktree with invalid submodule OIDs' '\n+\t# non-existent OID - ok\n+\tprintf \"160000 commit $(test_oid numeric)\\tA\\n\" >in &&\n+\tgit mktree <in >tree.actual &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\ttest_cmp in actual &&\n+\n+\t# existing OID, wrong type - error\n+\ttree_oid=\"$(cat tree)\" &&\n+\tprintf \"160000 commit $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"object $tree_oid is a tree but specified type was (commit)\" err\n+'\n+\n test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n \ttest_must_fail git mktree <all\n '\n@@ -67,4 +92,45 @@ test_expect_success 'mktree refuses to read ls-tree -r output (2)' '\n \ttest_must_fail git mktree <all.withsub\n '\n \n+test_expect_success 'mktree fails on malformed input' '\n+\t# empty line without --batch\n+\techo \"\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"blank line only valid in batch mode\" err &&\n+\n+\t# bad whitespace\n+\tprintf \"100644 blob $EMPTY_BLOB A\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"input format error\" err &&\n+\n+\t# invalid type\n+\tprintf \"100644 bad $EMPTY_BLOB\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"invalid object type\" err &&\n+\n+\t# invalid OID length\n+\tprintf \"100755 blob abc123\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"input format error\" err &&\n+\n+\t# bad quoting\n+\tprintf \"100644 blob $EMPTY_BLOB\\t\\\"A\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"bad quoting of path name\" err\n+'\n+\n+test_expect_success 'mktree fails on mode mismatch' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\n+\t# mode-type mismatch\n+\tprintf \"100644 tree $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"object type (tree) doesn${SQ}t match mode type (blob)\" err &&\n+\n+\t# mode-object mismatch (no --missing)\n+\tprintf \"100644 $tree_oid\\tA\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"object $tree_oid is a tree but specified type was (blob)\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497383","messageId":"e640a385b3d15a8c07d9ebd4c0872029cdc0d54f.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 09/17] mktree: add a --literally option","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:57Z","receivedAt":"2024-06-19T21:58:20Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nAdd the '--literally' option to 'git mktree' to allow constructing a tree\nwith invalid contents. For now, the only change this represents compared to\nthe normal 'git mktree' behavior is no longer sorting the inputs; in later\ncommits, deduplicaton and path validation will be added to the command and\n'--literally' will skip those as well.\n\nCertain tests use 'git mktree' to intentionally generate corrupt trees.\nUpdate these tests to use '--literally' so that they continue functioning\nproperly when additional input cleanup & validation is added to the base\ncommand. Note that, because 'mktree --literally' does not sort entries, some\nof the tests are updated to provide their inputs in tree order; otherwise,\nthe test would fail with an \"incorrect order\" error instead of the error the\ntest expects.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt       |  9 ++++++-\n builtin/mktree.c                   | 36 +++++++++++++++++++++++----\n t/t1010-mktree.sh                  | 40 ++++++++++++++++++++++++++++++\n t/t1014-read-tree-confusing.sh     |  6 ++---\n t/t1450-fsck.sh                    |  4 +--\n t/t1601-index-bogus.sh             |  2 +-\n t/t1700-split-index.sh             |  6 ++---\n t/t7008-filter-branch-null-sha1.sh |  6 ++---\n t/t7417-submodule-path-url.sh      |  2 +-\n t/t7450-bad-git-dotfiles.sh        |  8 +++---\n 10 files changed, 96 insertions(+), 23 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex c187403c6bd..5f3a6dfe38e 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -9,7 +9,7 @@ git-mktree - Build a tree-object from formatted tree entries\n SYNOPSIS\n --------\n [verse]\n-'git mktree' [-z] [--missing] [--batch]\n+'git mktree' [-z] [--missing] [--literally] [--batch]\n \n DESCRIPTION\n -----------\n@@ -28,6 +28,13 @@ OPTIONS\n \tobject.  This option has no effect on the treatment of gitlink entries\n \t(aka \"submodules\") which are always allowed to be missing.\n \n+--literally::\n+\tCreate the tree from the tree entries provided to stdin in the order\n+\tthey are provided without performing additional sorting,\n+\tdeduplication, or path validation on them. This option is primarily\n+\tuseful for creating invalid tree objects to use in tests of how Git\n+\tdeals with various forms of tree corruption.\n+\n --batch::\n \tAllow building of more than one tree object before exiting.  Each\n \ttree is separated by a single blank line. The final newline is\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex f509ed1a81f..4ff99d44d79 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -48,11 +48,11 @@ static void tree_entry_array_release(struct tree_entry_array *arr, int free_entr\n }\n \n static void append_to_tree(unsigned mode, struct object_id *oid, const char *path,\n-\t\t\t   struct tree_entry_array *arr)\n+\t\t\t   struct tree_entry_array *arr, int literally)\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n-\tif (strchr(path, '/'))\n+\tif (!literally && strchr(path, '/'))\n \t\tdie(\"path %s contains slash\", path);\n \n \tFLEX_ALLOC_MEM(ent, name, path, len);\n@@ -91,14 +91,35 @@ static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n \tstrbuf_release(&buf);\n }\n \n+static void write_tree_literally(struct tree_entry_array *arr,\n+\t\t\t\t struct object_id *oid)\n+{\n+\tstruct strbuf buf;\n+\tsize_t size = 0;\n+\n+\tfor (size_t i = 0; i < arr->nr; i++)\n+\t\tsize += 32 + arr->entries[i]->len;\n+\n+\tstrbuf_init(&buf, size);\n+\tfor (size_t i = 0; i < arr->nr; i++) {\n+\t\tstruct tree_entry *ent = arr->entries[i];\n+\t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n+\t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n+\t}\n+\n+\twrite_object_file(buf.buf, buf.len, OBJ_TREE, oid);\n+\tstrbuf_release(&buf);\n+}\n+\n static const char *mktree_usage[] = {\n-\t\"git mktree [-z] [--missing] [--batch]\",\n+\t\"git mktree [-z] [--missing] [--literally] [--batch]\",\n \tNULL\n };\n \n struct mktree_line_data {\n \tstruct tree_entry_array *arr;\n \tint allow_missing;\n+\tint literally;\n };\n \n static int mktree_line(unsigned int mode, struct object_id *oid,\n@@ -145,7 +166,7 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \t\t}\n \t}\n \n-\tappend_to_tree(mode, oid, path, data->arr);\n+\tappend_to_tree(mode, oid, path, data->arr, data->literally);\n \treturn 0;\n }\n \n@@ -162,6 +183,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tconst struct option option[] = {\n \t\tOPT_BOOL('z', NULL, &nul_term_line, N_(\"input is NUL terminated\")),\n \t\tOPT_BOOL(0, \"missing\", &mktree_line_data.allow_missing, N_(\"allow missing objects\")),\n+\t\tOPT_BOOL(0, \"literally\", &mktree_line_data.literally,\n+\t\t\t N_(\"do not sort, deduplicate, or validate paths of tree entries\")),\n \t\tOPT_BOOL(0, \"batch\", &is_batch_mode, N_(\"allow creation of more than one tree\")),\n \t\tOPT_END()\n \t};\n@@ -189,7 +212,10 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\t */\n \t\t\t; /* skip creating an empty tree */\n \t\t} else {\n-\t\t\twrite_tree(&arr, &oid);\n+\t\t\tif (mktree_line_data.literally)\n+\t\t\t\twrite_tree_literally(&arr, &oid);\n+\t\t\telse\n+\t\t\t\twrite_tree(&arr, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 48fc532e7af..961c0c3e55e 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -129,4 +129,44 @@ test_expect_success 'mktree fails on mode mismatch' '\n \ttest_grep \"object $tree_oid is a tree but specified type was (blob)\" err\n '\n \n+test_expect_success '--literally can create invalid trees' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse ${tree_oid}:one)\" &&\n+\n+\t# duplicate entries\n+\t{\n+\t\tprintf \"040000 tree $tree_oid\\tmy-tree\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest-file\\n\" &&\n+\t\tprintf \"100755 blob $blob_oid\\ttest-file\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\ttest_grep \"contains duplicate file entries\" err &&\n+\n+\t# disallowed path\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\t.git\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\ttest_grep \"contains ${SQ}.git${SQ}\" err &&\n+\n+\t# nested entry\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tdeeper/my-file\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\ttest_grep \"contains full pathnames\" err &&\n+\n+\t# bad entry ordering\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tB\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tA\\n\"\n+\t} | git mktree --literally >tree.bad &&\n+\tgit cat-file tree $(cat tree.bad) >top.bad &&\n+\ttest_must_fail git hash-object --stdin -t tree <top.bad 2>err &&\n+\ttest_grep \"not properly sorted\" err\n+'\n+\n test_done\ndiff --git a/t/t1014-read-tree-confusing.sh b/t/t1014-read-tree-confusing.sh\nindex 8ea8d36818b..762eb789704 100755\n--- a/t/t1014-read-tree-confusing.sh\n+++ b/t/t1014-read-tree-confusing.sh\n@@ -30,13 +30,13 @@ while read path pretty; do\n \tesac\n \ttest_expect_success \"reject $pretty at end of path\" '\n \t\tprintf \"100644 blob %s\\t%s\" \"$blob\" \"$path\" >tree &&\n-\t\tbogus=$(git mktree <tree) &&\n+\t\tbogus=$(git mktree --literally <tree) &&\n \t\ttest_must_fail git read-tree $bogus\n \t'\n \n \ttest_expect_success \"reject $pretty as subtree\" '\n \t\tprintf \"040000 tree %s\\t%s\" \"$tree\" \"$path\" >tree &&\n-\t\tbogus=$(git mktree <tree) &&\n+\t\tbogus=$(git mktree --literally <tree) &&\n \t\ttest_must_fail git read-tree $bogus\n \t'\n done <<-EOF\n@@ -58,7 +58,7 @@ test_expect_success 'utf-8 paths allowed with core.protectHFS off' '\n \ttest_when_finished \"git read-tree HEAD\" &&\n \ttest_config core.protectHFS false &&\n \tprintf \"100644 blob %s\\t%s\" \"$blob\" \".gi${u200c}t\" >tree &&\n-\tok=$(git mktree <tree) &&\n+\tok=$(git mktree --literally <tree) &&\n \tgit read-tree $ok\n '\n \ndiff --git a/t/t1450-fsck.sh b/t/t1450-fsck.sh\nindex 8a456b1142d..532d2770e88 100755\n--- a/t/t1450-fsck.sh\n+++ b/t/t1450-fsck.sh\n@@ -316,7 +316,7 @@ check_duplicate_names () {\n \t\t\t*)  printf \"100644 blob %s\\t%s\\n\" $blob \"$name\" ;;\n \t\t\tesac\n \t\tdone >badtree &&\n-\t\tbadtree=$(git mktree <badtree) &&\n+\t\tbadtree=$(git mktree --literally <badtree) &&\n \t\ttest_must_fail git fsck 2>out &&\n \t\ttest_grep \"$badtree\" out &&\n \t\ttest_grep \"error in tree .*contains duplicate file entries\" out\n@@ -614,7 +614,7 @@ while read name path pretty; do\n \t\t\ttree=$(git rev-parse HEAD^{tree}) &&\n \t\t\tvalue=$(eval \"echo \\$$type\") &&\n \t\t\tprintf \"$mode $type %s\\t%s\" \"$value\" \"$path\" >bad &&\n-\t\t\tbad_tree=$(git mktree <bad) &&\n+\t\t\tbad_tree=$(git mktree --literally <bad) &&\n \t\t\tgit fsck 2>out &&\n \t\t\ttest_grep \"warning.*tree $bad_tree\" out\n \t\t)'\ndiff --git a/t/t1601-index-bogus.sh b/t/t1601-index-bogus.sh\nindex 4171f1e1410..54e8ae038b7 100755\n--- a/t/t1601-index-bogus.sh\n+++ b/t/t1601-index-bogus.sh\n@@ -4,7 +4,7 @@ test_description='test handling of bogus index entries'\n . ./test-lib.sh\n \n test_expect_success 'create tree with null sha1' '\n-\ttree=$(printf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" | git mktree)\n+\ttree=$(printf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" | git mktree --literally)\n '\n \n test_expect_success 'read-tree refuses to read null sha1' '\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex ac4a5b2734c..97b58aa3cca 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -478,12 +478,12 @@ test_expect_success 'writing split index with null sha1 does not write cache tre\n \tgit config splitIndex.maxPercentChange 0 &&\n \tgit commit -m \"commit\" &&\n \t{\n-\t\tgit ls-tree HEAD &&\n-\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\"\n+\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" &&\n+\t\tgit ls-tree HEAD\n \t} >broken-tree &&\n \techo \"add broken entry\" >msg &&\n \n-\ttree=$(git mktree <broken-tree) &&\n+\ttree=$(git mktree --literally <broken-tree) &&\n \ttest_tick &&\n \tcommit=$(git commit-tree $tree -p HEAD <msg) &&\n \tgit update-ref HEAD \"$commit\" &&\ndiff --git a/t/t7008-filter-branch-null-sha1.sh b/t/t7008-filter-branch-null-sha1.sh\nindex 93fbc92b8db..a1b4c295c01 100755\n--- a/t/t7008-filter-branch-null-sha1.sh\n+++ b/t/t7008-filter-branch-null-sha1.sh\n@@ -12,12 +12,12 @@ test_expect_success 'setup: base commits' '\n \n test_expect_success 'setup: a commit with a bogus null sha1 in the tree' '\n \t{\n-\t\tgit ls-tree HEAD &&\n-\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\"\n+\t\tprintf \"160000 commit $ZERO_OID\\\\tbroken\\\\n\" &&\n+\t\tgit ls-tree HEAD\n \t} >broken-tree &&\n \techo \"add broken entry\" >msg &&\n \n-\ttree=$(git mktree <broken-tree) &&\n+\ttree=$(git mktree --literally <broken-tree) &&\n \ttest_tick &&\n \tcommit=$(git commit-tree $tree -p HEAD <msg) &&\n \tgit update-ref HEAD \"$commit\"\ndiff --git a/t/t7417-submodule-path-url.sh b/t/t7417-submodule-path-url.sh\nindex dbbb3853dc0..5d3c98e99a7 100755\n--- a/t/t7417-submodule-path-url.sh\n+++ b/t/t7417-submodule-path-url.sh\n@@ -42,7 +42,7 @@ test_expect_success MINGW 'submodule paths disallows trailing spaces' '\n \ttree=$(git -C super write-tree) &&\n \tgit -C super ls-tree $tree >tree &&\n \tsed \"s/sub/sub /\" <tree >tree.new &&\n-\ttree=$(git -C super mktree <tree.new) &&\n+\ttree=$(git -C super mktree --literally <tree.new) &&\n \tcommit=$(echo with space | git -C super commit-tree $tree) &&\n \tgit -C super update-ref refs/heads/main $commit &&\n \ndiff --git a/t/t7450-bad-git-dotfiles.sh b/t/t7450-bad-git-dotfiles.sh\nindex 4a9c22c9e2b..de2d45d2244 100755\n--- a/t/t7450-bad-git-dotfiles.sh\n+++ b/t/t7450-bad-git-dotfiles.sh\n@@ -203,11 +203,11 @@ check_dotx_symlink () {\n \t\t\tcontent=$(git hash-object -w ../.gitmodules) &&\n \t\t\ttarget=$(printf \"$tricky\" | git hash-object -w --stdin) &&\n \t\t\t{\n-\t\t\t\tprintf \"100644 blob $content\\t$tricky\\n\" &&\n-\t\t\t\tprintf \"120000 blob $target\\t$path\\n\"\n+\t\t\t\tprintf \"120000 blob $target\\t$path\\n\" &&\n+\t\t\t\tprintf \"100644 blob $content\\t$tricky\\n\"\n \t\t\t} >bad-tree\n \t\t) &&\n-\t\ttree=$(git -C $dir mktree <$dir/bad-tree)\n+\t\ttree=$(git -C $dir mktree --literally <$dir/bad-tree)\n \t'\n \n \ttest_expect_success \"fsck detects symlinked $name ($type)\" '\n@@ -261,7 +261,7 @@ test_expect_success 'fsck detects non-blob .gitmodules' '\n \t\tcp ../.gitmodules subdir/file &&\n \t\tgit add subdir/file &&\n \t\tgit commit -m ok &&\n-\t\tgit ls-tree HEAD | sed s/subdir/.gitmodules/ | git mktree &&\n+\t\tgit ls-tree HEAD | sed s/subdir/.gitmodules/ | git mktree --literally &&\n \n \t\ttest_must_fail git fsck 2>output &&\n \t\ttest_grep gitmodulesBlob output\n-- \ngitgitgadget\n\n"},{"id":"497384","messageId":"2eb207064f80d48a7db5617feea417a015bb6082.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 10/17] mktree: validate paths more carefully","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:58Z","receivedAt":"2024-06-19T21:58:22Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUse 'verify_path' to validate the paths provided as tree entries, ensuring\nwe do not create entries with paths not allowed in trees (e.g., .git). Also,\nremove trailing slashes on directories before validating, allowing users to\nprovide 'folder-name/' as the path for a tree object entry.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c  | 20 +++++++++++++++++---\n t/t1010-mktree.sh | 33 +++++++++++++++++++++++++++++++++\n 2 files changed, 50 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 4ff99d44d79..8f0af24b6b1 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -8,6 +8,7 @@\n #include \"hex.h\"\n #include \"index-info.h\"\n #include \"quote.h\"\n+#include \"read-cache-ll.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n #include \"parse-options.h\"\n@@ -52,10 +53,23 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n {\n \tstruct tree_entry *ent;\n \tsize_t len = strlen(path);\n-\tif (!literally && strchr(path, '/'))\n-\t\tdie(\"path %s contains slash\", path);\n \n-\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\tif (literally) {\n+\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\t} else {\n+\t\t/* Normalize and validate entry path */\n+\t\tif (S_ISDIR(mode)) {\n+\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n+\t\t\t\tlen--;\n+\t\t}\n+\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\n+\t\tif (!verify_path(ent->name, mode))\n+\t\t\tdie(_(\"invalid path '%s'\"), path);\n+\t\tif (strchr(ent->name, '/'))\n+\t\t\tdie(\"path %s contains slash\", path);\n+\t}\n+\n \tent->mode = mode;\n \tent->len = len;\n \toidcpy(&ent->oid, oid);\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 961c0c3e55e..7e750530455 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -169,4 +169,37 @@ test_expect_success '--literally can create invalid trees' '\n \ttest_grep \"not properly sorted\" err\n '\n \n+test_expect_success 'mktree validates path' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:a/one)\" &&\n+\thead_oid=\"$(git rev-parse HEAD)\" &&\n+\n+\t# Valid: tree with or without trailing slash, blob without trailing slash\n+\t{\n+\t\tprintf \"040000 tree $tree_oid\\tfolder1/\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tfolder2\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfile.txt\\n\"\n+\t} | git mktree >actual &&\n+\n+\t# Invalid: blob with trailing slash\n+\tprintf \"100644 blob $blob_oid\\ttest/\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"invalid path ${SQ}test/${SQ}\" err &&\n+\n+\t# Invalid: dotdot\n+\tprintf \"040000 tree $tree_oid\\t../\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"invalid path ${SQ}../${SQ}\" err &&\n+\n+\t# Invalid: dot\n+\tprintf \"040000 tree $tree_oid\\t.\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"invalid path ${SQ}.${SQ}\" err &&\n+\n+\t# Invalid: .git\n+\tprintf \"040000 tree $tree_oid\\t.git/\" |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"invalid path ${SQ}.git/${SQ}\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497385","messageId":"fb555658057f834d94f232f1d8b380a6304a3671.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 11/17] mktree: overwrite duplicate entries","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:57:59Z","receivedAt":"2024-06-19T21:58:23Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf multiple tree entries with the same name are provided as input to\n'mktree', only write the last one to the tree. Entries are considered\nduplicates if they have identical names (*not* considering mode); if a blob\nand a tree with the same name are provided, only the last one will be\nwritten to the tree. A tree with duplicate entries is invalid (per 'git\nfsck'), so that condition should be avoided wherever possible.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  3 ++-\n builtin/mktree.c             | 45 ++++++++++++++++++++++++++++++++----\n t/t1010-mktree.sh            | 36 +++++++++++++++++++++++++++--\n 3 files changed, 77 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 5f3a6dfe38e..cf1fd82f754 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -54,7 +54,8 @@ cannot be represented in a tree object. The command will fail without\n writing the tree if a higher order stage is specified for any entry.\n \n The order of the tree entries is normalized by `mktree` so pre-sorting the\n-input by path is not required.\n+input by path is not required. Multiple entries provided with the same path\n+are deduplicated, with only the last one specified added to the tree.\n \n GIT\n ---\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 8f0af24b6b1..a91d3a7b028 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -15,6 +15,9 @@\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n+\t/* Internal */\n+\tsize_t order;\n+\n \tunsigned mode;\n \tstruct object_id oid;\n \tint len;\n@@ -74,15 +77,49 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \tent->len = len;\n \toidcpy(&ent->oid, oid);\n \n+\tent->order = arr->nr;\n \ttree_entry_array_push(arr, ent);\n }\n \n-static int ent_compare(const void *a_, const void *b_)\n+static int ent_compare(const void *a_, const void *b_, void *ctx)\n {\n+\tint cmp;\n \tstruct tree_entry *a = *(struct tree_entry **)a_;\n \tstruct tree_entry *b = *(struct tree_entry **)b_;\n-\treturn base_name_compare(a->name, a->len, a->mode,\n-\t\t\t\t b->name, b->len, b->mode);\n+\tint ignore_mode = *((int *)ctx);\n+\n+\tif (ignore_mode)\n+\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n+\telse\n+\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n+\t\t\t\t\tb->name, b->len, b->mode);\n+\treturn cmp ? cmp : b->order - a->order;\n+}\n+\n+static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n+{\n+\tsize_t count = arr->nr;\n+\tstruct tree_entry *prev = NULL;\n+\n+\tint ignore_mode = 1;\n+\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n+\n+\tarr->nr = 0;\n+\tfor (size_t i = 0; i < count; i++) {\n+\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\tif (prev &&\n+\t\t    !name_compare(prev->name, prev->len,\n+\t\t\t\t  curr->name, curr->len)) {\n+\t\t\tFREE_AND_NULL(curr);\n+\t\t} else {\n+\t\t\tarr->entries[arr->nr++] = curr;\n+\t\t\tprev = curr;\n+\t\t}\n+\t}\n+\n+\t/* Sort again to order the entries for tree insertion */\n+\tignore_mode = 0;\n+\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n@@ -90,7 +127,7 @@ static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n \tstruct strbuf buf;\n \tsize_t size = 0;\n \n-\tQSORT(arr->entries, arr->nr, ent_compare);\n+\tsort_and_dedup_tree_entry_array(arr);\n \tfor (size_t i = 0; i < arr->nr; i++)\n \t\tsize += 32 + arr->entries[i]->len;\n \ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 7e750530455..08760141d6f 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -6,11 +6,16 @@ TEST_PASSES_SANITIZE_LEAK=true\n . ./test-lib.sh\n \n test_expect_success setup '\n-\tfor d in a a- a0\n+\tfor d in folder folder- folder0\n \tdo\n \t\tmkdir \"$d\" && echo \"$d/one\" >\"$d/one\" &&\n \t\tgit add \"$d\" || return 1\n \tdone &&\n+\tfor f in before folder.txt later\n+\tdo\n+\t\techo \"$f\" >\"$f\" &&\n+\t\tgit add \"$f\" || return 1\n+\tdone &&\n \techo zero >one &&\n \tgit update-index --add --info-only one &&\n \tgit write-tree --missing-ok >tree.missing &&\n@@ -171,7 +176,7 @@ test_expect_success '--literally can create invalid trees' '\n \n test_expect_success 'mktree validates path' '\n \ttree_oid=\"$(cat tree)\" &&\n-\tblob_oid=\"$(git rev-parse $tree_oid:a/one)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n \thead_oid=\"$(git rev-parse HEAD)\" &&\n \n \t# Valid: tree with or without trailing slash, blob without trailing slash\n@@ -202,4 +207,31 @@ test_expect_success 'mktree validates path' '\n \ttest_grep \"invalid path ${SQ}.git/${SQ}\" err\n '\n \n+test_expect_success 'mktree with duplicate entries' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\ttest\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest0\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest-\\n\"\n+\t} >top.dup &&\n+\tgit mktree <top.dup >tree.actual &&\n+\n+\t{\n+\t\tprintf \"160000 commit $head_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest0\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497386","messageId":"2333775ba5bd71766a6aece87e39a6d189aeaead.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 12/17] mktree: create tree using an in-core index","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:00Z","receivedAt":"2024-06-19T21:58:25Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nRather than manually write out the contents of a tree object file, construct\nan in-memory sparse index from the provided tree entries and create the tree\nby writing out its corresponding cache tree.\n\nThis patch does not change the behavior of the 'mktree' command. However,\nconstructing the tree this way will substantially simplify future extensions\nto the command's functionality, including handling deeper-than-toplevel tree\nentries and applying the provided entries to an existing tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 74 +++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 55 insertions(+), 19 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex a91d3a7b028..3ce8d3dc524 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -4,6 +4,7 @@\n  * Copyright (c) Junio C Hamano, 2006, 2009\n  */\n #include \"builtin.h\"\n+#include \"cache-tree.h\"\n #include \"gettext.h\"\n #include \"hex.h\"\n #include \"index-info.h\"\n@@ -24,6 +25,11 @@ struct tree_entry {\n \tchar name[FLEX_ARRAY];\n };\n \n+static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n+{\n+\treturn S_ISDIR(mode) ? pathlen - 1 : pathlen;\n+}\n+\n struct tree_entry_array {\n \tsize_t nr, alloc;\n \tstruct tree_entry **entries;\n@@ -60,17 +66,25 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \tif (literally) {\n \t\tFLEX_ALLOC_MEM(ent, name, path, len);\n \t} else {\n+\t\tsize_t len_to_copy = len;\n+\n \t\t/* Normalize and validate entry path */\n \t\tif (S_ISDIR(mode)) {\n-\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n-\t\t\t\tlen--;\n+\t\t\twhile(len_to_copy > 0 && is_dir_sep(path[len_to_copy - 1]))\n+\t\t\t\tlen_to_copy--;\n+\t\t\tlen = len_to_copy + 1; /* add space for trailing slash */\n \t\t}\n-\t\tFLEX_ALLOC_MEM(ent, name, path, len);\n+\t\tent = xcalloc(1, st_add3(sizeof(struct tree_entry), len, 1));\n+\t\tmemcpy(ent->name, path, len_to_copy);\n \n \t\tif (!verify_path(ent->name, mode))\n \t\t\tdie(_(\"invalid path '%s'\"), path);\n \t\tif (strchr(ent->name, '/'))\n \t\t\tdie(\"path %s contains slash\", path);\n+\n+\t\t/* Add trailing slash to dir */\n+\t\tif (S_ISDIR(mode))\n+\t\t\tent->name[len - 1] = '/';\n \t}\n \n \tent->mode = mode;\n@@ -88,11 +102,14 @@ static int ent_compare(const void *a_, const void *b_, void *ctx)\n \tstruct tree_entry *b = *(struct tree_entry **)b_;\n \tint ignore_mode = *((int *)ctx);\n \n-\tif (ignore_mode)\n-\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n-\telse\n-\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n-\t\t\t\t\tb->name, b->len, b->mode);\n+\tsize_t a_len = a->len, b_len = b->len;\n+\n+\tif (ignore_mode) {\n+\t\ta_len = df_path_len(a_len, a->mode);\n+\t\tb_len = df_path_len(b_len, b->mode);\n+\t}\n+\n+\tcmp = name_compare(a->name, a_len, b->name, b_len);\n \treturn cmp ? cmp : b->order - a->order;\n }\n \n@@ -108,8 +125,8 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tfor (size_t i = 0; i < count; i++) {\n \t\tstruct tree_entry *curr = arr->entries[i];\n \t\tif (prev &&\n-\t\t    !name_compare(prev->name, prev->len,\n-\t\t\t\t  curr->name, curr->len)) {\n+\t\t    !name_compare(prev->name, df_path_len(prev->len, prev->mode),\n+\t\t\t\t  curr->name, df_path_len(curr->len, curr->mode))) {\n \t\t\tFREE_AND_NULL(curr);\n \t\t} else {\n \t\t\tarr->entries[arr->nr++] = curr;\n@@ -122,24 +139,43 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n+static int add_tree_entry_to_index(struct index_state *istate,\n+\t\t\t\t   struct tree_entry *ent)\n+{\n+\tstruct cache_entry *ce;\n+\tstruct strbuf ce_name = STRBUF_INIT;\n+\tstrbuf_add(&ce_name, ent->name, ent->len);\n+\n+\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n+\tif (!ce)\n+\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\n+\tadd_index_entry(istate, ce, ADD_CACHE_JUST_APPEND);\n+\tstrbuf_release(&ce_name);\n+\treturn 0;\n+}\n+\n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n-\tstruct strbuf buf;\n-\tsize_t size = 0;\n+\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n+\tistate.sparse_index = 1;\n \n \tsort_and_dedup_tree_entry_array(arr);\n-\tfor (size_t i = 0; i < arr->nr; i++)\n-\t\tsize += 32 + arr->entries[i]->len;\n \n-\tstrbuf_init(&buf, size);\n+\t/* Construct an in-memory index from the provided entries */\n \tfor (size_t i = 0; i < arr->nr; i++) {\n \t\tstruct tree_entry *ent = arr->entries[i];\n-\t\tstrbuf_addf(&buf, \"%o %s%c\", ent->mode, ent->name, '\\0');\n-\t\tstrbuf_add(&buf, ent->oid.hash, the_hash_algo->rawsz);\n+\n+\t\tif (add_tree_entry_to_index(&istate, ent))\n+\t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n \t}\n \n-\twrite_object_file(buf.buf, buf.len, OBJ_TREE, oid);\n-\tstrbuf_release(&buf);\n+\t/* Write out new tree */\n+\tif (cache_tree_update(&istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n+\t\tdie(_(\"failed to write tree\"));\n+\toidcpy(oid, &istate.cache_tree->oid);\n+\n+\trelease_index(&istate);\n }\n \n static void write_tree_literally(struct tree_entry_array *arr,\n-- \ngitgitgadget\n\n"},{"id":"497387","messageId":"56f28efff5404a3fa22bd544d6de8ce2d919b78a.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 13/17] mktree: use iterator struct to add tree entries to index","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:01Z","receivedAt":"2024-06-19T21:58:26Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nCreate 'struct tree_entry_iterator' to manage iteration through a 'struct\ntree_entry_array'. Using an iterator allows for conditional iteration; this\nfunctionality will be necessary in later commits when performing parallel\niteration through multiple sets of tree entries.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 39 ++++++++++++++++++++++++++++++++++++---\n 1 file changed, 36 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 3ce8d3dc524..344c9b9b6fe 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -139,6 +139,35 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n }\n \n+struct tree_entry_iterator {\n+\tstruct tree_entry *current;\n+\n+\t/* private */\n+\tstruct {\n+\t\tstruct tree_entry_array *arr;\n+\t\tsize_t idx;\n+\t} priv;\n+};\n+\n+static void tree_entry_iterator_init(struct tree_entry_iterator *iter,\n+\t\t\t\t     struct tree_entry_array *arr)\n+{\n+\titer->priv.arr = arr;\n+\titer->priv.idx = 0;\n+\titer->current = 0 < arr->nr ? arr->entries[0] : NULL;\n+}\n+\n+/*\n+ * Advance the tree entry iterator to the next entry in the array. If no\n+ * entries remain, 'current' is set to NULL.\n+ */\n+static void tree_entry_iterator_advance(struct tree_entry_iterator *iter)\n+{\n+\titer->current = (iter->priv.idx + 1) < iter->priv.arr->nr\n+\t\t\t? iter->priv.arr->entries[++iter->priv.idx]\n+\t\t\t: NULL;\n+}\n+\n static int add_tree_entry_to_index(struct index_state *istate,\n \t\t\t\t   struct tree_entry *ent)\n {\n@@ -157,14 +186,18 @@ static int add_tree_entry_to_index(struct index_state *istate,\n \n static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n {\n+\tstruct tree_entry_iterator iter = { NULL };\n \tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n \tistate.sparse_index = 1;\n \n \tsort_and_dedup_tree_entry_array(arr);\n \n-\t/* Construct an in-memory index from the provided entries */\n-\tfor (size_t i = 0; i < arr->nr; i++) {\n-\t\tstruct tree_entry *ent = arr->entries[i];\n+\ttree_entry_iterator_init(&iter, arr);\n+\n+\t/* Construct an in-memory index from the provided entries & base tree */\n+\twhile (iter.current) {\n+\t\tstruct tree_entry *ent = iter.current;\n+\t\ttree_entry_iterator_advance(&iter);\n \n \t\tif (add_tree_entry_to_index(&istate, ent))\n \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n-- \ngitgitgadget\n\n"},{"id":"497388","messageId":"6f6d78ae7acb35991afbeaef9b61af892af93ca1.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 14/17] mktree: add directory-file conflict hashmap","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:02Z","receivedAt":"2024-06-19T21:58:27Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nCreate a hashmap member of a 'struct tree_entry_array' that contains all of\nthe (de-duplicated) provided tree entries, indexed by the hash of their path\nwith *no* trailing slash. This hashmap will be used in a later commit to\navoid adding a file to an existing tree that has the same path as a\ndirectory, or vice versa.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n builtin/mktree.c | 38 ++++++++++++++++++++++++++++++++++++++\n 1 file changed, 38 insertions(+)\n\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 344c9b9b6fe..b4d71dcdd02 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -16,6 +16,8 @@\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n+\tstruct hashmap_entry ent;\n+\n \t/* Internal */\n \tsize_t order;\n \n@@ -33,8 +35,33 @@ static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n struct tree_entry_array {\n \tsize_t nr, alloc;\n \tstruct tree_entry **entries;\n+\n+\tstruct hashmap df_name_hash;\n };\n \n+static int df_name_hash_cmp(const void *cmp_data UNUSED,\n+\t\t\t    const struct hashmap_entry *eptr,\n+\t\t\t    const struct hashmap_entry *entry_or_key,\n+\t\t\t    const void *keydata UNUSED)\n+{\n+\tconst struct tree_entry *e1, *e2;\n+\tsize_t e1_len, e2_len;\n+\n+\te1 = container_of(eptr, const struct tree_entry, ent);\n+\te2 = container_of(entry_or_key, const struct tree_entry, ent);\n+\n+\te1_len = df_path_len(e1->len, e1->mode);\n+\te2_len = df_path_len(e2->len, e2->mode);\n+\n+\treturn e1_len != e2_len ||\n+\t       name_compare(e1->name, e1_len, e2->name, e2_len);\n+}\n+\n+static void tree_entry_array_init(struct tree_entry_array *arr)\n+{\n+\thashmap_init(&arr->df_name_hash, df_name_hash_cmp, NULL, 0);\n+}\n+\n static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entry *ent)\n {\n \tALLOC_GROW(arr->entries, arr->nr + 1, arr->alloc);\n@@ -48,6 +75,7 @@ static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entrie\n \t\t\tFREE_AND_NULL(arr->entries[i]);\n \t}\n \tarr->nr = 0;\n+\thashmap_clear(&arr->df_name_hash);\n }\n \n static void tree_entry_array_release(struct tree_entry_array *arr, int free_entries)\n@@ -137,6 +165,14 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \t/* Sort again to order the entries for tree insertion */\n \tignore_mode = 0;\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n+\n+\t/* Finally, initialize the directory-file conflict hash map */\n+\tfor (size_t i = 0; i < count; i++) {\n+\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\thashmap_entry_init(&curr->ent,\n+\t\t\t\t   memhash(curr->name, df_path_len(curr->len, curr->mode)));\n+\t\thashmap_put(&arr->df_name_hash, &curr->ent);\n+\t}\n }\n \n struct tree_entry_iterator {\n@@ -311,6 +347,8 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n \n+\ttree_entry_array_init(&arr);\n+\n \tdo {\n \t\tret = read_index_info(nul_term_line, mktree_line, &mktree_line_data, &line);\n \t\tif (ret < 0)\n-- \ngitgitgadget\n\n"},{"id":"497389","messageId":"4b88f84b933b1598d12e3620f0c9fb85c559e8fb.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 15/17] mktree: optionally add to an existing tree","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:03Z","receivedAt":"2024-06-19T21:58:29Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nAllow users to specify a single \"tree-ish\" value as a positional argument.\nIf provided, the contents of the given tree serve as the basis for the new\ntree (or trees, in --batch mode) created by 'mktree', on top of which all of\nthe stdin-provided tree entries are applied.\n\nAt a high level, the entries are \"applied\" to a base tree by iterating\nthrough the base tree using 'read_tree' in parallel with iterating through\nthe sorted & deduplicated stdin entries via their iterator. That is, for\neach call to the 'build_index_from_tree callback of 'read_tree':\n\n* If the iterator entry precedes the base tree entry, add it to the in-core\n  index, increment the iterator, and repeat.\n* If the iterator entry has the same name as the base tree entry, add the\n  iterator entry to the index, increment the iterator, and return from the\n  callback to continue the 'read_tree' iteration.\n* If the iterator entry follows the base tree entry, first check\n  'df_name_hash' to ensure we won't be adding an entry with the same name\n  later (with a different mode). If there's no directory/file conflict, add\n  the base tree entry to the index. In either case, return from the callback\n  to continue the 'read_tree' iteration.\n\nFinally, once 'read_tree' is complete, add the remaining entries in the\niterator to the index and write out the index as a tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |   7 +-\n builtin/mktree.c             | 138 +++++++++++++++++++++++++++++------\n t/t1010-mktree.sh            |  36 +++++++++\n 3 files changed, 159 insertions(+), 22 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex cf1fd82f754..260d0e0bd7b 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -9,7 +9,7 @@ git-mktree - Build a tree-object from formatted tree entries\n SYNOPSIS\n --------\n [verse]\n-'git mktree' [-z] [--missing] [--literally] [--batch]\n+'git mktree' [-z] [--missing] [--literally] [--batch] [<tree-ish>]\n \n DESCRIPTION\n -----------\n@@ -41,6 +41,11 @@ OPTIONS\n \toptional.  Note - if the `-z` option is used, lines are terminated\n \twith NUL.\n \n+<tree-ish>::\n+\tIf provided, the tree entries provided in stdin are added to this\n+\ttree rather than a new empty one, replacing existing entries with\n+\tidentical names. Not compatible with `--literally`.\n+\n INPUT FORMAT\n ------------\n Tree entries may be specified in any of the formats compatible with the\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex b4d71dcdd02..96f06547a2a 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -12,7 +12,9 @@\n #include \"read-cache-ll.h\"\n #include \"strbuf.h\"\n #include \"tree.h\"\n+#include \"object-name.h\"\n #include \"parse-options.h\"\n+#include \"pathspec.h\"\n #include \"object-store-ll.h\"\n \n struct tree_entry {\n@@ -204,47 +206,124 @@ static void tree_entry_iterator_advance(struct tree_entry_iterator *iter)\n \t\t\t: NULL;\n }\n \n-static int add_tree_entry_to_index(struct index_state *istate,\n+struct build_index_data {\n+\tstruct tree_entry_iterator iter;\n+\tstruct hashmap *df_name_hash;\n+\tstruct index_state istate;\n+};\n+\n+static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n \tstruct cache_entry *ce;\n-\tstruct strbuf ce_name = STRBUF_INIT;\n-\tstrbuf_add(&ce_name, ent->name, ent->len);\n-\n-\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n+\tce = make_cache_entry(&data->istate, ent->mode, &ent->oid, ent->name, 0, 0);\n \tif (!ce)\n \t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n \n-\tadd_index_entry(istate, ce, ADD_CACHE_JUST_APPEND);\n-\tstrbuf_release(&ce_name);\n+\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n \treturn 0;\n }\n \n-static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n+static int build_index_from_tree(const struct object_id *oid,\n+\t\t\t\t struct strbuf *base, const char *filename,\n+\t\t\t\t unsigned mode, void *context)\n {\n-\tstruct tree_entry_iterator iter = { NULL };\n-\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n-\tistate.sparse_index = 1;\n+\tint result;\n+\tstruct tree_entry *base_tree_ent;\n+\tstruct build_index_data *cbdata = context;\n+\tsize_t filename_len = strlen(filename);\n+\tsize_t path_len = S_ISDIR(mode) ? st_add3(filename_len, base->len, 1)\n+\t\t\t\t\t: st_add(filename_len, base->len);\n+\n+\t/* Create a tree entry from the current entry in read_tree iteration */\n+\tbase_tree_ent = xcalloc(1, st_add3(sizeof(struct tree_entry), path_len, 1));\n+\tbase_tree_ent->len = path_len;\n+\tbase_tree_ent->mode = mode;\n+\toidcpy(&base_tree_ent->oid, oid);\n+\n+\tmemcpy(base_tree_ent->name, base->buf, base->len);\n+\tmemcpy(base_tree_ent->name + base->len, filename, filename_len);\n+\tif (S_ISDIR(mode))\n+\t\tbase_tree_ent->name[base_tree_ent->len - 1] = '/';\n+\n+\twhile (cbdata->iter.current) {\n+\t\tstruct tree_entry *ent = cbdata->iter.current;\n+\n+\t\tint cmp = name_compare(ent->name, ent->len,\n+\t\t\t\t       base_tree_ent->name, base_tree_ent->len);\n+\t\tif (!cmp || cmp < 0) {\n+\t\t\ttree_entry_iterator_advance(&cbdata->iter);\n+\n+\t\t\tif (add_tree_entry_to_index(cbdata, ent) < 0) {\n+\t\t\t\tresult = error(_(\"failed to add tree entry '%s'\"), ent->name);\n+\t\t\t\tgoto cleanup_and_return;\n+\t\t\t}\n+\n+\t\t\tif (!cmp) {\n+\t\t\t\tresult = 0;\n+\t\t\t\tgoto cleanup_and_return;\n+\t\t\t} else\n+\t\t\t\tcontinue;\n+\t\t}\n+\n+\t\tbreak;\n+\t}\n+\n+\t/*\n+\t * If the tree entry should be replaced with an entry with the same name\n+\t * (but different mode), skip it.\n+\t */\n+\thashmap_entry_init(&base_tree_ent->ent,\n+\t\t\t   memhash(base_tree_ent->name, df_path_len(base_tree_ent->len, base_tree_ent->mode)));\n+\tif (hashmap_get_entry(cbdata->df_name_hash, base_tree_ent, ent, NULL)) {\n+\t\tresult = 0;\n+\t\tgoto cleanup_and_return;\n+\t}\n+\n+\tif (add_tree_entry_to_index(cbdata, base_tree_ent)) {\n+\t\tresult = -1;\n+\t\tgoto cleanup_and_return;\n+\t}\n+\n+\tresult = 0;\n+\n+cleanup_and_return:\n+\tFREE_AND_NULL(base_tree_ent);\n+\treturn result;\n+}\n+\n+static void write_tree(struct tree_entry_array *arr, struct tree *base_tree,\n+\t\t       struct object_id *oid)\n+{\n+\tstruct build_index_data cbdata = { 0 };\n+\tstruct pathspec ps = { 0 };\n \n \tsort_and_dedup_tree_entry_array(arr);\n \n-\ttree_entry_iterator_init(&iter, arr);\n+\tindex_state_init(&cbdata.istate, the_repository);\n+\tcbdata.istate.sparse_index = 1;\n+\ttree_entry_iterator_init(&cbdata.iter, arr);\n+\tcbdata.df_name_hash = &arr->df_name_hash;\n \n \t/* Construct an in-memory index from the provided entries & base tree */\n-\twhile (iter.current) {\n-\t\tstruct tree_entry *ent = iter.current;\n-\t\ttree_entry_iterator_advance(&iter);\n+\tif (base_tree &&\n+\t    read_tree(the_repository, base_tree, &ps, build_index_from_tree, &cbdata) < 0)\n+\t\tdie(_(\"failed to create tree\"));\n+\n+\twhile (cbdata.iter.current) {\n+\t\tstruct tree_entry *ent = cbdata.iter.current;\n+\t\ttree_entry_iterator_advance(&cbdata.iter);\n \n-\t\tif (add_tree_entry_to_index(&istate, ent))\n+\t\tif (add_tree_entry_to_index(&cbdata, ent))\n \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n \t}\n \n \t/* Write out new tree */\n-\tif (cache_tree_update(&istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n+\tif (cache_tree_update(&cbdata.istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n \t\tdie(_(\"failed to write tree\"));\n-\toidcpy(oid, &istate.cache_tree->oid);\n+\toidcpy(oid, &cbdata.istate.cache_tree->oid);\n \n-\trelease_index(&istate);\n+\trelease_index(&cbdata.istate);\n }\n \n static void write_tree_literally(struct tree_entry_array *arr,\n@@ -268,7 +347,7 @@ static void write_tree_literally(struct tree_entry_array *arr,\n }\n \n static const char *mktree_usage[] = {\n-\t\"git mktree [-z] [--missing] [--literally] [--batch]\",\n+\t\"git mktree [-z] [--missing] [--literally] [--batch] [<tree-ish>]\",\n \tNULL\n };\n \n@@ -334,6 +413,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \tstruct tree_entry_array arr = { 0 };\n \tstruct mktree_line_data mktree_line_data = { .arr = &arr };\n \tstruct strbuf line = STRBUF_INIT;\n+\tstruct tree *base_tree = NULL;\n \tint ret;\n \n \tconst struct option option[] = {\n@@ -346,6 +426,22 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t};\n \n \tac = parse_options(ac, av, prefix, option, mktree_usage, 0);\n+\tif (ac > 1)\n+\t\tusage_with_options(mktree_usage, option);\n+\n+\tif (ac) {\n+\t\tstruct object_id base_tree_oid;\n+\n+\t\tif (mktree_line_data.literally)\n+\t\t\tdie(_(\"option '%s' and tree-ish cannot be used together\"), \"--literally\");\n+\n+\t\tif (repo_get_oid(the_repository, av[0], &base_tree_oid))\n+\t\t\tdie(_(\"not a valid object name %s\"), av[0]);\n+\n+\t\tbase_tree = parse_tree_indirect(&base_tree_oid);\n+\t\tif (!base_tree)\n+\t\t\tdie(_(\"not a tree object: %s\"), oid_to_hex(&base_tree_oid));\n+\t}\n \n \ttree_entry_array_init(&arr);\n \n@@ -373,7 +469,7 @@ int cmd_mktree(int ac, const char **av, const char *prefix)\n \t\t\tif (mktree_line_data.literally)\n \t\t\t\twrite_tree_literally(&arr, &oid);\n \t\t\telse\n-\t\t\t\twrite_tree(&arr, &oid);\n+\t\t\t\twrite_tree(&arr, base_tree, &oid);\n \t\t\tputs(oid_to_hex(&oid));\n \t\t\tfflush(stdout);\n \t\t}\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 08760141d6f..435ac23bd50 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -234,4 +234,40 @@ test_expect_success 'mktree with duplicate entries' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree with base tree' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest0\\n\"\n+\t} >top.base &&\n+\tgit mktree <top.base >tree.base &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tz\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.xyz\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ta\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest\\n\"\n+\t} >top.append &&\n+\tgit mktree $(cat tree.base) <top.append >tree.actual &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\ta\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\ttest-\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest.xyz\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\ttest0\\n\" &&\n+\t\tprintf \"100755 blob $before_oid\\tz\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497390","messageId":"46756c4e3140d34838ad4cd5e7a070d1f9f46b53.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 16/17] mktree: allow deeper paths in input","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:04Z","receivedAt":"2024-06-19T21:58:30Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nUpdate 'git mktree' to handle entries nested inside of directories (e.g.\n'path/to/a/file.txt'). This functionality requires a series of changes:\n\n* In 'sort_and_dedup_tree_entry_array()', remove entries inside of\n  directories that come after them in input order.\n* Also in 'sort_and_dedup_tree_entry_array()', mark directories that contain\n  entries that come after them in input order (e.g., 'folder/' followed by\n  'folder/file.txt') as \"need to expand\".\n* In 'add_tree_entry_to_index()', if a tree entry is marked as \"need to\n  expand\", recurse into it with 'read_tree_at()' & 'build_index_from_tree'.\n* In 'build_index_from_tree()', if a user-specified tree entry is contained\n  within the current iterated entry, return 'READ_TREE_RECURSIVE' to recurse\n  into the iterated tree.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |   5 ++\n builtin/mktree.c             | 101 ++++++++++++++++++++++++++++++---\n t/t1010-mktree.sh            | 107 +++++++++++++++++++++++++++++++++--\n 3 files changed, 200 insertions(+), 13 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 260d0e0bd7b..43cd9b10cc7 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -58,6 +58,11 @@ Higher stages represent conflicted files in an index; this information\n cannot be represented in a tree object. The command will fail without\n writing the tree if a higher order stage is specified for any entry.\n \n+Entries may use full pathnames containing directory separators to specify\n+entries nested within one or more directories. These entries are inserted\n+into the appropriate tree in the base tree-ish if one exists. Otherwise,\n+empty parent trees are created to contain the entries.\n+\n The order of the tree entries is normalized by `mktree` so pre-sorting the\n input by path is not required. Multiple entries provided with the same path\n are deduplicated, with only the last one specified added to the tree.\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 96f06547a2a..74cec92a517 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -22,6 +22,7 @@ struct tree_entry {\n \n \t/* Internal */\n \tsize_t order;\n+\tint expand_dir;\n \n \tunsigned mode;\n \tstruct object_id oid;\n@@ -39,6 +40,7 @@ struct tree_entry_array {\n \tstruct tree_entry **entries;\n \n \tstruct hashmap df_name_hash;\n+\tint has_nested_entries;\n };\n \n static int df_name_hash_cmp(const void *cmp_data UNUSED,\n@@ -70,6 +72,13 @@ static void tree_entry_array_push(struct tree_entry_array *arr, struct tree_entr\n \tarr->entries[arr->nr++] = ent;\n }\n \n+static struct tree_entry *tree_entry_array_pop(struct tree_entry_array *arr)\n+{\n+\tif (!arr->nr)\n+\t\treturn NULL;\n+\treturn arr->entries[--arr->nr];\n+}\n+\n static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entries)\n {\n \tif (free_entries) {\n@@ -109,8 +118,10 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \n \t\tif (!verify_path(ent->name, mode))\n \t\t\tdie(_(\"invalid path '%s'\"), path);\n-\t\tif (strchr(ent->name, '/'))\n-\t\t\tdie(\"path %s contains slash\", path);\n+\n+\t\t/* mark has_nested_entries if needed */\n+\t\tif (!arr->has_nested_entries && strchr(ent->name, '/'))\n+\t\t\tarr->has_nested_entries = 1;\n \n \t\t/* Add trailing slash to dir */\n \t\tif (S_ISDIR(mode))\n@@ -168,6 +179,46 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \tignore_mode = 0;\n \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n \n+\tif (arr->has_nested_entries) {\n+\t\tstruct tree_entry_array parent_dir_ents = { 0 };\n+\n+\t\tcount = arr->nr;\n+\t\tarr->nr = 0;\n+\n+\t\t/* Remove any entries where one of its parent dirs has a higher 'order' */\n+\t\tfor (size_t i = 0; i < count; i++) {\n+\t\t\tconst char *skipped_prefix;\n+\t\t\tstruct tree_entry *parent;\n+\t\t\tstruct tree_entry *curr = arr->entries[i];\n+\t\t\tint skip_entry = 0;\n+\n+\t\t\twhile ((parent = tree_entry_array_pop(&parent_dir_ents))) {\n+\t\t\t\tif (!skip_prefix(curr->name, parent->name, &skipped_prefix))\n+\t\t\t\t\tcontinue;\n+\n+\t\t\t\t/* entry in dir, so we push the parent back onto the stack */\n+\t\t\t\ttree_entry_array_push(&parent_dir_ents, parent);\n+\n+\t\t\t\tif (parent->order > curr->order)\n+\t\t\t\t\tskip_entry = 1;\n+\t\t\t\telse\n+\t\t\t\t\tparent->expand_dir = 1;\n+\n+\t\t\t\tbreak;\n+\t\t\t}\n+\n+\t\t\tif (!skip_entry) {\n+\t\t\t\tarr->entries[arr->nr++] = curr;\n+\t\t\t\tif (S_ISDIR(curr->mode))\n+\t\t\t\t\ttree_entry_array_push(&parent_dir_ents, curr);\n+\t\t\t} else {\n+\t\t\t\tFREE_AND_NULL(curr);\n+\t\t\t}\n+\t\t}\n+\n+\t\ttree_entry_array_release(&parent_dir_ents, 0);\n+\t}\n+\n \t/* Finally, initialize the directory-file conflict hash map */\n \tfor (size_t i = 0; i < count; i++) {\n \t\tstruct tree_entry *curr = arr->entries[i];\n@@ -212,15 +263,40 @@ struct build_index_data {\n \tstruct index_state istate;\n };\n \n+static int build_index_from_tree(const struct object_id *oid,\n+\t\t\t\t struct strbuf *base, const char *filename,\n+\t\t\t\t unsigned mode, void *context);\n+\n static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n-\tstruct cache_entry *ce;\n-\tce = make_cache_entry(&data->istate, ent->mode, &ent->oid, ent->name, 0, 0);\n-\tif (!ce)\n-\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\tif (ent->expand_dir) {\n+\t\tint ret = 0;\n+\t\tstruct pathspec ps = { 0 };\n+\t\tstruct tree *subtree = parse_tree_indirect(&ent->oid);\n+\t\tstruct strbuf base_path = STRBUF_INIT;\n+\t\tstrbuf_add(&base_path, ent->name, ent->len);\n+\n+\t\tif (!subtree)\n+\t\t\tret = error(_(\"not a tree object: %s\"), oid_to_hex(&ent->oid));\n+\t\telse if (read_tree_at(the_repository, subtree, &base_path, 0, &ps,\n+\t\t\t\t build_index_from_tree, data) < 0)\n+\t\t\tret = -1;\n+\n+\t\tstrbuf_release(&base_path);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\n+\t} else {\n+\t\tstruct cache_entry *ce = make_cache_entry(&data->istate,\n+\t\t\t\t\t\t\t  ent->mode, &ent->oid,\n+\t\t\t\t\t\t\t  ent->name, 0, 0);\n+\t\tif (!ce)\n+\t\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n+\n+\t\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n+\t}\n \n-\tadd_index_entry(&data->istate, ce, ADD_CACHE_JUST_APPEND);\n \treturn 0;\n }\n \n@@ -247,10 +323,12 @@ static int build_index_from_tree(const struct object_id *oid,\n \t\tbase_tree_ent->name[base_tree_ent->len - 1] = '/';\n \n \twhile (cbdata->iter.current) {\n+\t\tconst char *skipped_prefix;\n \t\tstruct tree_entry *ent = cbdata->iter.current;\n+\t\tint cmp;\n \n-\t\tint cmp = name_compare(ent->name, ent->len,\n-\t\t\t\t       base_tree_ent->name, base_tree_ent->len);\n+\t\tcmp = name_compare(ent->name, ent->len,\n+\t\t\t\t   base_tree_ent->name, base_tree_ent->len);\n \t\tif (!cmp || cmp < 0) {\n \t\t\ttree_entry_iterator_advance(&cbdata->iter);\n \n@@ -264,6 +342,11 @@ static int build_index_from_tree(const struct object_id *oid,\n \t\t\t\tgoto cleanup_and_return;\n \t\t\t} else\n \t\t\t\tcontinue;\n+\t\t} else if (skip_prefix(ent->name, base_tree_ent->name, &skipped_prefix) &&\n+\t\t\t   S_ISDIR(base_tree_ent->mode)) {\n+\t\t\t/* The entry is in the current traversed tree entry, so we recurse */\n+\t\t\tresult = READ_TREE_RECURSIVE;\n+\t\t\tgoto cleanup_and_return;\n \t\t}\n \n \t\tbreak;\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 435ac23bd50..9b0e0cf302f 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -85,12 +85,21 @@ test_expect_success 'mktree with invalid submodule OIDs' '\n \tdone\n '\n \n-test_expect_success 'mktree refuses to read ls-tree -r output (1)' '\n-\ttest_must_fail git mktree <all\n+test_expect_success 'mktree reads ls-tree -r output (1)' '\n+\tgit mktree <all >actual &&\n+\ttest_cmp tree actual\n '\n \n-test_expect_success 'mktree refuses to read ls-tree -r output (2)' '\n-\ttest_must_fail git mktree <all.withsub\n+test_expect_success 'mktree reads ls-tree -r output (2)' '\n+\tgit mktree <all.withsub >actual &&\n+\ttest_cmp tree.withsub actual\n+'\n+\n+test_expect_success 'mktree de-duplicates files inside directories' '\n+\tgit ls-tree $(cat tree) >everything &&\n+\tcat <all >top_and_all &&\n+\tgit mktree <top_and_all >actual &&\n+\ttest_cmp tree actual\n '\n \n test_expect_success 'mktree fails on malformed input' '\n@@ -234,6 +243,50 @@ test_expect_success 'mktree with duplicate entries' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree adds entry after nested entry' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tone_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tearly\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tearly/one\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tlater\\n\" &&\n+\t\tprintf \"040000 tree $EMPTY_TREE\\tnew-tree\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tnew-tree/one\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tzzz\\n\"\n+\t} >top.rec &&\n+\tgit mktree <top.rec >tree.actual &&\n+\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tearly\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tlater\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tnew-tree\\n\" &&\n+\t\tprintf \"100644 blob $one_oid\\tzzz\\n\"\n+\t} >expect &&\n+\tgit ls-tree $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'mktree inserts entries into directories' '\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tone_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\tblob_oid=$(git rev-parse ${tree_oid}:before) &&\n+\t{\n+\t\tprintf \"040000 tree $folder_oid\\tfolder\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/two\\n\"\n+\t} | git mktree >actual &&\n+\n+\t{\n+\t\tprintf \"100644 blob $one_oid\\tfolder/one\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/two\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'mktree with base tree' '\n \ttree_oid=$(cat tree) &&\n \tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n@@ -270,4 +323,50 @@ test_expect_success 'mktree with base tree' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'mktree with base tree (deep)' '\n+\ttree_oid=$(cat tree) &&\n+\tfolder_oid=$(git rev-parse ${tree_oid}:folder) &&\n+\tbefore_oid=$(git rev-parse ${tree_oid}:before) &&\n+\tfolder_one_oid=$(git rev-parse ${tree_oid}:folder/one) &&\n+\thead_oid=$(git rev-parse HEAD) &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tfolder/before\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\tfolder/one.txt\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\tfolder/sub\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tfolder/one\\n\" &&\n+\t\tprintf \"040000 tree $folder_oid\\tfolder/one/deeper\\n\"\n+\t} >top.append &&\n+\tgit mktree <top.append $(cat tree) >tree.actual &&\n+\n+\t{\n+\t\tprintf \"100755 blob $before_oid\\tfolder/before\\n\" &&\n+\t\tprintf \"100644 blob $before_oid\\tfolder/one.txt\\n\" &&\n+\t\tprintf \"100644 blob $folder_one_oid\\tfolder/one/deeper/one\\n\" &&\n+\t\tprintf \"100644 blob $folder_one_oid\\tfolder/one/one\\n\" &&\n+\t\tprintf \"160000 commit $head_oid\\tfolder/sub\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat tree.actual) -- folder/ >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'mktree fails on directory-file conflict' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\ttest\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest/deeper\\n\"\n+\t} |\n+\ttest_must_fail git mktree 2>err &&\n+\ttest_grep \"You have both test and test/deeper\" err &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\tfolder/one/deeper/deep\\n\"\n+\t} |\n+\ttest_must_fail git mktree $tree_oid 2>err &&\n+\ttest_grep \"You have both folder/one and folder/one/deeper/deep\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"497391","messageId":"d392c440b8a243a9fa3e5b603c42a81f02c26e62.1718834285.git.gitgitgadget@gmail.com","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"[PATCH v2 17/17] mktree: remove entries when mode is 0","fromName":"Victoria Dye via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2024-06-19T21:58:05Z","receivedAt":"2024-06-19T21:58:31Z","isPatch":true,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"From: Victoria Dye <vdye@github.com>\n\nIf tree entries are specified with a mode with value '0', remove them from\nthe tree instead of adding/updating them. If the mode is '0', both the\nprovided type string (if specified) and the object ID of the entry are\nignored.\n\nNote that entries with mode '0' are added to the 'struct tree_ent_array'\nwith a trailing slash so that it's always treated like a directory. This is\na bit of a hack to ensure that the removal supercedes any preceding entries\nwith matching names, as well as any nested inside a directory matching its\nname.\n\nSigned-off-by: Victoria Dye <vdye@github.com>\n---\n Documentation/git-mktree.txt |  4 ++++\n builtin/mktree.c             | 16 +++++++++++----\n t/t1010-mktree.sh            | 38 ++++++++++++++++++++++++++++++++++++\n 3 files changed, 54 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\nindex 43cd9b10cc7..52e6005c1d3 100644\n--- a/Documentation/git-mktree.txt\n+++ b/Documentation/git-mktree.txt\n@@ -63,6 +63,10 @@ entries nested within one or more directories. These entries are inserted\n into the appropriate tree in the base tree-ish if one exists. Otherwise,\n empty parent trees are created to contain the entries.\n \n+An entry with a mode of \"0\" will remove an entry of the same name from the\n+base tree-ish. If no tree-ish argument is given, or the entry does not exist\n+in that tree, the entry is ignored.\n+\n The order of the tree entries is normalized by `mktree` so pre-sorting the\n input by path is not required. Multiple entries provided with the same path\n are deduplicated, with only the last one specified added to the tree.\ndiff --git a/builtin/mktree.c b/builtin/mktree.c\nindex 74cec92a517..e7adcb384c8 100644\n--- a/builtin/mktree.c\n+++ b/builtin/mktree.c\n@@ -32,7 +32,7 @@ struct tree_entry {\n \n static inline size_t df_path_len(size_t pathlen, unsigned int mode)\n {\n-\treturn S_ISDIR(mode) ? pathlen - 1 : pathlen;\n+\treturn (S_ISDIR(mode) || !mode) ? pathlen - 1 : pathlen;\n }\n \n struct tree_entry_array {\n@@ -108,7 +108,7 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \t\tsize_t len_to_copy = len;\n \n \t\t/* Normalize and validate entry path */\n-\t\tif (S_ISDIR(mode)) {\n+\t\tif (S_ISDIR(mode) || !mode) {\n \t\t\twhile(len_to_copy > 0 && is_dir_sep(path[len_to_copy - 1]))\n \t\t\t\tlen_to_copy--;\n \t\t\tlen = len_to_copy + 1; /* add space for trailing slash */\n@@ -124,7 +124,7 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n \t\t\tarr->has_nested_entries = 1;\n \n \t\t/* Add trailing slash to dir */\n-\t\tif (S_ISDIR(mode))\n+\t\tif (S_ISDIR(mode) || !mode)\n \t\t\tent->name[len - 1] = '/';\n \t}\n \n@@ -209,7 +209,7 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n \n \t\t\tif (!skip_entry) {\n \t\t\t\tarr->entries[arr->nr++] = curr;\n-\t\t\t\tif (S_ISDIR(curr->mode))\n+\t\t\t\tif (S_ISDIR(curr->mode) || !curr->mode)\n \t\t\t\t\ttree_entry_array_push(&parent_dir_ents, curr);\n \t\t\t} else {\n \t\t\t\tFREE_AND_NULL(curr);\n@@ -270,6 +270,9 @@ static int build_index_from_tree(const struct object_id *oid,\n static int add_tree_entry_to_index(struct build_index_data *data,\n \t\t\t\t   struct tree_entry *ent)\n {\n+\tif (!ent->mode)\n+\t\treturn 0;\n+\n \tif (ent->expand_dir) {\n \t\tint ret = 0;\n \t\tstruct pathspec ps = { 0 };\n@@ -450,6 +453,10 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \tif (stage)\n \t\tdie(_(\"path '%s' is unmerged\"), path);\n \n+\t/* OID ignored for zero-mode entries; append unconditionally */\n+\tif (!mode)\n+\t\tgoto append_entry;\n+\n \tif (obj_type != OBJ_ANY && mode_type != obj_type)\n \t\tdie(\"object type (%s) doesn't match mode type (%s)\",\n \t\t    type_name(obj_type), type_name(mode_type));\n@@ -484,6 +491,7 @@ static int mktree_line(unsigned int mode, struct object_id *oid,\n \t\t}\n \t}\n \n+append_entry:\n \tappend_to_tree(mode, oid, path, data->arr, data->literally);\n \treturn 0;\n }\ndiff --git a/t/t1010-mktree.sh b/t/t1010-mktree.sh\nindex 9b0e0cf302f..5ed4352054a 100755\n--- a/t/t1010-mktree.sh\n+++ b/t/t1010-mktree.sh\n@@ -369,4 +369,42 @@ test_expect_success 'mktree fails on directory-file conflict' '\n \ttest_grep \"You have both folder/one and folder/one/deeper/deep\" err\n '\n \n+test_expect_success 'mktree with remove entries' '\n+\ttree_oid=\"$(cat tree)\" &&\n+\tblob_oid=\"$(git rev-parse $tree_oid:folder.txt)\" &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\ttest/deeper/deep.txt\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest.txt\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.a/file\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.txt\\n\" &&\n+\t\tprintf \"040000 tree $tree_oid\\tfolder\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\tfolder\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\tmissing\\n\"\n+\t} | git mktree >tree.base &&\n+\n+\t{\n+\t\tprintf \"0 $ZERO_OID\\texample.txt\\n\" &&\n+\t\tprintf \"0 $ZERO_OID\\ttest/deeper\\n\"\n+\t} | git mktree $(cat tree.base) >tree.actual &&\n+\n+\t{\n+\t\tprintf \"100644 blob $blob_oid\\texample\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\texample.a/file\\n\" &&\n+\t\tprintf \"100644 blob $blob_oid\\ttest.txt\\n\"\n+\t} >expect &&\n+\tgit ls-tree -r $(cat tree.actual) >actual &&\n+\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'type and oid not checked if entry mode is 0' '\n+\t# type and oid do not match\n+\tprintf \"0 commit $EMPTY_TREE\\tfolder.txt\\n\" |\n+\tgit mktree >tree.actual &&\n+\n+\ttest \"$(cat tree.actual)\" = $EMPTY_TREE\n+'\n+\n test_done\n-- \ngitgitgadget\n"},{"id":"497429","messageId":"xmqqbk3vtll3.fsf@gitster.g","threadId":"61624","inReplyTo":"9dc8e16a7fca886ec378d74a8e2ac61921a7f6ea.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 07/17] mktree: use read_index_info to read stdin lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-20T20:18:00Z","receivedAt":"2024-06-20T20:18:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +INPUT FORMAT\n> +------------\n> +Tree entries may be specified in any of the formats compatible with the\n> +`--index-info` option to linkgit:git-update-index[1]:\n> +\n> +include::index-info-formats.txt[]\n> +\n> +Note that if the `stage` of a tree entry is given, the value must be 0.\n> +Higher stages represent conflicted files in an index; this information\n> +cannot be represented in a tree object. The command will fail without\n> +writing the tree if a higher order stage is specified for any entry.\n> +\n> +The order of the tree entries is normalized by `mktree` so pre-sorting the\n> +input by path is not required.\n\nNicely done.  I was wondering how the common/shared text that was\nmade more generic in 04/17 would be made to fit in the new context,\nand the \"Note that\" makes them mix very well.\n\nThe updated code is exactly as expected.\n"},{"id":"497435","messageId":"xmqqh6dns21m.fsf@gitster.g","threadId":"61624","inReplyTo":"fb555658057f834d94f232f1d8b380a6304a3671.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 11/17] mktree: overwrite duplicate entries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-20T22:05:25Z","receivedAt":"2024-06-20T22:05:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> If multiple tree entries with the same name are provided as input to\n> 'mktree', only write the last one to the tree. Entries are considered\n> duplicates if they have identical names (*not* considering mode); if a blob\n> and a tree with the same name are provided, only the last one will be\n> written to the tree. A tree with duplicate entries is invalid (per 'git\n> fsck'), so that condition should be avoided wherever possible.\n\nThe \"should be avoided\" in the last sentence can be satisified\neither by the callers being extra careful, or the callee ignoring\nearlier entries with the same path.  I do not have a strong\nobjection against allowing looser callers, but if that is what is\ngoing on, perhaps\n\n\tBy teaching \"mktree\" to ignore the earlier entries for the\n        same path in the input, the callers can be more casual about\n        sending duplicate entries in order to avoid creating an\n        invalid tree objects.\n\nis a more honest justification for this setp?\n\n> diff --git a/Documentation/git-mktree.txt b/Documentation/git-mktree.txt\n> index 5f3a6dfe38e..cf1fd82f754 100644\n> --- a/Documentation/git-mktree.txt\n> +++ b/Documentation/git-mktree.txt\n> @@ -54,7 +54,8 @@ cannot be represented in a tree object. The command will fail without\n>  writing the tree if a higher order stage is specified for any entry.\n>  \n>  The order of the tree entries is normalized by `mktree` so pre-sorting the\n> -input by path is not required.\n> +input by path is not required. Multiple entries provided with the same path\n> +are deduplicated, with only the last one specified added to the tree.\n\nOK.\n\n>  struct tree_entry {\n> +\t/* Internal */\n> +\tsize_t order;\n> +\n>  \tunsigned mode;\n>  \tstruct object_id oid;\n>  \tint len;\n> @@ -74,15 +77,49 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n>  \tent->len = len;\n>  \toidcpy(&ent->oid, oid);\n>  \n> +\tent->order = arr->nr;\n>  \ttree_entry_array_push(arr, ent);\n>  }\n>  \n> -static int ent_compare(const void *a_, const void *b_)\n> +static int ent_compare(const void *a_, const void *b_, void *ctx)\n>  {\n> +\tint cmp;\n>  \tstruct tree_entry *a = *(struct tree_entry **)a_;\n>  \tstruct tree_entry *b = *(struct tree_entry **)b_;\n> -\treturn base_name_compare(a->name, a->len, a->mode,\n> -\t\t\t\t b->name, b->len, b->mode);\n> +\tint ignore_mode = *((int *)ctx);\n> +\n> +\tif (ignore_mode)\n> +\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n> +\telse\n> +\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n> +\t\t\t\t\tb->name, b->len, b->mode);\n> +\treturn cmp ? cmp : b->order - a->order;\n> +}\n\nHaving two similar functions that could go out of sync has bothered\nme somewhat.  We could instead do\n\n\tint a_mode = ignore_mode ? 0 : a->mode;\n\tint b_mode = ignore_mode ? 0 : b->mode;\n\tcmp = base_name_compare(a->name, a->len, a_mode,\n\t\t\t\tb->name, b->len, b_mode);\n\nbut that should be done by rewriting name_compare() in terms of\nbase_name_compare(), which will help more callers, not just this\none.\n\n> +static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n> +{\n> +\tsize_t count = arr->nr;\n> +\tstruct tree_entry *prev = NULL;\n> +\n> +\tint ignore_mode = 1;\n> +\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n\nSwap the decl for ignore_mode and the blank line above it?\n\nIf the callback context only needs a single bit, ent_compare() could\njust use the NULL-ness of ctx as \"do we want to ignore mode?\" bit.\n\n> +\tarr->nr = 0;\n> +\tfor (size_t i = 0; i < count; i++) {\n> +\t\tstruct tree_entry *curr = arr->entries[i];\n> +\t\tif (prev &&\n> +\t\t    !name_compare(prev->name, prev->len,\n> +\t\t\t\t  curr->name, curr->len)) {\n> +\t\t\tFREE_AND_NULL(curr);\n> +\t\t} else {\n> +\t\t\tarr->entries[arr->nr++] = curr;\n> +\t\t\tprev = curr;\n> +\t\t}\n> +\t}\n\nAs long as this is done for a single tree (i.e. the paths do not\nhave any slashes in them), this \"sort them all and keep the last\none\" is a good strategy.\n\n> +\t/* Sort again to order the entries for tree insertion */\n> +\tignore_mode = 0;\n> +\tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n\nOK.  We from time to time find need to do this, and I always regret\nthat we didn't design the sort order of paths in a tree (and in the\nindex) like so [*].  But that is almost 20 years too late ;-).\n\nLooking good.\n\n\n[Footnote]\n\n * A directory entry $T should have sorted after a non-directory\n   entry $T but before any non-directory entry whose path has $T\n   as its prefix (e.g. even a blob whose path is $T + \"\\001\" should\n   sort after a tree $T).  That way we didn't have to worry about a\n   blob at ($T + '-') sorting before a tree at $T but a blob at ($T\n   + '0') sorting after that tree.\n\n"},{"id":"497437","messageId":"xmqqtthnqmim.fsf@gitster.g","threadId":"61624","inReplyTo":"2333775ba5bd71766a6aece87e39a6d189aeaead.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 12/17] mktree: create tree using an in-core index","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-20T22:26:09Z","receivedAt":"2024-06-20T22:26:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> @@ -60,17 +66,25 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n>  \tif (literally) {\n>  \t\tFLEX_ALLOC_MEM(ent, name, path, len);\n>  \t} else {\n> +\t\tsize_t len_to_copy = len;\n> +\n>  \t\t/* Normalize and validate entry path */\n>  \t\tif (S_ISDIR(mode)) {\n> -\t\t\twhile(len > 0 && is_dir_sep(path[len - 1]))\n> -\t\t\t\tlen--;\n> +\t\t\twhile(len_to_copy > 0 && is_dir_sep(path[len_to_copy - 1]))\n\nLet's fix the style issue while at it, as we are doing other changes\nin this step anyway.  \"while(\" -> \"while (\".\n\n> +\t\t\t\tlen_to_copy--;\n> +\t\t\tlen = len_to_copy + 1; /* add space for trailing slash */\n\nDo we need to do st_add() here?  Perhaps not, but I just noticed the\ncareful use of st_add3() below, so...\n\n> +\t\tent = xcalloc(1, st_add3(sizeof(struct tree_entry), len, 1));\n\n> +\t\tmemcpy(ent->name, path, len_to_copy);\n>  \n>  \t\tif (!verify_path(ent->name, mode))\n>  \t\t\tdie(_(\"invalid path '%s'\"), path);\n>  \t\tif (strchr(ent->name, '/'))\n>  \t\t\tdie(\"path %s contains slash\", path);\n> +\n> +\t\t/* Add trailing slash to dir */\n> +\t\tif (S_ISDIR(mode))\n> +\t\t\tent->name[len - 1] = '/';\n\nOK.\n\n> @@ -88,11 +102,14 @@ static int ent_compare(const void *a_, const void *b_, void *ctx)\n>  \tstruct tree_entry *b = *(struct tree_entry **)b_;\n>  \tint ignore_mode = *((int *)ctx);\n>  \n> -\tif (ignore_mode)\n> -\t\tcmp = name_compare(a->name, a->len, b->name, b->len);\n> -\telse\n> -\t\tcmp = base_name_compare(a->name, a->len, a->mode,\n> -\t\t\t\t\tb->name, b->len, b->mode);\n> +\tsize_t a_len = a->len, b_len = b->len;\n> +\n> +\tif (ignore_mode) {\n> +\t\ta_len = df_path_len(a_len, a->mode);\n> +\t\tb_len = df_path_len(b_len, b->mode);\n> +\t}\n> +\n> +\tcmp = name_compare(a->name, a_len, b->name, b_len);\n>  \treturn cmp ? cmp : b->order - a->order;\n>  }\n\nOK, now the \"mode\" is sort of \"encoded\" already in the \"name\" by the\nslash at the end, the way \"ignore-mode\" works needs to be redesigned.\n\nIf we are ignoring mode, we are dropping the trailing '/' and\notherwise we just feed the name with possible trailing '/', and the\nsame name_compare() can be used.  OK.\n\n> @@ -108,8 +125,8 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>  \tfor (size_t i = 0; i < count; i++) {\n>  \t\tstruct tree_entry *curr = arr->entries[i];\n>  \t\tif (prev &&\n> -\t\t    !name_compare(prev->name, prev->len,\n> -\t\t\t\t  curr->name, curr->len)) {\n> +\t\t    !name_compare(prev->name, df_path_len(prev->len, prev->mode),\n> +\t\t\t\t  curr->name, df_path_len(curr->len, curr->mode))) {\n>  \t\t\tFREE_AND_NULL(curr);\n\nAnd here is the matching adjustment for the dedup comparison, which\nmakes sense.\n\n> @@ -122,24 +139,43 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>  \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n>  }\n>  \n> +static int add_tree_entry_to_index(struct index_state *istate,\n> +\t\t\t\t   struct tree_entry *ent)\n> +{\n> +\tstruct cache_entry *ce;\n> +\tstruct strbuf ce_name = STRBUF_INIT;\n> +\tstrbuf_add(&ce_name, ent->name, ent->len);\n> +\n\nPerhaps swap the first statement (which is strbuf_add()) and the\nblank line that ought to separate the decls and the first statement?\n\n> +\tce = make_cache_entry(istate, ent->mode, &ent->oid, ent->name, 0, 0);\n> +\tif (!ce)\n> +\t\treturn error(_(\"make_cache_entry failed for path '%s'\"), ent->name);\n> +\n> +\tadd_index_entry(istate, ce, ADD_CACHE_JUST_APPEND);\n> +\tstrbuf_release(&ce_name);\n> +\treturn 0;\n> +}\n\nThis is only to append; presumably the caller drives this function\nout of a sorted list.\n\n>  static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n>  {\n> +\tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n> +\tistate.sparse_index = 1;\n>  \n>  \tsort_and_dedup_tree_entry_array(arr);\n>  \n> +\t/* Construct an in-memory index from the provided entries */\n>  \tfor (size_t i = 0; i < arr->nr; i++) {\n>  \t\tstruct tree_entry *ent = arr->entries[i];\n> +\n> +\t\tif (add_tree_entry_to_index(&istate, ent))\n> +\t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n>  \t}\n> +\t/* Write out new tree */\n> +\tif (cache_tree_update(&istate, WRITE_TREE_SILENT | WRITE_TREE_MISSING_OK))\n> +\t\tdie(_(\"failed to write tree\"));\n\nHmph.  Are we doing any run-time verification of what we produce\n(e.g., if sort_and_dedup_tree_entry_array() fails to dedup or sort\ncorrectly due to a bug or two, would cache_tree_update() notice that\nthe in-core index array is fishy)?  I am not suggesting to add an\nunconditional \"we appended to the index, so we should sort the\nentries in it\" step before cache_tree_update() call.  It is the\nopposite---if we have extra checks in cache_tree_udpate() to slow us\ndown and if we are confident that the loop that added tree entries\nto the index is correct, if we can bypass such checks.\n\n> +\toidcpy(oid, &istate.cache_tree->oid);\n> +\n> +\trelease_index(&istate);\n>  }\n\nThis is the gem of the whole series.  Clever.\n\nWhat is so satisfying is that it takes not that much of code to\nreplace the \"here is a flat buffer of what the contents of a single\ntree object ought to look like\" with \"let's build in-core index and\nwrite it out just like write-tree would\".  Nice.\n"},{"id":"497666","messageId":"xmqqzfr861un.fsf@gitster.g","threadId":"61624","inReplyTo":"pull.1746.v2.git.1718834285.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 00/17] mktree: support more flexible usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-25T23:26:24Z","receivedAt":"2024-06-25T23:26:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> The goal of this series is to make 'git mktree' a much more flexible and\n> powerful tool for constructing arbitrary trees in memory without the use of\n> an index or worktree.\n\nI've read earlier parts of this series carefully (but didn't manage\nto get to the end of it), and I saw Patrick and Eric gave reviews on\nthe earlier round, but otherwise this topic seems to have stalled.\n\nhttps://lore.kernel.org/git/pull.1746.v2.git.1718834285.gitgitgadget@gmail.com/\n\nAny more comments?\n\nOtherwise, if I find time to read through 13-17/17 and did not find\nanything glaringly wrong, I am planning to mark the topic ready for\n'next'.  Thanks.\n"},{"id":"497735","messageId":"xmqq8qyrxveo.fsf@gitster.g","threadId":"61624","inReplyTo":"56f28efff5404a3fa22bd544d6de8ce2d919b78a.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 13/17] mktree: use iterator struct to add tree entries to index","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-26T21:10:23Z","receivedAt":"2024-06-26T21:10:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> @@ -157,14 +186,18 @@ static int add_tree_entry_to_index(struct index_state *istate,\n>  \n>  static void write_tree(struct tree_entry_array *arr, struct object_id *oid)\n>  {\n> +\tstruct tree_entry_iterator iter = { NULL };\n>  \tstruct index_state istate = INDEX_STATE_INIT(the_repository);\n>  \tistate.sparse_index = 1;\n>  \n>  \tsort_and_dedup_tree_entry_array(arr);\n>  \n> -\t/* Construct an in-memory index from the provided entries */\n> -\tfor (size_t i = 0; i < arr->nr; i++) {\n> -\t\tstruct tree_entry *ent = arr->entries[i];\n> +\ttree_entry_iterator_init(&iter, arr);\n> +\n> +\t/* Construct an in-memory index from the provided entries & base tree */\n> +\twhile (iter.current) {\n> +\t\tstruct tree_entry *ent = iter.current;\n> +\t\ttree_entry_iterator_advance(&iter);\n>  \n>  \t\tif (add_tree_entry_to_index(&istate, ent))\n>  \t\t\tdie(_(\"failed to add tree entry '%s'\"), ent->name);\n\nOK, looking good.\n\nIf we make _iterator_init() and _iterator_advance to both return the\ncurrent, then the loop can still be like so:\n\n\tfor (ent = tree_entry_iterator_init(&iter, arr);\n             ent;\n\t     ent = tree_entry_iterator_advance(&iter)) {\n\t\t... use ent ...\n\t}\n\nand .current does not need to be a non-private member, if we wanted\nto (I am not convinced it is necessarily a better interface to make\n.current as private---especially if we end up needing _peek() method\nto learn its value, i.e. the value the most recent call to _init()\nor _advance() returned.  If we need a write access to .current from\noutside the interator interface, then what I outlined above would\nnot be a good match).\n\n\nThanks.\n"},{"id":"497737","messageId":"xmqqo77nwg83.fsf@gitster.g","threadId":"61624","inReplyTo":"4b88f84b933b1598d12e3620f0c9fb85c559e8fb.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 15/17] mktree: optionally add to an existing tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-26T21:23:40Z","receivedAt":"2024-06-26T21:23:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Allow users to specify a single \"tree-ish\" value as a positional argument.\n> If provided, the contents of the given tree serve as the basis for the new\n> tree (or trees, in --batch mode) created by 'mktree', on top of which all of\n> the stdin-provided tree entries are applied.\n>\n> At a high level, the entries are \"applied\" to a base tree by iterating\n> through the base tree using 'read_tree' in parallel with iterating through\n> the sorted & deduplicated stdin entries via their iterator. That is, for\n> each call to the 'build_index_from_tree callback of 'read_tree':\n>\n> * If the iterator entry precedes the base tree entry, add it to the in-core\n>   index, increment the iterator, and repeat.\n\n\"add it\" -> \"add the base tree entry\"?  The next bullet point\nexplicitly says it adds \"the iterator entry\", which makes it crystal\nclear what is going on.\n\n> * If the iterator entry has the same name as the base tree entry, add the\n>   iterator entry to the index, increment the iterator, and return from the\n>   callback to continue the 'read_tree' iteration.\n> * If the iterator entry follows the base tree entry, first check\n>   'df_name_hash' to ensure we won't be adding an entry with the same name\n>   later (with a different mode). If there's no directory/file conflict, add\n>   the base tree entry to the index. In either case, return from the callback\n>   to continue the 'read_tree' iteration.\n\nIOW, we take advantage of the fact that iteration over the base tree\nand iteration over the sorted-and-deduped entries from the standard\ninput are already sorted, and do a simple bog-standard \"merge\" of\ntwo lists?\n\nWe'd probably have many common pitfalls to avoid with the read-tree\nwalking the index and tree(s) in parallel (I still remember the pain\nof maintaining the cache_bottom for the side that walks the index).\nMakes me wonder if this opens a way to a future where somehow\nread-tree also shares code with this new code in mktree (or vice\nversa).\n\n> Finally, once 'read_tree' is complete, add the remaining entries in the\n> iterator to the index and write out the index as a tree.\n\nOr vice versa?  We may finish iterating over the entries read from\nthe standard input but there still are entries from the base tree\nside remaining, which would need to be added to complete the index,\nright?\n\n> +<tree-ish>::\n> +\tIf provided, the tree entries provided in stdin are added to this\n> +\ttree rather than a new empty one, replacing existing entries with\n> +\tidentical names. Not compatible with `--literally`.\n\n\"replacing\" might need a bit more clarification when we start\nreading paths with multiple pathname components concatenated with\nslashes.  In the base tree, we may have\n\n    100644 blob 536e55524db72bd2acf175208aef4f3dfc148d42    D\n\nand it can (indirectly) replaced by the standard input stream\nfeeding entries like these\n\n    100644 blob b0517166ae2ad92f3b17638cbdee0f04b8170d99    D/a\n    100644 blob 495a54bc1397e2fd3177c2733baf4899b48d30bd    D/b\n\n\nwhich also leads us to compute a tree entry\n\n    040000 tree eccdce44520aa3ef4ac5ba090df53eadb01229ef    D/\n\nin the top-level tree?\n\nThe code looks good to me.  Thanks.\n\n"},{"id":"497762","messageId":"xmqqed8itc9l.fsf@gitster.g","threadId":"61624","inReplyTo":"46756c4e3140d34838ad4cd5e7a070d1f9f46b53.1718834285.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 16/17] mktree: allow deeper paths in input","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-27T19:29:42Z","receivedAt":"2024-06-27T19:29:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Victoria Dye <vdye@github.com>\n>\n> Update 'git mktree' to handle entries nested inside of directories (e.g.\n> 'path/to/a/file.txt'). This functionality requires a series of changes:\n>\n> * In 'sort_and_dedup_tree_entry_array()', remove entries inside of\n>   directories that come after them in input order.\n\nSo if you feed \"folder/file.txt\" and then \"folder/\", then\n\"folder/file.txt\" gets removed?  It is unclear offhand why that is\nthe right thing to do.\n\n> * Also in 'sort_and_dedup_tree_entry_array()', mark directories that contain\n>   entries that come after them in input order (e.g., 'folder/' followed by\n>   'folder/file.txt') as \"need to expand\".\n\nMakes me wonder what happens to the object name recorded in the\ninput for \"folder/\" when something like this happens.  Ideally,\nadding (or replacing) \"folder/file.txt\" to the set of files we\ncollected out of the base tree and the input stream for \"folder/\"\nand writing that \"folder/\" out as a tree would result in a tree\nwhose object name exactly matches it (and we will error out if it\ndoes not)?  Or is \"need to expand\" a signal that we should ignore\nthe object name in the input and we need to recompute it ourselves?\nAgain, it is unclear offhand what we want the \"need to expand\" is\nused for.\n\n> * In 'add_tree_entry_to_index()', if a tree entry is marked as \"need to\n>   expand\", recurse into it with 'read_tree_at()' & 'build_index_from_tree'.\n> * In 'build_index_from_tree()', if a user-specified tree entry is contained\n>   within the current iterated entry, return 'READ_TREE_RECURSIVE' to recurse\n>   into the iterated tree.\n\nSurely, no matter what we choose to do to the object name given with\n\"folder/\" when the input stream also talks about \"folder/file.txt\",\nwe'd need to recurse into the subtree.  But I think we need a higher\nlevel description of what exactly we want to do to the multi-level\npathnames (i.e., \"we want to handle them this way\") before going into\nthe implementation details of how we do so (i.e., \"hence we deal with\na multi-level pathname this way at these places in the code\") in\nthese bullet points.\n\nEspecially, I do not quite understand what semantics the first\nbullet point is trying to achieve.\n\n> +Entries may use full pathnames containing directory separators to specify\n> +entries nested within one or more directories. These entries are inserted\n> +into the appropriate tree in the base tree-ish if one exists. Otherwise,\n> +empty parent trees are created to contain the entries.\n\nThis still does not answer \"how overlapping entries are handled?\",\nwhich is more complex than \"for two exactly the same paths, the last\none wins\", which is mentioned in the next paragraph.\n\n>  The order of the tree entries is normalized by `mktree` so pre-sorting the\n>  input by path is not required. Multiple entries provided with the same path\n>  are deduplicated, with only the last one specified added to the tree.\n> diff --git a/builtin/mktree.c b/builtin/mktree.c\n> index 96f06547a2a..74cec92a517 100644\n> --- a/builtin/mktree.c\n> +++ b/builtin/mktree.c\n> ...\n> +static struct tree_entry *tree_entry_array_pop(struct tree_entry_array *arr)\n> +{\n> +\tif (!arr->nr)\n> +\t\treturn NULL;\n> +\treturn arr->entries[--arr->nr];\n> +}\n> +\n>  static void tree_entry_array_clear(struct tree_entry_array *arr, int free_entries)\n>  {\n>  \tif (free_entries) {\n> @@ -109,8 +118,10 @@ static void append_to_tree(unsigned mode, struct object_id *oid, const char *pat\n>  \n>  \t\tif (!verify_path(ent->name, mode))\n>  \t\t\tdie(_(\"invalid path '%s'\"), path);\n> -\t\tif (strchr(ent->name, '/'))\n> -\t\t\tdie(\"path %s contains slash\", path);\n> +\n> +\t\t/* mark has_nested_entries if needed */\n> +\t\tif (!arr->has_nested_entries && strchr(ent->name, '/'))\n> +\t\t\tarr->has_nested_entries = 1;\n\nOK.\n\n> @@ -168,6 +179,46 @@ static void sort_and_dedup_tree_entry_array(struct tree_entry_array *arr)\n>  \tignore_mode = 0;\n>  \tQSORT_S(arr->entries, arr->nr, ent_compare, &ignore_mode);\n\nWe have already sorted the array twice (once before simple deduping,\nonce after).  So we now have a sorted array of \"last one won\" paths\nand their object names.\n\n> +\tif (arr->has_nested_entries) {\n\nWe need to deal with overlapping entries if \"has-nested-entries\" is\ntrue.  Even though our input here is sorted, we'd still pay\nattention to the original input \"order\", which may be different from\nthe order in which we find these entries in arr->entries[].\n\nOK.\n\n> +\t\tstruct tree_entry_array parent_dir_ents = { 0 };\n> +\n> +\t\tcount = arr->nr;\n> +\t\tarr->nr = 0;\n> +\n> +\t\t/* Remove any entries where one of its parent dirs has a higher 'order' */\n\nIs \"has a higher order\" equivalent to \"appears later in the input\"?\n\nMore importantly, can the reason why they need to be removed be\nclarified?  For simple deduping, we can say \"we will make the last\none of multiple entries talking about the same path be used\", and\nthat would be a sufficient explanation why we discard the one that\nwe have seen earlier and replace it with the newly seen one for the\nsame path.  Can a similar and simple explanation be given for the\nbehaviour this loop tries to achieve?  Is it \"children, which appear\nearlier in the input, of a directory, which appears later than these\nchildren, are discarded, because the entry for the directory has a\nconcrete object name, and there is no point talking about individual\npaths inside the directory.  We know what the tree object that would\ncontain these child paths hashes to in the end.  This is a natural\nextension of 'last one wins' rule---a directory that comes later\ntrumps paths contained within that come earlier\"?\n\n> +\t\tfor (size_t i = 0; i < count; i++) {\n> +\t\t\tconst char *skipped_prefix;\n> +\t\t\tstruct tree_entry *parent;\n> +\t\t\tstruct tree_entry *curr = arr->entries[i];\n> +\t\t\tint skip_entry = 0;\n> +\n> +\t\t\twhile ((parent = tree_entry_array_pop(&parent_dir_ents))) {\n> +\t\t\t\tif (!skip_prefix(curr->name, parent->name, &skipped_prefix))\n> +\t\t\t\t\tcontinue;\n> +\n> +\t\t\t\t/* entry in dir, so we push the parent back onto the stack */\n> +\t\t\t\ttree_entry_array_push(&parent_dir_ents, parent);\n> +\n> +\t\t\t\tif (parent->order > curr->order)\n> +\t\t\t\t\tskip_entry = 1;\n> +\t\t\t\telse\n> +\t\t\t\t\tparent->expand_dir = 1;\n> +\n> +\t\t\t\tbreak;\n> +\t\t\t}\n> +\n> +\t\t\tif (!skip_entry) {\n> +\t\t\t\tarr->entries[arr->nr++] = curr;\n> +\t\t\t\tif (S_ISDIR(curr->mode))\n> +\t\t\t\t\ttree_entry_array_push(&parent_dir_ents, curr);\n> +\t\t\t} else {\n> +\t\t\t\tFREE_AND_NULL(curr);\n> +\t\t\t}\n> +\t\t}\n> +\n> +\t\ttree_entry_array_release(&parent_dir_ents, 0);\n> +\t}\n> +\n>  \t/* Finally, initialize the directory-file conflict hash map */\n>  \tfor (size_t i = 0; i < count; i++) {\n>  \t\tstruct tree_entry *curr = arr->entries[i];\n"},{"id":"498481","messageId":"xmqq34ohudr5.fsf@gitster.g","threadId":"61624","inReplyTo":"xmqqzfr861un.fsf@gitster.g","subject":"Re: [PATCH v2 00/17] mktree: support more flexible usage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-07-10T21:40:46Z","receivedAt":"2024-07-10T21:40:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Victoria Dye via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> The goal of this series is to make 'git mktree' a much more flexible and\n>> powerful tool for constructing arbitrary trees in memory without the use of\n>> an index or worktree.\n>\n> I've read earlier parts of this series carefully (but didn't manage\n> to get to the end of it), and I saw Patrick and Eric gave reviews on\n> the earlier round, but otherwise this topic seems to have stalled.\n>\n> https://lore.kernel.org/git/pull.1746.v2.git.1718834285.gitgitgadget@gmail.com/\n>\n> Any more comments?\n>\n> Otherwise, if I find time to read through 13-17/17 and did not find\n> anything glaringly wrong, I am planning to mark the topic ready for\n> 'next'.  Thanks.\n\nAnd I had a few review comments myself.  Then the topic stalled.  I\nam not ready to mark the topic ready for 'next' in this state.\n\nTHanks.\n"}]}