{"thread":{"id":"60068","subject":"[PATCH 0/5] Trailer readability cleanups","startedAt":"2023-08-05T05:04:48Z","lastAt":"2023-12-29T21:03:29Z","messageCount":72,"participants":["Linus Arver via GitGitGadget","Glen Choo","Phillip Wood","Linus Arver","Junio C Hamano","Jonathan Tan"],"isPatch":true,"patchVersion":1,"patchTotal":5},"messages":[{"id":"480178","messageId":"pull.1563.git.1691211879.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":null,"subject":"[PATCH 0/5] Trailer readability cleanups","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:34Z","receivedAt":"2023-08-05T05:04:48Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"These patches were created while digging into the trailer code to better\nunderstand how it works, in preparation for making the trailer.{c,h} files\nas small as possible to make them available as a library for external users.\nThis series was originally created as part of [1], but are sent here\nseparately because the changes here are arguably more subjective in nature.\nI think Patch 1 is the most important in this series. The others can wait,\nif folks are opposed to adding them on their own merits at this point in\ntime.\n\nThese patches do not add or change any features. Instead, their goal is to\nmake the code easier to understand for new contributors (like myself), by\nmaking various cleanups and improvements. Ultimately, my hope is that with\nsuch cleanups, we are better positioned to make larger changes (especially\nthe broader libification effort, as in \"Introduce Git Standard Library\"\n[2]).\n\nPatch 1 was inspired by 576de3d956 (unpack_trees: start splitting internal\nfields from public API, 2023-02-27) [3], and is in preparation for a\nlibification effort in the future around the trailer code. Independent of\nlibification, it still makes sense to discourage callers from peeking into\nthese trailer-internal fields.\n\nPatches 2-3 aim to make some functions do a little less multitasking.\n\nPatch 4 makes the find_patch_start function care about the \"--no-divider\"\noption, because it that option matters for determining the start of the\n\"patch part\" of the input.\n\nPatch 5 is a renaming change to reduce overloaded language in the codebase.\nIt is inspired by 229d6ab6bf (doc: trailer: examples: avoid the word\n\"message\" by itself, 2023-06-15) [4], which did a similar thing for the\ninterpret-trailers documentation.\n\n[1]\nhttps://lore.kernel.org/git/pull.1564.git.1691210737.gitgitgadget@gmail.com/T/#mb044012670663d8eb7a548924bbcc933bef116de\n[2]\nhttps://lore.kernel.org/git/20230627195251.1973421-1-calvinwan@google.com/\n[3]\nhttps://lore.kernel.org/git/pull.1149.git.1677143700.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/6b4cb31b17077181a311ca87e82464a1e2ad67dd.1686797630.git.gitgitgadget@gmail.com/\n\nLinus Arver (5):\n  trailer: separate public from internal portion of trailer_iterator\n  trailer: split process_input_file into separate pieces\n  trailer: split process_command_line_args into separate functions\n  trailer: teach find_patch_start about --no-divider\n  trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n\n trailer.c | 113 ++++++++++++++++++++++++++++++------------------------\n trailer.h |  12 +++---\n 2 files changed, 69 insertions(+), 56 deletions(-)\n\n\nbase-commit: 1b0a5129563ebe720330fdc8f5c6843d27641137\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1563%2Flistx%2Ftrailer-libification-prep-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1563/listx/trailer-libification-prep-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1563\n-- \ngitgitgadget\n"},{"id":"480179","messageId":"0bce4d4b0d5650edf477cbbcc9f4e467b7981426.1691211879.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH 1/5] trailer: separate public from internal portion of trailer_iterator","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:35Z","receivedAt":"2023-08-05T05:04:58Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nThe fields here are not meant to be used by downstream callers, so put\nthem behind an anonymous struct named as\n\"__private_to_trailer_c__do_not_use\" to warn against their use.\n\nInternally, use a \"#define\" to keep the code tidy.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 12 +++++++-----\n trailer.h |  6 ++++--\n 2 files changed, 11 insertions(+), 7 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex f408f9b058d..dff3fafe865 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1214,20 +1214,22 @@ void format_trailers_from_commit(struct strbuf *out, const char *msg,\n \ttrailer_info_release(&info);\n }\n \n+#define private __private_to_trailer_c__do_not_use\n+\n void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n {\n \tstruct process_trailer_options opts = PROCESS_TRAILER_OPTIONS_INIT;\n \tstrbuf_init(&iter->key, 0);\n \tstrbuf_init(&iter->val, 0);\n \topts.no_divider = 1;\n-\ttrailer_info_get(&iter->info, msg, &opts);\n-\titer->cur = 0;\n+\ttrailer_info_get(&iter->private.info, msg, &opts);\n+\titer->private.cur = 0;\n }\n \n int trailer_iterator_advance(struct trailer_iterator *iter)\n {\n-\twhile (iter->cur < iter->info.trailer_nr) {\n-\t\tchar *trailer = iter->info.trailers[iter->cur++];\n+\twhile (iter->private.cur < iter->private.info.trailer_nr) {\n+\t\tchar *trailer = iter->private.info.trailers[iter->private.cur++];\n \t\tint separator_pos = find_separator(trailer, separators);\n \n \t\tif (separator_pos < 1)\n@@ -1245,7 +1247,7 @@ int trailer_iterator_advance(struct trailer_iterator *iter)\n \n void trailer_iterator_release(struct trailer_iterator *iter)\n {\n-\ttrailer_info_release(&iter->info);\n+\ttrailer_info_release(&iter->private.info);\n \tstrbuf_release(&iter->val);\n \tstrbuf_release(&iter->key);\n }\ndiff --git a/trailer.h b/trailer.h\nindex 795d2fccfd9..db57e028650 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -119,8 +119,10 @@ struct trailer_iterator {\n \tstruct strbuf val;\n \n \t/* private */\n-\tstruct trailer_info info;\n-\tsize_t cur;\n+\tstruct {\n+\t\tstruct trailer_info info;\n+\t\tsize_t cur;\n+\t} __private_to_trailer_c__do_not_use;\n };\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"480180","messageId":"d023c297dcac0bb96f681dc1fc0116a649c2efec.1691211879.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH 2/5] trailer: split process_input_file into separate pieces","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:36Z","receivedAt":"2023-08-05T05:05:00Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nCurrently, process_input_file does three things:\n\n    (1) parse the input string for trailers,\n    (2) print text before the trailers, and\n    (3) calculate the position of the input where the trailers end.\n\nRename this function to parse_trailers(), and make it only do\n(1). The caller of this function, process_trailers, becomes responsible\nfor (2) and (3). These items belong inside process_trailers because they\nare both concerned with printing the surrounding text around\ntrailers (which is already one of the immediate concerns of\nprocess_trailers).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 40 +++++++++++++++++++++-------------------\n 1 file changed, 21 insertions(+), 19 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex dff3fafe865..16fbba03d07 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -961,28 +961,23 @@ static void unfold_value(struct strbuf *val)\n \tstrbuf_release(&out);\n }\n \n-static size_t process_input_file(FILE *outfile,\n-\t\t\t\t const char *str,\n-\t\t\t\t struct list_head *head,\n-\t\t\t\t const struct process_trailer_options *opts)\n+/*\n+ * Parse trailers in \"str\" and populate the \"head\" linked list structure.\n+ */\n+static void parse_trailers(struct trailer_info *info,\n+\t\t\t     const char *str,\n+\t\t\t     struct list_head *head,\n+\t\t\t     const struct process_trailer_options *opts)\n {\n-\tstruct trailer_info info;\n \tstruct strbuf tok = STRBUF_INIT;\n \tstruct strbuf val = STRBUF_INIT;\n \tsize_t i;\n \n-\ttrailer_info_get(&info, str, opts);\n-\n-\t/* Print lines before the trailers as is */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(str, 1, info.trailer_start - str, outfile);\n+\ttrailer_info_get(info, str, opts);\n \n-\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n-\t\tfprintf(outfile, \"\\n\");\n-\n-\tfor (i = 0; i < info.trailer_nr; i++) {\n+\tfor (i = 0; i < info->trailer_nr; i++) {\n \t\tint separator_pos;\n-\t\tchar *trailer = info.trailers[i];\n+\t\tchar *trailer = info->trailers[i];\n \t\tif (trailer[0] == comment_line_char)\n \t\t\tcontinue;\n \t\tseparator_pos = find_separator(trailer, separators);\n@@ -1003,9 +998,7 @@ static size_t process_input_file(FILE *outfile,\n \t\t}\n \t}\n \n-\ttrailer_info_release(&info);\n-\n-\treturn info.trailer_end - str;\n+\ttrailer_info_release(info);\n }\n \n static void free_all(struct list_head *head)\n@@ -1054,6 +1047,7 @@ void process_trailers(const char *file,\n {\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n+\tstruct trailer_info info;\n \tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n@@ -1064,8 +1058,16 @@ void process_trailers(const char *file,\n \tif (opts->in_place)\n \t\toutfile = create_in_place_tempfile(file);\n \n+\tparse_trailers(&info, sb.buf, &head, opts);\n+\ttrailer_end = info.trailer_end - sb.buf;\n+\n \t/* Print the lines before the trailers */\n-\ttrailer_end = process_input_file(outfile, sb.buf, &head, opts);\n+\tif (!opts->only_trailers)\n+\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\n+\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n+\t\tfprintf(outfile, \"\\n\");\n+\n \n \tif (!opts->only_input) {\n \t\tLIST_HEAD(arg_head);\n-- \ngitgitgadget\n\n"},{"id":"480181","messageId":"1fc060041db11b3df881cb2c7bd60630dc011a15.1691211879.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH 4/5] trailer: teach find_patch_start about --no-divider","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:38Z","receivedAt":"2023-08-05T05:05:02Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nCurrently, find_patch_start only finds the start of the patch part of\nthe input (by looking at the \"---\" divider) for cases where the\n\"--no-divider\" flag has not been provided. If the user provides this\nflag, we do not rely on find_patch_start at all and just call strlen()\ndirectly on the input.\n\nInstead, make find_patch_start aware of \"--no-divider\" and make it\nhandle that case as well. This means we no longer need to call strlen at\nall and can just rely on the existing code in find_patch_start.\n\nThis patch will make unit testing a bit more pleasant in this area in\nthe future when we adopt a unit testing framework, because we would not\nhave to test multiple functions to check how finding the start of a\npatch part works (we would only need to test find_patch_start).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 10 +++-------\n 1 file changed, 3 insertions(+), 7 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 89246a0d395..3b9ce199636 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -812,14 +812,14 @@ static ssize_t last_line(const char *buf, size_t len)\n  * Return the position of the start of the patch or the length of str if there\n  * is no patch in the message.\n  */\n-static size_t find_patch_start(const char *str)\n+static size_t find_patch_start(const char *str, int no_divider)\n {\n \tconst char *s;\n \n \tfor (s = str; *s; s = next_line(s)) {\n \t\tconst char *v;\n \n-\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v))\n \t\t\treturn s - str;\n \t}\n \n@@ -1109,11 +1109,7 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tensure_configured();\n \n-\tif (opts->no_divider)\n-\t\tpatch_start = strlen(str);\n-\telse\n-\t\tpatch_start = find_patch_start(str);\n-\n+\tpatch_start = find_patch_start(str, opts->no_divider);\n \ttrailer_end = find_trailer_end(str, patch_start);\n \ttrailer_start = find_trailer_start(str, trailer_end);\n \n-- \ngitgitgadget\n\n"},{"id":"480182","messageId":"c8bb013662187e9239d4a2499a63ed76daa78d14.1691211879.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH 3/5] trailer: split process_command_line_args into separate functions","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:37Z","receivedAt":"2023-08-05T05:05:06Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, process_command_line_args did two things:\n\n    (1) parse trailers from the configuration, and\n    (2) parse trailers defined on the command line.\n\nSeparate these concerns into parse_trailers_from_config and\nparse_trailers_from_command_line_args, respectively. Remove (now\nredundant) process_command_line_args.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 34 +++++++++++++++++++++-------------\n 1 file changed, 21 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 16fbba03d07..89246a0d395 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -711,30 +711,35 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tlist_add_tail(&new_item->list, arg_head);\n }\n \n-static void process_command_line_args(struct list_head *arg_head,\n-\t\t\t\t      struct list_head *new_trailer_head)\n+static void parse_trailers_from_config(struct list_head *config_head)\n {\n \tstruct arg_item *item;\n-\tstruct strbuf tok = STRBUF_INIT;\n-\tstruct strbuf val = STRBUF_INIT;\n-\tconst struct conf_info *conf;\n \tstruct list_head *pos;\n \n-\t/*\n-\t * In command-line arguments, '=' is accepted (in addition to the\n-\t * separators that are defined).\n-\t */\n-\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n-\n \t/* Add an arg item for each configured trailer with a command */\n \tlist_for_each(pos, &conf_head) {\n \t\titem = list_entry(pos, struct arg_item, list);\n \t\tif (item->conf.command)\n-\t\t\tadd_arg_item(arg_head,\n+\t\t\tadd_arg_item(config_head,\n \t\t\t\t     xstrdup(token_from_item(item, NULL)),\n \t\t\t\t     xstrdup(\"\"),\n \t\t\t\t     &item->conf, NULL);\n \t}\n+}\n+\n+static void parse_trailers_from_command_line_args(struct list_head *arg_head,\n+\t\t\t\t\t\t  struct list_head *new_trailer_head)\n+{\n+\tstruct strbuf tok = STRBUF_INIT;\n+\tstruct strbuf val = STRBUF_INIT;\n+\tconst struct conf_info *conf;\n+\tstruct list_head *pos;\n+\n+\t/*\n+\t * In command-line arguments, '=' is accepted (in addition to the\n+\t * separators that are defined).\n+\t */\n+\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n \n \t/* Add an arg item for each trailer on the command line */\n \tlist_for_each(pos, new_trailer_head) {\n@@ -1070,8 +1075,11 @@ void process_trailers(const char *file,\n \n \n \tif (!opts->only_input) {\n+\t\tLIST_HEAD(config_head);\n \t\tLIST_HEAD(arg_head);\n-\t\tprocess_command_line_args(&arg_head, new_trailer_head);\n+\t\tparse_trailers_from_config(&config_head);\n+\t\tparse_trailers_from_command_line_args(&arg_head, new_trailer_head);\n+\t\tlist_splice(&config_head, &arg_head);\n \t\tprocess_trailers_lists(&head, &arg_head);\n \t}\n \n-- \ngitgitgadget\n\n"},{"id":"480183","messageId":"7c9b63c26164b037272fde689bb3150b30aa7528.1691211879.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH 5/5] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-08-05T05:04:39Z","receivedAt":"2023-08-05T05:05:08Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nDo not use *_DEFAULT as a suffix to the enums, because the word\n\"default\" is overloaded. The following are two examples of the ambiguity\nof the word \"default\":\n\n(1) \"Default\" can mean using the \"default\" values that are hardcoded\n    in trailer.c as\n\n        default_conf_info.where = WHERE_END;\n        default_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n        default_conf_info.if_missing = MISSING_ADD;\n\n    in ensure_configured(). These values are referred to as \"the\n    default\" in the docs for interpret-trailers. These defaults are used\n    if no \"trailer.*\" configurations are defined.\n\n(2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n    because these configurations are used by \"default\" (ahead of the\n    hardcoded defaults in (1)) if no command line arguments are\n    provided.\n\nIn addition, the corresponding *_DEFAULT values are chosen when the user\nprovides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\non the command line. These \"--no-*\" flags are used to clear previously\nprovided flags of the form \"--where\", \"--if-exists\", and \"--if-missing\".\nUsing these \"--no-*\" flags undoes the specifying of these flags (if\nany), so using the word \"UNSPECIFIED\" is more natural here.\n\nSo instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\nsignals to the reader that the *_UNSPECIFIED value by itself carries no\nmeaning (it's a zero value and by itself does not \"default\" to anything,\nnecessitating the need to have some other way of getting to a useful\nvalue).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 17 ++++++++++-------\n trailer.h |  6 +++---\n 2 files changed, 13 insertions(+), 10 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 3b9ce199636..c49826decae 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -388,7 +388,7 @@ static void process_trailers_lists(struct list_head *head,\n int trailer_set_where(enum trailer_where *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = WHERE_DEFAULT;\n+\t\t*item = WHERE_UNSPECIFIED;\n \telse if (!strcasecmp(\"after\", value))\n \t\t*item = WHERE_AFTER;\n \telse if (!strcasecmp(\"before\", value))\n@@ -405,7 +405,7 @@ int trailer_set_where(enum trailer_where *item, const char *value)\n int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = EXISTS_DEFAULT;\n+\t\t*item = EXISTS_UNSPECIFIED;\n \telse if (!strcasecmp(\"addIfDifferent\", value))\n \t\t*item = EXISTS_ADD_IF_DIFFERENT;\n \telse if (!strcasecmp(\"addIfDifferentNeighbor\", value))\n@@ -424,7 +424,7 @@ int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n int trailer_set_if_missing(enum trailer_if_missing *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = MISSING_DEFAULT;\n+\t\t*item = MISSING_UNSPECIFIED;\n \telse if (!strcasecmp(\"doNothing\", value))\n \t\t*item = MISSING_DO_NOTHING;\n \telse if (!strcasecmp(\"add\", value))\n@@ -586,7 +586,10 @@ static void ensure_configured(void)\n \tif (configured)\n \t\treturn;\n \n-\t/* Default config must be setup first */\n+\t/*\n+\t * Default config must be setup first. These defaults are used if there\n+\t * are no \"trailer.*\" or \"trailer.<token>.*\" options configured.\n+\t */\n \tdefault_conf_info.where = WHERE_END;\n \tdefault_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n \tdefault_conf_info.if_missing = MISSING_ADD;\n@@ -701,11 +704,11 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tnew_item->value = val;\n \tduplicate_conf(&new_item->conf, conf);\n \tif (new_trailer_item) {\n-\t\tif (new_trailer_item->where != WHERE_DEFAULT)\n+\t\tif (new_trailer_item->where != WHERE_UNSPECIFIED)\n \t\t\tnew_item->conf.where = new_trailer_item->where;\n-\t\tif (new_trailer_item->if_exists != EXISTS_DEFAULT)\n+\t\tif (new_trailer_item->if_exists != EXISTS_UNSPECIFIED)\n \t\t\tnew_item->conf.if_exists = new_trailer_item->if_exists;\n-\t\tif (new_trailer_item->if_missing != MISSING_DEFAULT)\n+\t\tif (new_trailer_item->if_missing != MISSING_UNSPECIFIED)\n \t\t\tnew_item->conf.if_missing = new_trailer_item->if_missing;\n \t}\n \tlist_add_tail(&new_item->list, arg_head);\ndiff --git a/trailer.h b/trailer.h\nindex db57e028650..924cf5405c6 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -5,14 +5,14 @@\n #include \"strbuf.h\"\n \n enum trailer_where {\n-\tWHERE_DEFAULT,\n+\tWHERE_UNSPECIFIED,\n \tWHERE_END,\n \tWHERE_AFTER,\n \tWHERE_BEFORE,\n \tWHERE_START\n };\n enum trailer_if_exists {\n-\tEXISTS_DEFAULT,\n+\tEXISTS_UNSPECIFIED,\n \tEXISTS_ADD_IF_DIFFERENT_NEIGHBOR,\n \tEXISTS_ADD_IF_DIFFERENT,\n \tEXISTS_ADD,\n@@ -20,7 +20,7 @@ enum trailer_if_exists {\n \tEXISTS_DO_NOTHING\n };\n enum trailer_if_missing {\n-\tMISSING_DEFAULT,\n+\tMISSING_UNSPECIFIED,\n \tMISSING_ADD,\n \tMISSING_DO_NOTHING\n };\n-- \ngitgitgadget\n"},{"id":"480262","messageId":"kl6l8rama6yj.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"0bce4d4b0d5650edf477cbbcc9f4e467b7981426.1691211879.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/5] trailer: separate public from internal portion of trailer_iterator","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-07T21:16:52Z","receivedAt":"2023-08-07T21:17:04Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"As someone who isn't that familiar with trailer code, and will have less\ntime for the ML soon, this is more of a quick drive-by..\n\n\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> +#define private __private_to_trailer_c__do_not_use\n> +\n>  void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n>  {\n>  \tstruct process_trailer_options opts = PROCESS_TRAILER_OPTIONS_INIT;\n>  \tstrbuf_init(&iter->key, 0);\n>  \tstrbuf_init(&iter->val, 0);\n>  \topts.no_divider = 1;\n> -\ttrailer_info_get(&iter->info, msg, &opts);\n> -\titer->cur = 0;\n> +\ttrailer_info_get(&iter->private.info, msg, &opts);\n> +\titer->private.cur = 0;\n>  }\n> --- a/trailer.h\n> +++ b/trailer.h\n> @@ -119,8 +119,10 @@ struct trailer_iterator {\n>  \tstruct strbuf val;\n>...\n>  \t/* private */\n> -\tstruct trailer_info info;\n> -\tsize_t cur;\n> +\tstruct {\n> +\t\tstruct trailer_info info;\n> +\t\tsize_t cur;\n> +\t} __private_to_trailer_c__do_not_use;\n>  };\n\nInteresting approach to \"private members\". I like that it's fairly\nlightweight and clear. On the other hand, I think this will fail to\nautocomplete on most people's development setups, and I don't think this\nis worth the tradeoff.\n\nThis is the first instance of this I could find in the codebase. I'm not\nreally opposed to having a new way of doing things, but it would be nice\nfor us to be consistent with how we handle private members. Other\napproaches I've seen are:\n\n- Using a \"larger\" struct to hold private members and \"downcasting\" for\n  public users (struct dir_iterator and struct dir_iterator_int). I\n  dislike this because I think this enables 'wrong' memory access too\n  easily.\n\n  (As an aside, if we really wanted to 'strictly' enforce privateness in\n  this patch, shouldn't we move the \"#define private\" into the .c file,\n  the way dir_iterator_int is in the .c file?)\n\n- Prefixing private members with \"__\" (khash.h and other header-only\n  libraries use this at least, not sure if we have this in the 'main\n  tree'). I think this works pretty well most of the time.\n- Just marking private members with a comment. IMO this is good enough\n  the vast majority of the time - if something is private for a good\n  reason, it's unlikely to get used accidentally anyway. But properly\n  enforcing \"privateness\" is worthy goal anyway.\n\nPersonally, I think a decent tradeoff between enforcement and ergonomics\nwould be to use an inner struct like you do here, but name it something\nautocomplete-friendly and obviously private, like \"private\" or\n\"_private\". I suspect self-regulation and code review should be enough\nto catch nearly all accidental uses of private members.\n"},{"id":"480269","messageId":"kl6l5y5qa34v.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"d023c297dcac0bb96f681dc1fc0116a649c2efec.1691211879.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/5] trailer: split process_input_file into separate pieces","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-07T22:39:28Z","receivedAt":"2023-08-07T22:39:36Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Currently, process_input_file does three things:\n>\n>     (1) parse the input string for trailers,\n>     (2) print text before the trailers, and\n>     (3) calculate the position of the input where the trailers end.\n>\n> Rename this function to parse_trailers(), and make it only do\n> (1).\n\nOkay, process_input_file() is a very unhelpful name (What does it mean\nto \"process a file\"?). In contrast, parse_trailers() is more\nself-descriptive (It parses trailers into some appropriate format, so it\nshouldn't do things like print.) Makes sense.\n\nIs there some additional, unstated purpose behind this change besides\n\"move things around for readability\"? E.g. do you intend to move\nparse_trailers() to a future trailer parsing library? If so, that would\nbe useful context to evaluate the goodness of this split.\n\n> The caller of this function, process_trailers, becomes responsible\n> for (2) and (3). These items belong inside process_trailers because they\n> are both concerned with printing the surrounding text around\n> trailers (which is already one of the immediate concerns of\n> process_trailers).\n\nI agree that (2) doesn't belong in parse_trailers(). OTOH, (3) sounds\nlike something that belongs in parse_trailers() - you have to parse\ntrailers in order to tell where the trailers start and end, so it makes\nsense for the parsing function to give those values.\n\n> diff --git a/trailer.c b/trailer.c\n> index dff3fafe865..16fbba03d07 100644\n> --- a/trailer.c\n> +++ b/trailer.c\n> @@ -961,28 +961,23 @@ static void unfold_value(struct strbuf *val)\n>  \tstrbuf_release(&out);\n>  }\n>  \n> -static size_t process_input_file(FILE *outfile,\n> -\t\t\t\t const char *str,\n> -\t\t\t\t struct list_head *head,\n> -\t\t\t\t const struct process_trailer_options *opts)\n> +/*\n> + * Parse trailers in \"str\" and populate the \"head\" linked list structure.\n> + */\n> +static void parse_trailers(struct trailer_info *info,\n\n\"info\" is an out parameter, and IIRC we typically put out parameters\ntowards the end. I didn't find a callout in CodingGuidelines, though, so\nidk if this is an ironclad rule or not.\n\n> +\t\t\t     const char *str,\n> +\t\t\t     struct list_head *head,\n> +\t\t\t     const struct process_trailer_options *opts)\n>  {\n> -\tstruct trailer_info info;\n>  \tstruct strbuf tok = STRBUF_INIT;\n>  \tstruct strbuf val = STRBUF_INIT;\n>  \tsize_t i;\n>  \n> -\ttrailer_info_get(&info, str, opts);\n> -\n> -\t/* Print lines before the trailers as is */\n> -\tif (!opts->only_trailers)\n> -\t\tfwrite(str, 1, info.trailer_start - str, outfile);\n\nWe no longer fwrite the contents before the trailer, okay.\n\n> +\ttrailer_info_get(info, str, opts);\n\nThis is where we actually get the start and end of trailers, and each\ntrailer string. This is parsing out the trailers from a string, so what\nother parsing is left? Reading ahead shows that we're actually parsing\nthe trailer string into a \"struct trailer_item\". Okay, so this function\nis basically a wrapper around trailer_info_get() that also \"returns\" the\nparsed trailer_items.\n\n> -\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n> -\t\tfprintf(outfile, \"\\n\");\n> -\n\nSo we don't print the trailing line. Also makes sense.\n\n> @@ -1003,9 +998,7 @@ static size_t process_input_file(FILE *outfile,\n>  \t\t}\n>  \t}\n>  \n> -\ttrailer_info_release(&info);\n> -\n> -\treturn info.trailer_end - str;\n> +\ttrailer_info_release(info);\n>  }\n>  \n\nEven though \"info\" is a pointer passed into this function, we are\n_release-ing it. This is not an umabiguously good change, IMO. Before,\n\"info\" was never used outside of this function, so we should obviously\nrelease it before returning. However, now that \"info\" is an out\nparameter, we should be more careful about releasing it. I don't think\nit's obvious that the caller will see the right values for\ninfo.trailer_end and info.trailer_start, but free()-d values for\ninfo.trailers, and a meaningless value for info.trailer_nr (since the\nitems were free()-d).\n\nI think it might be better to update the comment on parse_trailers()\nlike so:\n\n  /*\n   * Parse trailers in \"str\", populating the trailer info and \"head\"\n   * linked list structure.\n   */\n\nand make it the caller's responsibility to call trailer_info_release().\nWe could move this call to where we \"free_all(head)\".\n\n>  static void free_all(struct list_head *head)\n> @@ -1054,6 +1047,7 @@ void process_trailers(const char *file,\n>  {\n>  \tLIST_HEAD(head);\n>  \tstruct strbuf sb = STRBUF_INIT;\n> +\tstruct trailer_info info;\n>  \tsize_t trailer_end;\n>  \tFILE *outfile = stdout;\n>  \n> @@ -1064,8 +1058,16 @@ void process_trailers(const char *file,\n>  \tif (opts->in_place)\n>  \t\toutfile = create_in_place_tempfile(file);\n\nThinking out loud, should we move the creation of outfile next to where\nwe first use it?\n\n> +\tparse_trailers(&info, sb.buf, &head, opts);\n> +\ttrailer_end = info.trailer_end - sb.buf;\n> +\n>  \t/* Print the lines before the trailers */\n> -\ttrailer_end = process_input_file(outfile, sb.buf, &head, opts);\n> +\tif (!opts->only_trailers)\n> +\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n\nI'm not sure if it is an unambiguously good change for the caller to\nlearn how to compute the start and end of the trailer sections by doing\npointer arithmetic, but I guess format_trailer_info() does this anyway,\nso your proposal to move (3) outside of the parse_trailers() makes\nsense.\n\nIt feels a bit non-obvious that trailer_start and trailer_end are\npointing inside the input string. I wonder if we should just return the\n_start and _end offsets directly instead of returning pointers. I.e.:\n\n   struct trailer_info {\n     int blank_line_before_trailer;\n -  /*\n -   * Pointers to the start and end of the trailer block found. If there\n -   * is no trailer block found, these 2 pointers point to the end of the\n -   * input string.\n -   */\n -   const char *trailer_start, *trailer_end;\n +   /* Offsets to the trailer block start and end in the input string */\n +   size_t *trailer_start, *trailer_end;\n\nWhich makes their intended use fairly unambiguous. A quick grep suggests\nthat in trailer.c, we're roughly as likely to use the pointer directly\nvs using it to do pointer arithmetic, so converging on one use might be\na win for readability. The only other user outside of trailer.c is\nsequencer.c, which doesn't care about the return type - it only checks\nif there are trailers.\n"},{"id":"480271","messageId":"kl6l1qgea2k0.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"c8bb013662187e9239d4a2499a63ed76daa78d14.1691211879.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 3/5] trailer: split process_command_line_args into separate functions","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-07T22:51:59Z","receivedAt":"2023-08-07T22:52:05Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Previously, process_command_line_args did two things:\n>\n>     (1) parse trailers from the configuration, and\n>     (2) parse trailers defined on the command line.\n\nIt parses trailers from two places, but it still only does \"one thing\",\nin that it only parses trailers.\n\n> @@ -711,30 +711,35 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n>  \tlist_add_tail(&new_item->list, arg_head);\n>  }\n>  \n> -static void process_command_line_args(struct list_head *arg_head,\n> -\t\t\t\t      struct list_head *new_trailer_head)\n> +static void parse_trailers_from_config(struct list_head *config_head)\n>  {\n>  \tstruct arg_item *item;\n> -\tstruct strbuf tok = STRBUF_INIT;\n> -\tstruct strbuf val = STRBUF_INIT;\n> -\tconst struct conf_info *conf;\n>  \tstruct list_head *pos;\n>  \n> -\t/*\n> -\t * In command-line arguments, '=' is accepted (in addition to the\n> -\t * separators that are defined).\n> -\t */\n> -\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n> -\n>  \t/* Add an arg item for each configured trailer with a command */\n>  \tlist_for_each(pos, &conf_head) {\n>  \t\titem = list_entry(pos, struct arg_item, list);\n>  \t\tif (item->conf.command)\n> -\t\t\tadd_arg_item(arg_head,\n> +\t\t\tadd_arg_item(config_head,\n>  \t\t\t\t     xstrdup(token_from_item(item, NULL)),\n>  \t\t\t\t     xstrdup(\"\"),\n>  \t\t\t\t     &item->conf, NULL);\n>  \t}\n> +}\n> +\n> +static void parse_trailers_from_command_line_args(struct list_head *arg_head,\n> +\t\t\t\t\t\t  struct list_head *new_trailer_head)\n> +{\n> +\tstruct strbuf tok = STRBUF_INIT;\n> +\tstruct strbuf val = STRBUF_INIT;\n> +\tconst struct conf_info *conf;\n> +\tstruct list_head *pos;\n> +\n> +\t/*\n> +\t * In command-line arguments, '=' is accepted (in addition to the\n> +\t * separators that are defined).\n> +\t */\n> +\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n>  \n>  \t/* Add an arg item for each trailer on the command line */\n>  \tlist_for_each(pos, new_trailer_head) {\n\nI find this equally readable as the preimage, which IMO is adequately\nscoped and commented.\n\n> @@ -1070,8 +1075,11 @@ void process_trailers(const char *file,\n>  \n>  \n>  \tif (!opts->only_input) {\n> +\t\tLIST_HEAD(config_head);\n>  \t\tLIST_HEAD(arg_head);\n> -\t\tprocess_command_line_args(&arg_head, new_trailer_head);\n> +\t\tparse_trailers_from_config(&config_head);\n> +\t\tparse_trailers_from_command_line_args(&arg_head, new_trailer_head);\n> +\t\tlist_splice(&config_head, &arg_head);\n>  \t\tprocess_trailers_lists(&head, &arg_head);\n>  \t}\n\nBut now, we have to remember to call two functions instead of just one.\nThis, and the slight additional churn makes me lean negative on this\nchange. I would be really happy if we had a use case where we only\nwanted to call one function but not the other, but it seems like this\nisn't the case.\n"},{"id":"480273","messageId":"kl6ly1im8ma5.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"1fc060041db11b3df881cb2c7bd60630dc011a15.1691211879.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/5] trailer: teach find_patch_start about --no-divider","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-07T23:28:50Z","receivedAt":"2023-08-07T23:29:07Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> This patch will make unit testing a bit more pleasant in this area in\n> the future when we adopt a unit testing framework, because we would not\n> have to test multiple functions to check how finding the start of a\n> patch part works (we would only need to test find_patch_start).\n\nUnit tests typically only test external-facing interfaces, not\nimplementatino details, so without seeing the unit tests or library\nboundary, it's hard to tell whether find_patch_start() is something we\nwant to unit test or not. I would have assumed it's not, given that it's\ntiny and only has a single caller, so I'm hesitant to say that we should\ndefinitely handle no_divider inside find_patch_start().\n\n> @@ -812,14 +812,14 @@ static ssize_t last_line(const char *buf, size_t len)\n>   * Return the position of the start of the patch or the length of str if there\n>   * is no patch in the message.\n>   */\n> -static size_t find_patch_start(const char *str)\n> +static size_t find_patch_start(const char *str, int no_divider)\n>  {\n>  \tconst char *s;\n>  \n>  \tfor (s = str; *s; s = next_line(s)) {\n>  \t\tconst char *v;\n>  \n> -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n> +\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v))\n>  \t\t\treturn s - str;\n>  \t}\n\nAssuming we wanted to make this unit-testable anyway, could we just move\nthe strlen() call into this function? Performance aside (I wouldn't be\nsurprised if a smart enough compiler could optimize away the noops), I\ndon't find this easier to understand. Now the reader needs to read the\ncode to see \"if no_divider is given, noop until the end of the string,\nat which point str will point to the end, and s - str will give us the\nlength of str\", as opposed to \"there are no dividers, so just return\nstrlen(str)\".\n"},{"id":"480274","messageId":"kl6lv8dq8li0.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"7c9b63c26164b037272fde689bb3150b30aa7528.1691211879.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/5] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-07T23:45:43Z","receivedAt":"2023-08-07T23:45:50Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> (2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n>     because these configurations are used by \"default\" (ahead of the\n>     hardcoded defaults in (1)) if no command line arguments are\n>     provided.\n\nInteresting, I would have never thought of config as 'default'. In fact,\nI would have thought that this de facto behavior (which you also\nclarified in [1]) is a bug if not for the fact that in an internal\nversion of this series, you cited a commit message that describes this\nas expected behavior. That context would be very welcome in the ML, I\nthink.\n\n[1] https://lore.kernel.org/git/6b427b4b1e82b1f01640f1f49fe8d1c2fd02111e.1691210737.git.gitgitgadget@gmail.com\n\n> In addition, the corresponding *_DEFAULT values are chosen when the user\n> provides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\n> on the command line. These \"--no-*\" flags are used to clear previously\n> provided flags of the form \"--where\", \"--if-exists\", and \"--if-missing\".\n> Using these \"--no-*\" flags undoes the specifying of these flags (if\n> any), so using the word \"UNSPECIFIED\" is more natural here.\n>\n> So instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\n> signals to the reader that the *_UNSPECIFIED value by itself carries no\n> meaning (it's a zero value and by itself does not \"default\" to anything,\n> necessitating the need to have some other way of getting to a useful\n> value).\n\nMakse sense. This seems like a good change.\n\n> @@ -586,7 +586,10 @@ static void ensure_configured(void)\n>  \tif (configured)\n>  \t\treturn;\n>  \n> -\t/* Default config must be setup first */\n> +\t/*\n> +\t * Default config must be setup first. These defaults are used if there\n> +\t * are no \"trailer.*\" or \"trailer.<token>.*\" options configured.\n> +\t */\n>  \tdefault_conf_info.where = WHERE_END;\n>  \tdefault_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n>  \tdefault_conf_info.if_missing = MISSING_ADD;\n\nAs mentioned earlier, I find it a bit odd that we're calling config\n'default' (and also that we're calling CLI args config), but\nrenaming default_conf_info to config_conf_info sounds worse, so let's\nleave it as-is.\n"},{"id":"480316","messageId":"1fd1f22d-e0db-04f3-7235-899b10909c7a@gmail.com","threadId":"60068","inReplyTo":"kl6l8rama6yj.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 1/5] trailer: separate public from internal portion of trailer_iterator","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2023-08-08T12:19:13Z","receivedAt":"2023-08-08T19:46:47Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 07/08/2023 22:16, Glen Choo wrote:\n> As someone who isn't that familiar with trailer code, and will have less\n> time for the ML soon, this is more of a quick drive-by..\n\nThis is a bit of a drive-by comment as well ...\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> +#define private __private_to_trailer_c__do_not_use\n>> +\n>>   void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n>>   {\n>>   \tstruct process_trailer_options opts = PROCESS_TRAILER_OPTIONS_INIT;\n>>   \tstrbuf_init(&iter->key, 0);\n>>   \tstrbuf_init(&iter->val, 0);\n>>   \topts.no_divider = 1;\n>> -\ttrailer_info_get(&iter->info, msg, &opts);\n>> -\titer->cur = 0;\n>> +\ttrailer_info_get(&iter->private.info, msg, &opts);\n>> +\titer->private.cur = 0;\n>>   }\n>> --- a/trailer.h\n>> +++ b/trailer.h\n>> @@ -119,8 +119,10 @@ struct trailer_iterator {\n>>   \tstruct strbuf val;\n>> ...\n>>   \t/* private */\n>> -\tstruct trailer_info info;\n>> -\tsize_t cur;\n>> +\tstruct {\n>> +\t\tstruct trailer_info info;\n>> +\t\tsize_t cur;\n>> +\t} __private_to_trailer_c__do_not_use;\n>>   };\n> \n> Interesting approach to \"private members\". I like that it's fairly\n> lightweight and clear. On the other hand, I think this will fail to\n> autocomplete on most people's development setups, and I don't think this\n> is worth the tradeoff.\n> \n> This is the first instance of this I could find in the codebase. \n\nWe have something similar in unpack_trees.h see 576de3d9560 \n(unpack_trees: start splitting internal fields from public API, \n2023-02-27). That adds an \"internal\" member to \"sturct unpack_trees\" of \ntype \"struct unpack_trees_internal which seems to be a easier naming scheme.\n\n> I'm not\n> really opposed to having a new way of doing things, but it would be nice\n> for us to be consistent with how we handle private members. Other\n> approaches I've seen are:\n> \n> - Using a \"larger\" struct to hold private members and \"downcasting\" for\n>    public users (struct dir_iterator and struct dir_iterator_int). I\n>    dislike this because I think this enables 'wrong' memory access too\n>    easily.\n> \n>    (As an aside, if we really wanted to 'strictly' enforce privateness in\n>    this patch, shouldn't we move the \"#define private\" into the .c file,\n>    the way dir_iterator_int is in the .c file?)\n\nThat #define is pretty ugly\n\nAnother common scheme is to have an opaque pointer to the private struct \n  in the public struct (aka pimpl idiom). The merge machinery uses this \n- see merge-recursive.h. (I'm working on something similar for the \nsequencer so we can change the internals without having to re-compile \neverything that includes \"sequencer.h\")\n\n> - Prefixing private members with \"__\" (khash.h and other header-only\n>    libraries use this at least, not sure if we have this in the 'main\n>    tree'). I think this works pretty well most of the time.\n\nIt is common but I think the C standard reserves names beginning with \"__\"\n\n> - Just marking private members with a comment. IMO this is good enough\n>    the vast majority of the time - if something is private for a good\n>    reason, it's unlikely to get used accidentally anyway. But properly\n>    enforcing \"privateness\" is worthy goal anyway.\n>\n> Personally, I think a decent tradeoff between enforcement and ergonomics\n> would be to use an inner struct like you do here, but name it something\n> autocomplete-friendly and obviously private, like \"private\" or\n> \"_private\".\n\nI agree, something like that would match the unpack_trees example\n\n> I suspect self-regulation and code review should be enough\n> to catch nearly all accidental uses of private members.\n\nAgreed\n\nBest Wishes\n\nPhillip\n\n"},{"id":"480511","messageId":"owly1qgah5qk.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"kl6l8rama6yj.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 1/5] trailer: separate public from internal portion of trailer_iterator","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-10T22:50:27Z","receivedAt":"2023-08-10T22:50:34Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Glen Choo <chooglen@google.com> writes:\n\n> As someone who isn't that familiar with trailer code, and will have less\n> time for the ML soon, this is more of a quick drive-by..\n\nAren't you also going on vacation soon? ;-)\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> +#define private __private_to_trailer_c__do_not_use\n>> +\n>>  void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n>>  {\n>>  \tstruct process_trailer_options opts = PROCESS_TRAILER_OPTIONS_INIT;\n>>  \tstrbuf_init(&iter->key, 0);\n>>  \tstrbuf_init(&iter->val, 0);\n>>  \topts.no_divider = 1;\n>> -\ttrailer_info_get(&iter->info, msg, &opts);\n>> -\titer->cur = 0;\n>> +\ttrailer_info_get(&iter->private.info, msg, &opts);\n>> +\titer->private.cur = 0;\n>>  }\n>> --- a/trailer.h\n>> +++ b/trailer.h\n>> @@ -119,8 +119,10 @@ struct trailer_iterator {\n>>  \tstruct strbuf val;\n>>...\n>>  \t/* private */\n>> -\tstruct trailer_info info;\n>> -\tsize_t cur;\n>> +\tstruct {\n>> +\t\tstruct trailer_info info;\n>> +\t\tsize_t cur;\n>> +\t} __private_to_trailer_c__do_not_use;\n>>  };\n>\n> [...]\n>\n> This is the first instance of this I could find in the codebase. I'm not\n> really opposed to having a new way of doing things, but it would be nice\n> for us to be consistent with how we handle private members. Other\n> approaches I've seen are:\n>\n> - Using a \"larger\" struct to hold private members and \"downcasting\" for\n>   public users (struct dir_iterator and struct dir_iterator_int). I\n>   dislike this because I think this enables 'wrong' memory access too\n>   easily.\n>   [...]\n> - Prefixing private members with \"__\" (khash.h and other header-only\n>   libraries use this at least, not sure if we have this in the 'main\n>   tree'). I think this works pretty well most of the time.\n> - Just marking private members with a comment. IMO this is good enough\n>   the vast majority of the time - if something is private for a good\n>   reason, it's unlikely to get used accidentally anyway. But properly\n>   enforcing \"privateness\" is worthy goal anyway.\n\nThanks for documenting these other approaches.\n\nI prefer the \"larger\" struct to hold private members pattern. More\nspecifically I like the container_of approach pointed out by Jacob [2],\nbecause it is an established pattern in the Linux Kernel and because we\nalready sort of use the same idea in the list_head type we imported from\nthe Kernel in 94e99012fc (http-walker: reduce O(n) ops with\ndoubly-linked list, 2016-07-11). That is, for example for the\nnew_trailer_item struct we have\n\n    struct new_trailer_item {\n        struct list_head list;\n        <list item stuff>\n    };\n\nand to me this is symmetric to the container_of pattern described by Jacob:\n\n    struct dir_entry_private {\n        struct dir_entry entry;\n        <private stuff>\n    };\n\nAccordingly, we are already doing the \"structure pointer math\" (which\nJacob described in [2]) for list_head in list.h:\n\n    /* Get typed element from list at a given position. */\n    #define list_entry(ptr, type, member) \\\n        ((type *) ((char *) (ptr) - offsetof(type, member)))\n\nIn this patch series though, I decided to just stick with giving the\nstruct a private-sounding name, because I don't think we reached\nconsensus on what the preferred approach is for separating\npublic/private fields.\n\n>   (As an aside, if we really wanted to 'strictly' enforce privateness in\n>   this patch, shouldn't we move the \"#define private\" into the .c file,\n>   the way dir_iterator_int is in the .c file?)\n\nI think you meant moving the struct into the .c file (the \"#define\" is\nalready in the .c file).\n\n> Personally, I think a decent tradeoff between enforcement and ergonomics\n> would be to use an inner struct like you do here, but name it something\n> autocomplete-friendly and obviously private, like \"private\" or\n> \"_private\".\n\nSGTM. I think I'll go with \"internal\" though, to align with 576de3d956\n(unpack_trees: start splitting internal fields from public API,\n2023-02-27) which Phillip pointed out. Will reroll.\n\n> I suspect self-regulation and code review should be enough\n> to catch nearly all accidental uses of private members.\n\nAck. In the future, if and when we want compiler-level guarantees to\nmake it impossible (this was the discussion at [1]), we can revisit this\narea.\n\n[1] https://lore.kernel.org/git/16ff5069-0408-21cd-995c-8b47afb9810d@github.com/\n[2] https://lore.kernel.org/git/CA+P7+xo02dGkjb5DwJ1Af_hoQ5HiuxASheZxoFz+r6B-6cQMug@mail.gmail.com/\n"},{"id":"480512","messageId":"owlyy1iifq0n.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"1fd1f22d-e0db-04f3-7235-899b10909c7a@gmail.com","subject":"Re: [PATCH 1/5] trailer: separate public from internal portion of trailer_iterator","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-10T23:15:20Z","receivedAt":"2023-08-10T23:15:25Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> We have something similar in unpack_trees.h see 576de3d9560 \n> (unpack_trees: start splitting internal fields from public API, \n> 2023-02-27). That adds an \"internal\" member to \"sturct unpack_trees\" of \n> type \"struct unpack_trees_internal which seems to be a easier naming scheme.\n\nAck, I will use \"internal\" as the member name in the next reroll.\n\n>>> +#define private __private_to_trailer_c__do_not_use\n> [...]\n> That #define is pretty ugly\n\nHaha, indeed. But I think that's the point though (i.e., the degree of\nugliness matches the strength of our codebase's posture to discourage\nits use by external users).\n\nI will drop the #define in the next reroll though, so, I guess it's a\nmoot point anyway.\n\n> Another common scheme is to have an opaque pointer to the private struct \n>   in the public struct (aka pimpl idiom). The merge machinery uses this \n> - see merge-recursive.h. (I'm working on something similar for the \n> sequencer so we can change the internals without having to re-compile \n> everything that includes \"sequencer.h\")\n\nVery interesting! I look forward to seeing your work. :)\n\n>> - Prefixing private members with \"__\" (khash.h and other header-only\n>>    libraries use this at least, not sure if we have this in the 'main\n>>    tree'). I think this works pretty well most of the time.\n>\n> It is common but I think the C standard reserves names beginning with \"__\"\n\nIndeed (see [1]).\n\n[1] https://devblogs.microsoft.com/oldnewthing/20230109-00/?p=107685\n"},{"id":"480517","messageId":"owlyttt6fm0v.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"kl6l5y5qa34v.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 2/5] trailer: split process_input_file into separate pieces","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-11T00:41:36Z","receivedAt":"2023-08-11T00:41:41Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Glen Choo <chooglen@google.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> Currently, process_input_file does three things:\n>>\n>>     (1) parse the input string for trailers,\n>>     (2) print text before the trailers, and\n>>     (3) calculate the position of the input where the trailers end.\n>>\n>> Rename this function to parse_trailers(), and make it only do\n>> (1).\n>\n> [...]\n>\n> Is there some additional, unstated purpose behind this change besides\n> \"move things around for readability\"? E.g. do you intend to move\n> parse_trailers() to a future trailer parsing library? If so, that would\n> be useful context to evaluate the goodness of this split.\n\nI think it's still too early to say whether certain functions will make\nit (unmodified) into the public, libified API. So currently \"move things\naround for readability\" is the most concrete reason.\n\n>> The caller of this function, process_trailers, becomes responsible\n>> for (2) and (3). These items belong inside process_trailers because they\n>> are both concerned with printing the surrounding text around\n>> trailers (which is already one of the immediate concerns of\n>> process_trailers).\n>\n> I agree that (2) doesn't belong in parse_trailers(). OTOH, (3) sounds\n> like something that belongs in parse_trailers() - you have to parse\n> trailers in order to tell where the trailers start and end, so it makes\n> sense for the parsing function to give those values.\n\nI don't think (3) should belong in parse_trailers, mainly because the\n\"info\" struct we pass into it gets this information populated by\nparse_trailers already. Which is why we can do (3) in the caller with\n\n    parse_trailers(&info, sb.buf, &head, opts);\n    trailer_end = info.trailer_end - sb.buf;\n\nto get the same information. Also, the endpoint of the trailers is no\nmore inherently special than, for example, the following other possible\nreturn values:\n\n- the number of trailers that were recognized and parsed\n- whether we encountered any trailers or not\n- the start position of when we first encountered a trailer in the input\n\nwhich makes me want to avoid returning this \"trailer_end\" value from\nparse_trailers.\n\nOne more thing: we already have a function named \"find_trailer_end\"\nwhich is supposed to do this already. But it uses \"ignore_non_trailer\"\nfrom commit.c (that function should probably use the trailer API later\non to figure this out...). I wanted to clean that part up in the future\nas part of libifcation.\n\n>> -static size_t process_input_file(FILE *outfile,\n>> -\t\t\t\t const char *str,\n>> -\t\t\t\t struct list_head *head,\n>> -\t\t\t\t const struct process_trailer_options *opts)\n>> +/*\n>> + * Parse trailers in \"str\" and populate the \"head\" linked list structure.\n>> + */\n>> +static void parse_trailers(struct trailer_info *info,\n>\n> \"info\" is an out parameter, and IIRC we typically put out parameters\n> towards the end. I didn't find a callout in CodingGuidelines, though, so\n> idk if this is an ironclad rule or not.\n\nI wanted to minimize churn as much as possible (hence my hesitation with\nchanging around the order of these parameters). But also,\ntrailer_info_get uses \"info\" as the first parameter, so I wanted to\nalign with that usage.\n\n>> @@ -1003,9 +998,7 @@ static size_t process_input_file(FILE *outfile,\n>>  \t\t}\n>>  \t}\n>>\n>> -\ttrailer_info_release(&info);\n>> -\n>> -\treturn info.trailer_end - str;\n>> +\ttrailer_info_release(info);\n>>  }\n>>\n>\n> Even though \"info\" is a pointer passed into this function, we are\n> _release-ing it. This is not an umabiguously good change, IMO. Before,\n> \"info\" was never used outside of this function, so we should obviously\n> release it before returning. However, now that \"info\" is an out\n> parameter, we should be more careful about releasing it.\n\nHmm, agreed.\n\n> I don't think\n> it's obvious that the caller will see the right values for\n> info.trailer_end and info.trailer_start, but free()-d values for\n> info.trailers, and a meaningless value for info.trailer_nr (since the\n> items were free()-d).\n\nAgreed. Will update to avoid calling trailer_info_release() inside\nparse_trailers() because the caller might still need that information. I\nthink the fix is to move the trailer_info_get outside to the caller,\nmuch like how format_trailers_from_commit() does it.\n\n> I think it might be better to update the comment on parse_trailers()\n> like so:\n>\n>   /*\n>    * Parse trailers in \"str\", populating the trailer info and \"head\"\n>    * linked list structure.\n>    */\n>\n> and make it the caller's responsibility to call trailer_info_release().\n> We could move this call to where we \"free_all(head)\".\n\nSGTM. (I regret not reading this text before drafting my response above.)\n\n>>  static void free_all(struct list_head *head)\n>> @@ -1054,6 +1047,7 @@ void process_trailers(const char *file,\n>>  {\n>>  \tLIST_HEAD(head);\n>>  \tstruct strbuf sb = STRBUF_INIT;\n>> +\tstruct trailer_info info;\n>>  \tsize_t trailer_end;\n>>  \tFILE *outfile = stdout;\n>>\n>> @@ -1064,8 +1058,16 @@ void process_trailers(const char *file,\n>>  \tif (opts->in_place)\n>>  \t\toutfile = create_in_place_tempfile(file);\n>\n> Thinking out loud, should we move the creation of outfile next to where\n> we first use it?\n\nNot sure what you mean here. Can you clarify?\n\n>> +\tparse_trailers(&info, sb.buf, &head, opts);\n>> +\ttrailer_end = info.trailer_end - sb.buf;\n>> +\n>>  \t/* Print the lines before the trailers */\n>> -\ttrailer_end = process_input_file(outfile, sb.buf, &head, opts);\n>> +\tif (!opts->only_trailers)\n>> +\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n>\n> I'm not sure if it is an unambiguously good change for the caller to\n> learn how to compute the start and end of the trailer sections by doing\n> pointer arithmetic,\n\nI think a future cleanup (in a follow-up series) involving\nfind_trailer_end should simplify this area and avoid the need for\npointer arithmetic in the caller.\n\n> It feels a bit non-obvious that trailer_start and trailer_end are\n> pointing inside the input string. I wonder if we should just return the\n> _start and _end offsets directly instead of returning pointers. I.e.:\n>\n>    struct trailer_info {\n>      int blank_line_before_trailer;\n>  -  /*\n>  -   * Pointers to the start and end of the trailer block found. If there\n>  -   * is no trailer block found, these 2 pointers point to the end of the\n>  -   * input string.\n>  -   */\n>  -   const char *trailer_start, *trailer_end;\n>  +   /* Offsets to the trailer block start and end in the input string */\n>  +   size_t *trailer_start, *trailer_end;\n>\n> Which makes their intended use fairly unambiguous. A quick grep suggests\n> that in trailer.c, we're roughly as likely to use the pointer directly\n> vs using it to do pointer arithmetic, so converging on one use might be\n> a win for readability.\n\nAgreed! I would prefer to use offsets everywhere, as I think that is\nmore direct (because we are concerned with locations in the input).\n"},{"id":"480518","messageId":"owlypm3ufl7j.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"kl6l1qgea2k0.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 3/5] trailer: split process_command_line_args into separate functions","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-11T00:59:12Z","receivedAt":"2023-08-11T00:59:21Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Glen Choo <chooglen@google.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> Previously, process_command_line_args did two things:\n>>\n>>     (1) parse trailers from the configuration, and\n>>     (2) parse trailers defined on the command line.\n>\n> It parses trailers from two places, but it still only does \"one thing\",\n> in that it only parses trailers.\n\nMore precisely, it parses trailers from the command line by first\nparsing trailers from the configuration. In other words, parsing\ntrailers from the configuration (independent of the input string!) is a\nrequired dependency for parsing trailers coming from the command line.\n\nIf we take a step back, we need to do 3 things:\n\n   (1) parse trailers from the configuration\n   (2) parse trailers from the command line\n   (3) parse trailers from the input\n\nI think these three should be separated into separate functions. I think\nno one wants to combine all three into one function. And I can't think\nof a good enough reason to combine (1) and (2) together either. Hence\nthis patch.\n\n> I find this equally readable as the preimage, which IMO is adequately\n> scoped and commented.\n\nAside: is \"preimage\" the status quo before applying the patch?\n\n>> @@ -1070,8 +1075,11 @@ void process_trailers(const char *file,\n>>  \n>>  \n>>  \tif (!opts->only_input) {\n>> +\t\tLIST_HEAD(config_head);\n>>  \t\tLIST_HEAD(arg_head);\n>> -\t\tprocess_command_line_args(&arg_head, new_trailer_head);\n>> +\t\tparse_trailers_from_config(&config_head);\n>> +\t\tparse_trailers_from_command_line_args(&arg_head, new_trailer_head);\n>> +\t\tlist_splice(&config_head, &arg_head);\n>>  \t\tprocess_trailers_lists(&head, &arg_head);\n>>  \t}\n>\n> But now, we have to remember to call two functions instead of just one.\n\nBut only inside interpret-trailers.c, right?\n"},{"id":"480520","messageId":"owlymsyyfl26.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"owlypm3ufl7j.fsf@fine.c.googlers.com","subject":"Re: [PATCH 3/5] trailer: split process_command_line_args into separate functions","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-11T01:02:25Z","receivedAt":"2023-08-11T01:02:31Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n>>\n>> But now, we have to remember to call two functions instead of just one.\n>\n> But only inside interpret-trailers.c, right?\n\nOops, I meant trailer.c.\n\nIn the future I expect to move this to interpret-trailers.c.\n"},{"id":"480522","messageId":"owlyjzu2fjz9.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"kl6ly1im8ma5.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 4/5] trailer: teach find_patch_start about --no-divider","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-11T01:25:46Z","receivedAt":"2023-08-11T01:25:53Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Glen Choo <chooglen@google.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> @@ -812,14 +812,14 @@ static ssize_t last_line(const char *buf, size_t len)\n>>   * Return the position of the start of the patch or the length of str if there\n>>   * is no patch in the message.\n>>   */\n>> -static size_t find_patch_start(const char *str)\n>> +static size_t find_patch_start(const char *str, int no_divider)\n>>  {\n>>  \tconst char *s;\n>>  \n>>  \tfor (s = str; *s; s = next_line(s)) {\n>>  \t\tconst char *v;\n>>  \n>> -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n>> +\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v))\n>>  \t\t\treturn s - str;\n>>  \t}\n>\n> Assuming we wanted to make this unit-testable anyway, could we just move\n> the strlen() call into this function?\n\nI don't see why we should preserve the if-statement and associated\nstrlen() call if we can just get rid of it.\n\n> [...] I\n> don't find this easier to understand. Now the reader needs to read the\n> code to see \"if no_divider is given, noop until the end of the string,\n> at which point str will point to the end, and s - str will give us the\n> length of str\", as opposed to \"there are no dividers, so just return\n> strlen(str)\".\n\nThe main idea behind this patch is to make find_patch_start() return the\ncorrect answer. Currently it does not in all cases (whether --no-divider\nis provided), and so the caller has to calculate the\nstart of the patch with strlen manually. By moving the --no-divider flag\ninto this function, we force all callers to consider this important\noption.\n\nFor additional context we recently had to fix a bug where we weren't\npassing in this flag to the interpret-trailers builtin. See be3d654343\n(commit: pass --no-divider to interpret-trailers, 2023-06-17). There we\nacknowledged that some callers forgot to pass in --no-divider to\ninterpret-trailers (such as the caller that was fixed up in that\ncommit).\n\nI mention the above example because although it's not the exact same\nthing as here, I think the scenario of \"sometimes callers can forget\nabout --no-divider\" is an important one to prevent wherever possible.\nThat's why I like this patch (in addition to the reasons cited in the\ncommit message).\n"},{"id":"480568","messageId":"owlycyztfohb.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"kl6lv8dq8li0.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: [PATCH 5/5] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-08-11T18:00:48Z","receivedAt":"2023-08-11T18:00:51Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Glen Choo <chooglen@google.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> (2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n>>     because these configurations are used by \"default\" (ahead of the\n>>     hardcoded defaults in (1)) if no command line arguments are\n>>     provided.\n>\n> Interesting, I would have never thought of config as 'default'. In fact,\n> I would have thought that this de facto behavior (which you also\n> clarified in [1]) is a bug if not for the fact that in an internal\n> version of this series, you cited a commit message that describes this\n> as expected behavior. That context would be very welcome in the ML, I\n> think.\n>\n> [1] https://lore.kernel.org/git/6b427b4b1e82b1f01640f1f49fe8d1c2fd02111e.1691210737.git.gitgitgadget@gmail.com\n\nI forget the internal version/discussion, but I assume you're thinking\nof 0ea5292e6b (interpret-trailers: add options for actions, 2017-08-01).\nI will reroll and mention it to the commit message.\n\n>> In addition, the corresponding *_DEFAULT values are chosen when the user\n>> provides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\n>> on the command line. These \"--no-*\" flags are used to clear previously\n>> provided flags of the form \"--where\", \"--if-exists\", and \"--if-missing\".\n>> Using these \"--no-*\" flags undoes the specifying of these flags (if\n>> any), so using the word \"UNSPECIFIED\" is more natural here.\n>>\n>> So instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\n>> signals to the reader that the *_UNSPECIFIED value by itself carries no\n>> meaning (it's a zero value and by itself does not \"default\" to anything,\n>> necessitating the need to have some other way of getting to a useful\n>> value).\n>\n> Makse sense. This seems like a good change.\n>\n>> @@ -586,7 +586,10 @@ static void ensure_configured(void)\n>>  \tif (configured)\n>>  \t\treturn;\n>>  \n>> -\t/* Default config must be setup first */\n>> +\t/*\n>> +\t * Default config must be setup first. These defaults are used if there\n>> +\t * are no \"trailer.*\" or \"trailer.<token>.*\" options configured.\n>> +\t */\n>>  \tdefault_conf_info.where = WHERE_END;\n>>  \tdefault_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n>>  \tdefault_conf_info.if_missing = MISSING_ADD;\n>\n> As mentioned earlier, I find it a bit odd that we're calling config\n> 'default' (and also that we're calling CLI args config), but\n> renaming default_conf_info to config_conf_info sounds worse, so let's\n> leave it as-is.\n\nSGTM. Although, we could also just rename it to \"conf_info\" (same name\nas the struct name). Unless such same-variable-name-as-the-struct is\ndiscouraged in the codebase.\n"},{"id":"480588","messageId":"kl6l8rah8fq6.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"owlyjzu2fjz9.fsf@fine.c.googlers.com","subject":"Re: [PATCH 4/5] trailer: teach find_patch_start about --no-divider","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-11T20:51:45Z","receivedAt":"2023-08-11T20:52:02Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n> I don't see why we should preserve the if-statement and associated\n> strlen() call if we can just get rid of it.\n\nHere are some reasons:\n\n- Without compiler optimizations, it is faster.\n- Subjectively, I find the early return easier to understand.\n\nI don't think I need to nitpick over such a tiny issue though, so I'm\nokay either way.\n"},{"id":"480589","messageId":"kl6l5y5l8etg.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"60068","inReplyTo":"owlypm3ufl7j.fsf@fine.c.googlers.com","subject":"Re: [PATCH 3/5] trailer: split process_command_line_args into separate functions","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2023-08-11T21:11:23Z","receivedAt":"2023-08-11T21:11:28Z","isPatch":true,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n>> I find this equally readable as the preimage, which IMO is adequately\n>> scoped and commented.\n>\n> Aside: is \"preimage\" the status quo before applying the patch?\n\nYup.\n"},{"id":"481616","messageId":"4f116d2550f6cf218477560a9e25dbe4c384a2a6.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 1/6] trailer: separate public from internal portion of trailer_iterator","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:12Z","receivedAt":"2023-09-09T06:16:24Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nThe fields here are not meant to be used by downstream callers, so put\nthem behind an anonymous struct named as \"internal\" to warn against\ntheir use. This follows the pattern in 576de3d956 (unpack_trees: start\nsplitting internal fields from public API, 2023-02-27).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 10 +++++-----\n trailer.h |  6 ++++--\n 2 files changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex f408f9b058d..de4bdece847 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1220,14 +1220,14 @@ void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n \tstrbuf_init(&iter->key, 0);\n \tstrbuf_init(&iter->val, 0);\n \topts.no_divider = 1;\n-\ttrailer_info_get(&iter->info, msg, &opts);\n-\titer->cur = 0;\n+\ttrailer_info_get(&iter->internal.info, msg, &opts);\n+\titer->internal.cur = 0;\n }\n \n int trailer_iterator_advance(struct trailer_iterator *iter)\n {\n-\twhile (iter->cur < iter->info.trailer_nr) {\n-\t\tchar *trailer = iter->info.trailers[iter->cur++];\n+\twhile (iter->internal.cur < iter->internal.info.trailer_nr) {\n+\t\tchar *trailer = iter->internal.info.trailers[iter->internal.cur++];\n \t\tint separator_pos = find_separator(trailer, separators);\n \n \t\tif (separator_pos < 1)\n@@ -1245,7 +1245,7 @@ int trailer_iterator_advance(struct trailer_iterator *iter)\n \n void trailer_iterator_release(struct trailer_iterator *iter)\n {\n-\ttrailer_info_release(&iter->info);\n+\ttrailer_info_release(&iter->internal.info);\n \tstrbuf_release(&iter->val);\n \tstrbuf_release(&iter->key);\n }\ndiff --git a/trailer.h b/trailer.h\nindex 795d2fccfd9..ab2cd017567 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -119,8 +119,10 @@ struct trailer_iterator {\n \tstruct strbuf val;\n \n \t/* private */\n-\tstruct trailer_info info;\n-\tsize_t cur;\n+\tstruct {\n+\t\tstruct trailer_info info;\n+\t\tsize_t cur;\n+\t} internal;\n };\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"481617","messageId":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.git.1691211879.gitgitgadget@gmail.com","subject":"[PATCH v2 0/6] Trailer readability cleanups","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:11Z","receivedAt":"2023-09-09T06:16:25Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"These patches were created while digging into the trailer code to better\nunderstand how it works, in preparation for making the trailer.{c,h} files\nas small as possible to make them available as a library for external users.\nThis series was originally created as part of [1], but are sent here\nseparately because the changes here are arguably more subjective in nature.\nI think Patch 1 is the most important in this series. The others can wait,\nif folks are opposed to adding them on their own merits at this point in\ntime.\n\nThese patches do not add or change any features. Instead, their goal is to\nmake the code easier to understand for new contributors (like myself), by\nmaking various cleanups and improvements. Ultimately, my hope is that with\nsuch cleanups, we are better positioned to make larger changes (especially\nthe broader libification effort, as in \"Introduce Git Standard Library\"\n[2]).\n\nPatch 1 was inspired by 576de3d956 (unpack_trees: start splitting internal\nfields from public API, 2023-02-27) [3], and is in preparation for a\nlibification effort in the future around the trailer code. Independent of\nlibification, it still makes sense to discourage callers from peeking into\nthese trailer-internal fields.\n\nPatches 2-3 aim to make some functions do a little less multitasking.\n\nPatch 4 makes the find_patch_start function care about the \"--no-divider\"\noption, because it that option matters for determining the start of the\n\"patch part\" of the input.\n\nPatch 5 is a renaming change to reduce overloaded language in the codebase.\nIt is inspired by 229d6ab6bf (doc: trailer: examples: avoid the word\n\"message\" by itself, 2023-06-15) [4], which did a similar thing for the\ninterpret-trailers documentation.\n\nPatch 6 makes trailer_info use offsets for trailer_start and trailer_end.\n\n\nUpdates in v2\n=============\n\n * Patch 1: Drop the use of a #define. Instead just use an anonymous struct\n   named internal.\n * Patch 2: Don't free info out parameter inside parse_trailers(). Instead\n   free it from the caller, process_trailers(). Update comment in\n   parse_trailers().\n * Patch 3: Reword commit message.\n * Patch 4: Mention be3d654343 (commit: pass --no-divider to\n   interpret-trailers, 2023-06-17) in commit message.\n * Added Patch 6 to make trailer_info use offsets for trailer_start and\n   trailer_end (thanks to Glen Choo for the suggestion).\n\n[1]\nhttps://lore.kernel.org/git/pull.1564.git.1691210737.gitgitgadget@gmail.com/T/#mb044012670663d8eb7a548924bbcc933bef116de\n[2]\nhttps://lore.kernel.org/git/20230627195251.1973421-1-calvinwan@google.com/\n[3]\nhttps://lore.kernel.org/git/pull.1149.git.1677143700.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/6b4cb31b17077181a311ca87e82464a1e2ad67dd.1686797630.git.gitgitgadget@gmail.com/\n\nLinus Arver (6):\n  trailer: separate public from internal portion of trailer_iterator\n  trailer: split process_input_file into separate pieces\n  trailer: split process_command_line_args into separate functions\n  trailer: teach find_patch_start about --no-divider\n  trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n  trailer: use offsets for trailer_start/trailer_end\n\n trailer.c | 126 +++++++++++++++++++++++++++++-------------------------\n trailer.h |  19 ++++----\n 2 files changed, 77 insertions(+), 68 deletions(-)\n\n\nbase-commit: 1b0a5129563ebe720330fdc8f5c6843d27641137\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1563%2Flistx%2Ftrailer-libification-prep-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1563/listx/trailer-libification-prep-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1563\n\nRange-diff vs v1:\n\n 1:  0bce4d4b0d5 ! 1:  4f116d2550f trailer: separate public from internal portion of trailer_iterator\n     @@ Commit message\n          trailer: separate public from internal portion of trailer_iterator\n      \n          The fields here are not meant to be used by downstream callers, so put\n     -    them behind an anonymous struct named as\n     -    \"__private_to_trailer_c__do_not_use\" to warn against their use.\n     +    them behind an anonymous struct named as \"internal\" to warn against\n     +    their use. This follows the pattern in 576de3d956 (unpack_trees: start\n     +    splitting internal fields from public API, 2023-02-27).\n      \n     -    Internally, use a \"#define\" to keep the code tidy.\n     -\n     -    Helped-by: Junio C Hamano <gitster@pobox.com>\n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n       ## trailer.c ##\n     -@@ trailer.c: void format_trailers_from_commit(struct strbuf *out, const char *msg,\n     - \ttrailer_info_release(&info);\n     - }\n     - \n     -+#define private __private_to_trailer_c__do_not_use\n     -+\n     - void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n     - {\n     - \tstruct process_trailer_options opts = PROCESS_TRAILER_OPTIONS_INIT;\n     +@@ trailer.c: void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n       \tstrbuf_init(&iter->key, 0);\n       \tstrbuf_init(&iter->val, 0);\n       \topts.no_divider = 1;\n      -\ttrailer_info_get(&iter->info, msg, &opts);\n      -\titer->cur = 0;\n     -+\ttrailer_info_get(&iter->private.info, msg, &opts);\n     -+\titer->private.cur = 0;\n     ++\ttrailer_info_get(&iter->internal.info, msg, &opts);\n     ++\titer->internal.cur = 0;\n       }\n       \n       int trailer_iterator_advance(struct trailer_iterator *iter)\n       {\n      -\twhile (iter->cur < iter->info.trailer_nr) {\n      -\t\tchar *trailer = iter->info.trailers[iter->cur++];\n     -+\twhile (iter->private.cur < iter->private.info.trailer_nr) {\n     -+\t\tchar *trailer = iter->private.info.trailers[iter->private.cur++];\n     ++\twhile (iter->internal.cur < iter->internal.info.trailer_nr) {\n     ++\t\tchar *trailer = iter->internal.info.trailers[iter->internal.cur++];\n       \t\tint separator_pos = find_separator(trailer, separators);\n       \n       \t\tif (separator_pos < 1)\n     @@ trailer.c: int trailer_iterator_advance(struct trailer_iterator *iter)\n       void trailer_iterator_release(struct trailer_iterator *iter)\n       {\n      -\ttrailer_info_release(&iter->info);\n     -+\ttrailer_info_release(&iter->private.info);\n     ++\ttrailer_info_release(&iter->internal.info);\n       \tstrbuf_release(&iter->val);\n       \tstrbuf_release(&iter->key);\n       }\n     @@ trailer.h: struct trailer_iterator {\n      +\tstruct {\n      +\t\tstruct trailer_info info;\n      +\t\tsize_t cur;\n     -+\t} __private_to_trailer_c__do_not_use;\n     ++\t} internal;\n       };\n       \n       /*\n 2:  d023c297dca ! 2:  c00f4623d0b trailer: split process_input_file into separate pieces\n     @@ trailer.c: static void unfold_value(struct strbuf *val)\n      -\t\t\t\t struct list_head *head,\n      -\t\t\t\t const struct process_trailer_options *opts)\n      +/*\n     -+ * Parse trailers in \"str\" and populate the \"head\" linked list structure.\n     ++ * Parse trailers in \"str\", populating the trailer info and \"head\"\n     ++ * linked list structure.\n      + */\n      +static void parse_trailers(struct trailer_info *info,\n      +\t\t\t     const char *str,\n     @@ trailer.c: static void unfold_value(struct strbuf *val)\n       \t\t\tcontinue;\n       \t\tseparator_pos = find_separator(trailer, separators);\n      @@ trailer.c: static size_t process_input_file(FILE *outfile,\n     + \t\t\t\t\t strbuf_detach(&val, NULL));\n       \t\t}\n       \t}\n     - \n     +-\n      -\ttrailer_info_release(&info);\n      -\n      -\treturn info.trailer_end - str;\n     -+\ttrailer_info_release(info);\n       }\n       \n       static void free_all(struct list_head *head)\n     @@ trailer.c: void process_trailers(const char *file,\n       \n       \tif (!opts->only_input) {\n       \t\tLIST_HEAD(arg_head);\n     +@@ trailer.c: void process_trailers(const char *file,\n     + \tprint_all(outfile, &head, opts);\n     + \n     + \tfree_all(&head);\n     ++\ttrailer_info_release(&info);\n     + \n     + \t/* Print the lines after the trailers as is */\n     + \tif (!opts->only_trailers)\n 3:  c8bb0136621 ! 3:  f78c2345fad trailer: split process_command_line_args into separate functions\n     @@ Commit message\n              (1) parse trailers from the configuration, and\n              (2) parse trailers defined on the command line.\n      \n     -    Separate these concerns into parse_trailers_from_config and\n     -    parse_trailers_from_command_line_args, respectively. Remove (now\n     -    redundant) process_command_line_args.\n     +    Separate (1) outside to a new function, parse_trailers_from_config.\n     +    Rename the remaining logic to parse_trailers_from_command_line_args.\n      \n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n 4:  1fc060041db ! 4:  f5f507c4c6c trailer: teach find_patch_start about --no-divider\n     @@ Commit message\n      \n          Instead, make find_patch_start aware of \"--no-divider\" and make it\n          handle that case as well. This means we no longer need to call strlen at\n     -    all and can just rely on the existing code in find_patch_start.\n     +    all and can just rely on the existing code in find_patch_start. By\n     +    forcing callers to consider this important option, we avoid the kind of\n     +    mistake described in be3d654343 (commit: pass --no-divider to\n     +    interpret-trailers, 2023-06-17).\n      \n          This patch will make unit testing a bit more pleasant in this area in\n          the future when we adopt a unit testing framework, because we would not\n 5:  7c9b63c2616 ! 5:  52958c3557c trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n     @@ Commit message\n          (2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n              because these configurations are used by \"default\" (ahead of the\n              hardcoded defaults in (1)) if no command line arguments are\n     -        provided.\n     +        provided. This concept of defaulting back to the configurations was\n     +        introduced in 0ea5292e6b (interpret-trailers: add options for\n     +        actions, 2017-08-01).\n      \n          In addition, the corresponding *_DEFAULT values are chosen when the user\n          provides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\n -:  ----------- > 6:  0463066ebe0 trailer: use offsets for trailer_start/trailer_end\n\n-- \ngitgitgadget\n"},{"id":"481618","messageId":"c00f4623d0b97cc8ed71ea018e6ecf6e21739b53.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 2/6] trailer: split process_input_file into separate pieces","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:13Z","receivedAt":"2023-09-09T06:16:25Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nCurrently, process_input_file does three things:\n\n    (1) parse the input string for trailers,\n    (2) print text before the trailers, and\n    (3) calculate the position of the input where the trailers end.\n\nRename this function to parse_trailers(), and make it only do\n(1). The caller of this function, process_trailers, becomes responsible\nfor (2) and (3). These items belong inside process_trailers because they\nare both concerned with printing the surrounding text around\ntrailers (which is already one of the immediate concerns of\nprocess_trailers).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 42 ++++++++++++++++++++++--------------------\n 1 file changed, 22 insertions(+), 20 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex de4bdece847..2c56cbc4a2e 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -961,28 +961,24 @@ static void unfold_value(struct strbuf *val)\n \tstrbuf_release(&out);\n }\n \n-static size_t process_input_file(FILE *outfile,\n-\t\t\t\t const char *str,\n-\t\t\t\t struct list_head *head,\n-\t\t\t\t const struct process_trailer_options *opts)\n+/*\n+ * Parse trailers in \"str\", populating the trailer info and \"head\"\n+ * linked list structure.\n+ */\n+static void parse_trailers(struct trailer_info *info,\n+\t\t\t     const char *str,\n+\t\t\t     struct list_head *head,\n+\t\t\t     const struct process_trailer_options *opts)\n {\n-\tstruct trailer_info info;\n \tstruct strbuf tok = STRBUF_INIT;\n \tstruct strbuf val = STRBUF_INIT;\n \tsize_t i;\n \n-\ttrailer_info_get(&info, str, opts);\n-\n-\t/* Print lines before the trailers as is */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(str, 1, info.trailer_start - str, outfile);\n+\ttrailer_info_get(info, str, opts);\n \n-\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n-\t\tfprintf(outfile, \"\\n\");\n-\n-\tfor (i = 0; i < info.trailer_nr; i++) {\n+\tfor (i = 0; i < info->trailer_nr; i++) {\n \t\tint separator_pos;\n-\t\tchar *trailer = info.trailers[i];\n+\t\tchar *trailer = info->trailers[i];\n \t\tif (trailer[0] == comment_line_char)\n \t\t\tcontinue;\n \t\tseparator_pos = find_separator(trailer, separators);\n@@ -1002,10 +998,6 @@ static size_t process_input_file(FILE *outfile,\n \t\t\t\t\t strbuf_detach(&val, NULL));\n \t\t}\n \t}\n-\n-\ttrailer_info_release(&info);\n-\n-\treturn info.trailer_end - str;\n }\n \n static void free_all(struct list_head *head)\n@@ -1054,6 +1046,7 @@ void process_trailers(const char *file,\n {\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n+\tstruct trailer_info info;\n \tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n@@ -1064,8 +1057,16 @@ void process_trailers(const char *file,\n \tif (opts->in_place)\n \t\toutfile = create_in_place_tempfile(file);\n \n+\tparse_trailers(&info, sb.buf, &head, opts);\n+\ttrailer_end = info.trailer_end - sb.buf;\n+\n \t/* Print the lines before the trailers */\n-\ttrailer_end = process_input_file(outfile, sb.buf, &head, opts);\n+\tif (!opts->only_trailers)\n+\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\n+\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n+\t\tfprintf(outfile, \"\\n\");\n+\n \n \tif (!opts->only_input) {\n \t\tLIST_HEAD(arg_head);\n@@ -1076,6 +1077,7 @@ void process_trailers(const char *file,\n \tprint_all(outfile, &head, opts);\n \n \tfree_all(&head);\n+\ttrailer_info_release(&info);\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-- \ngitgitgadget\n\n"},{"id":"481619","messageId":"f78c2345fadb37e10feb3a18aabb536357549790.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 3/6] trailer: split process_command_line_args into separate functions","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:14Z","receivedAt":"2023-09-09T06:16:29Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, process_command_line_args did two things:\n\n    (1) parse trailers from the configuration, and\n    (2) parse trailers defined on the command line.\n\nSeparate (1) outside to a new function, parse_trailers_from_config.\nRename the remaining logic to parse_trailers_from_command_line_args.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 34 +++++++++++++++++++++-------------\n 1 file changed, 21 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 2c56cbc4a2e..b6de5d9cb2d 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -711,30 +711,35 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tlist_add_tail(&new_item->list, arg_head);\n }\n \n-static void process_command_line_args(struct list_head *arg_head,\n-\t\t\t\t      struct list_head *new_trailer_head)\n+static void parse_trailers_from_config(struct list_head *config_head)\n {\n \tstruct arg_item *item;\n-\tstruct strbuf tok = STRBUF_INIT;\n-\tstruct strbuf val = STRBUF_INIT;\n-\tconst struct conf_info *conf;\n \tstruct list_head *pos;\n \n-\t/*\n-\t * In command-line arguments, '=' is accepted (in addition to the\n-\t * separators that are defined).\n-\t */\n-\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n-\n \t/* Add an arg item for each configured trailer with a command */\n \tlist_for_each(pos, &conf_head) {\n \t\titem = list_entry(pos, struct arg_item, list);\n \t\tif (item->conf.command)\n-\t\t\tadd_arg_item(arg_head,\n+\t\t\tadd_arg_item(config_head,\n \t\t\t\t     xstrdup(token_from_item(item, NULL)),\n \t\t\t\t     xstrdup(\"\"),\n \t\t\t\t     &item->conf, NULL);\n \t}\n+}\n+\n+static void parse_trailers_from_command_line_args(struct list_head *arg_head,\n+\t\t\t\t\t\t  struct list_head *new_trailer_head)\n+{\n+\tstruct strbuf tok = STRBUF_INIT;\n+\tstruct strbuf val = STRBUF_INIT;\n+\tconst struct conf_info *conf;\n+\tstruct list_head *pos;\n+\n+\t/*\n+\t * In command-line arguments, '=' is accepted (in addition to the\n+\t * separators that are defined).\n+\t */\n+\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n \n \t/* Add an arg item for each trailer on the command line */\n \tlist_for_each(pos, new_trailer_head) {\n@@ -1069,8 +1074,11 @@ void process_trailers(const char *file,\n \n \n \tif (!opts->only_input) {\n+\t\tLIST_HEAD(config_head);\n \t\tLIST_HEAD(arg_head);\n-\t\tprocess_command_line_args(&arg_head, new_trailer_head);\n+\t\tparse_trailers_from_config(&config_head);\n+\t\tparse_trailers_from_command_line_args(&arg_head, new_trailer_head);\n+\t\tlist_splice(&config_head, &arg_head);\n \t\tprocess_trailers_lists(&head, &arg_head);\n \t}\n \n-- \ngitgitgadget\n\n"},{"id":"481620","messageId":"f5f507c4c6c4514af7dca35e307ca68e72435afb.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 4/6] trailer: teach find_patch_start about --no-divider","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:15Z","receivedAt":"2023-09-09T06:16:32Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nCurrently, find_patch_start only finds the start of the patch part of\nthe input (by looking at the \"---\" divider) for cases where the\n\"--no-divider\" flag has not been provided. If the user provides this\nflag, we do not rely on find_patch_start at all and just call strlen()\ndirectly on the input.\n\nInstead, make find_patch_start aware of \"--no-divider\" and make it\nhandle that case as well. This means we no longer need to call strlen at\nall and can just rely on the existing code in find_patch_start. By\nforcing callers to consider this important option, we avoid the kind of\nmistake described in be3d654343 (commit: pass --no-divider to\ninterpret-trailers, 2023-06-17).\n\nThis patch will make unit testing a bit more pleasant in this area in\nthe future when we adopt a unit testing framework, because we would not\nhave to test multiple functions to check how finding the start of a\npatch part works (we would only need to test find_patch_start).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 10 +++-------\n 1 file changed, 3 insertions(+), 7 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex b6de5d9cb2d..f646e484a23 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -812,14 +812,14 @@ static ssize_t last_line(const char *buf, size_t len)\n  * Return the position of the start of the patch or the length of str if there\n  * is no patch in the message.\n  */\n-static size_t find_patch_start(const char *str)\n+static size_t find_patch_start(const char *str, int no_divider)\n {\n \tconst char *s;\n \n \tfor (s = str; *s; s = next_line(s)) {\n \t\tconst char *v;\n \n-\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v))\n \t\t\treturn s - str;\n \t}\n \n@@ -1109,11 +1109,7 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tensure_configured();\n \n-\tif (opts->no_divider)\n-\t\tpatch_start = strlen(str);\n-\telse\n-\t\tpatch_start = find_patch_start(str);\n-\n+\tpatch_start = find_patch_start(str, opts->no_divider);\n \ttrailer_end = find_trailer_end(str, patch_start);\n \ttrailer_start = find_trailer_start(str, trailer_end);\n \n-- \ngitgitgadget\n\n"},{"id":"481621","messageId":"0463066ebe0889b72b6a1f6c344f2de127458391.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 6/6] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:17Z","receivedAt":"2023-09-09T06:16:34Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously these fields in the trailer_info struct were of type \"const\nchar *\" and pointed to positions in the input string directly (to the\nstart and end positions of the trailer block).\n\nUse offsets to make the intended usage less ambiguous. We only need to\nreference the input string in format_trailer_info(), so update that\nfunction to take a pointer to the input.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 17 ++++++++---------\n trailer.h |  7 +++----\n 2 files changed, 11 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 6ad2fbca942..00326720e81 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1055,7 +1055,6 @@ void process_trailers(const char *file,\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n \tstruct trailer_info info;\n-\tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n \tensure_configured();\n@@ -1066,11 +1065,10 @@ void process_trailers(const char *file,\n \t\toutfile = create_in_place_tempfile(file);\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n-\ttrailer_end = info.trailer_end - sb.buf;\n \n \t/* Print the lines before the trailers */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\t\tfwrite(sb.buf, 1, info.trailer_start, outfile);\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1092,7 +1090,7 @@ void process_trailers(const char *file,\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + trailer_end, 1, sb.len - trailer_end, outfile);\n+\t\tfwrite(sb.buf + info.trailer_end, 1, sb.len - info.trailer_end, outfile);\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1104,7 +1102,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint patch_start, trailer_end, trailer_start;\n+\tsize_t patch_start, trailer_end = 0, trailer_start = 0;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1139,8 +1137,8 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_start);\n-\tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + trailer_end;\n+\tinfo->trailer_start = trailer_start;\n+\tinfo->trailer_end = trailer_end;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n@@ -1155,6 +1153,7 @@ void trailer_info_release(struct trailer_info *info)\n \n static void format_trailer_info(struct strbuf *out,\n \t\t\t\tconst struct trailer_info *info,\n+\t\t\t\tconst char *msg,\n \t\t\t\tconst struct process_trailer_options *opts)\n {\n \tsize_t origlen = out->len;\n@@ -1164,7 +1163,7 @@ static void format_trailer_info(struct strbuf *out,\n \tif (!opts->only_trailers && !opts->unfold && !opts->filter &&\n \t    !opts->separator && !opts->key_only && !opts->value_only &&\n \t    !opts->key_value_separator) {\n-\t\tstrbuf_add(out, info->trailer_start,\n+\t\tstrbuf_add(out, msg + info->trailer_start,\n \t\t\t   info->trailer_end - info->trailer_start);\n \t\treturn;\n \t}\n@@ -1219,7 +1218,7 @@ void format_trailers_from_commit(struct strbuf *out, const char *msg,\n \tstruct trailer_info info;\n \n \ttrailer_info_get(&info, msg, opts);\n-\tformat_trailer_info(out, &info, opts);\n+\tformat_trailer_info(out, &info, msg, opts);\n \ttrailer_info_release(&info);\n }\n \ndiff --git a/trailer.h b/trailer.h\nindex a689d768c79..13fbf0dcd12 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -37,11 +37,10 @@ struct trailer_info {\n \tint blank_line_before_trailer;\n \n \t/*\n-\t * Pointers to the start and end of the trailer block found. If there\n-\t * is no trailer block found, these 2 pointers point to the end of the\n-\t * input string.\n+\t * Offsets to the trailer block start and end positions in the input\n+\t * string. If no trailer block is found, these are set to 0.\n \t */\n-\tconst char *trailer_start, *trailer_end;\n+\tsize_t trailer_start, trailer_end;\n \n \t/*\n \t * Array of trailers found.\n-- \ngitgitgadget\n"},{"id":"481622","messageId":"52958c3557c34992df59e9c10f098f457526702c.1694240177.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-09T06:16:16Z","receivedAt":"2023-09-09T06:16:34Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nDo not use *_DEFAULT as a suffix to the enums, because the word\n\"default\" is overloaded. The following are two examples of the ambiguity\nof the word \"default\":\n\n(1) \"Default\" can mean using the \"default\" values that are hardcoded\n    in trailer.c as\n\n        default_conf_info.where = WHERE_END;\n        default_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n        default_conf_info.if_missing = MISSING_ADD;\n\n    in ensure_configured(). These values are referred to as \"the\n    default\" in the docs for interpret-trailers. These defaults are used\n    if no \"trailer.*\" configurations are defined.\n\n(2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n    because these configurations are used by \"default\" (ahead of the\n    hardcoded defaults in (1)) if no command line arguments are\n    provided. This concept of defaulting back to the configurations was\n    introduced in 0ea5292e6b (interpret-trailers: add options for\n    actions, 2017-08-01).\n\nIn addition, the corresponding *_DEFAULT values are chosen when the user\nprovides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\non the command line. These \"--no-*\" flags are used to clear previously\nprovided flags of the form \"--where\", \"--if-exists\", and \"--if-missing\".\nUsing these \"--no-*\" flags undoes the specifying of these flags (if\nany), so using the word \"UNSPECIFIED\" is more natural here.\n\nSo instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\nsignals to the reader that the *_UNSPECIFIED value by itself carries no\nmeaning (it's a zero value and by itself does not \"default\" to anything,\nnecessitating the need to have some other way of getting to a useful\nvalue).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 17 ++++++++++-------\n trailer.h |  6 +++---\n 2 files changed, 13 insertions(+), 10 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex f646e484a23..6ad2fbca942 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -388,7 +388,7 @@ static void process_trailers_lists(struct list_head *head,\n int trailer_set_where(enum trailer_where *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = WHERE_DEFAULT;\n+\t\t*item = WHERE_UNSPECIFIED;\n \telse if (!strcasecmp(\"after\", value))\n \t\t*item = WHERE_AFTER;\n \telse if (!strcasecmp(\"before\", value))\n@@ -405,7 +405,7 @@ int trailer_set_where(enum trailer_where *item, const char *value)\n int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = EXISTS_DEFAULT;\n+\t\t*item = EXISTS_UNSPECIFIED;\n \telse if (!strcasecmp(\"addIfDifferent\", value))\n \t\t*item = EXISTS_ADD_IF_DIFFERENT;\n \telse if (!strcasecmp(\"addIfDifferentNeighbor\", value))\n@@ -424,7 +424,7 @@ int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n int trailer_set_if_missing(enum trailer_if_missing *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = MISSING_DEFAULT;\n+\t\t*item = MISSING_UNSPECIFIED;\n \telse if (!strcasecmp(\"doNothing\", value))\n \t\t*item = MISSING_DO_NOTHING;\n \telse if (!strcasecmp(\"add\", value))\n@@ -586,7 +586,10 @@ static void ensure_configured(void)\n \tif (configured)\n \t\treturn;\n \n-\t/* Default config must be setup first */\n+\t/*\n+\t * Default config must be setup first. These defaults are used if there\n+\t * are no \"trailer.*\" or \"trailer.<token>.*\" options configured.\n+\t */\n \tdefault_conf_info.where = WHERE_END;\n \tdefault_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n \tdefault_conf_info.if_missing = MISSING_ADD;\n@@ -701,11 +704,11 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tnew_item->value = val;\n \tduplicate_conf(&new_item->conf, conf);\n \tif (new_trailer_item) {\n-\t\tif (new_trailer_item->where != WHERE_DEFAULT)\n+\t\tif (new_trailer_item->where != WHERE_UNSPECIFIED)\n \t\t\tnew_item->conf.where = new_trailer_item->where;\n-\t\tif (new_trailer_item->if_exists != EXISTS_DEFAULT)\n+\t\tif (new_trailer_item->if_exists != EXISTS_UNSPECIFIED)\n \t\t\tnew_item->conf.if_exists = new_trailer_item->if_exists;\n-\t\tif (new_trailer_item->if_missing != MISSING_DEFAULT)\n+\t\tif (new_trailer_item->if_missing != MISSING_UNSPECIFIED)\n \t\t\tnew_item->conf.if_missing = new_trailer_item->if_missing;\n \t}\n \tlist_add_tail(&new_item->list, arg_head);\ndiff --git a/trailer.h b/trailer.h\nindex ab2cd017567..a689d768c79 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -5,14 +5,14 @@\n #include \"strbuf.h\"\n \n enum trailer_where {\n-\tWHERE_DEFAULT,\n+\tWHERE_UNSPECIFIED,\n \tWHERE_END,\n \tWHERE_AFTER,\n \tWHERE_BEFORE,\n \tWHERE_START\n };\n enum trailer_if_exists {\n-\tEXISTS_DEFAULT,\n+\tEXISTS_UNSPECIFIED,\n \tEXISTS_ADD_IF_DIFFERENT_NEIGHBOR,\n \tEXISTS_ADD_IF_DIFFERENT,\n \tEXISTS_ADD,\n@@ -20,7 +20,7 @@ enum trailer_if_exists {\n \tEXISTS_DO_NOTHING\n };\n enum trailer_if_missing {\n-\tMISSING_DEFAULT,\n+\tMISSING_UNSPECIFIED,\n \tMISSING_ADD,\n \tMISSING_DO_NOTHING\n };\n-- \ngitgitgadget\n\n"},{"id":"481686","messageId":"xmqq7cowws8m.fsf@gitster.g","threadId":"60068","inReplyTo":"c00f4623d0b97cc8ed71ea018e6ecf6e21739b53.1694240177.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/6] trailer: split process_input_file into separate pieces","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T17:10:33Z","receivedAt":"2023-09-11T21:38:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> Currently, process_input_file does three things:\n>\n>     (1) parse the input string for trailers,\n>     (2) print text before the trailers, and\n>     (3) calculate the position of the input where the trailers end.\n>\n> Rename this function to parse_trailers(), and make it only do\n> (1). The caller of this function, process_trailers, becomes responsible\n> for (2) and (3). These items belong inside process_trailers because they\n> are both concerned with printing the surrounding text around\n> trailers (which is already one of the immediate concerns of\n> process_trailers).\n\nNicely explained and the resulting code reads well.\n\nThanks.\n"},{"id":"481701","messageId":"xmqqv8cgvd02.fsf@gitster.g","threadId":"60068","inReplyTo":"f5f507c4c6c4514af7dca35e307ca68e72435afb.1694240177.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 4/6] trailer: teach find_patch_start about --no-divider","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T17:25:01Z","receivedAt":"2023-09-11T21:39:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> Currently, find_patch_start only finds the start of the patch part of\n> the input (by looking at the \"---\" divider) for cases where the\n> \"--no-divider\" flag has not been provided. If the user provides this\n> flag, we do not rely on find_patch_start at all and just call strlen()\n> directly on the input.\n>\n> Instead, make find_patch_start aware of \"--no-divider\" and make it\n> handle that case as well. This means we no longer need to call strlen at\n> all and can just rely on the existing code in find_patch_start. By\n> forcing callers to consider this important option, we avoid the kind of\n> mistake described in be3d654343 (commit: pass --no-divider to\n> interpret-trailers, 2023-06-17).\n\nOK.  The code pays attention to \"---\" so making it stop doing so\nwhen the \"--no-*\" option is given will make the function responsible\nfor finding the beginning of the patch.\n\nI wonder if we should rename this function while we are at it,\nthough.  When \"--no-divider\" is given, the expected use case is\n*not* to have a patch at all, and it is dubious that a function\nwhose name is find_patch_start() can possibly do anything useful.\n\nThe real purpose of this function is to find the end of the log\nmessage, isn't it?  And the caller uses the end of the log message\nit found and gives it to find_trailer_start() and find_trailer_end()\nas the upper boundary of the search for the trailer block.\n"},{"id":"481706","messageId":"xmqqr0n4v8ul.fsf@gitster.g","threadId":"60068","inReplyTo":"52958c3557c34992df59e9c10f098f457526702c.1694240177.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T18:54:42Z","receivedAt":"2023-09-11T21:39:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> Do not use *_DEFAULT as a suffix to the enums, because the word\n> \"default\" is overloaded. The following are two examples of the ambiguity\n> of the word \"default\":\n\nIn this case these are left unspecified to use the default; while it\nis not wrong per-se to say *_DEFAULT, using *_UNSPECIFIED makes it\nmore obvious.\n\n> So instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\n> signals to the reader that the *_UNSPECIFIED value by itself carries no\n> meaning (it's a zero value and by itself does not \"default\" to anything,\n> necessitating the need to have some other way of getting to a useful\n> value).\n\nIt gets tempting to initialize a variable to the default and arrange\nthe rest of the system so that the variable set to the default\ntriggers the default activity.  Such an obvious solution however\ncannot be used when (1) being left unspecified to use the default\nvalue and (2) explicitly set by the user to a value that happens to\nbe the same as the default have to behave differently.  I am not\nsure if that applies to the trailers system, though.\n\nThanks.\n\n\nPS.  Glen's old e-mail address is no longer valid and there is no\nforwarding done by @google.com mailservers, it seems.  Can you tell\nGGG to drop the address (optionally replace it with his new address)?\n"},{"id":"481709","messageId":"xmqqbke8ws9g.fsf@gitster.g","threadId":"60068","inReplyTo":"4f116d2550f6cf218477560a9e25dbe4c384a2a6.1694240177.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 1/6] trailer: separate public from internal portion of trailer_iterator","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T17:10:03Z","receivedAt":"2023-09-11T23:02:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> The fields here are not meant to be used by downstream callers, so put\n> them behind an anonymous struct named as \"internal\" to warn against\n> their use. This follows the pattern in 576de3d956 (unpack_trees: start\n> splitting internal fields from public API, 2023-02-27).\n\nOK.  The patch shows that there exist no code external to this file\nthat touch these members that are marked \"private\", so it has some\nauditing value from that point of view, which is nice.\n\nBut that is only about today's code and does not protect us from\nfuture breakage.  In other words, \"git grep internal\\\\.\" would not\nbe an effective way to find misuses of these members from the\nsidelines.  But that is OK, as \"git grep -E '([.]|->)info'\" would\nnot be an effective way in today's code, either, and the patch is\nnot making things worse.\n\nQueued.  Thanks.\n\n> Signed-off-by: Linus Arver <linusa@google.com>\n> ---\n>  trailer.c | 10 +++++-----\n>  trailer.h |  6 ++++--\n>  2 files changed, 9 insertions(+), 7 deletions(-)\n>\n> diff --git a/trailer.c b/trailer.c\n> index f408f9b058d..de4bdece847 100644\n> --- a/trailer.c\n> +++ b/trailer.c\n> @@ -1220,14 +1220,14 @@ void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n>  \tstrbuf_init(&iter->key, 0);\n>  \tstrbuf_init(&iter->val, 0);\n>  \topts.no_divider = 1;\n> -\ttrailer_info_get(&iter->info, msg, &opts);\n> -\titer->cur = 0;\n> +\ttrailer_info_get(&iter->internal.info, msg, &opts);\n> +\titer->internal.cur = 0;\n>  }\n>  \n>  int trailer_iterator_advance(struct trailer_iterator *iter)\n>  {\n> -\twhile (iter->cur < iter->info.trailer_nr) {\n> -\t\tchar *trailer = iter->info.trailers[iter->cur++];\n> +\twhile (iter->internal.cur < iter->internal.info.trailer_nr) {\n> +\t\tchar *trailer = iter->internal.info.trailers[iter->internal.cur++];\n>  \t\tint separator_pos = find_separator(trailer, separators);\n>  \n>  \t\tif (separator_pos < 1)\n> @@ -1245,7 +1245,7 @@ int trailer_iterator_advance(struct trailer_iterator *iter)\n>  \n>  void trailer_iterator_release(struct trailer_iterator *iter)\n>  {\n> -\ttrailer_info_release(&iter->info);\n> +\ttrailer_info_release(&iter->internal.info);\n>  \tstrbuf_release(&iter->val);\n>  \tstrbuf_release(&iter->key);\n>  }\n> diff --git a/trailer.h b/trailer.h\n> index 795d2fccfd9..ab2cd017567 100644\n> --- a/trailer.h\n> +++ b/trailer.h\n> @@ -119,8 +119,10 @@ struct trailer_iterator {\n>  \tstruct strbuf val;\n>  \n>  \t/* private */\n> -\tstruct trailer_info info;\n> -\tsize_t cur;\n> +\tstruct {\n> +\t\tstruct trailer_info info;\n> +\t\tsize_t cur;\n> +\t} internal;\n>  };\n>  \n>  /*\n"},{"id":"481714","messageId":"xmqqmsxsv8ik.fsf@gitster.g","threadId":"60068","inReplyTo":"0463066ebe0889b72b6a1f6c344f2de127458391.1694240177.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 6/6] trailer: use offsets for trailer_start/trailer_end","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T19:01:55Z","receivedAt":"2023-09-11T23:02:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> Previously these fields in the trailer_info struct were of type \"const\n> char *\" and pointed to positions in the input string directly (to the\n> start and end positions of the trailer block).\n>\n> Use offsets to make the intended usage less ambiguous. We only need to\n> reference the input string in format_trailer_info(), so update that\n> function to take a pointer to the input.\n\nHmm, I am not sure if this is an improvement.  If the underlying\nbuffer can be reallocated (to grow), the approach to use the offsets\ncertainly is easier to deal with, as they will stay valid even after\nsuch a reallocation.  But you lose the obvious sentinel value NULL\nthat can mean something special, and have to make the readers aware\nof the local convention you happened to have picked with a comment\nlike ...\n\n> Signed-off-by: Linus Arver <linusa@google.com>\n> ---\n>  trailer.c | 17 ++++++++---------\n>  trailer.h |  7 +++----\n>  2 files changed, 11 insertions(+), 13 deletions(-)\n> ...\n>  \t/*\n> -\t * Pointers to the start and end of the trailer block found. If there\n> -\t * is no trailer block found, these 2 pointers point to the end of the\n> -\t * input string.\n> +\t * Offsets to the trailer block start and end positions in the input\n> +\t * string. If no trailer block is found, these are set to 0.\n>  \t */\n\n... this, simply because there is no obvious sentinel value for an\nunsigned integral type; even if you count MAX_ULONG and its friends,\nthey are not as obvious as NULL for pointer types.\n\nSo, I dunno.\n"},{"id":"481823","messageId":"owly5y4dlfcg.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqmsxsv8ik.fsf@gitster.g","subject":"Re: [PATCH v2 6/6] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-14T01:21:19Z","receivedAt":"2023-09-14T01:21:22Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Linus Arver <linusa@google.com>\n>>\n>> Previously these fields in the trailer_info struct were of type \"const\n>> char *\" and pointed to positions in the input string directly (to the\n>> start and end positions of the trailer block).\n>>\n>> Use offsets to make the intended usage less ambiguous. We only need to\n>> reference the input string in format_trailer_info(), so update that\n>> function to take a pointer to the input.\n>\n> Hmm, I am not sure if this is an improvement.  If the underlying\n> buffer can be reallocated (to grow), the approach to use the offsets\n> certainly is easier to deal with, as they will stay valid even after\n> such a reallocation.  But you lose the obvious sentinel value NULL\n> that can mean something special\n\nTrue.\n\n> and have to make the readers aware\n> of the local convention you happened to have picked with a comment\n> like ...\n>\n>> Signed-off-by: Linus Arver <linusa@google.com>\n>> ---\n>>  trailer.c | 17 ++++++++---------\n>>  trailer.h |  7 +++----\n>>  2 files changed, 11 insertions(+), 13 deletions(-)\n>> ...\n>>  \t/*\n>> -\t * Pointers to the start and end of the trailer block found. If there\n>> -\t * is no trailer block found, these 2 pointers point to the end of the\n>> -\t * input string.\n>> +\t * Offsets to the trailer block start and end positions in the input\n>> +\t * string. If no trailer block is found, these are set to 0.\n>>  \t */\n>\n> ... this, simply because there is no obvious sentinel value for an\n> unsigned integral type; even if you count MAX_ULONG and its friends,\n> they are not as obvious as NULL for pointer types.\n\nI agree that there is no trustworthy sentinel value for an unsigned\nintegral type.\n\nOn the other hand, we never used NULL as a sentinel value before even\nwhen they were const char pointers --- the current comment for these\nfields which say ...\n\n    If there is no trailer block found, these 2 pointers point to the end of the\n    input string.\n\n... sounds somewhat arbitrary to me (and I don't think we care about\nthis property in trailer.c, and AFAICS it's also not something that\nconsumers should be aware of). Consumers of the trailer_info struct\ncould also just see if \"info->trailer_nr > 0\" to check whether any\ntrailers were found, although if we're merging Patch 1 [1] of this\nseries the consumers will not have easy access any more to any of the\ntrailer_info fields, and they should instead be using a public-facing\nfunction that does the \"were trailers found\" check.\n\n> So, I dunno.\n\nIf the \"we don't use NULL sentinel values currently anyway\" argument is\nconvincing enough, I'm happy to add this to the commit message on a\nreroll. But I'm also OK with dropping this patch. Thoughts?\n\n[1] https://lore.kernel.org/git/pull.1563.git.1691211879.gitgitgadget@gmail.com/T/#m8f1b5f1eb346331658c8c7b3e057a4ee31223664\n"},{"id":"481824","messageId":"owly34zhlcoa.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqv8cgvd02.fsf@gitster.g","subject":"Re: [PATCH v2 4/6] trailer: teach find_patch_start about --no-divider","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-14T02:19:01Z","receivedAt":"2023-09-14T02:19:05Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Linus Arver <linusa@google.com>\n>>\n>> Currently, find_patch_start only finds the start of the patch part of\n>> the input (by looking at the \"---\" divider) for cases where the\n>> \"--no-divider\" flag has not been provided. If the user provides this\n>> flag, we do not rely on find_patch_start at all and just call strlen()\n>> directly on the input.\n>>\n>> Instead, make find_patch_start aware of \"--no-divider\" and make it\n>> handle that case as well. This means we no longer need to call strlen at\n>> all and can just rely on the existing code in find_patch_start. By\n>> forcing callers to consider this important option, we avoid the kind of\n>> mistake described in be3d654343 (commit: pass --no-divider to\n>> interpret-trailers, 2023-06-17).\n>\n> OK.  The code pays attention to \"---\" so making it stop doing so\n> when the \"--no-*\" option is given will make the function responsible\n> for finding the beginning of the patch.\n>\n> I wonder if we should rename this function while we are at it,\n> though.  When \"--no-divider\" is given, the expected use case is\n> *not* to have a patch at all, and it is dubious that a function\n> whose name is find_patch_start() can possibly do anything useful.\n\nIOW, saying as the caller of this function, \"find the patch start in\nthis input I'm giving you, but also FYI the input has no patch in it\"\nsounds wrong. Agreed.\n\n> The real purpose of this function is to find the end of the log\n> message, isn't it?\n\nIndeed.\n\n> And the caller uses the end of the log message\n> it found and gives it to find_trailer_start() and find_trailer_end()\n> as the upper boundary of the search for the trailer block.\n\nRight! So a better name might be something like\n\"find_trailer_search_boundary\" with a comment like\n\n    Find the point at which we should stop searching for trailers (to\n    parse them). This is either the end of the input string (obviously),\n    or the point when we see \"---\" indicating the start of the \"patch\n    part\".\n\nWill update.\n"},{"id":"481825","messageId":"owlyzg1pjx2f.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqr0n4v8ul.fsf@gitster.g","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-14T02:41:28Z","receivedAt":"2023-09-14T02:41:32Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Linus Arver <linusa@google.com>\n>>\n>> [...]\n>> So instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\n>> signals to the reader that the *_UNSPECIFIED value by itself carries no\n>> meaning (it's a zero value and by itself does not \"default\" to anything,\n>> necessitating the need to have some other way of getting to a useful\n>> value).\n>\n> It gets tempting to initialize a variable to the default and arrange\n> the rest of the system so that the variable set to the default\n> triggers the default activity.  Such an obvious solution however\n> cannot be used when (1) being left unspecified to use the default\n> value and (2) explicitly set by the user to a value that happens to\n> be the same as the default have to behave differently.  I am not\n> sure if that applies to the trailers system, though.\n>\n> Thanks.\n\nI get the feeling that you wrote the \"Such an obvious ... differently\"\nsentence after writing the last sentence in that paragraph, because when\nyou say\n\n    I am not\n    sure if that applies to the trailers system, though.\n\nI read the \"that\" (emphasis added) in there as referring to the solution\ndescribed in the first sentence, and not the conditions (1) and (2) you\nenumerated. IOW, you are OK with this patch.\n\nAm I parsing your paragraph correctly?\n\n> PS.  Glen's old e-mail address is no longer valid and there is no\n> forwarding done by @google.com mailservers, it seems.  Can you tell\n> GGG to drop the address (optionally replace it with his new address)?\n\nThanks for catching this. I've updated the GGG PR [1] to use Glen's new\naddress.\n\n[1] https://github.com/gitgitgadget/git/pull/1563#issue-1837440424\n"},{"id":"481826","messageId":"xmqq5y4dla6p.fsf@gitster.g","threadId":"60068","inReplyTo":"owly34zhlcoa.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 4/6] trailer: teach find_patch_start about --no-divider","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-14T03:12:46Z","receivedAt":"2023-09-14T03:12:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n>> The real purpose of this function is to find the end of the log\n>> message, isn't it?\n>\n> Indeed.\n> ...\n> Right! So a better name might be something like\n> \"find_trailer_search_boundary\" with a comment like\n\nOr \"find_end_of_log_message()\", which we agreed to be the real\npurpose of this function ;-)\n"},{"id":"481827","messageId":"xmqq1qf1la0q.fsf@gitster.g","threadId":"60068","inReplyTo":"owlyzg1pjx2f.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-14T03:16:21Z","receivedAt":"2023-09-14T03:16:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n>> It gets tempting to initialize a variable to the default and arrange\n>> the rest of the system so that the variable set to the default\n>> triggers the default activity.  Such an obvious solution however\n>> cannot be used when (1) being left unspecified to use the default\n>> value and (2) explicitly set by the user to a value that happens to\n>> be the same as the default have to behave differently.  I am not\n>> sure if that applies to the trailers system, though.\n>>\n>> Thanks.\n>\n> I get the feeling that you wrote the \"Such an obvious ... differently\"\n> sentence after writing the last sentence in that paragraph, because when\n> you say\n>\n>     I am not\n>     sure if that applies to the trailers system, though.\n>\n> I read the \"that\" (emphasis added) in there as referring to the solution\n> described in the first sentence, and not the conditions (1) and (2) you\n> enumerated. IOW, you are OK with this patch.\n\n\"that\" refers to \"the reason not to use such an obvious solution\".\nI do not know if trailer subsystem wants to treat \"left unspecified\"\nand \"set to the value that happens to be the same as the default\" in\na different way.  If it does want to do so, then I do not see a\nstrong reason not to use the \"obvious solution\".\n\nThanks.\n"},{"id":"481828","messageId":"owlywmwtjvcu.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"owly5y4dlfcg.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 6/6] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-14T03:18:25Z","receivedAt":"2023-09-14T03:18:58Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>>> From: Linus Arver <linusa@google.com>\n>>>\n>>> Previously these fields in the trailer_info struct were of type \"const\n>>> char *\" and pointed to positions in the input string directly (to the\n>>> start and end positions of the trailer block).\n>>>\n>>> Use offsets to make the intended usage less ambiguous. We only need to\n>>> reference the input string in format_trailer_info(), so update that\n>>> function to take a pointer to the input.\n>>\n>> Hmm, I am not sure if this is an improvement.  If the underlying\n>> buffer can be reallocated (to grow), the approach to use the offsets\n>> certainly is easier to deal with, as they will stay valid even after\n>> such a reallocation.  But you lose the obvious sentinel value NULL\n>> that can mean something special\n>\n> True.\n>\n>> and have to make the readers aware\n>> of the local convention you happened to have picked with a comment\n>> like ...\n>>\n>>> Signed-off-by: Linus Arver <linusa@google.com>\n>>> ---\n>>>  trailer.c | 17 ++++++++---------\n>>>  trailer.h |  7 +++----\n>>>  2 files changed, 11 insertions(+), 13 deletions(-)\n>>> ...\n>>>  \t/*\n>>> -\t * Pointers to the start and end of the trailer block found. If there\n>>> -\t * is no trailer block found, these 2 pointers point to the end of the\n>>> -\t * input string.\n>>> +\t * Offsets to the trailer block start and end positions in the input\n>>> +\t * string. If no trailer block is found, these are set to 0.\n>>>  \t */\n\nI've just realized that the new comment \"If no trailer block is found,\nthese are set to 0\" is perhaps not always true in the current version\nof this patch. This is because this patch did a mechanical type\nconversion of the old fields from \"const char *\" to \"size_t\", and we\nstill run\n\n    trailer_end = find_trailer_end(str, trailer_search_boundary);\n    trailer_start = find_trailer_start(str, trailer_end);\n\ninside trailer_info_get() (not visible in the patch context lines). I\nwill update this patch to make the comment \"If no trailer block is found,\nthese are set to 0\" true.\n"},{"id":"481830","messageId":"owlyttrxjp6r.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqq5y4dla6p.fsf@gitster.g","subject":"Re: [PATCH v2 4/6] trailer: teach find_patch_start about --no-divider","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-14T05:31:40Z","receivedAt":"2023-09-14T05:31:44Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>>> The real purpose of this function is to find the end of the log\n>>> message, isn't it?\n>>\n>> Indeed.\n>> ...\n>> Right! So a better name might be something like\n>> \"find_trailer_search_boundary\" with a comment like\n>\n> Or \"find_end_of_log_message()\", which we agreed to be the real\n> purpose of this function ;-)\n\nI did this locally, but in doing so realized that we have in trailer.c\n\n    /* Return the position of the end of the trailers. */\n    static size_t find_trailer_end(const char *buf, size_t len)\n    {\n        return len - ignore_non_trailer(buf, len);\n    }\n\nand the ignore_non_trailer() comes from commit.c, which says\n\n    /*\n    * Inspect the given string and determine the true \"end\" of the log message, in\n    * order to find where to put a new Signed-off-by trailer.  Ignored are\n    * trailing comment lines and blank lines.  To support \"git commit -s\n    * --amend\" on an existing commit, we also ignore \"Conflicts:\".  To\n    * support \"git commit -v\", we truncate at cut lines.\n    *\n    * Returns the number of bytes from the tail to ignore, to be fed as\n    * the second parameter to append_signoff().\n    */\n    size_t ignore_non_trailer(const char *buf, size_t len)\n    {\n\n...and I am not so sure the new \"find_end_of_log_message\" name for\nfind_patch_start() is desirable because of the overlap in meaning with\nthe comment for ignore_non_trailer(). To recap, we have in\ntrailer_info_get() in trailer.c which (without this patch) has\n\n    if (opts->no_divider)\n        patch_start = strlen(str);\n    else\n        patch_start = find_patch_start(str);\n\n    trailer_end = find_trailer_end(str, patch_start);\n    trailer_start = find_trailer_start(str, trailer_end);\n                \nto find the trailer_end and trailer_start positions in the input. These\npositions are the boundaries for parsing for actual trailers (if any).\nMore precisely, the \"patch_start\" variable helps us _skip_ the \"patch\npart\" of the input (if any, denoted by \"---\"). The call to\nfind_trailer_end() helps us (again) _skip_ any parts that are not part\nof the actual log message (per the comment in ignore_non_trailer()). So\nas far as trailer_info_get() is concerned, we are just trying to skip\nover areas where we shouldn't search/parse for trailers.\n\nThe above analysis leads me to some new ideas:\n\n(1) For \"find_end_of_log_message()\", I think this name should really\n    belong to ignore_non_trailer() in commit.c instead of\n    find_patch_start() in trailer.c. The existing comment for\n    ignore_non_trailer() already says that its job is to determine the\n    true end of the log message (and does a lot of the necessary work to\n    do this job).\n\n(2) find_patch_start() should be named something like\n    \"shrink_trailer_search_space\" (although this meaning belongs equally\n    well to find_trailer_end()), because the point is to reduce the\n    search space for parsing trailers.\n\n(3) \"find_trailer_end\" is not a great function name because on first\n    glance it implies that it found the end of trailers, but this is\n    only true for the \"happy path\" of actually finding trailers.\n\nI need to consider all of the above ideas and reroll this patch. Because\nthe ideas here closely relate to the \"trailer_end\" and \"trailer_start\"\nvariables, I will probably reorder the series so that this patch and the\nlast patch are closer together.\n"},{"id":"482161","messageId":"owly8r8yt6cr.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqq1qf1la0q.fsf@gitster.g","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-22T18:23:16Z","receivedAt":"2023-09-22T18:23:24Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>>> It gets tempting to initialize a variable to the default and arrange\n>>> the rest of the system so that the variable set to the default\n>>> triggers the default activity.  Such an obvious solution however\n>>> cannot be used when (1) being left unspecified to use the default\n>>> value and (2) explicitly set by the user to a value that happens to\n>>> be the same as the default have to behave differently.  I am not\n>>> sure if that applies to the trailers system, though.\n>>>\n>>> Thanks.\n>>\n>> I get the feeling that you wrote the \"Such an obvious ... differently\"\n>> sentence after writing the last sentence in that paragraph, because when\n>> you say\n>>\n>>     I am not\n>>     sure if that applies to the trailers system, though.\n>>\n>> I read the \"that\" (emphasis added) in there as referring to the solution\n>> described in the first sentence, and not the conditions (1) and (2) you\n>> enumerated. IOW, you are OK with this patch.\n>\n> \"that\" refers to \"the reason not to use such an obvious solution\".\n> I do not know if trailer subsystem wants to treat \"left unspecified\"\n> and \"set to the value that happens to be the same as the default\" in\n> a different way.  If it does want to do so, then I do not see a\n> strong reason not to use the \"obvious solution\".\n\nCurrently we set the defaults (these take effect absent any\nconfiguration or CLI options) in trailer.c like this:\n\n    static void ensure_configured(void)\n    {\n            if (configured)\n                    return;\n\n            /* Default config must be setup first */\n            default_conf_info.where = WHERE_END;\n            default_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n            default_conf_info.if_missing = MISSING_ADD;\n            git_config(git_trailer_default_config, NULL);\n            git_config(git_trailer_config, NULL);\n            configured = 1;\n    }\n\nSo technically we already sort of do the \"obvious solution\". And then\nthese get overriden by configuration (if any), and finally any CLI\noptions that are passed in (e.g., \"--where after\"). The reason why I\nprefer the *_UNSPECIFIED style in this patch for these enums though is\nbecause of this (and other similar functions) in trailer.c:\n\n    int trailer_set_where(enum trailer_where *item, const char *value)\n    {\n            if (!value)\n                    *item = WHERE_DEFAULT;\n            else if (!strcasecmp(\"after\", value))\n                    *item = WHERE_AFTER;\n            else if (!strcasecmp(\"before\", value))\n                    *item = WHERE_BEFORE;\n            else if (!strcasecmp(\"end\", value))\n                    *item = WHERE_END;\n            else if (!strcasecmp(\"start\", value))\n                    *item = WHERE_START;\n            else\n                    return -1;\n            return 0;\n    }\n\nand this function is used as a callback to the \"--where\" flag, such that\nthe WHERE_DEFAULT gets chosen if \"--no-where\" is where. I prefer the\nWHERE_UNSPECIFIED as in this patch because the WHERE_DEFAULT is\nambiguous on its own (i.e., WHERE_DEFAULT could mean that we either use\nthe default value WHERE_END in default_conf_info, or it could mean that\nwe fall back to the configuration variables (where it could be something\nelse)).\n"},{"id":"482169","messageId":"xmqqil820z25.fsf@gitster.g","threadId":"60068","inReplyTo":"owly8r8yt6cr.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-22T19:48:18Z","receivedAt":"2023-09-22T19:48:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n> ... I prefer the\n> WHERE_UNSPECIFIED as in this patch because the WHERE_DEFAULT is\n> ambiguous on its own (i.e., WHERE_DEFAULT could mean that we either use\n> the default value WHERE_END in default_conf_info, or it could mean that\n> we fall back to the configuration variables (where it could be something\n> else)).\n\nYup.  \"Turning something that is left UNSPECIFIED after command line\noptions and configuration files are processed into the hardcoded\nDEFAULT\" is one mental model that is easy to explain.\n\nI however am not sure if it is easier than \"Setting something to\nhardcoded DEFAULT before command line options and configuration\nfiles are allowed to tweak it, and if nobody touches it, then it\ngets the hardcoded DEFAULT value in the end\", which is another valid\nmental model, though.  If both can be used, I'd personally prefer\nthe latter, and reserve the \"UNSPECIFIED to DEFAULT\" pattern to\nsignal that we are dealing with a case where the simpler pattern\nwithout UNSPECIFIED cannot solve.\n\nThe simpler pattern would not work, when the default is defined\ndepending on a larger context.  Imagine we have two Boolean\nvariables, A and B, where A defaults to false, and B defaults to\nsome value derived from the value of A (say, opposite of A).\n\nIn the most natural implementation, you'd initialize A to false and\nB to unspecified, let command line options and configuration\nvariables to set them to true or false, and after all that, you do\nnot have to tweak A's value (it will be left to false that is the\ndefault unless the user or the configuration gave an explicit\nvalue), but you need to check if B is left unspecified and tweak it\nto true or false using the final value of A.\n\nFor a variable with such a need like B, we cannot avoid having\n\"unspecified\".  If you initialize it to false (or true), after the\ncommand line and the configuration files are read and you find B is\nset to false (or true), you cannot tell if the user or the\nconfiguration explicitly set B to false (or true), in which case you\ndo not want to futz with its value based on what is in A, or it is\nfalse (or true) only because nobody touched it, in which case you\nneed to compute its value based on what is in A.\n\nAnd that is why I asked if we need to special case \"the user did not\ntouch and the variable is left untouched\" in the trailer subsystem.\n\nThanks.\n"},{"id":"482171","messageId":"4f116d2550f6cf218477560a9e25dbe4c384a2a6.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 1/9] trailer: separate public from internal portion of trailer_iterator","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:37Z","receivedAt":"2023-09-22T19:50:52Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nThe fields here are not meant to be used by downstream callers, so put\nthem behind an anonymous struct named as \"internal\" to warn against\ntheir use. This follows the pattern in 576de3d956 (unpack_trees: start\nsplitting internal fields from public API, 2023-02-27).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 10 +++++-----\n trailer.h |  6 ++++--\n 2 files changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex f408f9b058d..de4bdece847 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1220,14 +1220,14 @@ void trailer_iterator_init(struct trailer_iterator *iter, const char *msg)\n \tstrbuf_init(&iter->key, 0);\n \tstrbuf_init(&iter->val, 0);\n \topts.no_divider = 1;\n-\ttrailer_info_get(&iter->info, msg, &opts);\n-\titer->cur = 0;\n+\ttrailer_info_get(&iter->internal.info, msg, &opts);\n+\titer->internal.cur = 0;\n }\n \n int trailer_iterator_advance(struct trailer_iterator *iter)\n {\n-\twhile (iter->cur < iter->info.trailer_nr) {\n-\t\tchar *trailer = iter->info.trailers[iter->cur++];\n+\twhile (iter->internal.cur < iter->internal.info.trailer_nr) {\n+\t\tchar *trailer = iter->internal.info.trailers[iter->internal.cur++];\n \t\tint separator_pos = find_separator(trailer, separators);\n \n \t\tif (separator_pos < 1)\n@@ -1245,7 +1245,7 @@ int trailer_iterator_advance(struct trailer_iterator *iter)\n \n void trailer_iterator_release(struct trailer_iterator *iter)\n {\n-\ttrailer_info_release(&iter->info);\n+\ttrailer_info_release(&iter->internal.info);\n \tstrbuf_release(&iter->val);\n \tstrbuf_release(&iter->key);\n }\ndiff --git a/trailer.h b/trailer.h\nindex 795d2fccfd9..ab2cd017567 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -119,8 +119,10 @@ struct trailer_iterator {\n \tstruct strbuf val;\n \n \t/* private */\n-\tstruct trailer_info info;\n-\tsize_t cur;\n+\tstruct {\n+\t\tstruct trailer_info info;\n+\t\tsize_t cur;\n+\t} internal;\n };\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"482172","messageId":"c00f4623d0b97cc8ed71ea018e6ecf6e21739b53.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 2/9] trailer: split process_input_file into separate pieces","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:38Z","receivedAt":"2023-09-22T19:50:54Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nCurrently, process_input_file does three things:\n\n    (1) parse the input string for trailers,\n    (2) print text before the trailers, and\n    (3) calculate the position of the input where the trailers end.\n\nRename this function to parse_trailers(), and make it only do\n(1). The caller of this function, process_trailers, becomes responsible\nfor (2) and (3). These items belong inside process_trailers because they\nare both concerned with printing the surrounding text around\ntrailers (which is already one of the immediate concerns of\nprocess_trailers).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 42 ++++++++++++++++++++++--------------------\n 1 file changed, 22 insertions(+), 20 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex de4bdece847..2c56cbc4a2e 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -961,28 +961,24 @@ static void unfold_value(struct strbuf *val)\n \tstrbuf_release(&out);\n }\n \n-static size_t process_input_file(FILE *outfile,\n-\t\t\t\t const char *str,\n-\t\t\t\t struct list_head *head,\n-\t\t\t\t const struct process_trailer_options *opts)\n+/*\n+ * Parse trailers in \"str\", populating the trailer info and \"head\"\n+ * linked list structure.\n+ */\n+static void parse_trailers(struct trailer_info *info,\n+\t\t\t     const char *str,\n+\t\t\t     struct list_head *head,\n+\t\t\t     const struct process_trailer_options *opts)\n {\n-\tstruct trailer_info info;\n \tstruct strbuf tok = STRBUF_INIT;\n \tstruct strbuf val = STRBUF_INIT;\n \tsize_t i;\n \n-\ttrailer_info_get(&info, str, opts);\n-\n-\t/* Print lines before the trailers as is */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(str, 1, info.trailer_start - str, outfile);\n+\ttrailer_info_get(info, str, opts);\n \n-\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n-\t\tfprintf(outfile, \"\\n\");\n-\n-\tfor (i = 0; i < info.trailer_nr; i++) {\n+\tfor (i = 0; i < info->trailer_nr; i++) {\n \t\tint separator_pos;\n-\t\tchar *trailer = info.trailers[i];\n+\t\tchar *trailer = info->trailers[i];\n \t\tif (trailer[0] == comment_line_char)\n \t\t\tcontinue;\n \t\tseparator_pos = find_separator(trailer, separators);\n@@ -1002,10 +998,6 @@ static size_t process_input_file(FILE *outfile,\n \t\t\t\t\t strbuf_detach(&val, NULL));\n \t\t}\n \t}\n-\n-\ttrailer_info_release(&info);\n-\n-\treturn info.trailer_end - str;\n }\n \n static void free_all(struct list_head *head)\n@@ -1054,6 +1046,7 @@ void process_trailers(const char *file,\n {\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n+\tstruct trailer_info info;\n \tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n@@ -1064,8 +1057,16 @@ void process_trailers(const char *file,\n \tif (opts->in_place)\n \t\toutfile = create_in_place_tempfile(file);\n \n+\tparse_trailers(&info, sb.buf, &head, opts);\n+\ttrailer_end = info.trailer_end - sb.buf;\n+\n \t/* Print the lines before the trailers */\n-\ttrailer_end = process_input_file(outfile, sb.buf, &head, opts);\n+\tif (!opts->only_trailers)\n+\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\n+\tif (!opts->only_trailers && !info.blank_line_before_trailer)\n+\t\tfprintf(outfile, \"\\n\");\n+\n \n \tif (!opts->only_input) {\n \t\tLIST_HEAD(arg_head);\n@@ -1076,6 +1077,7 @@ void process_trailers(const char *file,\n \tprint_all(outfile, &head, opts);\n \n \tfree_all(&head);\n+\ttrailer_info_release(&info);\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-- \ngitgitgadget\n\n"},{"id":"482173","messageId":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v2.git.1694240177.gitgitgadget@gmail.com","subject":"[PATCH v3 0/9] Trailer readability cleanups","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:36Z","receivedAt":"2023-09-22T19:50:56Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"These patches were created while digging into the trailer code to better\nunderstand how it works, in preparation for making the trailer.{c,h} files\nas small as possible to make them available as a library for external users.\nThis series was originally created as part of [1], but are sent here\nseparately because the changes here are arguably more subjective in nature.\nI think Patch 1 is the most important in this series. The others can wait,\nif folks are opposed to adding them on their own merits at this point in\ntime.\n\nThese patches do not add or change any features. Instead, their goal is to\nmake the code easier to understand for new contributors (like myself), by\nmaking various cleanups and improvements. Ultimately, my hope is that with\nsuch cleanups, we are better positioned to make larger changes (especially\nthe broader libification effort, as in \"Introduce Git Standard Library\"\n[2]).\n\nPatch 1 was inspired by 576de3d956 (unpack_trees: start splitting internal\nfields from public API, 2023-02-27) [3], and is in preparation for a\nlibification effort in the future around the trailer code. Independent of\nlibification, it still makes sense to discourage callers from peeking into\nthese trailer-internal fields.\n\nPatches 2-3 aim to make some functions do a little less multitasking.\n\nPatch 4 is a renaming change to reduce overloaded language in the codebase.\nIt is inspired by 229d6ab6bf (doc: trailer: examples: avoid the word\n\"message\" by itself, 2023-06-15) [4], which did a similar thing for the\ninterpret-trailers documentation.\n\nPatches 5-8 clean up the area around handling the trailer block start and\nend of the input. In particular we rename find_patch_start() to\nfind_end_of_log_message(). These patches address the new approach I cited in\n[5].\n\n\nUpdates in v3\n=============\n\n * Patches 4 and 6 (--no-divider and trailer block start/end cleanups) have\n   been reorganized to Patches 5-8. This ended up touching commit.c in a\n   minor way, but otherwise all of the changes here are cleanups and do not\n   change any behavior.\n\n\nUpdates in v2\n=============\n\n * Patch 1: Drop the use of a #define. Instead just use an anonymous struct\n   named internal.\n * Patch 2: Don't free info out parameter inside parse_trailers(). Instead\n   free it from the caller, process_trailers(). Update comment in\n   parse_trailers().\n * Patch 3: Reword commit message.\n * Patch 4: Mention be3d654343 (commit: pass --no-divider to\n   interpret-trailers, 2023-06-17) in commit message.\n * Added Patch 6 to make trailer_info use offsets for trailer_start and\n   trailer_end (thanks to Glen Choo for the suggestion).\n\n[1]\nhttps://lore.kernel.org/git/pull.1564.git.1691210737.gitgitgadget@gmail.com/T/#mb044012670663d8eb7a548924bbcc933bef116de\n[2]\nhttps://lore.kernel.org/git/20230627195251.1973421-1-calvinwan@google.com/\n[3]\nhttps://lore.kernel.org/git/pull.1149.git.1677143700.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/6b4cb31b17077181a311ca87e82464a1e2ad67dd.1686797630.git.gitgitgadget@gmail.com/\n[5]\nhttps://lore.kernel.org/git/pull.1563.git.1691211879.gitgitgadget@gmail.com/T/#m0131f9829c35d8e0103ffa88f07d8e0e43dd732c\n\nLinus Arver (9):\n  trailer: separate public from internal portion of trailer_iterator\n  trailer: split process_input_file into separate pieces\n  trailer: split process_command_line_args into separate functions\n  trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n  commit: ignore_non_trailer computes number of bytes to ignore\n  trailer: find the end of the log message\n  trailer: use offsets for trailer_start/trailer_end\n  trailer: only use trailer_block_* variables if trailers were found\n  trailer: make stack variable names match field names\n\n builtin/commit.c |   2 +-\n builtin/merge.c  |   2 +-\n commit.c         |   2 +-\n commit.h         |   4 +-\n sequencer.c      |   2 +-\n trailer.c        | 220 ++++++++++++++++++++++++++++-------------------\n trailer.h        |  27 +++---\n 7 files changed, 154 insertions(+), 105 deletions(-)\n\n\nbase-commit: 1b0a5129563ebe720330fdc8f5c6843d27641137\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1563%2Flistx%2Ftrailer-libification-prep-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1563/listx/trailer-libification-prep-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/1563\n\nRange-diff vs v2:\n\n  1:  4f116d2550f =  1:  4f116d2550f trailer: separate public from internal portion of trailer_iterator\n  2:  c00f4623d0b =  2:  c00f4623d0b trailer: split process_input_file into separate pieces\n  3:  f78c2345fad =  3:  f78c2345fad trailer: split process_command_line_args into separate functions\n  5:  52958c3557c =  4:  47186a09b24 trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n  -:  ----------- >  5:  da52cec42e1 commit: ignore_non_trailer computes number of bytes to ignore\n  4:  f5f507c4c6c !  6:  ab8a6ced143 trailer: teach find_patch_start about --no-divider\n     @@ Metadata\n      Author: Linus Arver <linusa@google.com>\n      \n       ## Commit message ##\n     -    trailer: teach find_patch_start about --no-divider\n     +    trailer: find the end of the log message\n      \n     -    Currently, find_patch_start only finds the start of the patch part of\n     -    the input (by looking at the \"---\" divider) for cases where the\n     -    \"--no-divider\" flag has not been provided. If the user provides this\n     -    flag, we do not rely on find_patch_start at all and just call strlen()\n     -    directly on the input.\n     +    Previously, trailer_info_get() computed the trailer block end position\n     +    by\n      \n     -    Instead, make find_patch_start aware of \"--no-divider\" and make it\n     -    handle that case as well. This means we no longer need to call strlen at\n     -    all and can just rely on the existing code in find_patch_start. By\n     -    forcing callers to consider this important option, we avoid the kind of\n     -    mistake described in be3d654343 (commit: pass --no-divider to\n     -    interpret-trailers, 2023-06-17).\n     +    (1) checking for the opts->no_divider flag and optionally calling\n     +        find_patch_start() to find the \"patch start\" location (patch_start), and\n     +    (2) calling find_trailer_end() to find the end of the trailer block\n     +        using patch_start as a guide, saving the return value into\n     +        \"trailer_end\".\n      \n     -    This patch will make unit testing a bit more pleasant in this area in\n     -    the future when we adopt a unit testing framework, because we would not\n     -    have to test multiple functions to check how finding the start of a\n     -    patch part works (we would only need to test find_patch_start).\n     +    The logic in (1) was awkward because the variable \"patch_start\" is\n     +    misleading if there is no patch in the input. The logic in (2) was\n     +    misleading because it could be the case that no trailers are in the\n     +    input (yet we are setting a \"trailer_end\" variable before even searching\n     +    for trailers, which happens later in find_trailer_start()). The name\n     +    \"find_trailer_end\" was misleading because that function did not look for\n     +    any trailer block itself --- instead it just computed the end position\n     +    of the log message in the input where the end of the trailer block (if\n     +    it exists) would be (because trailer blocks must always come after the\n     +    end of the log message).\n      \n     +    Combine the logic in (1) and (2) together into find_patch_start() by\n     +    renaming it to find_end_of_log_message(). The end of the log message is\n     +    the starting point which find_trailer_start() needs to start searching\n     +    backward to parse individual trailers (if any).\n     +\n     +    Helped-by: Junio C Hamano <gitster@pobox.com>\n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n       ## trailer.c ##\n      @@ trailer.c: static ssize_t last_line(const char *buf, size_t len)\n     -  * Return the position of the start of the patch or the length of str if there\n     -  * is no patch in the message.\n     + }\n     + \n     + /*\n     +- * Return the position of the start of the patch or the length of str if there\n     +- * is no patch in the message.\n     ++ * Find the end of the log message as an offset from the start of the input\n     ++ * (where callers of this function are interested in looking for a trailers\n     ++ * block in the same input). We have to consider two categories of content that\n     ++ * can come at the end of the input which we want to ignore (because they don't\n     ++ * belong in the log message):\n     ++ *\n     ++ * (1) the \"patch part\" which begins with a \"---\" divider and has patch\n     ++ * information (like the output of git-format-patch), and\n     ++ *\n     ++ * (2) any trailing comment lines, blank lines like in the output of \"git\n     ++ * commit -v\", or stuff below the \"cut\" (scissor) line.\n     ++ *\n     ++ * As a formula, the situation looks like this:\n     ++ *\n     ++ *     INPUT = LOG MESSAGE + IGNORED\n     ++ *\n     ++ * where IGNORED can be either of the two categories described above. It may be\n     ++ * that there is nothing to ignore. Now it may be the case that the LOG MESSAGE\n     ++ * contains a trailer block, but that's not the concern of this function.\n        */\n      -static size_t find_patch_start(const char *str)\n     -+static size_t find_patch_start(const char *str, int no_divider)\n     ++static size_t find_end_of_log_message(const char *input, int no_divider)\n       {\n     ++\tsize_t end;\n     ++\n       \tconst char *s;\n       \n     - \tfor (s = str; *s; s = next_line(s)) {\n     +-\tfor (s = str; *s; s = next_line(s)) {\n     ++\t/* Assume the naive end of the input is already what we want. */\n     ++\tend = strlen(input);\n     ++\n     ++\t/* Optionally skip over any patch part (\"---\" line and below). */\n     ++\tfor (s = input; *s; s = next_line(s)) {\n       \t\tconst char *v;\n       \n      -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n     -+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v))\n     - \t\t\treturn s - str;\n     +-\t\t\treturn s - str;\n     ++\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n     ++\t\t\tend = s - input;\n     ++\t\t\tbreak;\n     ++\t\t}\n       \t}\n       \n     +-\treturn s - str;\n     ++\t/* Skip over other ignorable bits. */\n     ++\treturn end - ignored_log_message_bytes(input, end);\n     + }\n     + \n     + /*\n     +@@ trailer.c: continue_outer_loop:\n     + \treturn len;\n     + }\n     + \n     +-/* Return the position of the end of the trailers. */\n     +-static size_t find_trailer_end(const char *buf, size_t len)\n     +-{\n     +-\treturn len - ignored_log_message_bytes(buf, len);\n     +-}\n     +-\n     + static int ends_with_blank_line(const char *buf, size_t len)\n     + {\n     + \tssize_t ll = last_line(buf, len);\n     +@@ trailer.c: void process_trailers(const char *file,\n     + void trailer_info_get(struct trailer_info *info, const char *str,\n     + \t\t      const struct process_trailer_options *opts)\n     + {\n     +-\tint patch_start, trailer_end, trailer_start;\n     ++\tint end_of_log_message, trailer_start;\n     + \tstruct strbuf **trailer_lines, **ptr;\n     + \tchar **trailer_strings = NULL;\n     + \tsize_t nr = 0, alloc = 0;\n      @@ trailer.c: void trailer_info_get(struct trailer_info *info, const char *str,\n       \n       \tensure_configured();\n     @@ trailer.c: void trailer_info_get(struct trailer_info *info, const char *str,\n      -\telse\n      -\t\tpatch_start = find_patch_start(str);\n      -\n     -+\tpatch_start = find_patch_start(str, opts->no_divider);\n     - \ttrailer_end = find_trailer_end(str, patch_start);\n     - \ttrailer_start = find_trailer_start(str, trailer_end);\n     +-\ttrailer_end = find_trailer_end(str, patch_start);\n     +-\ttrailer_start = find_trailer_start(str, trailer_end);\n     ++\tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n     ++\ttrailer_start = find_trailer_start(str, end_of_log_message);\n       \n     + \ttrailer_lines = strbuf_split_buf(str + trailer_start,\n     +-\t\t\t\t\t trailer_end - trailer_start,\n     ++\t\t\t\t\t end_of_log_message - trailer_start,\n     + \t\t\t\t\t '\\n',\n     + \t\t\t\t\t 0);\n     + \tfor (ptr = trailer_lines; *ptr; ptr++) {\n     +@@ trailer.c: void trailer_info_get(struct trailer_info *info, const char *str,\n     + \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n     + \t\t\t\t\t\t\t       trailer_start);\n     + \tinfo->trailer_start = str + trailer_start;\n     +-\tinfo->trailer_end = str + trailer_end;\n     ++\tinfo->trailer_end = str + end_of_log_message;\n     + \tinfo->trailers = trailer_strings;\n     + \tinfo->trailer_nr = nr;\n     + }\n  6:  0463066ebe0 !  7:  091805eb7d9 trailer: use offsets for trailer_start/trailer_end\n     @@ Commit message\n          reference the input string in format_trailer_info(), so update that\n          function to take a pointer to the input.\n      \n     +    While we're at it, rename trailer_start to trailer_block_start to be\n     +    more explicit about these offsets (that they are for the entire trailer\n     +    block including other trailers). Ditto for trailer_end.\n     +\n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n     + ## sequencer.c ##\n     +@@ sequencer.c: static int has_conforming_footer(struct strbuf *sb, struct strbuf *sob,\n     + \tif (ignore_footer)\n     + \t\tsb->buf[sb->len - ignore_footer] = saved_char;\n     + \n     +-\tif (info.trailer_start == info.trailer_end)\n     ++\tif (info.trailer_block_start == info.trailer_block_end)\n     + \t\treturn 0;\n     + \n     + \tfor (i = 0; i < info.trailer_nr; i++)\n     +\n       ## trailer.c ##\n     +@@ trailer.c: static size_t find_end_of_log_message(const char *input, int no_divider)\n     +  * Return the position of the first trailer line or len if there are no\n     +  * trailers.\n     +  */\n     +-static size_t find_trailer_start(const char *buf, size_t len)\n     ++static size_t find_trailer_block_start(const char *buf, size_t len)\n     + {\n     + \tconst char *s;\n     + \tssize_t end_of_title, l;\n      @@ trailer.c: void process_trailers(const char *file,\n       \tLIST_HEAD(head);\n       \tstruct strbuf sb = STRBUF_INIT;\n     @@ trailer.c: void process_trailers(const char *file,\n       \t/* Print the lines before the trailers */\n       \tif (!opts->only_trailers)\n      -\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n     -+\t\tfwrite(sb.buf, 1, info.trailer_start, outfile);\n     ++\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n       \n       \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n       \t\tfprintf(outfile, \"\\n\");\n     @@ trailer.c: void process_trailers(const char *file,\n       \t/* Print the lines after the trailers as is */\n       \tif (!opts->only_trailers)\n      -\t\tfwrite(sb.buf + trailer_end, 1, sb.len - trailer_end, outfile);\n     -+\t\tfwrite(sb.buf + info.trailer_end, 1, sb.len - info.trailer_end, outfile);\n     ++\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n       \n       \tif (opts->in_place)\n       \t\tif (rename_tempfile(&trailers_tempfile, file))\n     @@ trailer.c: void process_trailers(const char *file,\n       void trailer_info_get(struct trailer_info *info, const char *str,\n       \t\t      const struct process_trailer_options *opts)\n       {\n     --\tint patch_start, trailer_end, trailer_start;\n     -+\tsize_t patch_start, trailer_end = 0, trailer_start = 0;\n     +-\tint end_of_log_message, trailer_start;\n     ++\tsize_t end_of_log_message = 0, trailer_block_start = 0;\n       \tstruct strbuf **trailer_lines, **ptr;\n       \tchar **trailer_strings = NULL;\n       \tsize_t nr = 0, alloc = 0;\n      @@ trailer.c: void trailer_info_get(struct trailer_info *info, const char *str,\n     + \tensure_configured();\n     + \n     + \tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n     +-\ttrailer_start = find_trailer_start(str, end_of_log_message);\n     ++\ttrailer_block_start = find_trailer_block_start(str, end_of_log_message);\n     + \n     +-\ttrailer_lines = strbuf_split_buf(str + trailer_start,\n     +-\t\t\t\t\t end_of_log_message - trailer_start,\n     ++\ttrailer_lines = strbuf_split_buf(str + trailer_block_start,\n     ++\t\t\t\t\t end_of_log_message - trailer_block_start,\n     + \t\t\t\t\t '\\n',\n     + \t\t\t\t\t 0);\n     + \tfor (ptr = trailer_lines; *ptr; ptr++) {\n     +@@ trailer.c: void trailer_info_get(struct trailer_info *info, const char *str,\n     + \tstrbuf_list_free(trailer_lines);\n       \n       \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n     - \t\t\t\t\t\t\t       trailer_start);\n     +-\t\t\t\t\t\t\t       trailer_start);\n      -\tinfo->trailer_start = str + trailer_start;\n     --\tinfo->trailer_end = str + trailer_end;\n     -+\tinfo->trailer_start = trailer_start;\n     -+\tinfo->trailer_end = trailer_end;\n     +-\tinfo->trailer_end = str + end_of_log_message;\n     ++\t\t\t\t\t\t\t       trailer_block_start);\n     ++\tinfo->trailer_block_start = trailer_block_start;\n     ++\tinfo->trailer_block_end = end_of_log_message;\n       \tinfo->trailers = trailer_strings;\n       \tinfo->trailer_nr = nr;\n       }\n     @@ trailer.c: static void format_trailer_info(struct strbuf *out,\n       \t    !opts->separator && !opts->key_only && !opts->value_only &&\n       \t    !opts->key_value_separator) {\n      -\t\tstrbuf_add(out, info->trailer_start,\n     -+\t\tstrbuf_add(out, msg + info->trailer_start,\n     - \t\t\t   info->trailer_end - info->trailer_start);\n     +-\t\t\t   info->trailer_end - info->trailer_start);\n     ++\t\tstrbuf_add(out, msg + info->trailer_block_start,\n     ++\t\t\t   info->trailer_block_end - info->trailer_block_start);\n       \t\treturn;\n       \t}\n     + \n      @@ trailer.c: void format_trailers_from_commit(struct strbuf *out, const char *msg,\n       \tstruct trailer_info info;\n       \n     @@ trailer.c: void format_trailers_from_commit(struct strbuf *out, const char *msg,\n       \n      \n       ## trailer.h ##\n     -@@ trailer.h: struct trailer_info {\n     +@@ trailer.h: int trailer_set_if_missing(enum trailer_if_missing *item, const char *value);\n     + struct trailer_info {\n     + \t/*\n     + \t * True if there is a blank line before the location pointed to by\n     +-\t * trailer_start.\n     ++\t * trailer_block_start.\n     + \t */\n       \tint blank_line_before_trailer;\n       \n       \t/*\n     @@ trailer.h: struct trailer_info {\n      -\t * is no trailer block found, these 2 pointers point to the end of the\n      -\t * input string.\n      +\t * Offsets to the trailer block start and end positions in the input\n     -+\t * string. If no trailer block is found, these are set to 0.\n     ++\t * string. If no trailer block is found, these are both set to the\n     ++\t * \"true\" end of the input, per find_true_end_of_input().\n     ++\t *\n     ++\t * NOTE: This will be changed so that these point to 0 in the next\n     ++\t * patch if no trailers are found.\n       \t */\n      -\tconst char *trailer_start, *trailer_end;\n     -+\tsize_t trailer_start, trailer_end;\n     ++\tsize_t trailer_block_start, trailer_block_end;\n       \n       \t/*\n       \t * Array of trailers found.\n  -:  ----------- >  8:  1762f78a613 trailer: only use trailer_block_* variables if trailers were found\n  -:  ----------- >  9:  a784c45ed71 trailer: make stack variable names match field names\n\n-- \ngitgitgadget\n"},{"id":"482174","messageId":"f78c2345fadb37e10feb3a18aabb536357549790.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 3/9] trailer: split process_command_line_args into separate functions","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:39Z","receivedAt":"2023-09-22T19:50:57Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, process_command_line_args did two things:\n\n    (1) parse trailers from the configuration, and\n    (2) parse trailers defined on the command line.\n\nSeparate (1) outside to a new function, parse_trailers_from_config.\nRename the remaining logic to parse_trailers_from_command_line_args.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 34 +++++++++++++++++++++-------------\n 1 file changed, 21 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 2c56cbc4a2e..b6de5d9cb2d 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -711,30 +711,35 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tlist_add_tail(&new_item->list, arg_head);\n }\n \n-static void process_command_line_args(struct list_head *arg_head,\n-\t\t\t\t      struct list_head *new_trailer_head)\n+static void parse_trailers_from_config(struct list_head *config_head)\n {\n \tstruct arg_item *item;\n-\tstruct strbuf tok = STRBUF_INIT;\n-\tstruct strbuf val = STRBUF_INIT;\n-\tconst struct conf_info *conf;\n \tstruct list_head *pos;\n \n-\t/*\n-\t * In command-line arguments, '=' is accepted (in addition to the\n-\t * separators that are defined).\n-\t */\n-\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n-\n \t/* Add an arg item for each configured trailer with a command */\n \tlist_for_each(pos, &conf_head) {\n \t\titem = list_entry(pos, struct arg_item, list);\n \t\tif (item->conf.command)\n-\t\t\tadd_arg_item(arg_head,\n+\t\t\tadd_arg_item(config_head,\n \t\t\t\t     xstrdup(token_from_item(item, NULL)),\n \t\t\t\t     xstrdup(\"\"),\n \t\t\t\t     &item->conf, NULL);\n \t}\n+}\n+\n+static void parse_trailers_from_command_line_args(struct list_head *arg_head,\n+\t\t\t\t\t\t  struct list_head *new_trailer_head)\n+{\n+\tstruct strbuf tok = STRBUF_INIT;\n+\tstruct strbuf val = STRBUF_INIT;\n+\tconst struct conf_info *conf;\n+\tstruct list_head *pos;\n+\n+\t/*\n+\t * In command-line arguments, '=' is accepted (in addition to the\n+\t * separators that are defined).\n+\t */\n+\tchar *cl_separators = xstrfmt(\"=%s\", separators);\n \n \t/* Add an arg item for each trailer on the command line */\n \tlist_for_each(pos, new_trailer_head) {\n@@ -1069,8 +1074,11 @@ void process_trailers(const char *file,\n \n \n \tif (!opts->only_input) {\n+\t\tLIST_HEAD(config_head);\n \t\tLIST_HEAD(arg_head);\n-\t\tprocess_command_line_args(&arg_head, new_trailer_head);\n+\t\tparse_trailers_from_config(&config_head);\n+\t\tparse_trailers_from_command_line_args(&arg_head, new_trailer_head);\n+\t\tlist_splice(&config_head, &arg_head);\n \t\tprocess_trailers_lists(&head, &arg_head);\n \t}\n \n-- \ngitgitgadget\n\n"},{"id":"482175","messageId":"47186a09b24522bacf459006330fd469766072f2.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 4/9] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:40Z","receivedAt":"2023-09-22T19:50:59Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nDo not use *_DEFAULT as a suffix to the enums, because the word\n\"default\" is overloaded. The following are two examples of the ambiguity\nof the word \"default\":\n\n(1) \"Default\" can mean using the \"default\" values that are hardcoded\n    in trailer.c as\n\n        default_conf_info.where = WHERE_END;\n        default_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n        default_conf_info.if_missing = MISSING_ADD;\n\n    in ensure_configured(). These values are referred to as \"the\n    default\" in the docs for interpret-trailers. These defaults are used\n    if no \"trailer.*\" configurations are defined.\n\n(2) \"Default\" can also mean the \"trailer.*\" configurations themselves,\n    because these configurations are used by \"default\" (ahead of the\n    hardcoded defaults in (1)) if no command line arguments are\n    provided. This concept of defaulting back to the configurations was\n    introduced in 0ea5292e6b (interpret-trailers: add options for\n    actions, 2017-08-01).\n\nIn addition, the corresponding *_DEFAULT values are chosen when the user\nprovides the \"--no-where\", \"--no-if-exists\", or \"--no-if-missing\" flags\non the command line. These \"--no-*\" flags are used to clear previously\nprovided flags of the form \"--where\", \"--if-exists\", and \"--if-missing\".\nUsing these \"--no-*\" flags undoes the specifying of these flags (if\nany), so using the word \"UNSPECIFIED\" is more natural here.\n\nSo instead of using \"*_DEFAULT\", use \"*_UNSPECIFIED\" because this\nsignals to the reader that the *_UNSPECIFIED value by itself carries no\nmeaning (it's a zero value and by itself does not \"default\" to anything,\nnecessitating the need to have some other way of getting to a useful\nvalue).\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 17 ++++++++++-------\n trailer.h |  6 +++---\n 2 files changed, 13 insertions(+), 10 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex b6de5d9cb2d..0b66effceb5 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -388,7 +388,7 @@ static void process_trailers_lists(struct list_head *head,\n int trailer_set_where(enum trailer_where *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = WHERE_DEFAULT;\n+\t\t*item = WHERE_UNSPECIFIED;\n \telse if (!strcasecmp(\"after\", value))\n \t\t*item = WHERE_AFTER;\n \telse if (!strcasecmp(\"before\", value))\n@@ -405,7 +405,7 @@ int trailer_set_where(enum trailer_where *item, const char *value)\n int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = EXISTS_DEFAULT;\n+\t\t*item = EXISTS_UNSPECIFIED;\n \telse if (!strcasecmp(\"addIfDifferent\", value))\n \t\t*item = EXISTS_ADD_IF_DIFFERENT;\n \telse if (!strcasecmp(\"addIfDifferentNeighbor\", value))\n@@ -424,7 +424,7 @@ int trailer_set_if_exists(enum trailer_if_exists *item, const char *value)\n int trailer_set_if_missing(enum trailer_if_missing *item, const char *value)\n {\n \tif (!value)\n-\t\t*item = MISSING_DEFAULT;\n+\t\t*item = MISSING_UNSPECIFIED;\n \telse if (!strcasecmp(\"doNothing\", value))\n \t\t*item = MISSING_DO_NOTHING;\n \telse if (!strcasecmp(\"add\", value))\n@@ -586,7 +586,10 @@ static void ensure_configured(void)\n \tif (configured)\n \t\treturn;\n \n-\t/* Default config must be setup first */\n+\t/*\n+\t * Default config must be setup first. These defaults are used if there\n+\t * are no \"trailer.*\" or \"trailer.<token>.*\" options configured.\n+\t */\n \tdefault_conf_info.where = WHERE_END;\n \tdefault_conf_info.if_exists = EXISTS_ADD_IF_DIFFERENT_NEIGHBOR;\n \tdefault_conf_info.if_missing = MISSING_ADD;\n@@ -701,11 +704,11 @@ static void add_arg_item(struct list_head *arg_head, char *tok, char *val,\n \tnew_item->value = val;\n \tduplicate_conf(&new_item->conf, conf);\n \tif (new_trailer_item) {\n-\t\tif (new_trailer_item->where != WHERE_DEFAULT)\n+\t\tif (new_trailer_item->where != WHERE_UNSPECIFIED)\n \t\t\tnew_item->conf.where = new_trailer_item->where;\n-\t\tif (new_trailer_item->if_exists != EXISTS_DEFAULT)\n+\t\tif (new_trailer_item->if_exists != EXISTS_UNSPECIFIED)\n \t\t\tnew_item->conf.if_exists = new_trailer_item->if_exists;\n-\t\tif (new_trailer_item->if_missing != MISSING_DEFAULT)\n+\t\tif (new_trailer_item->if_missing != MISSING_UNSPECIFIED)\n \t\t\tnew_item->conf.if_missing = new_trailer_item->if_missing;\n \t}\n \tlist_add_tail(&new_item->list, arg_head);\ndiff --git a/trailer.h b/trailer.h\nindex ab2cd017567..a689d768c79 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -5,14 +5,14 @@\n #include \"strbuf.h\"\n \n enum trailer_where {\n-\tWHERE_DEFAULT,\n+\tWHERE_UNSPECIFIED,\n \tWHERE_END,\n \tWHERE_AFTER,\n \tWHERE_BEFORE,\n \tWHERE_START\n };\n enum trailer_if_exists {\n-\tEXISTS_DEFAULT,\n+\tEXISTS_UNSPECIFIED,\n \tEXISTS_ADD_IF_DIFFERENT_NEIGHBOR,\n \tEXISTS_ADD_IF_DIFFERENT,\n \tEXISTS_ADD,\n@@ -20,7 +20,7 @@ enum trailer_if_exists {\n \tEXISTS_DO_NOTHING\n };\n enum trailer_if_missing {\n-\tMISSING_DEFAULT,\n+\tMISSING_UNSPECIFIED,\n \tMISSING_ADD,\n \tMISSING_DO_NOTHING\n };\n-- \ngitgitgadget\n\n"},{"id":"482176","messageId":"da52cec42e1a64221a7daa958f841ab5e5bd304e.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 5/9] commit: ignore_non_trailer computes number of bytes to ignore","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:41Z","receivedAt":"2023-09-22T19:51:02Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nignore_non_trailer() returns the _number of bytes_ that should be\nignored from the end of the log message. It does not by itself \"ignore\"\nanything.\n\nRename this function to remove the leading \"ignore\" verb, to sound more\nlike a quantity than an action.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n builtin/commit.c | 2 +-\n builtin/merge.c  | 2 +-\n commit.c         | 2 +-\n commit.h         | 4 ++--\n trailer.c        | 2 +-\n 5 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484..d1785d32db1 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -900,7 +900,7 @@ static int prepare_to_commit(const char *index_file, const char *prefix,\n \t\tstrbuf_stripspace(&sb, '\\0');\n \n \tif (signoff)\n-\t\tappend_signoff(&sb, ignore_non_trailer(sb.buf, sb.len), 0);\n+\t\tappend_signoff(&sb, ignored_log_message_bytes(sb.buf, sb.len), 0);\n \n \tif (fwrite(sb.buf, 1, sb.len, s->fp) < sb.len)\n \t\tdie_errno(_(\"could not write commit template\"));\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex de68910177f..6cbbebca13d 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -891,7 +891,7 @@ static void prepare_to_commit(struct commit_list *remoteheads)\n \t\t\t\t_(no_scissors_editor_comment), comment_line_char);\n \t}\n \tif (signoff)\n-\t\tappend_signoff(&msg, ignore_non_trailer(msg.buf, msg.len), 0);\n+\t\tappend_signoff(&msg, ignored_log_message_bytes(msg.buf, msg.len), 0);\n \twrite_merge_heads(remoteheads);\n \twrite_file_buf(git_path_merge_msg(the_repository), msg.buf, msg.len);\n \tif (run_commit_hook(0 < option_edit, get_index_file(), NULL,\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2..4440fbabb83 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1769,7 +1769,7 @@ const char *find_commit_header(const char *msg, const char *key, size_t *out_len\n  * Returns the number of bytes from the tail to ignore, to be fed as\n  * the second parameter to append_signoff().\n  */\n-size_t ignore_non_trailer(const char *buf, size_t len)\n+size_t ignored_log_message_bytes(const char *buf, size_t len)\n {\n \tsize_t boc = 0;\n \tsize_t bol = 0;\ndiff --git a/commit.h b/commit.h\nindex 28928833c54..1cc872f225f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -294,8 +294,8 @@ const char *find_header_mem(const char *msg, size_t len,\n const char *find_commit_header(const char *msg, const char *key,\n \t\t\t       size_t *out_len);\n \n-/* Find the end of the log message, the right place for a new trailer. */\n-size_t ignore_non_trailer(const char *buf, size_t len);\n+/* Find the number of bytes to ignore from the end of a log message. */\n+size_t ignored_log_message_bytes(const char *buf, size_t len);\n \n typedef int (*each_mergetag_fn)(struct commit *commit, struct commit_extra_header *extra,\n \t\t\t\tvoid *cb_data);\ndiff --git a/trailer.c b/trailer.c\nindex 0b66effceb5..185b3e2707f 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -931,7 +931,7 @@ continue_outer_loop:\n /* Return the position of the end of the trailers. */\n static size_t find_trailer_end(const char *buf, size_t len)\n {\n-\treturn len - ignore_non_trailer(buf, len);\n+\treturn len - ignored_log_message_bytes(buf, len);\n }\n \n static int ends_with_blank_line(const char *buf, size_t len)\n-- \ngitgitgadget\n\n"},{"id":"482177","messageId":"ab8a6ced1435580ad3e8585e464e447d40805232.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 6/9] trailer: find the end of the log message","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:42Z","receivedAt":"2023-09-22T19:51:05Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, trailer_info_get() computed the trailer block end position\nby\n\n(1) checking for the opts->no_divider flag and optionally calling\n    find_patch_start() to find the \"patch start\" location (patch_start), and\n(2) calling find_trailer_end() to find the end of the trailer block\n    using patch_start as a guide, saving the return value into\n    \"trailer_end\".\n\nThe logic in (1) was awkward because the variable \"patch_start\" is\nmisleading if there is no patch in the input. The logic in (2) was\nmisleading because it could be the case that no trailers are in the\ninput (yet we are setting a \"trailer_end\" variable before even searching\nfor trailers, which happens later in find_trailer_start()). The name\n\"find_trailer_end\" was misleading because that function did not look for\nany trailer block itself --- instead it just computed the end position\nof the log message in the input where the end of the trailer block (if\nit exists) would be (because trailer blocks must always come after the\nend of the log message).\n\nCombine the logic in (1) and (2) together into find_patch_start() by\nrenaming it to find_end_of_log_message(). The end of the log message is\nthe starting point which find_trailer_start() needs to start searching\nbackward to parse individual trailers (if any).\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 61 ++++++++++++++++++++++++++++++++++---------------------\n 1 file changed, 38 insertions(+), 23 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 185b3e2707f..9da89df9d8a 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -812,21 +812,47 @@ static ssize_t last_line(const char *buf, size_t len)\n }\n \n /*\n- * Return the position of the start of the patch or the length of str if there\n- * is no patch in the message.\n+ * Find the end of the log message as an offset from the start of the input\n+ * (where callers of this function are interested in looking for a trailers\n+ * block in the same input). We have to consider two categories of content that\n+ * can come at the end of the input which we want to ignore (because they don't\n+ * belong in the log message):\n+ *\n+ * (1) the \"patch part\" which begins with a \"---\" divider and has patch\n+ * information (like the output of git-format-patch), and\n+ *\n+ * (2) any trailing comment lines, blank lines like in the output of \"git\n+ * commit -v\", or stuff below the \"cut\" (scissor) line.\n+ *\n+ * As a formula, the situation looks like this:\n+ *\n+ *     INPUT = LOG MESSAGE + IGNORED\n+ *\n+ * where IGNORED can be either of the two categories described above. It may be\n+ * that there is nothing to ignore. Now it may be the case that the LOG MESSAGE\n+ * contains a trailer block, but that's not the concern of this function.\n  */\n-static size_t find_patch_start(const char *str)\n+static size_t find_end_of_log_message(const char *input, int no_divider)\n {\n+\tsize_t end;\n+\n \tconst char *s;\n \n-\tfor (s = str; *s; s = next_line(s)) {\n+\t/* Assume the naive end of the input is already what we want. */\n+\tend = strlen(input);\n+\n+\t/* Optionally skip over any patch part (\"---\" line and below). */\n+\tfor (s = input; *s; s = next_line(s)) {\n \t\tconst char *v;\n \n-\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n-\t\t\treturn s - str;\n+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n+\t\t\tend = s - input;\n+\t\t\tbreak;\n+\t\t}\n \t}\n \n-\treturn s - str;\n+\t/* Skip over other ignorable bits. */\n+\treturn end - ignored_log_message_bytes(input, end);\n }\n \n /*\n@@ -928,12 +954,6 @@ continue_outer_loop:\n \treturn len;\n }\n \n-/* Return the position of the end of the trailers. */\n-static size_t find_trailer_end(const char *buf, size_t len)\n-{\n-\treturn len - ignored_log_message_bytes(buf, len);\n-}\n-\n static int ends_with_blank_line(const char *buf, size_t len)\n {\n \tssize_t ll = last_line(buf, len);\n@@ -1104,7 +1124,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint patch_start, trailer_end, trailer_start;\n+\tint end_of_log_message, trailer_start;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1112,16 +1132,11 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tensure_configured();\n \n-\tif (opts->no_divider)\n-\t\tpatch_start = strlen(str);\n-\telse\n-\t\tpatch_start = find_patch_start(str);\n-\n-\ttrailer_end = find_trailer_end(str, patch_start);\n-\ttrailer_start = find_trailer_start(str, trailer_end);\n+\tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n+\ttrailer_start = find_trailer_start(str, end_of_log_message);\n \n \ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t trailer_end - trailer_start,\n+\t\t\t\t\t end_of_log_message - trailer_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1144,7 +1159,7 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_start);\n \tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + trailer_end;\n+\tinfo->trailer_end = str + end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n-- \ngitgitgadget\n\n"},{"id":"482178","messageId":"a784c45ed715c5066961f1566f2e91eb597d89a3.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 9/9] trailer: make stack variable names match field names","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:45Z","receivedAt":"2023-09-22T19:51:07Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 20 ++++++++++----------\n 1 file changed, 10 insertions(+), 10 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 9a3837be770..739acafc4e9 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1134,8 +1134,8 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n {\n \tsize_t end_of_log_message = 0, trailer_block_start = 0;\n \tstruct strbuf **trailer_lines, **ptr;\n-\tchar **trailer_strings = NULL;\n-\tsize_t nr = 0, alloc = 0;\n+\tchar **trailers = NULL;\n+\tsize_t trailer_nr = 0, alloc = 0;\n \tchar **last = NULL;\n \n \tensure_configured();\n@@ -1155,12 +1155,12 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t\t*last = strbuf_detach(&sb, NULL);\n \t\t\tcontinue;\n \t\t}\n-\t\tALLOC_GROW(trailer_strings, nr + 1, alloc);\n-\t\ttrailer_strings[nr] = strbuf_detach(*ptr, NULL);\n-\t\tlast = find_separator(trailer_strings[nr], separators) >= 1\n-\t\t\t? &trailer_strings[nr]\n+\t\tALLOC_GROW(trailers, trailer_nr + 1, alloc);\n+\t\ttrailers[trailer_nr] = strbuf_detach(*ptr, NULL);\n+\t\tlast = find_separator(trailers[trailer_nr], separators) >= 1\n+\t\t\t? &trailers[trailer_nr]\n \t\t\t: NULL;\n-\t\tnr++;\n+\t\ttrailer_nr++;\n \t}\n \tstrbuf_list_free(trailer_lines);\n \n@@ -1168,13 +1168,13 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t\t\t\t\t\t       trailer_block_start);\n \tinfo->trailer_block_start = 0;\n \tinfo->trailer_block_end = 0;\n-\tif (nr) {\n+\tif (trailer_nr) {\n \t\tinfo->trailer_block_start = trailer_block_start;\n \t\tinfo->trailer_block_end = end_of_log_message;\n \t}\n \tinfo->end_of_log_message = end_of_log_message;\n-\tinfo->trailers = trailer_strings;\n-\tinfo->trailer_nr = nr;\n+\tinfo->trailers = trailers;\n+\tinfo->trailer_nr = trailer_nr;\n }\n \n void trailer_info_release(struct trailer_info *info)\n-- \ngitgitgadget\n"},{"id":"482179","messageId":"091805eb7d93efa6fbe3831bcddd2a6fdc033388.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 7/9] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:43Z","receivedAt":"2023-09-22T19:51:23Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously these fields in the trailer_info struct were of type \"const\nchar *\" and pointed to positions in the input string directly (to the\nstart and end positions of the trailer block).\n\nUse offsets to make the intended usage less ambiguous. We only need to\nreference the input string in format_trailer_info(), so update that\nfunction to take a pointer to the input.\n\nWhile we're at it, rename trailer_start to trailer_block_start to be\nmore explicit about these offsets (that they are for the entire trailer\nblock including other trailers). Ditto for trailer_end.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n sequencer.c |  2 +-\n trailer.c   | 29 ++++++++++++++---------------\n trailer.h   | 13 ++++++++-----\n 3 files changed, 23 insertions(+), 21 deletions(-)\n\ndiff --git a/sequencer.c b/sequencer.c\nindex adc9cfb4df3..77362f5cd5d 100644\n--- a/sequencer.c\n+++ b/sequencer.c\n@@ -331,7 +331,7 @@ static int has_conforming_footer(struct strbuf *sb, struct strbuf *sob,\n \tif (ignore_footer)\n \t\tsb->buf[sb->len - ignore_footer] = saved_char;\n \n-\tif (info.trailer_start == info.trailer_end)\n+\tif (info.trailer_block_start == info.trailer_block_end)\n \t\treturn 0;\n \n \tfor (i = 0; i < info.trailer_nr; i++)\ndiff --git a/trailer.c b/trailer.c\nindex 9da89df9d8a..471c2722536 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -859,7 +859,7 @@ static size_t find_end_of_log_message(const char *input, int no_divider)\n  * Return the position of the first trailer line or len if there are no\n  * trailers.\n  */\n-static size_t find_trailer_start(const char *buf, size_t len)\n+static size_t find_trailer_block_start(const char *buf, size_t len)\n {\n \tconst char *s;\n \tssize_t end_of_title, l;\n@@ -1075,7 +1075,6 @@ void process_trailers(const char *file,\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n \tstruct trailer_info info;\n-\tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n \tensure_configured();\n@@ -1086,11 +1085,10 @@ void process_trailers(const char *file,\n \t\toutfile = create_in_place_tempfile(file);\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n-\ttrailer_end = info.trailer_end - sb.buf;\n \n \t/* Print the lines before the trailers */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1112,7 +1110,7 @@ void process_trailers(const char *file,\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + trailer_end, 1, sb.len - trailer_end, outfile);\n+\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1124,7 +1122,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint end_of_log_message, trailer_start;\n+\tsize_t end_of_log_message = 0, trailer_block_start = 0;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1133,10 +1131,10 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tensure_configured();\n \n \tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n-\ttrailer_start = find_trailer_start(str, end_of_log_message);\n+\ttrailer_block_start = find_trailer_block_start(str, end_of_log_message);\n \n-\ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t end_of_log_message - trailer_start,\n+\ttrailer_lines = strbuf_split_buf(str + trailer_block_start,\n+\t\t\t\t\t end_of_log_message - trailer_block_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1157,9 +1155,9 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tstrbuf_list_free(trailer_lines);\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n-\t\t\t\t\t\t\t       trailer_start);\n-\tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + end_of_log_message;\n+\t\t\t\t\t\t\t       trailer_block_start);\n+\tinfo->trailer_block_start = trailer_block_start;\n+\tinfo->trailer_block_end = end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n@@ -1174,6 +1172,7 @@ void trailer_info_release(struct trailer_info *info)\n \n static void format_trailer_info(struct strbuf *out,\n \t\t\t\tconst struct trailer_info *info,\n+\t\t\t\tconst char *msg,\n \t\t\t\tconst struct process_trailer_options *opts)\n {\n \tsize_t origlen = out->len;\n@@ -1183,8 +1182,8 @@ static void format_trailer_info(struct strbuf *out,\n \tif (!opts->only_trailers && !opts->unfold && !opts->filter &&\n \t    !opts->separator && !opts->key_only && !opts->value_only &&\n \t    !opts->key_value_separator) {\n-\t\tstrbuf_add(out, info->trailer_start,\n-\t\t\t   info->trailer_end - info->trailer_start);\n+\t\tstrbuf_add(out, msg + info->trailer_block_start,\n+\t\t\t   info->trailer_block_end - info->trailer_block_start);\n \t\treturn;\n \t}\n \n@@ -1238,7 +1237,7 @@ void format_trailers_from_commit(struct strbuf *out, const char *msg,\n \tstruct trailer_info info;\n \n \ttrailer_info_get(&info, msg, opts);\n-\tformat_trailer_info(out, &info, opts);\n+\tformat_trailer_info(out, &info, msg, opts);\n \ttrailer_info_release(&info);\n }\n \ndiff --git a/trailer.h b/trailer.h\nindex a689d768c79..4dcb9080327 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -32,16 +32,19 @@ int trailer_set_if_missing(enum trailer_if_missing *item, const char *value);\n struct trailer_info {\n \t/*\n \t * True if there is a blank line before the location pointed to by\n-\t * trailer_start.\n+\t * trailer_block_start.\n \t */\n \tint blank_line_before_trailer;\n \n \t/*\n-\t * Pointers to the start and end of the trailer block found. If there\n-\t * is no trailer block found, these 2 pointers point to the end of the\n-\t * input string.\n+\t * Offsets to the trailer block start and end positions in the input\n+\t * string. If no trailer block is found, these are both set to the\n+\t * \"true\" end of the input, per find_true_end_of_input().\n+\t *\n+\t * NOTE: This will be changed so that these point to 0 in the next\n+\t * patch if no trailers are found.\n \t */\n-\tconst char *trailer_start, *trailer_end;\n+\tsize_t trailer_block_start, trailer_block_end;\n \n \t/*\n \t * Array of trailers found.\n-- \ngitgitgadget\n\n"},{"id":"482180","messageId":"1762f78a613f4a744e76ad515b6d27ca9bea47ed.1695412245.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v3 8/9] trailer: only use trailer_block_* variables if trailers were found","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-22T19:50:44Z","receivedAt":"2023-09-22T19:51:25Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, these variables were overloaded to act as the end of the log\nmessage even if no trailers were found.\n\nRemove the overloaded meaning by adding a new end_of_log_message field\nto the trailer_info struct. The trailer_info struct consumers now only\nrefer to the trailer_block_start and trailer_block_end fields if\ntrailers were found (trailer_nr > 0), and otherwise refer to the\nend_of_log_message.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 31 +++++++++++++++++++++++--------\n trailer.h | 12 +++++++-----\n 2 files changed, 30 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 471c2722536..9a3837be770 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1086,9 +1086,14 @@ void process_trailers(const char *file,\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n \n-\t/* Print the lines before the trailers */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n+\t/* Print the lines before the trailers (if any) as is. */\n+\tif (!opts->only_trailers) {\n+\t\tif (info.trailer_nr) {\n+\t\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n+\t\t} else {\n+\t\t\tfwrite(sb.buf, 1, info.end_of_log_message, outfile);\n+\t\t}\n+\t}\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1108,9 +1113,14 @@ void process_trailers(const char *file,\n \tfree_all(&head);\n \ttrailer_info_release(&info);\n \n-\t/* Print the lines after the trailers as is */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n+\t/* Print the lines after the trailers (if any) as is. */\n+\tif (!opts->only_trailers) {\n+\t\tif (info.trailer_nr) {\n+\t\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n+\t\t} else {\n+\t\t\tfwrite(sb.buf + info.end_of_log_message, 1, sb.len - info.end_of_log_message, outfile);\n+\t\t}\n+\t}\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1156,8 +1166,13 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_block_start);\n-\tinfo->trailer_block_start = trailer_block_start;\n-\tinfo->trailer_block_end = end_of_log_message;\n+\tinfo->trailer_block_start = 0;\n+\tinfo->trailer_block_end = 0;\n+\tif (nr) {\n+\t\tinfo->trailer_block_start = trailer_block_start;\n+\t\tinfo->trailer_block_end = end_of_log_message;\n+\t}\n+\tinfo->end_of_log_message = end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\ndiff --git a/trailer.h b/trailer.h\nindex 4dcb9080327..5e2843d320a 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -38,14 +38,16 @@ struct trailer_info {\n \n \t/*\n \t * Offsets to the trailer block start and end positions in the input\n-\t * string. If no trailer block is found, these are both set to the\n-\t * \"true\" end of the input, per find_true_end_of_input().\n-\t *\n-\t * NOTE: This will be changed so that these point to 0 in the next\n-\t * patch if no trailers are found.\n+\t * string. If no trailer block is found, these are set to 0.\n \t */\n \tsize_t trailer_block_start, trailer_block_end;\n \n+\t/*\n+\t * Offset to the end of the log message in the input (may not be the\n+\t * same as the end of the input).\n+\t */\n+\tsize_t end_of_log_message;\n+\n \t/*\n \t * Array of trailers found.\n \t */\n-- \ngitgitgadget\n\n"},{"id":"482192","messageId":"xmqq8r8xyge6.fsf@gitster.g","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 0/9] Trailer readability cleanups","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-22T22:47:29Z","receivedAt":"2023-09-22T22:47:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> These patches were created while digging into the trailer code to better\n> understand how it works, in preparation for making the trailer.{c,h} files\n> as small as possible to make them available as a library for external users.\n> This series was originally created as part of [1], but are sent here\n> separately because the changes here are arguably more subjective in nature.\n> I think Patch 1 is the most important in this series. The others can wait,\n> if folks are opposed to adding them on their own merits at this point in\n> time.\n\nHmph, as we discussed, these changes have already been cooking in\n'next' for some time:\n\n    13211ae23f trailer: separate public from internal portion of trailer_iterator\n    c2a8edf997 trailer: split process_input_file into separate pieces\n    94430d03df trailer: split process_command_line_args into separate functions\n    ee8c5ee08c trailer: teach find_patch_start about --no-divider\n    d2be104085 trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n    b5e75f87b5 trailer: use offsets for trailer_start/trailer_end\n\nand I thought we agreed that we'll park them in 'next' and do\nwhatever necessary fix-up on top as incremental patches?  The first\nthree patches in this iteration seems to be identical to the\nprevious round, so I can ignore them and keep the old iteration, but\nthe remainder of this series are replacements that are suitable when\nthe series was still out of 'next', but they are no longer usable\nonce the series is in 'next'.\n\nI could revert and discard [4-6/6] of the previous iteration out of\n'next' and have only the first three (which I thought have been\nadequately reviewed without remaining issues) graduate to 'master',\nif it makes it easier to fix this update on top, but I'd rather not\nto encourage people to form a habit of reverting changes out of\n'next'.\n\nThanks.\n"},{"id":"482193","messageId":"owlyzg1dsswr.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqq8r8xyge6.fsf@gitster.g","subject":"Re: [PATCH v3 0/9] Trailer readability cleanups","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-22T23:13:40Z","receivedAt":"2023-09-22T23:13:44Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> These patches were created while digging into the trailer code to better\n>> understand how it works, in preparation for making the trailer.{c,h} files\n>> as small as possible to make them available as a library for external users.\n>> This series was originally created as part of [1], but are sent here\n>> separately because the changes here are arguably more subjective in nature.\n>> I think Patch 1 is the most important in this series. The others can wait,\n>> if folks are opposed to adding them on their own merits at this point in\n>> time.\n>\n> Hmph, as we discussed, these changes have already been cooking in\n> 'next' for some time:\n>\n>     13211ae23f trailer: separate public from internal portion of trailer_iterator\n>     c2a8edf997 trailer: split process_input_file into separate pieces\n>     94430d03df trailer: split process_command_line_args into separate functions\n>     ee8c5ee08c trailer: teach find_patch_start about --no-divider\n>     d2be104085 trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n>     b5e75f87b5 trailer: use offsets for trailer_start/trailer_end\n>\n> and I thought we agreed that we'll park them in 'next' and do\n> whatever necessary fix-up on top as incremental patches?\n\nAhhh yes! I completely forgot. So sorry for the noise...\n\n> The first\n> three patches in this iteration seems to be identical to the\n> previous round, so I can ignore them and keep the old iteration, but\n> the remainder of this series are replacements that are suitable when\n> the series was still out of 'next', but they are no longer usable\n> once the series is in 'next'.\n\nRight.\n\n> I could revert and discard [4-6/6] of the previous iteration out of\n> 'next' and have only the first three (which I thought have been\n> adequately reviewed without remaining issues) graduate to 'master',\n> if it makes it easier to fix this update on top, but I'd rather not\n> to encourage people to form a habit of reverting changes out of\n> 'next'.\n>\n> Thanks.\n\nI totally agree that reverting changes out of next is undesirable. I\nwill do a reroll on top of 'next' with only those incremental (new)\npatches.\n"},{"id":"482194","messageId":"xmqqo7htww7k.fsf@gitster.g","threadId":"60068","inReplyTo":"owlyzg1dsswr.fsf@fine.c.googlers.com","subject":"Re: [PATCH v3 0/9] Trailer readability cleanups","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-23T00:48:47Z","receivedAt":"2023-09-23T00:48:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n>> I could revert and discard [4-6/6] of the previous iteration out of\n>> 'next' and have only the first three (which I thought have been\n>> adequately reviewed without remaining issues) graduate to 'master',\n>> if it makes it easier to fix this update on top, but I'd rather not\n>> to encourage people to form a habit of reverting changes out of\n>> 'next'.\n>>\n>> Thanks.\n>\n> I totally agree that reverting changes out of next is undesirable. I\n> will do a reroll on top of 'next' with only those incremental (new)\n> patches.\n\nOK, so the first 3 patches are now in 'master', and the remainder of\nthe previous series have been discarded.\n\nThanks.\n"},{"id":"482318","messageId":"owlywmwdsdr0.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqil820z25.fsf@gitster.g","subject":"Re: [PATCH v2 5/6] trailer: rename *_DEFAULT enums to *_UNSPECIFIED","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-26T05:30:11Z","receivedAt":"2023-09-26T05:30:18Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>> ... I prefer the\n>> WHERE_UNSPECIFIED as in this patch because the WHERE_DEFAULT is\n>> ambiguous on its own (i.e., WHERE_DEFAULT could mean that we either use\n>> the default value WHERE_END in default_conf_info, or it could mean that\n>> we fall back to the configuration variables (where it could be something\n>> else)).\n>\n> Yup.  \"Turning something that is left UNSPECIFIED after command line\n> options and configuration files are processed into the hardcoded\n> DEFAULT\" is one mental model that is easy to explain.\n>\n> I however am not sure if it is easier than \"Setting something to\n> hardcoded DEFAULT before command line options and configuration\n> files are allowed to tweak it, and if nobody touches it, then it\n> gets the hardcoded DEFAULT value in the end\", which is another valid\n> mental model, though.\n\nTrue.\n\n> If both can be used, I'd personally prefer\n> the latter, and reserve the \"UNSPECIFIED to DEFAULT\" pattern to\n> signal that we are dealing with a case where the simpler pattern\n> without UNSPECIFIED cannot solve.\n\nSGTM.\n\n> The simpler pattern would not work, when the default is defined\n> depending on a larger context.  Imagine we have two Boolean\n> variables, A and B, where A defaults to false, and B defaults to\n> some value derived from the value of A (say, opposite of A).\n>\n> In the most natural implementation, you'd initialize A to false and\n> B to unspecified, let command line options and configuration\n> variables to set them to true or false, and after all that, you do\n> not have to tweak A's value (it will be left to false that is the\n> default unless the user or the configuration gave an explicit\n> value), but you need to check if B is left unspecified and tweak it\n> to true or false using the final value of A.\n>\n> For a variable with such a need like B, we cannot avoid having\n> \"unspecified\".  If you initialize it to false (or true), after the\n> command line and the configuration files are read and you find B is\n> set to false (or true), you cannot tell if the user or the\n> configuration explicitly set B to false (or true), in which case you\n> do not want to futz with its value based on what is in A, or it is\n> false (or true) only because nobody touched it, in which case you\n> need to compute its value based on what is in A.\n\nThanks for the illustrative example! I don't think we have a case of a\n\"B\" variable here for trailers.\n\n> And that is why I asked if we need to special case \"the user did not\n> touch and the variable is left untouched\" in the trailer subsystem.\n\nI think the answer is \"no, we don't need to special case\". I'll be\ndropping this patch in the next re-roll.\n\n> Thanks.\n"},{"id":"482319","messageId":"owlyttrhsdaf.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqo7htww7k.fsf@gitster.g","subject":"Re: [PATCH v3 0/9] Trailer readability cleanups","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-09-26T05:40:08Z","receivedAt":"2023-09-26T05:40:13Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>>> I could revert and discard [4-6/6] of the previous iteration out of\n>>> 'next' and have only the first three (which I thought have been\n>>> adequately reviewed without remaining issues) graduate to 'master',\n>>> if it makes it easier to fix this update on top, but I'd rather not\n>>> to encourage people to form a habit of reverting changes out of\n>>> 'next'.\n>>>\n>>> Thanks.\n>>\n>> I totally agree that reverting changes out of next is undesirable. I\n>> will do a reroll on top of 'next' with only those incremental (new)\n>> patches.\n>\n> OK, so the first 3 patches are now in 'master', and the remainder of\n> the previous series have been discarded.\n>\n> Thanks.\n\nOh, that simplifies things. I will re-roll on top of 'master' as the\nstarting point for the other remaining patches instead of using 'next'\nas I suggested earlier.\n"},{"id":"482327","messageId":"4ce5cf77005eb8c6da243777b3c29103add7ddbd.1695709372.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","subject":"[PATCH v4 1/4] commit: ignore_non_trailer computes number of bytes to ignore","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-26T06:22:49Z","receivedAt":"2023-09-26T06:23:02Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nignore_non_trailer() returns the _number of bytes_ that should be\nignored from the end of the log message. It does not by itself \"ignore\"\nanything.\n\nRename this function to remove the leading \"ignore\" verb, to sound more\nlike a quantity than an action.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n builtin/commit.c | 2 +-\n builtin/merge.c  | 2 +-\n commit.c         | 2 +-\n commit.h         | 4 ++--\n trailer.c        | 2 +-\n 5 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484..d1785d32db1 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -900,7 +900,7 @@ static int prepare_to_commit(const char *index_file, const char *prefix,\n \t\tstrbuf_stripspace(&sb, '\\0');\n \n \tif (signoff)\n-\t\tappend_signoff(&sb, ignore_non_trailer(sb.buf, sb.len), 0);\n+\t\tappend_signoff(&sb, ignored_log_message_bytes(sb.buf, sb.len), 0);\n \n \tif (fwrite(sb.buf, 1, sb.len, s->fp) < sb.len)\n \t\tdie_errno(_(\"could not write commit template\"));\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex 545da0c8a11..c654a29fe85 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -870,7 +870,7 @@ static void prepare_to_commit(struct commit_list *remoteheads)\n \t\t\t\t_(no_scissors_editor_comment), comment_line_char);\n \t}\n \tif (signoff)\n-\t\tappend_signoff(&msg, ignore_non_trailer(msg.buf, msg.len), 0);\n+\t\tappend_signoff(&msg, ignored_log_message_bytes(msg.buf, msg.len), 0);\n \twrite_merge_heads(remoteheads);\n \twrite_file_buf(git_path_merge_msg(the_repository), msg.buf, msg.len);\n \tif (run_commit_hook(0 < option_edit, get_index_file(), NULL,\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2..4440fbabb83 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1769,7 +1769,7 @@ const char *find_commit_header(const char *msg, const char *key, size_t *out_len\n  * Returns the number of bytes from the tail to ignore, to be fed as\n  * the second parameter to append_signoff().\n  */\n-size_t ignore_non_trailer(const char *buf, size_t len)\n+size_t ignored_log_message_bytes(const char *buf, size_t len)\n {\n \tsize_t boc = 0;\n \tsize_t bol = 0;\ndiff --git a/commit.h b/commit.h\nindex 28928833c54..1cc872f225f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -294,8 +294,8 @@ const char *find_header_mem(const char *msg, size_t len,\n const char *find_commit_header(const char *msg, const char *key,\n \t\t\t       size_t *out_len);\n \n-/* Find the end of the log message, the right place for a new trailer. */\n-size_t ignore_non_trailer(const char *buf, size_t len);\n+/* Find the number of bytes to ignore from the end of a log message. */\n+size_t ignored_log_message_bytes(const char *buf, size_t len);\n \n typedef int (*each_mergetag_fn)(struct commit *commit, struct commit_extra_header *extra,\n \t\t\t\tvoid *cb_data);\ndiff --git a/trailer.c b/trailer.c\nindex b6de5d9cb2d..3c54b38a85a 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -928,7 +928,7 @@ continue_outer_loop:\n /* Return the position of the end of the trailers. */\n static size_t find_trailer_end(const char *buf, size_t len)\n {\n-\treturn len - ignore_non_trailer(buf, len);\n+\treturn len - ignored_log_message_bytes(buf, len);\n }\n \n static int ends_with_blank_line(const char *buf, size_t len)\n-- \ngitgitgadget\n\n"},{"id":"482328","messageId":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v3.git.1695412245.gitgitgadget@gmail.com","subject":"[PATCH v4 0/4] Trailer readability cleanups","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-26T06:22:48Z","receivedAt":"2023-09-26T06:23:04Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"These patches were created while digging into the trailer code to better\nunderstand how it works, in preparation for making the trailer.{c,h} files\nas small as possible to make them available as a library for external users.\nThis series was originally created as part of [1], but are sent here\nseparately because the changes here are arguably more subjective in nature.\n\nThese patches do not add or change any features. Instead, their goal is to\nmake the code easier to understand for new contributors (like myself), by\nmaking various cleanups and improvements. Ultimately, my hope is that with\nsuch cleanups, we are better positioned to make larger changes (especially\nthe broader libification effort, as in \"Introduce Git Standard Library\"\n[2]).\n\n\nUpdates in v4\n=============\n\n * The first 3 patches in v3 were merged into 'master'. Necessarily, those 3\n   patches have been dropped.\n * Patch 4 in v3 (\"trailer: rename *_DEFAULT enums to *_UNSPECIFIED\") has\n   been dropped, as well as Patch 9 in v3 (\"trailer: make stack variable\n   names match field names\"). These were dropped to simplify this series for\n   what I think is the more immediate, important change (see next bullet\n   point).\n * Patches 5-8 in v3 are the only ones remaining in this series. They still\n   solely deal with --no-divider and trailer block start/end cleanups.\n\n\nUpdates in v3\n=============\n\n * Patches 4 and 6 (--no-divider and trailer block start/end cleanups) have\n   been reorganized to Patches 5-8. This ended up touching commit.c in a\n   minor way, but otherwise all of the changes here are cleanups and do not\n   change any behavior.\n\n\nUpdates in v2\n=============\n\n * Patch 1: Drop the use of a #define. Instead just use an anonymous struct\n   named internal.\n * Patch 2: Don't free info out parameter inside parse_trailers(). Instead\n   free it from the caller, process_trailers(). Update comment in\n   parse_trailers().\n * Patch 3: Reword commit message.\n * Patch 4: Mention be3d654343 (commit: pass --no-divider to\n   interpret-trailers, 2023-06-17) in commit message.\n * Added Patch 6 to make trailer_info use offsets for trailer_start and\n   trailer_end (thanks to Glen Choo for the suggestion).\n\n[1]\nhttps://lore.kernel.org/git/pull.1564.git.1691210737.gitgitgadget@gmail.com/T/#mb044012670663d8eb7a548924bbcc933bef116de\n[2]\nhttps://lore.kernel.org/git/20230627195251.1973421-1-calvinwan@google.com/\n[3]\nhttps://lore.kernel.org/git/pull.1149.git.1677143700.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/6b4cb31b17077181a311ca87e82464a1e2ad67dd.1686797630.git.gitgitgadget@gmail.com/\n[5]\nhttps://lore.kernel.org/git/pull.1563.git.1691211879.gitgitgadget@gmail.com/T/#m0131f9829c35d8e0103ffa88f07d8e0e43dd732c\n\nLinus Arver (4):\n  commit: ignore_non_trailer computes number of bytes to ignore\n  trailer: find the end of the log message\n  trailer: use offsets for trailer_start/trailer_end\n  trailer: only use trailer_block_* variables if trailers were found\n\n builtin/commit.c |   2 +-\n builtin/merge.c  |   2 +-\n commit.c         |   2 +-\n commit.h         |   4 +-\n sequencer.c      |   2 +-\n trailer.c        | 105 ++++++++++++++++++++++++++++++-----------------\n trailer.h        |  15 ++++---\n 7 files changed, 83 insertions(+), 49 deletions(-)\n\n\nbase-commit: bcb6cae2966cc407ca1afc77413b3ef11103c175\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1563%2Flistx%2Ftrailer-libification-prep-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1563/listx/trailer-libification-prep-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/1563\n\nRange-diff vs v3:\n\n  1:  4f116d2550f <  -:  ----------- trailer: separate public from internal portion of trailer_iterator\n  2:  c00f4623d0b <  -:  ----------- trailer: split process_input_file into separate pieces\n  3:  f78c2345fad <  -:  ----------- trailer: split process_command_line_args into separate functions\n  4:  47186a09b24 <  -:  ----------- trailer: rename *_DEFAULT enums to *_UNSPECIFIED\n  5:  da52cec42e1 =  1:  4ce5cf77005 commit: ignore_non_trailer computes number of bytes to ignore\n  6:  ab8a6ced143 =  2:  c904caba7e1 trailer: find the end of the log message\n  7:  091805eb7d9 =  3:  796e47c1e5f trailer: use offsets for trailer_start/trailer_end\n  8:  1762f78a613 =  4:  64e1bd4e4be trailer: only use trailer_block_* variables if trailers were found\n  9:  a784c45ed71 <  -:  ----------- trailer: make stack variable names match field names\n\n-- \ngitgitgadget\n"},{"id":"482329","messageId":"c904caba7e17b6f2784933e9f18634ea66f28537.1695709372.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","subject":"[PATCH v4 2/4] trailer: find the end of the log message","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-26T06:22:50Z","receivedAt":"2023-09-26T06:23:06Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, trailer_info_get() computed the trailer block end position\nby\n\n(1) checking for the opts->no_divider flag and optionally calling\n    find_patch_start() to find the \"patch start\" location (patch_start), and\n(2) calling find_trailer_end() to find the end of the trailer block\n    using patch_start as a guide, saving the return value into\n    \"trailer_end\".\n\nThe logic in (1) was awkward because the variable \"patch_start\" is\nmisleading if there is no patch in the input. The logic in (2) was\nmisleading because it could be the case that no trailers are in the\ninput (yet we are setting a \"trailer_end\" variable before even searching\nfor trailers, which happens later in find_trailer_start()). The name\n\"find_trailer_end\" was misleading because that function did not look for\nany trailer block itself --- instead it just computed the end position\nof the log message in the input where the end of the trailer block (if\nit exists) would be (because trailer blocks must always come after the\nend of the log message).\n\nCombine the logic in (1) and (2) together into find_patch_start() by\nrenaming it to find_end_of_log_message(). The end of the log message is\nthe starting point which find_trailer_start() needs to start searching\nbackward to parse individual trailers (if any).\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 61 ++++++++++++++++++++++++++++++++++---------------------\n 1 file changed, 38 insertions(+), 23 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 3c54b38a85a..96cb285a4ea 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -809,21 +809,47 @@ static ssize_t last_line(const char *buf, size_t len)\n }\n \n /*\n- * Return the position of the start of the patch or the length of str if there\n- * is no patch in the message.\n+ * Find the end of the log message as an offset from the start of the input\n+ * (where callers of this function are interested in looking for a trailers\n+ * block in the same input). We have to consider two categories of content that\n+ * can come at the end of the input which we want to ignore (because they don't\n+ * belong in the log message):\n+ *\n+ * (1) the \"patch part\" which begins with a \"---\" divider and has patch\n+ * information (like the output of git-format-patch), and\n+ *\n+ * (2) any trailing comment lines, blank lines like in the output of \"git\n+ * commit -v\", or stuff below the \"cut\" (scissor) line.\n+ *\n+ * As a formula, the situation looks like this:\n+ *\n+ *     INPUT = LOG MESSAGE + IGNORED\n+ *\n+ * where IGNORED can be either of the two categories described above. It may be\n+ * that there is nothing to ignore. Now it may be the case that the LOG MESSAGE\n+ * contains a trailer block, but that's not the concern of this function.\n  */\n-static size_t find_patch_start(const char *str)\n+static size_t find_end_of_log_message(const char *input, int no_divider)\n {\n+\tsize_t end;\n+\n \tconst char *s;\n \n-\tfor (s = str; *s; s = next_line(s)) {\n+\t/* Assume the naive end of the input is already what we want. */\n+\tend = strlen(input);\n+\n+\t/* Optionally skip over any patch part (\"---\" line and below). */\n+\tfor (s = input; *s; s = next_line(s)) {\n \t\tconst char *v;\n \n-\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n-\t\t\treturn s - str;\n+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n+\t\t\tend = s - input;\n+\t\t\tbreak;\n+\t\t}\n \t}\n \n-\treturn s - str;\n+\t/* Skip over other ignorable bits. */\n+\treturn end - ignored_log_message_bytes(input, end);\n }\n \n /*\n@@ -925,12 +951,6 @@ continue_outer_loop:\n \treturn len;\n }\n \n-/* Return the position of the end of the trailers. */\n-static size_t find_trailer_end(const char *buf, size_t len)\n-{\n-\treturn len - ignored_log_message_bytes(buf, len);\n-}\n-\n static int ends_with_blank_line(const char *buf, size_t len)\n {\n \tssize_t ll = last_line(buf, len);\n@@ -1101,7 +1121,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint patch_start, trailer_end, trailer_start;\n+\tint end_of_log_message, trailer_start;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1109,16 +1129,11 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tensure_configured();\n \n-\tif (opts->no_divider)\n-\t\tpatch_start = strlen(str);\n-\telse\n-\t\tpatch_start = find_patch_start(str);\n-\n-\ttrailer_end = find_trailer_end(str, patch_start);\n-\ttrailer_start = find_trailer_start(str, trailer_end);\n+\tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n+\ttrailer_start = find_trailer_start(str, end_of_log_message);\n \n \ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t trailer_end - trailer_start,\n+\t\t\t\t\t end_of_log_message - trailer_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1141,7 +1156,7 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_start);\n \tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + trailer_end;\n+\tinfo->trailer_end = str + end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n-- \ngitgitgadget\n\n"},{"id":"482330","messageId":"796e47c1e5fb50a8adb4cf803320de926912a8ad.1695709372.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","subject":"[PATCH v4 3/4] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-26T06:22:51Z","receivedAt":"2023-09-26T06:23:09Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously these fields in the trailer_info struct were of type \"const\nchar *\" and pointed to positions in the input string directly (to the\nstart and end positions of the trailer block).\n\nUse offsets to make the intended usage less ambiguous. We only need to\nreference the input string in format_trailer_info(), so update that\nfunction to take a pointer to the input.\n\nWhile we're at it, rename trailer_start to trailer_block_start to be\nmore explicit about these offsets (that they are for the entire trailer\nblock including other trailers). Ditto for trailer_end.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n sequencer.c |  2 +-\n trailer.c   | 29 ++++++++++++++---------------\n trailer.h   | 13 ++++++++-----\n 3 files changed, 23 insertions(+), 21 deletions(-)\n\ndiff --git a/sequencer.c b/sequencer.c\nindex d584cac8ed9..8707a92204f 100644\n--- a/sequencer.c\n+++ b/sequencer.c\n@@ -345,7 +345,7 @@ static int has_conforming_footer(struct strbuf *sb, struct strbuf *sob,\n \tif (ignore_footer)\n \t\tsb->buf[sb->len - ignore_footer] = saved_char;\n \n-\tif (info.trailer_start == info.trailer_end)\n+\tif (info.trailer_block_start == info.trailer_block_end)\n \t\treturn 0;\n \n \tfor (i = 0; i < info.trailer_nr; i++)\ndiff --git a/trailer.c b/trailer.c\nindex 96cb285a4ea..3dc2faa969c 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -856,7 +856,7 @@ static size_t find_end_of_log_message(const char *input, int no_divider)\n  * Return the position of the first trailer line or len if there are no\n  * trailers.\n  */\n-static size_t find_trailer_start(const char *buf, size_t len)\n+static size_t find_trailer_block_start(const char *buf, size_t len)\n {\n \tconst char *s;\n \tssize_t end_of_title, l;\n@@ -1072,7 +1072,6 @@ void process_trailers(const char *file,\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n \tstruct trailer_info info;\n-\tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n \tensure_configured();\n@@ -1083,11 +1082,10 @@ void process_trailers(const char *file,\n \t\toutfile = create_in_place_tempfile(file);\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n-\ttrailer_end = info.trailer_end - sb.buf;\n \n \t/* Print the lines before the trailers */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1109,7 +1107,7 @@ void process_trailers(const char *file,\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + trailer_end, 1, sb.len - trailer_end, outfile);\n+\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1121,7 +1119,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint end_of_log_message, trailer_start;\n+\tsize_t end_of_log_message = 0, trailer_block_start = 0;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1130,10 +1128,10 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tensure_configured();\n \n \tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n-\ttrailer_start = find_trailer_start(str, end_of_log_message);\n+\ttrailer_block_start = find_trailer_block_start(str, end_of_log_message);\n \n-\ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t end_of_log_message - trailer_start,\n+\ttrailer_lines = strbuf_split_buf(str + trailer_block_start,\n+\t\t\t\t\t end_of_log_message - trailer_block_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1154,9 +1152,9 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tstrbuf_list_free(trailer_lines);\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n-\t\t\t\t\t\t\t       trailer_start);\n-\tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + end_of_log_message;\n+\t\t\t\t\t\t\t       trailer_block_start);\n+\tinfo->trailer_block_start = trailer_block_start;\n+\tinfo->trailer_block_end = end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n@@ -1171,6 +1169,7 @@ void trailer_info_release(struct trailer_info *info)\n \n static void format_trailer_info(struct strbuf *out,\n \t\t\t\tconst struct trailer_info *info,\n+\t\t\t\tconst char *msg,\n \t\t\t\tconst struct process_trailer_options *opts)\n {\n \tsize_t origlen = out->len;\n@@ -1180,8 +1179,8 @@ static void format_trailer_info(struct strbuf *out,\n \tif (!opts->only_trailers && !opts->unfold && !opts->filter &&\n \t    !opts->separator && !opts->key_only && !opts->value_only &&\n \t    !opts->key_value_separator) {\n-\t\tstrbuf_add(out, info->trailer_start,\n-\t\t\t   info->trailer_end - info->trailer_start);\n+\t\tstrbuf_add(out, msg + info->trailer_block_start,\n+\t\t\t   info->trailer_block_end - info->trailer_block_start);\n \t\treturn;\n \t}\n \n@@ -1235,7 +1234,7 @@ void format_trailers_from_commit(struct strbuf *out, const char *msg,\n \tstruct trailer_info info;\n \n \ttrailer_info_get(&info, msg, opts);\n-\tformat_trailer_info(out, &info, opts);\n+\tformat_trailer_info(out, &info, msg, opts);\n \ttrailer_info_release(&info);\n }\n \ndiff --git a/trailer.h b/trailer.h\nindex ab2cd017567..70d7b8bf1d8 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -32,16 +32,19 @@ int trailer_set_if_missing(enum trailer_if_missing *item, const char *value);\n struct trailer_info {\n \t/*\n \t * True if there is a blank line before the location pointed to by\n-\t * trailer_start.\n+\t * trailer_block_start.\n \t */\n \tint blank_line_before_trailer;\n \n \t/*\n-\t * Pointers to the start and end of the trailer block found. If there\n-\t * is no trailer block found, these 2 pointers point to the end of the\n-\t * input string.\n+\t * Offsets to the trailer block start and end positions in the input\n+\t * string. If no trailer block is found, these are both set to the\n+\t * \"true\" end of the input, per find_true_end_of_input().\n+\t *\n+\t * NOTE: This will be changed so that these point to 0 in the next\n+\t * patch if no trailers are found.\n \t */\n-\tconst char *trailer_start, *trailer_end;\n+\tsize_t trailer_block_start, trailer_block_end;\n \n \t/*\n \t * Array of trailers found.\n-- \ngitgitgadget\n\n"},{"id":"482331","messageId":"64e1bd4e4be6d5f59b17986601aa2b7285362937.1695709372.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","subject":"[PATCH v4 4/4] trailer: only use trailer_block_* variables if trailers were found","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-09-26T06:22:52Z","receivedAt":"2023-09-26T06:23:18Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, these variables were overloaded to act as the end of the log\nmessage even if no trailers were found.\n\nRemove the overloaded meaning by adding a new end_of_log_message field\nto the trailer_info struct. The trailer_info struct consumers now only\nrefer to the trailer_block_start and trailer_block_end fields if\ntrailers were found (trailer_nr > 0), and otherwise refer to the\nend_of_log_message.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 31 +++++++++++++++++++++++--------\n trailer.h | 12 +++++++-----\n 2 files changed, 30 insertions(+), 13 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 3dc2faa969c..c11839ae365 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -1083,9 +1083,14 @@ void process_trailers(const char *file,\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n \n-\t/* Print the lines before the trailers */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n+\t/* Print the lines before the trailers (if any) as is. */\n+\tif (!opts->only_trailers) {\n+\t\tif (info.trailer_nr) {\n+\t\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n+\t\t} else {\n+\t\t\tfwrite(sb.buf, 1, info.end_of_log_message, outfile);\n+\t\t}\n+\t}\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1105,9 +1110,14 @@ void process_trailers(const char *file,\n \tfree_all(&head);\n \ttrailer_info_release(&info);\n \n-\t/* Print the lines after the trailers as is */\n-\tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n+\t/* Print the lines after the trailers (if any) as is. */\n+\tif (!opts->only_trailers) {\n+\t\tif (info.trailer_nr) {\n+\t\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n+\t\t} else {\n+\t\t\tfwrite(sb.buf + info.end_of_log_message, 1, sb.len - info.end_of_log_message, outfile);\n+\t\t}\n+\t}\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1153,8 +1163,13 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_block_start);\n-\tinfo->trailer_block_start = trailer_block_start;\n-\tinfo->trailer_block_end = end_of_log_message;\n+\tinfo->trailer_block_start = 0;\n+\tinfo->trailer_block_end = 0;\n+\tif (nr) {\n+\t\tinfo->trailer_block_start = trailer_block_start;\n+\t\tinfo->trailer_block_end = end_of_log_message;\n+\t}\n+\tinfo->end_of_log_message = end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\ndiff --git a/trailer.h b/trailer.h\nindex 70d7b8bf1d8..d1e8751952b 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -38,14 +38,16 @@ struct trailer_info {\n \n \t/*\n \t * Offsets to the trailer block start and end positions in the input\n-\t * string. If no trailer block is found, these are both set to the\n-\t * \"true\" end of the input, per find_true_end_of_input().\n-\t *\n-\t * NOTE: This will be changed so that these point to 0 in the next\n-\t * patch if no trailers are found.\n+\t * string. If no trailer block is found, these are set to 0.\n \t */\n \tsize_t trailer_block_start, trailer_block_end;\n \n+\t/*\n+\t * Offset to the end of the log message in the input (may not be the\n+\t * same as the end of the input).\n+\t */\n+\tsize_t end_of_log_message;\n+\n \t/*\n \t * Array of trailers found.\n \t */\n-- \ngitgitgadget\n"},{"id":"482445","messageId":"20230928231644.3529127-1-jonathantanmy@google.com","threadId":"60068","inReplyTo":"c904caba7e17b6f2784933e9f18634ea66f28537.1695709372.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 2/4] trailer: find the end of the log message","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2023-09-28T23:16:44Z","receivedAt":"2023-09-28T23:16:56Z","isPatch":true,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> From: Linus Arver <linusa@google.com>\n> \n> Previously, trailer_info_get() computed the trailer block end position\n> by\n> \n> (1) checking for the opts->no_divider flag and optionally calling\n>     find_patch_start() to find the \"patch start\" location (patch_start), and\n> (2) calling find_trailer_end() to find the end of the trailer block\n>     using patch_start as a guide, saving the return value into\n>     \"trailer_end\".\n> \n> The logic in (1) was awkward because the variable \"patch_start\" is\n> misleading if there is no patch in the input. The logic in (2) was\n> misleading because it could be the case that no trailers are in the\n> input (yet we are setting a \"trailer_end\" variable before even searching\n> for trailers, which happens later in find_trailer_start()). The name\n> \"find_trailer_end\" was misleading because that function did not look for\n> any trailer block itself --- instead it just computed the end position\n> of the log message in the input where the end of the trailer block (if\n> it exists) would be (because trailer blocks must always come after the\n> end of the log message).\n\nI might be biased since I wrote the text in question in 022349c3b0\n(trailer: avoid unnecessary splitting on lines, 2016-11-02), but the\nconcept of patch_start and trailer_end being where the patch would start\nand where the trailer would end (if they were present) goes back to\n2013d8505d (trailer: parse trailers from file or stdin, 2014-10-13). I\ndon't remember exactly my thoughts in 2016, but today, this makes sense\nto me.\n\nAs it is, the reader still needs to know that trailer_start is where\nthe trailer would start if it was present, and then I think it's quite\nnatural to have trailer_end be where the trailer would end if it was\npresent.\n\nI believe the code is simpler this way, because trailer absence now no\nlonger needs to be special-cased when we use these variables (or maybe\nthey sometimes do, but not as often, since code that writes to the end\nof the trailers, for example, can now just write at trailer_end instead\nof having to check whether trailers exist). Same comment for patch 4\nregarding using the special value 0 if no trailers are found (I think\nthe existing code makes more sense).\n\n> Combine the logic in (1) and (2) together into find_patch_start() by\n> renaming it to find_end_of_log_message(). The end of the log message is\n> the starting point which find_trailer_start() needs to start searching\n> backward to parse individual trailers (if any).\n\nHaving said that, if patch_start is too confusing for whatever reason,\nthis refactoring makes sense. (Avoid the confusing name by avoiding\nneeding to name it in the first place.)\n\n> -static size_t find_patch_start(const char *str)\n> +static size_t find_end_of_log_message(const char *input, int no_divider)\n>  {\n> +\tsize_t end;\n> +\n>  \tconst char *s;\n>  \n> -\tfor (s = str; *s; s = next_line(s)) {\n> +\t/* Assume the naive end of the input is already what we want. */\n> +\tend = strlen(input);\n> +\n> +\t/* Optionally skip over any patch part (\"---\" line and below). */\n> +\tfor (s = input; *s; s = next_line(s)) {\n>  \t\tconst char *v;\n>  \n> -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n> -\t\t\treturn s - str;\n> +\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n> +\t\t\tend = s - input;\n> +\t\t\tbreak;\n> +\t\t}\n>  \t}\n>  \n> -\treturn s - str;\n> +\t/* Skip over other ignorable bits. */\n> +\treturn end - ignored_log_message_bytes(input, end);\n>  }\n\nThis sometimes redundantly calls strlen and sometimes redundantly loops.\nI think it's better to do as the code currently does - so, have a big\nif/else at the beginning of this function that checks no_divider.\n\n"},{"id":"483534","messageId":"owlymsweqgx4.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"20230928231644.3529127-1-jonathantanmy@google.com","subject":"Re: [PATCH v4 2/4] trailer: find the end of the log message","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-10-20T00:24:55Z","receivedAt":"2023-10-20T00:24:59Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Hi Jonathan, it's been a while because I was on vacation. I've forgotten\nabout most of the intricacies of this patch (I think this was a good\nthing, read on below).\n\nJonathan Tan <jonathantanmy@google.com> writes:\n\n> \"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> From: Linus Arver <linusa@google.com>\n>> \n>> Previously, trailer_info_get() computed the trailer block end position\n>> by\n>> \n>> (1) checking for the opts->no_divider flag and optionally calling\n>>     find_patch_start() to find the \"patch start\" location (patch_start), and\n>> (2) calling find_trailer_end() to find the end of the trailer block\n>>     using patch_start as a guide, saving the return value into\n>>     \"trailer_end\".\n>> \n>> The logic in (1) was awkward because the variable \"patch_start\" is\n>> misleading if there is no patch in the input. The logic in (2) was\n>> misleading because it could be the case that no trailers are in the\n>> input (yet we are setting a \"trailer_end\" variable before even searching\n>> for trailers, which happens later in find_trailer_start()). The name\n>> \"find_trailer_end\" was misleading because that function did not look for\n>> any trailer block itself --- instead it just computed the end position\n>> of the log message in the input where the end of the trailer block (if\n>> it exists) would be (because trailer blocks must always come after the\n>> end of the log message).\n>\n> [...]\n>\n> As it is, the reader still needs to know that trailer_start is where\n> the trailer would start if it was present, and then I think it's quite\n> natural to have trailer_end be where the trailer would end if it was\n> present.\n>\n> I believe the code is simpler this way, because trailer absence now no\n> longer needs to be special-cased when we use these variables (or maybe\n> they sometimes do, but not as often, since code that writes to the end\n> of the trailers, for example, can now just write at trailer_end instead\n> of having to check whether trailers exist). Same comment for patch 4\n> regarding using the special value 0 if no trailers are found (I think\n> the existing code makes more sense).\n\nI think the root cause of my confusion with this codebase is due to the\nvariables being named as if the things they refer to exist, but without\nany guarantees that they indeed exist. This applies to \"patch_start\"\n(the patch part might not be present in the input),\n\"trailer_{start,end}\" (trailers block might not exist (yet)). IOW these\nvariables are named as if the intent is to always add new trailers into\nthe input, which may not be the case (we have \"--parse\", after all).\n\nLooking again at patch 4, I'm now leaning toward dropping it. Other\nthan the reasons you cited, we also add a new struct field which by\nitself does not add new information (the information can already be\ndeduced from the other fields). In the near future I want to simplify\nthe data structures as much as possible, and adding a new field seems to\ngo against this desire of mine.\n\n>> Combine the logic in (1) and (2) together into find_patch_start() by\n>> renaming it to find_end_of_log_message(). The end of the log message is\n>> the starting point which find_trailer_start() needs to start searching\n>> backward to parse individual trailers (if any).\n>\n> Having said that, if patch_start is too confusing for whatever reason,\n> this refactoring makes sense. (Avoid the confusing name by avoiding\n> needing to name it in the first place.)\n\nI think the existing code is confusing, and would prefer to keep this\npatch.\n\n>> -static size_t find_patch_start(const char *str)\n>> +static size_t find_end_of_log_message(const char *input, int no_divider)\n>>  {\n>> +\tsize_t end;\n>> +\n>>  \tconst char *s;\n>>  \n>> -\tfor (s = str; *s; s = next_line(s)) {\n>> +\t/* Assume the naive end of the input is already what we want. */\n>> +\tend = strlen(input);\n>> +\n>> +\t/* Optionally skip over any patch part (\"---\" line and below). */\n>> +\tfor (s = input; *s; s = next_line(s)) {\n>>  \t\tconst char *v;\n>>  \n>> -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n>> -\t\t\treturn s - str;\n>> +\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n>> +\t\t\tend = s - input;\n>> +\t\t\tbreak;\n>> +\t\t}\n>>  \t}\n>>  \n>> -\treturn s - str;\n>> +\t/* Skip over other ignorable bits. */\n>> +\treturn end - ignored_log_message_bytes(input, end);\n>>  }\n>\n> This sometimes redundantly calls strlen and sometimes redundantly loops.\n> I think it's better to do as the code currently does - so, have a big\n> if/else at the beginning of this function that checks no_divider.\n\nWill update, thanks.\n"},{"id":"483536","messageId":"xmqqpm1at9il.fsf@gitster.g","threadId":"60068","inReplyTo":"owlymsweqgx4.fsf@fine.c.googlers.com","subject":"Re: [PATCH v4 2/4] trailer: find the end of the log message","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-20T00:36:34Z","receivedAt":"2023-10-20T00:36:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Arver <linusa@google.com> writes:\n\n> Hi Jonathan, it's been a while because I was on vacation. I've forgotten\n> about most of the intricacies of this patch (I think this was a good\n> thing, read on below).\n\nWelcome back ;-).\n\n> Will update, thanks.\n\nThanks.\n\n"},{"id":"483592","messageId":"pull.1563.v5.git.1697828495.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v4.git.1695709372.gitgitgadget@gmail.com","subject":"[PATCH v5 0/3] Trailer readability cleanups","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-10-20T19:01:32Z","receivedAt":"2023-10-20T19:01:40Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"These patches were created while digging into the trailer code to better\nunderstand how it works, in preparation for making the trailer.{c,h} files\nas small as possible to make them available as a library for external users.\nThis series was originally created as part of [1], but are sent here\nseparately because the changes here are arguably more subjective in nature.\n\nThese patches do not add or change any features. Instead, their goal is to\nmake the code easier to understand for new contributors (like myself), by\nmaking various cleanups and improvements. Ultimately, my hope is that with\nsuch cleanups, we are better positioned to make larger changes (especially\nthe broader libification effort, as in \"Introduce Git Standard Library\"\n[2]).\n\n\nUpdates in v5\n=============\n\n * Patch 4 (\"trailer: only use trailer_block_* variables if trailers were\n   found\") has been dropped.\n * Patch 2 returns early if \"--no-divider\" is true, avoiding unnecessary\n   loop iterations (thanks Jonathan).\n * Added missing Reported-by trailer for Patch 3 (it was originally Glen's\n   idea to use offsets).\n * Patch 3: Fixed typo in \"trailer.h\" that referred to an obsolete function\n   name (\"find_true_end_of_input()\", instead of\n   \"find_end_of_log_message()\").\n\n\nUpdates in v4\n=============\n\n * The first 3 patches in v3 were merged into 'master'. Necessarily, those 3\n   patches have been dropped.\n * Patch 4 in v3 (\"trailer: rename *_DEFAULT enums to *_UNSPECIFIED\") has\n   been dropped, as well as Patch 9 in v3 (\"trailer: make stack variable\n   names match field names\"). These were dropped to simplify this series for\n   what I think is the more immediate, important change (see next bullet\n   point).\n * Patches 5-8 in v3 are the only ones remaining in this series. They still\n   solely deal with --no-divider and trailer block start/end cleanups.\n\n\nUpdates in v3\n=============\n\n * Patches 4 and 6 (--no-divider and trailer block start/end cleanups) have\n   been reorganized to Patches 5-8. This ended up touching commit.c in a\n   minor way, but otherwise all of the changes here are cleanups and do not\n   change any behavior.\n\n\nUpdates in v2\n=============\n\n * Patch 1: Drop the use of a #define. Instead just use an anonymous struct\n   named internal.\n * Patch 2: Don't free info out parameter inside parse_trailers(). Instead\n   free it from the caller, process_trailers(). Update comment in\n   parse_trailers().\n * Patch 3: Reword commit message.\n * Patch 4: Mention be3d654343 (commit: pass --no-divider to\n   interpret-trailers, 2023-06-17) in commit message.\n * Added Patch 6 to make trailer_info use offsets for trailer_start and\n   trailer_end (thanks to Glen Choo for the suggestion).\n\n[1]\nhttps://lore.kernel.org/git/pull.1564.git.1691210737.gitgitgadget@gmail.com/T/#mb044012670663d8eb7a548924bbcc933bef116de\n[2]\nhttps://lore.kernel.org/git/20230627195251.1973421-1-calvinwan@google.com/\n[3]\nhttps://lore.kernel.org/git/pull.1149.git.1677143700.gitgitgadget@gmail.com/\n[4]\nhttps://lore.kernel.org/git/6b4cb31b17077181a311ca87e82464a1e2ad67dd.1686797630.git.gitgitgadget@gmail.com/\n[5]\nhttps://lore.kernel.org/git/pull.1563.git.1691211879.gitgitgadget@gmail.com/T/#m0131f9829c35d8e0103ffa88f07d8e0e43dd732c\n\nLinus Arver (3):\n  commit: ignore_non_trailer computes number of bytes to ignore\n  trailer: find the end of the log message\n  trailer: use offsets for trailer_start/trailer_end\n\n builtin/commit.c |  2 +-\n builtin/merge.c  |  2 +-\n commit.c         |  2 +-\n commit.h         |  4 +--\n sequencer.c      |  2 +-\n trailer.c        | 85 +++++++++++++++++++++++++++++-------------------\n trailer.h        | 10 +++---\n 7 files changed, 62 insertions(+), 45 deletions(-)\n\n\nbase-commit: bcb6cae2966cc407ca1afc77413b3ef11103c175\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1563%2Flistx%2Ftrailer-libification-prep-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1563/listx/trailer-libification-prep-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/1563\n\nRange-diff vs v4:\n\n 1:  4ce5cf77005 = 1:  4ce5cf77005 commit: ignore_non_trailer computes number of bytes to ignore\n 2:  c904caba7e1 ! 2:  ce25420db29 trailer: find the end of the log message\n     @@ Commit message\n          the starting point which find_trailer_start() needs to start searching\n          backward to parse individual trailers (if any).\n      \n     +    Helped-by: Jonathan Tan <jonathantanmy@google.com>\n          Helped-by: Junio C Hamano <gitster@pobox.com>\n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n     @@ trailer.c: static ssize_t last_line(const char *buf, size_t len)\n      +static size_t find_end_of_log_message(const char *input, int no_divider)\n       {\n      +\tsize_t end;\n     -+\n       \tconst char *s;\n       \n      -\tfor (s = str; *s; s = next_line(s)) {\n      +\t/* Assume the naive end of the input is already what we want. */\n      +\tend = strlen(input);\n      +\n     ++\tif (no_divider) {\n     ++\t\treturn end;\n     ++\t}\n     ++\n      +\t/* Optionally skip over any patch part (\"---\" line and below). */\n      +\tfor (s = input; *s; s = next_line(s)) {\n       \t\tconst char *v;\n       \n      -\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n      -\t\t\treturn s - str;\n     -+\t\tif (!no_divider && skip_prefix(s, \"---\", &v) && isspace(*v)) {\n     ++\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v)) {\n      +\t\t\tend = s - input;\n      +\t\t\tbreak;\n      +\t\t}\n 3:  796e47c1e5f ! 3:  e3a7b150241 trailer: use offsets for trailer_start/trailer_end\n     @@ Commit message\n          more explicit about these offsets (that they are for the entire trailer\n          block including other trailers). Ditto for trailer_end.\n      \n     +    Reported-by: Glen Choo <glencbz@gmail.com>\n          Signed-off-by: Linus Arver <linusa@google.com>\n      \n       ## sequencer.c ##\n     @@ trailer.h: int trailer_set_if_missing(enum trailer_if_missing *item, const char\n      -\t * input string.\n      +\t * Offsets to the trailer block start and end positions in the input\n      +\t * string. If no trailer block is found, these are both set to the\n     -+\t * \"true\" end of the input, per find_true_end_of_input().\n     -+\t *\n     -+\t * NOTE: This will be changed so that these point to 0 in the next\n     -+\t * patch if no trailers are found.\n     ++\t * \"true\" end of the input (find_end_of_log_message()).\n       \t */\n      -\tconst char *trailer_start, *trailer_end;\n      +\tsize_t trailer_block_start, trailer_block_end;\n 4:  64e1bd4e4be < -:  ----------- trailer: only use trailer_block_* variables if trailers were found\n\n-- \ngitgitgadget\n"},{"id":"483593","messageId":"4ce5cf77005eb8c6da243777b3c29103add7ddbd.1697828495.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v5.git.1697828495.gitgitgadget@gmail.com","subject":"[PATCH v5 1/3] commit: ignore_non_trailer computes number of bytes to ignore","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-10-20T19:01:33Z","receivedAt":"2023-10-20T19:01:41Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nignore_non_trailer() returns the _number of bytes_ that should be\nignored from the end of the log message. It does not by itself \"ignore\"\nanything.\n\nRename this function to remove the leading \"ignore\" verb, to sound more\nlike a quantity than an action.\n\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n builtin/commit.c | 2 +-\n builtin/merge.c  | 2 +-\n commit.c         | 2 +-\n commit.h         | 4 ++--\n trailer.c        | 2 +-\n 5 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484..d1785d32db1 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -900,7 +900,7 @@ static int prepare_to_commit(const char *index_file, const char *prefix,\n \t\tstrbuf_stripspace(&sb, '\\0');\n \n \tif (signoff)\n-\t\tappend_signoff(&sb, ignore_non_trailer(sb.buf, sb.len), 0);\n+\t\tappend_signoff(&sb, ignored_log_message_bytes(sb.buf, sb.len), 0);\n \n \tif (fwrite(sb.buf, 1, sb.len, s->fp) < sb.len)\n \t\tdie_errno(_(\"could not write commit template\"));\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex 545da0c8a11..c654a29fe85 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -870,7 +870,7 @@ static void prepare_to_commit(struct commit_list *remoteheads)\n \t\t\t\t_(no_scissors_editor_comment), comment_line_char);\n \t}\n \tif (signoff)\n-\t\tappend_signoff(&msg, ignore_non_trailer(msg.buf, msg.len), 0);\n+\t\tappend_signoff(&msg, ignored_log_message_bytes(msg.buf, msg.len), 0);\n \twrite_merge_heads(remoteheads);\n \twrite_file_buf(git_path_merge_msg(the_repository), msg.buf, msg.len);\n \tif (run_commit_hook(0 < option_edit, get_index_file(), NULL,\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2..4440fbabb83 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1769,7 +1769,7 @@ const char *find_commit_header(const char *msg, const char *key, size_t *out_len\n  * Returns the number of bytes from the tail to ignore, to be fed as\n  * the second parameter to append_signoff().\n  */\n-size_t ignore_non_trailer(const char *buf, size_t len)\n+size_t ignored_log_message_bytes(const char *buf, size_t len)\n {\n \tsize_t boc = 0;\n \tsize_t bol = 0;\ndiff --git a/commit.h b/commit.h\nindex 28928833c54..1cc872f225f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -294,8 +294,8 @@ const char *find_header_mem(const char *msg, size_t len,\n const char *find_commit_header(const char *msg, const char *key,\n \t\t\t       size_t *out_len);\n \n-/* Find the end of the log message, the right place for a new trailer. */\n-size_t ignore_non_trailer(const char *buf, size_t len);\n+/* Find the number of bytes to ignore from the end of a log message. */\n+size_t ignored_log_message_bytes(const char *buf, size_t len);\n \n typedef int (*each_mergetag_fn)(struct commit *commit, struct commit_extra_header *extra,\n \t\t\t\tvoid *cb_data);\ndiff --git a/trailer.c b/trailer.c\nindex b6de5d9cb2d..3c54b38a85a 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -928,7 +928,7 @@ continue_outer_loop:\n /* Return the position of the end of the trailers. */\n static size_t find_trailer_end(const char *buf, size_t len)\n {\n-\treturn len - ignore_non_trailer(buf, len);\n+\treturn len - ignored_log_message_bytes(buf, len);\n }\n \n static int ends_with_blank_line(const char *buf, size_t len)\n-- \ngitgitgadget\n\n"},{"id":"483594","messageId":"ce25420db29c9953095db652584dbed4e35d67ad.1697828495.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v5.git.1697828495.gitgitgadget@gmail.com","subject":"[PATCH v5 2/3] trailer: find the end of the log message","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-10-20T19:01:34Z","receivedAt":"2023-10-20T19:01:42Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously, trailer_info_get() computed the trailer block end position\nby\n\n(1) checking for the opts->no_divider flag and optionally calling\n    find_patch_start() to find the \"patch start\" location (patch_start), and\n(2) calling find_trailer_end() to find the end of the trailer block\n    using patch_start as a guide, saving the return value into\n    \"trailer_end\".\n\nThe logic in (1) was awkward because the variable \"patch_start\" is\nmisleading if there is no patch in the input. The logic in (2) was\nmisleading because it could be the case that no trailers are in the\ninput (yet we are setting a \"trailer_end\" variable before even searching\nfor trailers, which happens later in find_trailer_start()). The name\n\"find_trailer_end\" was misleading because that function did not look for\nany trailer block itself --- instead it just computed the end position\nof the log message in the input where the end of the trailer block (if\nit exists) would be (because trailer blocks must always come after the\nend of the log message).\n\nCombine the logic in (1) and (2) together into find_patch_start() by\nrenaming it to find_end_of_log_message(). The end of the log message is\nthe starting point which find_trailer_start() needs to start searching\nbackward to parse individual trailers (if any).\n\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n trailer.c | 64 +++++++++++++++++++++++++++++++++++--------------------\n 1 file changed, 41 insertions(+), 23 deletions(-)\n\ndiff --git a/trailer.c b/trailer.c\nindex 3c54b38a85a..70c81fda710 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -809,21 +809,50 @@ static ssize_t last_line(const char *buf, size_t len)\n }\n \n /*\n- * Return the position of the start of the patch or the length of str if there\n- * is no patch in the message.\n+ * Find the end of the log message as an offset from the start of the input\n+ * (where callers of this function are interested in looking for a trailers\n+ * block in the same input). We have to consider two categories of content that\n+ * can come at the end of the input which we want to ignore (because they don't\n+ * belong in the log message):\n+ *\n+ * (1) the \"patch part\" which begins with a \"---\" divider and has patch\n+ * information (like the output of git-format-patch), and\n+ *\n+ * (2) any trailing comment lines, blank lines like in the output of \"git\n+ * commit -v\", or stuff below the \"cut\" (scissor) line.\n+ *\n+ * As a formula, the situation looks like this:\n+ *\n+ *     INPUT = LOG MESSAGE + IGNORED\n+ *\n+ * where IGNORED can be either of the two categories described above. It may be\n+ * that there is nothing to ignore. Now it may be the case that the LOG MESSAGE\n+ * contains a trailer block, but that's not the concern of this function.\n  */\n-static size_t find_patch_start(const char *str)\n+static size_t find_end_of_log_message(const char *input, int no_divider)\n {\n+\tsize_t end;\n \tconst char *s;\n \n-\tfor (s = str; *s; s = next_line(s)) {\n+\t/* Assume the naive end of the input is already what we want. */\n+\tend = strlen(input);\n+\n+\tif (no_divider) {\n+\t\treturn end;\n+\t}\n+\n+\t/* Optionally skip over any patch part (\"---\" line and below). */\n+\tfor (s = input; *s; s = next_line(s)) {\n \t\tconst char *v;\n \n-\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v))\n-\t\t\treturn s - str;\n+\t\tif (skip_prefix(s, \"---\", &v) && isspace(*v)) {\n+\t\t\tend = s - input;\n+\t\t\tbreak;\n+\t\t}\n \t}\n \n-\treturn s - str;\n+\t/* Skip over other ignorable bits. */\n+\treturn end - ignored_log_message_bytes(input, end);\n }\n \n /*\n@@ -925,12 +954,6 @@ continue_outer_loop:\n \treturn len;\n }\n \n-/* Return the position of the end of the trailers. */\n-static size_t find_trailer_end(const char *buf, size_t len)\n-{\n-\treturn len - ignored_log_message_bytes(buf, len);\n-}\n-\n static int ends_with_blank_line(const char *buf, size_t len)\n {\n \tssize_t ll = last_line(buf, len);\n@@ -1101,7 +1124,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint patch_start, trailer_end, trailer_start;\n+\tint end_of_log_message, trailer_start;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1109,16 +1132,11 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \n \tensure_configured();\n \n-\tif (opts->no_divider)\n-\t\tpatch_start = strlen(str);\n-\telse\n-\t\tpatch_start = find_patch_start(str);\n-\n-\ttrailer_end = find_trailer_end(str, patch_start);\n-\ttrailer_start = find_trailer_start(str, trailer_end);\n+\tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n+\ttrailer_start = find_trailer_start(str, end_of_log_message);\n \n \ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t trailer_end - trailer_start,\n+\t\t\t\t\t end_of_log_message - trailer_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1141,7 +1159,7 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n \t\t\t\t\t\t\t       trailer_start);\n \tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + trailer_end;\n+\tinfo->trailer_end = str + end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n-- \ngitgitgadget\n\n"},{"id":"483595","messageId":"e3a7b150241c2d997026b8ccf7c88ecfecdfe4e7.1697828495.git.gitgitgadget@gmail.com","threadId":"60068","inReplyTo":"pull.1563.v5.git.1697828495.gitgitgadget@gmail.com","subject":"[PATCH v5 3/3] trailer: use offsets for trailer_start/trailer_end","fromName":"Linus Arver via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2023-10-20T19:01:35Z","receivedAt":"2023-10-20T19:01:43Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"From: Linus Arver <linusa@google.com>\n\nPreviously these fields in the trailer_info struct were of type \"const\nchar *\" and pointed to positions in the input string directly (to the\nstart and end positions of the trailer block).\n\nUse offsets to make the intended usage less ambiguous. We only need to\nreference the input string in format_trailer_info(), so update that\nfunction to take a pointer to the input.\n\nWhile we're at it, rename trailer_start to trailer_block_start to be\nmore explicit about these offsets (that they are for the entire trailer\nblock including other trailers). Ditto for trailer_end.\n\nReported-by: Glen Choo <glencbz@gmail.com>\nSigned-off-by: Linus Arver <linusa@google.com>\n---\n sequencer.c |  2 +-\n trailer.c   | 29 ++++++++++++++---------------\n trailer.h   | 10 +++++-----\n 3 files changed, 20 insertions(+), 21 deletions(-)\n\ndiff --git a/sequencer.c b/sequencer.c\nindex d584cac8ed9..8707a92204f 100644\n--- a/sequencer.c\n+++ b/sequencer.c\n@@ -345,7 +345,7 @@ static int has_conforming_footer(struct strbuf *sb, struct strbuf *sob,\n \tif (ignore_footer)\n \t\tsb->buf[sb->len - ignore_footer] = saved_char;\n \n-\tif (info.trailer_start == info.trailer_end)\n+\tif (info.trailer_block_start == info.trailer_block_end)\n \t\treturn 0;\n \n \tfor (i = 0; i < info.trailer_nr; i++)\ndiff --git a/trailer.c b/trailer.c\nindex 70c81fda710..f7dc7c4c008 100644\n--- a/trailer.c\n+++ b/trailer.c\n@@ -859,7 +859,7 @@ static size_t find_end_of_log_message(const char *input, int no_divider)\n  * Return the position of the first trailer line or len if there are no\n  * trailers.\n  */\n-static size_t find_trailer_start(const char *buf, size_t len)\n+static size_t find_trailer_block_start(const char *buf, size_t len)\n {\n \tconst char *s;\n \tssize_t end_of_title, l;\n@@ -1075,7 +1075,6 @@ void process_trailers(const char *file,\n \tLIST_HEAD(head);\n \tstruct strbuf sb = STRBUF_INIT;\n \tstruct trailer_info info;\n-\tsize_t trailer_end;\n \tFILE *outfile = stdout;\n \n \tensure_configured();\n@@ -1086,11 +1085,10 @@ void process_trailers(const char *file,\n \t\toutfile = create_in_place_tempfile(file);\n \n \tparse_trailers(&info, sb.buf, &head, opts);\n-\ttrailer_end = info.trailer_end - sb.buf;\n \n \t/* Print the lines before the trailers */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf, 1, info.trailer_start - sb.buf, outfile);\n+\t\tfwrite(sb.buf, 1, info.trailer_block_start, outfile);\n \n \tif (!opts->only_trailers && !info.blank_line_before_trailer)\n \t\tfprintf(outfile, \"\\n\");\n@@ -1112,7 +1110,7 @@ void process_trailers(const char *file,\n \n \t/* Print the lines after the trailers as is */\n \tif (!opts->only_trailers)\n-\t\tfwrite(sb.buf + trailer_end, 1, sb.len - trailer_end, outfile);\n+\t\tfwrite(sb.buf + info.trailer_block_end, 1, sb.len - info.trailer_block_end, outfile);\n \n \tif (opts->in_place)\n \t\tif (rename_tempfile(&trailers_tempfile, file))\n@@ -1124,7 +1122,7 @@ void process_trailers(const char *file,\n void trailer_info_get(struct trailer_info *info, const char *str,\n \t\t      const struct process_trailer_options *opts)\n {\n-\tint end_of_log_message, trailer_start;\n+\tsize_t end_of_log_message = 0, trailer_block_start = 0;\n \tstruct strbuf **trailer_lines, **ptr;\n \tchar **trailer_strings = NULL;\n \tsize_t nr = 0, alloc = 0;\n@@ -1133,10 +1131,10 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tensure_configured();\n \n \tend_of_log_message = find_end_of_log_message(str, opts->no_divider);\n-\ttrailer_start = find_trailer_start(str, end_of_log_message);\n+\ttrailer_block_start = find_trailer_block_start(str, end_of_log_message);\n \n-\ttrailer_lines = strbuf_split_buf(str + trailer_start,\n-\t\t\t\t\t end_of_log_message - trailer_start,\n+\ttrailer_lines = strbuf_split_buf(str + trailer_block_start,\n+\t\t\t\t\t end_of_log_message - trailer_block_start,\n \t\t\t\t\t '\\n',\n \t\t\t\t\t 0);\n \tfor (ptr = trailer_lines; *ptr; ptr++) {\n@@ -1157,9 +1155,9 @@ void trailer_info_get(struct trailer_info *info, const char *str,\n \tstrbuf_list_free(trailer_lines);\n \n \tinfo->blank_line_before_trailer = ends_with_blank_line(str,\n-\t\t\t\t\t\t\t       trailer_start);\n-\tinfo->trailer_start = str + trailer_start;\n-\tinfo->trailer_end = str + end_of_log_message;\n+\t\t\t\t\t\t\t       trailer_block_start);\n+\tinfo->trailer_block_start = trailer_block_start;\n+\tinfo->trailer_block_end = end_of_log_message;\n \tinfo->trailers = trailer_strings;\n \tinfo->trailer_nr = nr;\n }\n@@ -1174,6 +1172,7 @@ void trailer_info_release(struct trailer_info *info)\n \n static void format_trailer_info(struct strbuf *out,\n \t\t\t\tconst struct trailer_info *info,\n+\t\t\t\tconst char *msg,\n \t\t\t\tconst struct process_trailer_options *opts)\n {\n \tsize_t origlen = out->len;\n@@ -1183,8 +1182,8 @@ static void format_trailer_info(struct strbuf *out,\n \tif (!opts->only_trailers && !opts->unfold && !opts->filter &&\n \t    !opts->separator && !opts->key_only && !opts->value_only &&\n \t    !opts->key_value_separator) {\n-\t\tstrbuf_add(out, info->trailer_start,\n-\t\t\t   info->trailer_end - info->trailer_start);\n+\t\tstrbuf_add(out, msg + info->trailer_block_start,\n+\t\t\t   info->trailer_block_end - info->trailer_block_start);\n \t\treturn;\n \t}\n \n@@ -1238,7 +1237,7 @@ void format_trailers_from_commit(struct strbuf *out, const char *msg,\n \tstruct trailer_info info;\n \n \ttrailer_info_get(&info, msg, opts);\n-\tformat_trailer_info(out, &info, opts);\n+\tformat_trailer_info(out, &info, msg, opts);\n \ttrailer_info_release(&info);\n }\n \ndiff --git a/trailer.h b/trailer.h\nindex ab2cd017567..1644cd05f60 100644\n--- a/trailer.h\n+++ b/trailer.h\n@@ -32,16 +32,16 @@ int trailer_set_if_missing(enum trailer_if_missing *item, const char *value);\n struct trailer_info {\n \t/*\n \t * True if there is a blank line before the location pointed to by\n-\t * trailer_start.\n+\t * trailer_block_start.\n \t */\n \tint blank_line_before_trailer;\n \n \t/*\n-\t * Pointers to the start and end of the trailer block found. If there\n-\t * is no trailer block found, these 2 pointers point to the end of the\n-\t * input string.\n+\t * Offsets to the trailer block start and end positions in the input\n+\t * string. If no trailer block is found, these are both set to the\n+\t * \"true\" end of the input (find_end_of_log_message()).\n \t */\n-\tconst char *trailer_start, *trailer_end;\n+\tsize_t trailer_block_start, trailer_block_end;\n \n \t/*\n \t * Array of trailers found.\n-- \ngitgitgadget\n"},{"id":"483601","messageId":"xmqqr0lpoue3.fsf@gitster.g","threadId":"60068","inReplyTo":"ce25420db29c9953095db652584dbed4e35d67ad.1697828495.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 2/3] trailer: find the end of the log message","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-10-20T21:29:08Z","receivedAt":"2023-10-20T21:29:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Linus Arver via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Linus Arver <linusa@google.com>\n>\n> Previously, trailer_info_get() computed the trailer block end position\n> by\n>\n> (1) checking for the opts->no_divider flag and optionally calling\n>     find_patch_start() to find the \"patch start\" location (patch_start), and\n> (2) calling find_trailer_end() to find the end of the trailer block\n>     using patch_start as a guide, saving the return value into\n>     \"trailer_end\".\n>\n> The logic in (1) was awkward because the variable \"patch_start\" is\n> misleading if there is no patch in the input. The logic in (2) was\n> misleading because it could be the case that no trailers are in the\n> input (yet we are setting a \"trailer_end\" variable before even searching\n> for trailers, which happens later in find_trailer_start()). The name\n> \"find_trailer_end\" was misleading because that function did not look for\n> any trailer block itself --- instead it just computed the end position\n> of the log message in the input where the end of the trailer block (if\n> it exists) would be (because trailer blocks must always come after the\n> end of the log message).\n>\n> Combine the logic in (1) and (2) together into find_patch_start() by\n> renaming it to find_end_of_log_message(). The end of the log message is\n> the starting point which find_trailer_start() needs to start searching\n> backward to parse individual trailers (if any).\n>\n> Helped-by: Jonathan Tan <jonathantanmy@google.com>\n> Helped-by: Junio C Hamano <gitster@pobox.com>\n> Signed-off-by: Linus Arver <linusa@google.com>\n> ---\n>  trailer.c | 64 +++++++++++++++++++++++++++++++++++--------------------\n>  1 file changed, 41 insertions(+), 23 deletions(-)\n>\n> diff --git a/trailer.c b/trailer.c\n> index 3c54b38a85a..70c81fda710 100644\n> --- a/trailer.c\n> +++ b/trailer.c\n> @@ -809,21 +809,50 @@ static ssize_t last_line(const char *buf, size_t len)\n>  }\n>  \n>  /*\n> - * Return the position of the start of the patch or the length of str if there\n> - * is no patch in the message.\n> + * Find the end of the log message as an offset from the start of the input\n> + * (where callers of this function are interested in looking for a trailers\n> + * block in the same input). We have to consider two categories of content that\n> + * can come at the end of the input which we want to ignore (because they don't\n> + * belong in the log message):\n> + *\n> + * (1) the \"patch part\" which begins with a \"---\" divider and has patch\n> + * information (like the output of git-format-patch), and\n> + *\n> + * (2) any trailing comment lines, blank lines like in the output of \"git\n> + * commit -v\", or stuff below the \"cut\" (scissor) line.\n> + *\n> + * As a formula, the situation looks like this:\n> + *\n> + *     INPUT = LOG MESSAGE + IGNORED\n> + *\n> + * where IGNORED can be either of the two categories described above. It may be\n> + * that there is nothing to ignore. Now it may be the case that the LOG MESSAGE\n> + * contains a trailer block, but that's not the concern of this function.\n>   */\n> -static size_t find_patch_start(const char *str)\n> +static size_t find_end_of_log_message(const char *input, int no_divider)\n>  {\n> +\tsize_t end;\n>  \tconst char *s;\n>  \n> -\tfor (s = str; *s; s = next_line(s)) {\n> +\t/* Assume the naive end of the input is already what we want. */\n> +\tend = strlen(input);\n> +\n> +\tif (no_divider) {\n> +\t\treturn end;\n> +\t}\n\nOK.  The early return may make sense, as we are essentially\ndeclaring that everything is the \"INPUT (= message + ignored)\".\n\nYou do not need {braces} around a single-statement block, though.\n\nOther than that, I didn't find anything quesionable in any of the\npatches in this round.  Looking good.\n\nThanks.\n"},{"id":"486145","messageId":"owlyv88hqzlu.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"xmqqr0lpoue3.fsf@gitster.g","subject":"Re: [PATCH v5 2/3] trailer: find the end of the log message","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-12-29T06:42:05Z","receivedAt":"2023-12-29T06:42:08Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\nTL;DR: I'm working on a new approach.\n\nJunio C Hamano <gitster@pobox.com> writes:\n> Other than that, I didn't find anything quesionable in any of the\n> patches in this round.  Looking good.\n\nSo actually, I'm now taking a much more aggressive approach to libifying\nthe trailer subsystem. Instead of incrementally simplifying/improving\nthings as in this series, I think I need to get to the root problem,\nwhich is that the trailer.h API isn't rich enough to make it pleasant\nfor clients to use, including our own builtin/interpret-trailers.c\nclient. That is, the problem we have today is that the trailer subsystem\nis not very ergonomic for internal use, much less external use (outside\nof Git itself).\n\nAs an example, the current API exposes process_trailers() which does a\nwhole bunch of things that only builtin/interpret-trailers.c cares\nabout. Multiple other clients of trailer.h exist in our codebase (e.g.,\nsequencer.c, pretty.c, ref-filter.c) but none of them use\nprocess_trailers().\n\nOne really useful data structure is the trailer_iterator that was\nintroduced in f0939a0eb1 (trailer: add interface for iterating over\ncommit trailers, 2020-09-27). The only problem is that it is not generic\nenough such that interpret-trailers.c can use it.\n\nMy new goal is to introduce a new API in trailer.h so that\ninterpret-trailers.c and everyone else can start using these new data\nstructures and associated functions (while preserving the\ntrailer_iterator interface). So the order of operations should be:\n\n(1) enrich the trailer API (make trailer.h have simpler data structures\n    and practical functions that clients can readily use), and\n(2) make builtin/interpret-trailers.c, and other clients in the Git\n    codebase use this new API.\n\nThis way when the unit test framework selection process is finalized we\ncan\n\n(3) write unit tests for the functions in the (enriched) trailer API,\n\nwhich is one of the major goals for my efforts around this area.\n\nThe work I've started locally for (1) does not depend on this series,\nand I think it'll be cleaner (less churn) that way. So, feel free to\ndrop this series in favor of the forthcoming work described in this\nmessage.\n\nThanks.\n"},{"id":"486164","messageId":"owlysf3kraao.fsf@fine.c.googlers.com","threadId":"60068","inReplyTo":"owlyv88hqzlu.fsf@fine.c.googlers.com","subject":"Re: [PATCH v5 2/3] trailer: find the end of the log message","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2023-12-29T21:03:27Z","receivedAt":"2023-12-29T21:03:29Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\n> (1) enrich the trailer API (make trailer.h have simpler data structures\n>     and practical functions that clients can readily use), and\n> (2) make builtin/interpret-trailers.c, and other clients in the Git\n>     codebase use this new API.\n\nI've done some more thinking/hacking and I'm realizing that changing the\ndata structures significantly as a first step will be difficult to get\nright without being able to unit-test things as we go. As we don't have\nunit tests (sorry, I keep forgetting...), I think changing the shape of\ndata structures right now is a showstopper.\n\nStep (1) will have to be done without changing any of the existing data\nstructures.\n"}]}