{"thread":{"id":"55909","subject":"[PATCH 3/8] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","startedAt":"2021-06-12T11:14:24Z","lastAt":"2021-07-09T10:04:07Z","messageCount":121,"participants":["ZheNing Hu via GitGitGadget","Christian Couder","ZheNing Hu","Junio C Hamano","Ævar Arnfjörð Bjarmason","Bagas Sanjaya","Hariom verma"],"isPatch":true,"patchVersion":1,"patchTotal":8},"messages":[{"id":"427176","messageId":"c99d1d070a182d09013792d724d6a62bb9a7c0a2.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 3/8] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:12Z","receivedAt":"2021-06-12T11:14:24Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc098242..452558ec9575 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be903071..af8c15aef44d 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2546,9 +2546,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2593,7 +2593,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a718965..74fb423fc89f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"427177","messageId":"44ebf75e2e937b76a7c2887cd98da2912240811c.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 6/8] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:15Z","receivedAt":"2021-06-12T11:14:27Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, the following\nmethods are used:\n\n1. Add `cat_file_mode` member in struct `ref_format`, this can\nhelp us reject atoms in verify_ref_format() which cat-file\ncannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n2. Change the type of member `format` in struct `batch_options`\nto `ref_format`, We can add format data in it.\n3. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n4. Use `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists.\n5. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue less than zero, use `die()` to print the error message\nand exit; else return value greater than zero, only print the\nerror message, but not exit.\n6. Let get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\",\nthis can help `format_ref_array_item()` just report that the\nobject is missing without letting Git exit.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have refname.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 250 +++++++-------------------------\n ref-filter.c                   |  15 +-\n ref-filter.h                   |   3 +-\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n t/t6301-for-each-ref-errors.sh |   2 +-\n 6 files changed, 324 insertions(+), 204 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd9..ef8ab952b2fa 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e83..0bc524e656e1 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,86 +209,31 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret = 0;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (!ret) {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n+\t\tstrbuf_release(&err);\n+\t} else if (ret < 0) {\n+\t\tdie(\"%s\\n\", err.buf);\n+\t\tstrbuf_release(&err);\n+\t} else {\n+\t\t/* when ret > 0 , don't call die and print the err to stdout*/\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n+\t\tstrbuf_release(&err);\n \t}\n }\n \n@@ -428,6 +285,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n@@ -488,42 +352,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *opt, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (opt->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (!opt->format.format)\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n+\telse\n+\t\tstrbuf_addstr(&format, opt->format.format);\n \tif (opt->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\topt->format.format = format.buf;\n+\tif (verify_ref_format(&opt->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (opt->cmdmode || opt->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (opt->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -556,6 +412,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -587,19 +444,13 @@ static int batch_objects(struct batch_options *opt)\n \n \t\tbatch_one_object(input.buf, &output, opt, &data);\n \t}\n-\n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -622,7 +473,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -631,7 +482,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -670,6 +523,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -713,7 +567,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 420c0bf9384f..d4c88d496698 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1017,8 +1017,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n@@ -1735,8 +1742,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 9dc07476a584..bece9583cf18 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n@@ -86,7 +87,7 @@ struct ref_format {\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74ad..5efa7397cfbc 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 mast failed\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab534..3553f84a00c1 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"427178","messageId":"0004d5b24a0fb735d7fa9cb9a8e214d6e838baeb.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 8/8] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:17Z","receivedAt":"2021-06-12T11:14:28Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  5 ++++\n ref-filter.c       | 66 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  5 ++--\n 3 files changed, 72 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 1a73c3d23dde..3fde2587201b 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -376,6 +376,11 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \tif (opt->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \topt->format.format = format.buf;\n+\tif (opt->cmdmode == 'c')\n+\t\topt->format.use_textconv = 1;\n+\tif (opt->cmdmode == 'w')\n+\t\topt->format.use_filters = 1;\n+\n \tif (verify_ref_format(&opt->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex d4c88d496698..8264ef7d2786 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1027,6 +1031,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n@@ -1735,10 +1742,41 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv) {\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tif(!ref->rest)\n+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\t\tif (act_oi.type == OBJ_BLOB) {\n+\t\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\t\tgoto success;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1748,19 +1786,43 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters) {\n+\t\t\tif(!ref->rest)\n+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\t\tif (oi->type == OBJ_BLOB) {\n+\t\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\t\tstruct checkout_metadata meta;\n+\t\t\t\tact_oi = *oi;\n+\n+\t\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\t\tif (convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta)) {\n+\t\t\t\t\tact_oi.size = strbuf.len;\n+\t\t\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\t} else {\n+\t\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\t    oid_to_hex(&oi->oid), ref->rest);\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex bece9583cf18..cf7bad4e8b49 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,14 +80,15 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n-\n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n-- \ngitgitgadget\n"},{"id":"427179","messageId":"5a5b5f78aeeac1f541852dc219d617530fbe87ea.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 4/8] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:13Z","receivedAt":"2021-06-12T11:14:40Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex af8c15aef44d..8868cf98f090 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1006,6 +1017,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n@@ -1920,6 +1934,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2137,6 +2157,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89f..9dc07476a584 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a00..2780ec8803fd 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must failed' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex e2867de791e7..8c97c3b877c6 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1187,6 +1187,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883e..9fc4c4323949 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must failed' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602bf..785b32eb88f9 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must failed' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"427180","messageId":"abee6a03becb929ffb292648d1ef64e61b66d53d.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:11Z","receivedAt":"2021-06-12T11:14:40Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, add a new member in `struct atom_value`: `s_size`, which\ncan record raw object size, it can help us add raw object data to\nthe buffer or compare two buffers which contain raw object data.\n\nBeyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, `--perl` because if our binary raw data is passed to a variable\nin the host language, the host language may not support arbitrary binary\ndata in the variables of its string type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 207 +++++++++++++++++++++++++++++\n 3 files changed, 328 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de706..7f1f0a1ca3b6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the host language may not support arbitrary binary data in the\n+variables of its string type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fbaf..7822be903071 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4d..e2867de791e7 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo \"\" >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,165 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo \"\" >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tgit hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n+\tgit hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n+\tgit hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n+\tgit hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n+\tgit hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n+\tgit hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n+\tgit hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n+\tgit hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"427181","messageId":"pull.980.git.1623496458.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":null,"subject":"[PATCH 0/8] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:09Z","receivedAt":"2021-06-12T11:15:22Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic, which based on\n5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n\n 1. Modified the logic of cat-file --batch, use verify_ref_format() and\n    format_ref_array_item() to get object data.\n 2. Re-implement --textconv, --filters.\n\nNow cat-file can support most ref-filter atoms, like %(tree), %(parent),\n%(if)...\n\nThere is still an unresolved issue: performance overhead is very large, so\nthat when we use:\n\ngit cat-file --batch --batch-all-objects >/dev/null\n\non git.git, it may fail.\n\nZheNing Hu (8):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: teach get_object() return useful value\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_objet_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 267 ++++++------------------\n builtin/tag.c                      |   2 +-\n ref-filter.c                       | 320 +++++++++++++++++++++++------\n ref-filter.h                       |  13 +-\n t/t1006-cat-file.sh                | 252 +++++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 211 +++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 12 files changed, 818 insertions(+), 276 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n-- \ngitgitgadget\n"},{"id":"427182","messageId":"48d256db5c349c1fa0615bb60d74039c78a831fd.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 1/8] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:10Z","receivedAt":"2021-06-12T11:15:23Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c6..5cee6512fbaf 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"427183","messageId":"c208b8a45d66556a3f905063bc7c5026ac4f1e82.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 5/8] [GSOC] ref-filter: teach get_object() return useful value","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:14Z","receivedAt":"2021-06-12T11:15:25Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nLet `populate_value()`, `get_ref_atom_value()` and\n`format_ref_array_item()` get the return value of `get_value()`\ncorrectly. This can help us later let `cat-file --batch` get the\ncorrect error message and return value of `get_value()`.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 19 +++++++++++--------\n 1 file changed, 11 insertions(+), 8 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 8868cf98f090..420c0bf9384f 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1808,7 +1808,7 @@ static char *get_worktree_path(const struct used_atom *atom, const struct ref_ar\n static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n-\tint i;\n+\tint i, ret = 0;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1965,8 +1965,8 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tif ((ret = get_object(ref, 0, &obj, &oi, err)))\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -1997,9 +1997,11 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n+\tint ret = 0;\n+\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tif ((ret = populate_value(ref, err)))\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2573,6 +2575,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret = 0;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n@@ -2610,7 +2613,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tstrbuf_addbuf(final_buf, &state.stack->output);\n \tpop_stack_element(&state.stack);\n-\treturn 0;\n+\treturn ret;\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-- \ngitgitgadget\n\n"},{"id":"427184","messageId":"d31059c391d0c3f40ba45be0803a5ac6d49d5c6f.1623496458.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH 7/8] [GSOC] cat-file: reuse err buf in batch_objet_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-12T11:14:16Z","receivedAt":"2021-06-12T11:15:41Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 24 ++++++++++++++----------\n 1 file changed, 14 insertions(+), 10 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 0bc524e656e1..1a73c3d23dde 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,33 +212,32 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret = 0;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (!ret) {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n-\t\tstrbuf_release(&err);\n \t} else if (ret < 0) {\n-\t\tdie(\"%s\\n\", err.buf);\n-\t\tstrbuf_release(&err);\n+\t\tdie(\"%s\\n\", err->buf);\n \t} else {\n \t\t/* when ret > 0 , don't call die and print the err to stdout*/\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n-\t\tstrbuf_release(&err);\n \t}\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -292,7 +291,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -300,13 +299,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -362,6 +362,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -390,6 +391,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\tcb.opt = opt;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (opt->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -414,6 +416,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -442,11 +445,12 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, opt, &data);\n \t}\n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"427193","messageId":"CAP8UFD0zVLF4ZKorv-8TaEh16-5K1e-=tFafGknucQGU14ZYGg@mail.gmail.com","threadId":"55909","inReplyTo":"c208b8a45d66556a3f905063bc7c5026ac4f1e82.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/8] [GSOC] ref-filter: teach get_object() return useful value","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-06-12T20:09:18Z","receivedAt":"2021-06-12T20:10:37Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Sat, Jun 12, 2021 at 1:14 PM ZheNing Hu via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Let `populate_value()`, `get_ref_atom_value()` and\n> `format_ref_array_item()` get the return value of `get_value()`\n> correctly. This can help us later let `cat-file --batch` get the\n> correct error message and return value of `get_value()`.\n\nIs it get_object() or get_value()?\n"},{"id":"427228","messageId":"CAOLTT8QJh==+ji3dvb4VpTKfuWp5CC9fbNSdWTvux7gk0-Hvhw@mail.gmail.com","threadId":"55909","inReplyTo":"CAP8UFD0zVLF4ZKorv-8TaEh16-5K1e-=tFafGknucQGU14ZYGg@mail.gmail.com","subject":"Re: [PATCH 5/8] [GSOC] ref-filter: teach get_object() return useful value","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-13T10:11:36Z","receivedAt":"2021-06-13T10:14:53Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Christian Couder <christian.couder@gmail.com> 于2021年6月13日周日 上午4:09写道：\n>\n> On Sat, Jun 12, 2021 at 1:14 PM ZheNing Hu via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> > Let `populate_value()`, `get_ref_atom_value()` and\n> > `format_ref_array_item()` get the return value of `get_value()`\n> > correctly. This can help us later let `cat-file --batch` get the\n> > correct error message and return value of `get_value()`.\n>\n> Is it get_object() or get_value()?\n\nOh, it's get_object().\n\nThanks for pointing out:)\n--\nZheNing Hu\n"},{"id":"427459","messageId":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.git.1623496458.gitgitgadget@gmail.com","subject":"[PATCH v2 0/9] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:28:56Z","receivedAt":"2021-06-15T13:29:34Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic, which based on\n5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n\nChange from last version:\n\n 1. Use free_array_item_internal() to solve the memory leak problem.\n 2. Change commit message of ([GSOC] ref-filter: teach get_object() return\n    useful value).\n\nZheNing Hu (9):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: teach get_object() return useful value\n  [GSOC] ref-filter: introduce free_array_item_internal() function\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_objet_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 267 ++++++-----------------\n builtin/tag.c                      |   2 +-\n ref-filter.c                       | 331 ++++++++++++++++++++++-------\n ref-filter.h                       |  14 +-\n t/t1006-cat-file.sh                | 252 ++++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 211 ++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 12 files changed, 829 insertions(+), 277 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n\nRange-diff vs v1:\n\n  1:  48d256db5c34 =  1:  48d256db5c34 [GSOC] ref-filter: add obj-type check in grab contents\n  2:  abee6a03becb =  2:  abee6a03becb [GSOC] ref-filter: add %(raw) atom\n  3:  c99d1d070a18 =  3:  c99d1d070a18 [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  4:  5a5b5f78aeea =  4:  5a5b5f78aeea [GSOC] ref-filter: add %(rest) atom\n  5:  c208b8a45d66 !  5:  49063372e003 [GSOC] ref-filter: teach get_object() return useful value\n     @@ Commit message\n          [GSOC] ref-filter: teach get_object() return useful value\n      \n          Let `populate_value()`, `get_ref_atom_value()` and\n     -    `format_ref_array_item()` get the return value of `get_value()`\n     +    `format_ref_array_item()` get the return value of `get_object()`\n          correctly. This can help us later let `cat-file --batch` get the\n     -    correct error message and return value of `get_value()`.\n     +    correct error message and return value of `get_object()`.\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n  -:  ------------ >  6:  d2f2563eb76a [GSOC] ref-filter: introduce free_array_item_internal() function\n  6:  44ebf75e2e93 !  7:  765337a46ab0 [GSOC] cat-file: reuse ref-filter logic\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      +\tif (!ret) {\n      +\t\tstrbuf_addch(scratch, '\\n');\n      +\t\tbatch_write(opt, scratch->buf, scratch->len);\n     -+\t\tstrbuf_release(&err);\n      +\t} else if (ret < 0) {\n      +\t\tdie(\"%s\\n\", err.buf);\n     -+\t\tstrbuf_release(&err);\n      +\t} else {\n      +\t\t/* when ret > 0 , don't call die and print the err to stdout*/\n      +\t\tprintf(\"%s\\n\", err.buf);\n      +\t\tfflush(stdout);\n     -+\t\tstrbuf_release(&err);\n       \t}\n     ++\tfree_array_item_internal(&item);\n     ++\tstrbuf_release(&err);\n       }\n       \n     + static void batch_one_object(const char *obj_name,\n      @@ builtin/cat-file.c: static void batch_one_object(const char *obj_name,\n       \t\treturn;\n       \t}\n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt)\n       \t\treturn 0;\n       \t}\n      @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt)\n     - \n       \t\tbatch_one_object(input.buf, &output, opt, &data);\n       \t}\n     --\n     + \n      +\tstrbuf_release(&format);\n       \tstrbuf_release(&input);\n       \tstrbuf_release(&output);\n  7:  d31059c391d0 !  8:  058b304686fd [GSOC] cat-file: reuse err buf in batch_objet_write()\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n       \tif (!ret) {\n       \t\tstrbuf_addch(scratch, '\\n');\n       \t\tbatch_write(opt, scratch->buf, scratch->len);\n     --\t\tstrbuf_release(&err);\n       \t} else if (ret < 0) {\n      -\t\tdie(\"%s\\n\", err.buf);\n     --\t\tstrbuf_release(&err);\n      +\t\tdie(\"%s\\n\", err->buf);\n       \t} else {\n       \t\t/* when ret > 0 , don't call die and print the err to stdout*/\n      -\t\tprintf(\"%s\\n\", err.buf);\n      +\t\tprintf(\"%s\\n\", err->buf);\n       \t\tfflush(stdout);\n     --\t\tstrbuf_release(&err);\n       \t}\n     + \tfree_array_item_internal(&item);\n     +-\tstrbuf_release(&err);\n       }\n       \n       static void batch_one_object(const char *obj_name,\n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const st\n      -\t\tbatch_one_object(input.buf, &output, opt, &data);\n      +\t\tbatch_one_object(input.buf, &output, &err, opt, &data);\n       \t}\n     + \n       \tstrbuf_release(&format);\n       \tstrbuf_release(&input);\n       \tstrbuf_release(&output);\n  8:  0004d5b24a0f !  9:  cbf7d51933ea [GSOC] cat-file: re-implement --textconv, --filters options\n     @@ ref-filter.h: struct ref_format {\n      +\tint use_filters;\n       \tint use_rest;\n       \tint use_color;\n     --\n     - \t/* Internal state to ref-filter */\n     + \n     +@@ ref-filter.h: struct ref_format {\n       \tint need_color_reset_at_eol;\n       };\n       \n\n-- \ngitgitgadget\n"},{"id":"427460","messageId":"48d256db5c349c1fa0615bb60d74039c78a831fd.1623763746.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 1/9] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:28:57Z","receivedAt":"2021-06-15T13:29:34Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c6..5cee6512fbaf 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"427461","messageId":"abee6a03becb929ffb292648d1ef64e61b66d53d.1623763746.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 2/9] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:28:58Z","receivedAt":"2021-06-15T13:29:38Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, add a new member in `struct atom_value`: `s_size`, which\ncan record raw object size, it can help us add raw object data to\nthe buffer or compare two buffers which contain raw object data.\n\nBeyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, `--perl` because if our binary raw data is passed to a variable\nin the host language, the host language may not support arbitrary binary\ndata in the variables of its string type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 207 +++++++++++++++++++++++++++++\n 3 files changed, 328 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de706..7f1f0a1ca3b6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the host language may not support arbitrary binary data in the\n+variables of its string type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fbaf..7822be903071 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4d..e2867de791e7 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\techo \"\" >>expected.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo \"\" >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,165 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo \"\" >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tgit hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n+\tgit hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n+\tgit hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n+\tgit hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n+\tgit hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n+\tgit hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n+\tgit hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n+\tgit hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"427462","messageId":"5a5b5f78aeeac1f541852dc219d617530fbe87ea.1623763746.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 4/9] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:00Z","receivedAt":"2021-06-15T13:29:38Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex af8c15aef44d..8868cf98f090 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1006,6 +1017,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n@@ -1920,6 +1934,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2137,6 +2157,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89f..9dc07476a584 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a00..2780ec8803fd 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must failed' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex e2867de791e7..8c97c3b877c6 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1187,6 +1187,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must failed' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883e..9fc4c4323949 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must failed' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602bf..785b32eb88f9 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must failed' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"427463","messageId":"c99d1d070a182d09013792d724d6a62bb9a7c0a2.1623763746.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 3/9] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:28:59Z","receivedAt":"2021-06-15T13:29:41Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc098242..452558ec9575 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be903071..af8c15aef44d 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2546,9 +2546,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2593,7 +2593,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a718965..74fb423fc89f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"427464","messageId":"d2f2563eb76ac2e2c88a76edfac7353284407ad2.1623763747.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 6/9] [GSOC] ref-filter: introduce free_array_item_internal() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:02Z","receivedAt":"2021-06-15T13:29:43Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIntroduce free_array_item_internal() for freeing ref_array_item value.\nIt will be called internally by free_array_item(), and it will help\n`cat-file --batch` free ref_array_item's memory later.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 ++++++++---\n ref-filter.h |  2 ++\n 2 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 420c0bf9384f..aa2ce634106b 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -2282,16 +2282,21 @@ static int ref_filter_handler(const char *refname, const struct object_id *oid,\n \treturn 0;\n }\n \n-/*  Free memory allocated for a ref_array_item */\n-static void free_array_item(struct ref_array_item *item)\n+void free_array_item_internal(struct ref_array_item *item)\n {\n-\tfree((char *)item->symref);\n \tif (item->value) {\n \t\tint i;\n \t\tfor (i = 0; i < used_atom_cnt; i++)\n \t\t\tfree((char *)item->value[i].s);\n \t\tfree(item->value);\n \t}\n+}\n+\n+/*  Free memory allocated for a ref_array_item */\n+static void free_array_item(struct ref_array_item *item)\n+{\n+\tfree((char *)item->symref);\n+\tfree_array_item_internal(item);\n \tfree(item);\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 9dc07476a584..d4531fef5f91 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -111,6 +111,8 @@ struct ref_format {\n int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n /*  Clear all memory allocated to ref_array */\n void ref_array_clear(struct ref_array *array);\n+/* Free array item's value */\n+void free_array_item_internal(struct ref_array_item *item);\n /*  Used to verify if the given format is correct and to parse out the used atoms */\n int verify_ref_format(struct ref_format *format);\n /*  Sort the given ref_array as per the ref_sorting provided */\n-- \ngitgitgadget\n\n"},{"id":"427465","messageId":"49063372e0035c5384f834d78854da56f5726d13.1623763747.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 5/9] [GSOC] ref-filter: teach get_object() return useful value","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:01Z","receivedAt":"2021-06-15T13:29:44Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nLet `populate_value()`, `get_ref_atom_value()` and\n`format_ref_array_item()` get the return value of `get_object()`\ncorrectly. This can help us later let `cat-file --batch` get the\ncorrect error message and return value of `get_object()`.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 19 +++++++++++--------\n 1 file changed, 11 insertions(+), 8 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 8868cf98f090..420c0bf9384f 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1808,7 +1808,7 @@ static char *get_worktree_path(const struct used_atom *atom, const struct ref_ar\n static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n-\tint i;\n+\tint i, ret = 0;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1965,8 +1965,8 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tif ((ret = get_object(ref, 0, &obj, &oi, err)))\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -1997,9 +1997,11 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n+\tint ret = 0;\n+\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tif ((ret = populate_value(ref, err)))\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2573,6 +2575,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret = 0;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n@@ -2610,7 +2613,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tstrbuf_addbuf(final_buf, &state.stack->output);\n \tpop_stack_element(&state.stack);\n-\treturn 0;\n+\treturn ret;\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-- \ngitgitgadget\n\n"},{"id":"427467","messageId":"765337a46ab0f6c8ad38007a78c850f5dd4cfa38.1623763747.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 7/9] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:03Z","receivedAt":"2021-06-15T13:29:47Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, the following\nmethods are used:\n\n1. Add `cat_file_mode` member in struct `ref_format`, this can\nhelp us reject atoms in verify_ref_format() which cat-file\ncannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n2. Change the type of member `format` in struct `batch_options`\nto `ref_format`, We can add format data in it.\n3. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n4. Use `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists.\n5. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue less than zero, use `die()` to print the error message\nand exit; else return value greater than zero, only print the\nerror message, but not exit.\n6. Let get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\",\nthis can help `format_ref_array_item()` just report that the\nobject is missing without letting Git exit.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have refname.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 248 +++++++-------------------------\n ref-filter.c                   |  15 +-\n ref-filter.h                   |   3 +-\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n t/t6301-for-each-ref-errors.sh |   2 +-\n 6 files changed, 323 insertions(+), 203 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd9..ef8ab952b2fa 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e83..026d9405636a 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,87 +209,31 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret = 0;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (!ret) {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n+\t} else if (ret < 0) {\n+\t\tdie(\"%s\\n\", err.buf);\n+\t} else {\n+\t\t/* when ret > 0 , don't call die and print the err to stdout*/\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n \t}\n+\tfree_array_item_internal(&item);\n+\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n@@ -428,6 +284,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n@@ -488,42 +351,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *opt, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (opt->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (!opt->format.format)\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n+\telse\n+\t\tstrbuf_addstr(&format, opt->format.format);\n \tif (opt->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\topt->format.format = format.buf;\n+\tif (verify_ref_format(&opt->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (opt->cmdmode || opt->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (opt->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -556,6 +411,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -588,18 +444,13 @@ static int batch_objects(struct batch_options *opt)\n \t\tbatch_one_object(input.buf, &output, opt, &data);\n \t}\n \n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -622,7 +473,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -631,7 +482,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -670,6 +523,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -713,7 +567,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/ref-filter.c b/ref-filter.c\nindex aa2ce634106b..660e52f97f79 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1017,8 +1017,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n@@ -1735,8 +1742,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex d4531fef5f91..3db07729f46d 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n@@ -86,7 +87,7 @@ struct ref_format {\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74ad..5efa7397cfbc 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 mast failed\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab534..3553f84a00c1 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"427466","messageId":"058b304686fdd97f1cdd81f4d5eb22f2b08ce987.1623763747.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 8/9] [GSOC] cat-file: reuse err buf in batch_objet_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:04Z","receivedAt":"2021-06-15T13:29:48Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 026d9405636a..6766efb06f41 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,32 +212,33 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret = 0;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (!ret) {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n \t} else if (ret < 0) {\n-\t\tdie(\"%s\\n\", err.buf);\n+\t\tdie(\"%s\\n\", err->buf);\n \t} else {\n \t\t/* when ret > 0 , don't call die and print the err to stdout*/\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n \t}\n \tfree_array_item_internal(&item);\n-\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -291,7 +292,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -299,13 +300,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -361,6 +363,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -389,6 +392,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\tcb.opt = opt;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (opt->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -413,6 +417,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -441,12 +446,13 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, opt, &data);\n \t}\n \n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"427468","messageId":"cbf7d51933ea562e1f75b386caebce878dc3e2c1.1623763747.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v2 9/9] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-15T13:29:05Z","receivedAt":"2021-06-15T13:29:52Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  5 ++++\n ref-filter.c       | 66 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  4 ++-\n 3 files changed, 72 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 6766efb06f41..2978b8990e8d 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -377,6 +377,11 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \tif (opt->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \topt->format.format = format.buf;\n+\tif (opt->cmdmode == 'c')\n+\t\topt->format.use_textconv = 1;\n+\tif (opt->cmdmode == 'w')\n+\t\topt->format.use_filters = 1;\n+\n \tif (verify_ref_format(&opt->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 660e52f97f79..5e4aaa94db8a 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1027,6 +1031,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n@@ -1735,10 +1742,41 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv) {\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tif(!ref->rest)\n+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\t\tif (act_oi.type == OBJ_BLOB) {\n+\t\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\t\tgoto success;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1748,19 +1786,43 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters) {\n+\t\t\tif(!ref->rest)\n+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\t\tif (oi->type == OBJ_BLOB) {\n+\t\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\t\tstruct checkout_metadata meta;\n+\t\t\t\tact_oi = *oi;\n+\n+\t\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\t\tif (convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta)) {\n+\t\t\t\t\tact_oi.size = strbuf.len;\n+\t\t\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\t} else {\n+\t\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\t    oid_to_hex(&oi->oid), ref->rest);\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 3db07729f46d..56a5ec4c9eaa 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,6 +80,8 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n \n@@ -87,7 +89,7 @@ struct ref_format {\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n-- \ngitgitgadget\n"},{"id":"427581","messageId":"xmqq5yyeqncn.fsf@gitster.g","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/9] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-16T07:29:12Z","receivedAt":"2021-06-16T07:29:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> This patch series make cat-file reuse ref-filter logic, which based on\n> 5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n\nHmph, does anybody have 5a5b5f78?\n\nThe way to deal with this and avoid resending the same patches\n(assuming that this is not a 9-patch series, but only 5 of them are\nnew) is to rebase your topic on 723bc66d (ref-filter: add %(rest)\natom, 2021-06-09), which will allow you to discard the 4 earlier\npatches, and force push, with base set to 723bc66d, I think, but I\nam not a GGG user, so there may need an extra step or two on top of\nthat.\n"},{"id":"427582","messageId":"xmqq1r92qn18.fsf@gitster.g","threadId":"55909","inReplyTo":"49063372e0035c5384f834d78854da56f5726d13.1623763747.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 5/9] [GSOC] ref-filter: teach get_object() return useful value","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-16T07:36:03Z","receivedAt":"2021-06-16T07:36:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Let `populate_value()`, `get_ref_atom_value()` and\n> `format_ref_array_item()` get the return value of `get_object()`\n> correctly.\n\nThe \"get\" the value correctly, I think.  What you are teaching them\nis to pass the return value from get_object() through the callchain\nto their callers.\n\nThe readers will be helped if you say what kind of errors\nget_object() wants to tell its callers, not just \"-1\" is for error,\nwhich is what populate_value() assumes to be sufficient.  In other\nwords, which non-zero returns from get_object() are interesting and\nwhy?\n\n> @@ -1997,9 +1997,11 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n>  static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n>  \t\t\t      struct atom_value **v, struct strbuf *err)\n>  {\n> +\tint ret = 0;\n> +\n>  \tif (!ref->value) {\n> -\t\tif (populate_value(ref, err))\n> -\t\t\treturn -1;\n> +\t\tif ((ret = populate_value(ref, err)))\n> +\t\t\treturn ret;\n\nThe new variable only needs to be in this scope, and does not have\nto be shown to the entire function.\n\n> @@ -2573,6 +2575,7 @@ int format_ref_array_item(struct ref_array_item *info,\n>  {\n>  \tconst char *cp, *sp, *ep;\n>  \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n> +\tint ret = 0;\n\nThis is dubious...\n\n>  \tstate.quote_style = format->quote_style;\n>  \tpush_stack_element(&state.stack);\n> @@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n>  \t\tif (cp < sp)\n>  \t\t\tappend_literal(cp, sp, &state);\n>  \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n> -\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n> +\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n\nHere, if \"ret\" gets assigned any non-zero value, the condition is\nsatisfied, and ...\n\n>  \t\t    atomv->handler(atomv, &state, error_buf)) {\n>  \t\t\tpop_stack_element(&state.stack);\n> -\t\t\treturn -1;\n> +\t\t\treturn ret ? ret : -1;\n\n... the control flow will leave this function.  Therefore, ...\n\n>  \t\t}\n>  \t}\n>  \tif (*cp) {\n> @@ -2610,7 +2613,7 @@ int format_ref_array_item(struct ref_array_item *info,\n>  \t}\n>  \tstrbuf_addbuf(final_buf, &state.stack->output);\n>  \tpop_stack_element(&state.stack);\n> -\treturn 0;\n> +\treturn ret;\n\n... at this point, \"ret\" can never be anything other than zero.  Am\nI misreading the patch?\n\nIf I am not misreading the patch, then \"ret\" does not have to be\nglobally visible in this function---it can have the same scope as\n\"pos\".\n\n>  }\n>  \n>  void pretty_print_ref(const char *name, const struct object_id *oid,\n"},{"id":"427584","messageId":"xmqqv96ep7u6.fsf@gitster.g","threadId":"55909","inReplyTo":"d2f2563eb76ac2e2c88a76edfac7353284407ad2.1623763747.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 6/9] [GSOC] ref-filter: introduce free_array_item_internal() function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-16T07:49:37Z","receivedAt":"2021-06-16T07:49:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Introduce free_array_item_internal() for freeing ref_array_item value.\n> It will be called internally by free_array_item(), and it will help\n> `cat-file --batch` free ref_array_item's memory later.\n\nAs a file local static function, the horrible name free_array_item()\nwas tolerable.  But before exposing a name like that to the outside\nworld, think twice if that is specific enough, and it is not.  There\nare 47 different kinds of \"array\"s we use in the system, but this\nnew helper function only works with ref_array_item and not on items\nin any other kinds of arrays.\n\n> -/*  Free memory allocated for a ref_array_item */\n> -static void free_array_item(struct ref_array_item *item)\n> +void free_array_item_internal(struct ref_array_item *item)\n>  {\n\nAnd \"internal\" is a horrible name to have as an external name.  You\nprobably can come up with a more appropriate name when you imagine\nyourself explaining to somebody who is relatively new to this part\nof the codebase what the difference between free_array_item() and\nthis new helper is, where the difference comes from, why the symref\nmember (and no other member) is so special, etc.\n\nI _think_ what is special is not the .symref but is the .value\nfield, IOW, you are trying to come up with an interface to free the\nvalue part of ref_array_item without touching other things.  But it\nis not helpful at all to readers if you do not explain why you want\nto do so.  Why is the .value member so special?  The ability to\nclear only the .value member without touching other members is useful\nbecause ...?\n\nIn any case, assuming that you'd establish why the .value member is\nso special to deserve an externally callable function, when external\ncallers do not have to be able to free the item as a whole (i.e.\nfree_array_item() is still file-scope static), in the proposed log\nmessage in an updated patch, I would imagine that \n\n    free_ref_array_item_value()\n\nwould be a more suitable name than the _internal thing.  When it\nhappens, you might want to rename the static one to\nfree_ref_array_item() to match, even if it does not have external\ncallers.\n"},{"id":"427665","messageId":"CAOLTT8R8CBF6XSO0gx2P_UaSCf1ansXS1Y+Y7Kuft6AKmxTQLA@mail.gmail.com","threadId":"55909","inReplyTo":"xmqq5yyeqncn.fsf@gitster.g","subject":"Re: [PATCH v2 0/9] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T06:07:37Z","receivedAt":"2021-06-17T06:07:57Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2021年6月16日周三 下午3:29写道：\n>\n> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > This patch series make cat-file reuse ref-filter logic, which based on\n> > 5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n>\n> Hmph, does anybody have 5a5b5f78?\n>\n> The way to deal with this and avoid resending the same patches\n> (assuming that this is not a 9-patch series, but only 5 of them are\n> new) is to rebase your topic on 723bc66d (ref-filter: add %(rest)\n> atom, 2021-06-09), which will allow you to discard the 4 earlier\n> patches, and force push, with base set to 723bc66d, I think, but I\n> am not a GGG user, so there may need an extra step or two on top of\n> that.\n\nOh, 5a5b5f78 is in gitgitgadget.git and 723bc66d in git.git, their commit\nmessages are different. Later I will rebase my new patch to zh/xxx.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"427666","messageId":"878s39x95m.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"48d256db5c349c1fa0615bb60d74039c78a831fd.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/8] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:04:57Z","receivedAt":"2021-06-17T07:06:21Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Only tag and commit objects use `grab_sub_body_contents()` to grab\n> object contents in the current codebase.  We want to teach the\n> function to also handle blobs and trees to get their raw data,\n> without parsing a blob (whose contents looks like a commit or a tag)\n> incorrectly as a commit or a tag.\n>\n> Skip the block of code that is specific to handling commits and tags\n> early when the given object is of a wrong type to help later\n> addition to handle other types of objects in this function.\n>\n> Mentored-by: Christian Couder <christian.couder@gmail.com>\n> Mentored-by: Hariom Verma <hariom18599@gmail.com>\n> Helped-by: Junio C Hamano <gitster@pobox.com>\n> Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n> ---\n>  ref-filter.c | 24 +++++++++++++++---------\n>  1 file changed, 15 insertions(+), 9 deletions(-)\n>\n> diff --git a/ref-filter.c b/ref-filter.c\n> index 4db0e40ff4c6..5cee6512fbaf 100644\n> --- a/ref-filter.c\n> +++ b/ref-filter.c\n> @@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n>  }\n>  \n>  /* See grab_values */\n> -static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n> +static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n>  {\n>  \tint i;\n>  \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n>  \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n> +\tvoid *buf = data->content;\n>  \n>  \tfor (i = 0; i < used_atom_cnt; i++) {\n>  \t\tstruct used_atom *atom = &used_atom[i];\n> @@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n>  \t\t\tcontinue;\n>  \t\tif (deref)\n>  \t\t\tname++;\n> -\t\tif (strcmp(name, \"body\") &&\n> -\t\t    !starts_with(name, \"subject\") &&\n> -\t\t    !starts_with(name, \"trailers\") &&\n> -\t\t    !starts_with(name, \"contents\"))\n> +\n> +\t\tif ((data->type != OBJ_TAG &&\n> +\t\t     data->type != OBJ_COMMIT) ||\n> +\t\t    (strcmp(name, \"body\") &&\n> +\t\t     !starts_with(name, \"subject\") &&\n> +\t\t     !starts_with(name, \"trailers\") &&\n> +\t\t     !starts_with(name, \"contents\")))\n\nWe have 4 \"real\" object types, commit, tree, blob, tag. Do you really\nmean \"not tag or commit\" here, don't you mean \"is tree or blob\" instead?\nI.e. do we really want to pass OBJ_NONE etc. here?\n\n>  \t\t\tcontinue;\n>  \t\tif (!subpos)\n>  \t\t\tfind_subpos(buf,\n> @@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n>   * pointed at by the ref itself; otherwise it is the object the\n>   * ref (which is a tag) refers to.\n>   */\n> -static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n> +static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n>  {\n> +\tvoid *buf = data->content;\n> +\n>  \tswitch (obj->type) {\n>  \tcase OBJ_TAG:\n>  \t\tgrab_tag_values(val, deref, obj);\n> -\t\tgrab_sub_body_contents(val, deref, buf);\n> +\t\tgrab_sub_body_contents(val, deref, data);\n>  \t\tgrab_person(\"tagger\", val, deref, buf);\n>  \t\tbreak;\n>  \tcase OBJ_COMMIT:\n>  \t\tgrab_commit_values(val, deref, obj);\n> -\t\tgrab_sub_body_contents(val, deref, buf);\n> +\t\tgrab_sub_body_contents(val, deref, data);\n>  \t\tgrab_person(\"author\", val, deref, buf);\n>  \t\tgrab_person(\"committer\", val, deref, buf);\n>  \t\tbreak;\n> @@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n>  \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n>  \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n>  \t\t}\n> -\t\tgrab_values(ref->value, deref, *obj, oi->content);\n> +\t\tgrab_values(ref->value, deref, *obj, oi);\n>  \t}\n>  \n>  \tgrab_common_values(ref->value, deref, oi);\n\n"},{"id":"427667","messageId":"875yydx8oo.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"abee6a03becb929ffb292648d1ef64e61b66d53d.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:10:34Z","receivedAt":"2021-06-17T07:16:29Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Add new formatting option `%(raw)`, which will print the raw\n> object data without any changes. It will help further to migrate\n> all cat-file formatting logic from cat-file to ref-filter.\n\nNice goal and feature to have.\n\n> The raw data of blob, tree objects may contain '\\0', but most of\n> the logic in `ref-filter` depends on the output of the atom being\n> text (specifically, no embedded NULs in it).\n>\n> E.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\n> add the data to the buffer. The raw data of a tree object is\n> `100644 one\\0...`, only the `100644 one` will be added to the buffer,\n> which is incorrect.\n>\n> Therefore, add a new member in `struct atom_value`: `s_size`, which\n> can record raw object size, it can help us add raw object data to\n> the buffer or compare two buffers which contain raw object data.\n\nMost of the functions that deal with this already use a strbuf in some\nway, before we had a const char *, now there's a size_t to go along with\nit, why not simply use a strbuf in the struct for the data? You'll then\nget the size and \\0 handling for free, and any functions to deal with\nconversion can stick to the strbuf API, there seems to be a lot of back\nand forth now.\n\n> Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> `--tcl`, `--perl` because if our binary raw data is passed to a variable\n> in the host language, the host language may not support arbitrary binary\n> data in the variables of its string type.\n\nPerl at least deals with that just fine, and to the extent that it\ndoesn't any new problems here would have nothing to do with \\0 being in\nthe data. Perl doesn't have a notion of \"binary has \\0 in it\", it always\nsupports \\0, it has a notion of \"is it utf-8 or not?\", so any encoding\nproblems wouldn't be new. I'd think that the same would be true of\nPython, but I'm not sure.\n\n\n> +test_expect_success 'basic atom: refs/tags/testtag *raw' '\n> +\tgit cat-file commit refs/tags/testtag^{} >expected &&\n> +\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n> +\tsanitize_pgp <expected >expected.clean &&\n> +\tsanitize_pgp <actual >actual.clean &&\n> +\techo \"\" >>expected.clean &&\n\nJust \"echo\" will do, ditto for the rest. Also odd to go back and forth\nbetween populating expected.clean & actual.clean.\n\n\n> +test_expect_success 'set up refs pointing to binary blob' '\n> +\tprintf \"a\\0b\\0c\" >blob1 &&\n> +\tprintf \"a\\0c\\0b\" >blob2 &&\n> +\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n> +\tprintf \"abc\" >blob4 &&\n> +\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n> +\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n> +\tprintf \"  \" >blob7 &&\n> +\t>blob8 &&\n> +\tgit hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n> +\tgit hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n> +\tgit hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n> +\tgit hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n> +\tgit hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n> +\tgit hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n> +\tgit hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n> +\tgit hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n\nHrm, xargs just to avoid:\n\n    git update-ref ... $(git hash-object) ?\n\n> +test_expect_success '%(raw) with --python must failed' '\n> +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n> +'\n> +\n> +test_expect_success '%(raw) with --tcl must failed' '\n> +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n> +'\n> +\n> +test_expect_success '%(raw) with --perl must failed' '\n> +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n> +'\n> +\n> +test_expect_success '%(raw) with --shell must failed' '\n> +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n> +'\n> +\n> +test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n> +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n> +'\n\ns/must failed/must fail/, but see question above about encoding in these\nlanguages...\n"},{"id":"427668","messageId":"8735thx8nn.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"d31059c391d0c3f40ba45be0803a5ac6d49d5c6f.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 7/8] [GSOC] cat-file: reuse err buf in batch_objet_write()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:16:39Z","receivedAt":"2021-06-17T07:17:06Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\n\n\nOn Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n\nSubject: s/objet/object/\n"},{"id":"427670","messageId":"87zgvpvtu4.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"0004d5b24a0fb735d7fa9cb9a8e214d6e838baeb.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 8/8] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:18:18Z","receivedAt":"2021-06-17T07:22:55Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n\n>  \topt->format.format = format.buf;\n> +\tif (opt->cmdmode == 'c')\n> +\t\topt->format.use_textconv = 1;\n> +\tif (opt->cmdmode == 'w')\n> +\t\topt->format.use_filters = 1;\n> +\n\n\nStyle nit: if/if -> if/else if, both can't be true.\n\n> +\t\t/* get the type and size */\n> +\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n> +\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n> +\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n> +\t\t\t\t\t       oid_to_hex(&oi->oid));\n> +\n> +\t\toi->info.sizep = NULL;\n> +\t\toi->info.typep = NULL;\n> +\t\toi->info.contentp = temp_contentp;\n\nHere we have a call function and error if bad, then proceed to populate\nstuff (good)...\n\n> +\n> +\t\tif (use_textconv) {\n> +\t\t\tact_oi = *oi;\n> +\n> +\t\t\tif(!ref->rest)\n> +\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n> +\t\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n> +\t\t\tif (act_oi.type == OBJ_BLOB) {\n> +\t\t\t\tif (textconv_object(the_repository,\n> +\t\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n> +\t\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n> +\t\t\t\t\tactual_oi = &act_oi;\n> +\t\t\t\t\tgoto success;\n> +\t\t\t\t}\n> +\t\t\t}\n> +\t\t}\n\n\nMaybe change the if (A) { if (B) {} ) to if (A && B) {} ?\n\n>  \t}\n>  \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n>  \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n> @@ -1748,19 +1786,43 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n>  \t\tBUG(\"Object size is less than zero.\");\n>  \n>  \tif (oi->info.contentp) {\n> -\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n> +\t\tif (use_filters) {\n> +\t\t\tif(!ref->rest)\n> +\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n> +\t\t\t\t\t\t       oid_to_hex(&oi->oid));\n> +\t\t\tif (oi->type == OBJ_BLOB) {\n> +\t\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n> +\t\t\t\tstruct checkout_metadata meta;\n> +\t\t\t\tact_oi = *oi;\n> +\n> +\t\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n> +\t\t\t\tif (convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta)) {\n> +\t\t\t\t\tact_oi.size = strbuf.len;\n> +\t\t\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n> +\t\t\t\t\tactual_oi = &act_oi;\n> +\t\t\t\t} else {\n> +\t\t\t\t\tdie(\"could not convert '%s' %s\",\n> +\t\t\t\t\t    oid_to_hex(&oi->oid), ref->rest);\n> +\t\t\t\t}\n\n... but here instead of \"if (!x) { bad } do stuff\" we have if (x) {do\nstuff} else { bad }\". Better to get the die out of the way, and avoid\nthe indentation on the \"do stuff\" IMO.\n\n\n> -#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n> +#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, 0, 0, -1 }\n\nNot a new problem, but an earlier cleanup to simply change this to\ndesignated initializers would be welcome, see recent work of mine in\nfsck.h for an example.\n\nI.e. we keep churning on changing this *_INIT just to populate the one\n-1 field at the end, can also be simply:\n\n    #define FOO_INIT { .that_field = 1 }\n\n"},{"id":"427671","messageId":"CAOLTT8Q3_X3kZ5ZdUYe8LKc2WuctWFPCg-akJEpQAKY9_rZv3g@mail.gmail.com","threadId":"55909","inReplyTo":"xmqq1r92qn18.fsf@gitster.g","subject":"Re: [PATCH v2 5/9] [GSOC] ref-filter: teach get_object() return useful value","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T07:23:20Z","receivedAt":"2021-06-17T07:23:37Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2021年6月16日周三 下午3:36写道：\n>\n> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> > Let `populate_value()`, `get_ref_atom_value()` and\n> > `format_ref_array_item()` get the return value of `get_object()`\n> > correctly.\n>\n> The \"get\" the value correctly, I think.  What you are teaching them\n> is to pass the return value from get_object() through the callchain\n> to their callers.\n>\n\nYes, this is exactly what I meant.\n\n> The readers will be helped if you say what kind of errors\n> get_object() wants to tell its callers, not just \"-1\" is for error,\n> which is what populate_value() assumes to be sufficient.  In other\n> words, which non-zero returns from get_object() are interesting and\n> why?\n>\n\nAs stated in 765337a, We can just print the error without exiting if the\nreturn value of format_ref_array_item() is greater than 0. Therefore,\nthe current patch is to make get_object() return a value other than\n-1 when an error occurs.\n\n> > @@ -1997,9 +1997,11 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n> >  static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n> >                             struct atom_value **v, struct strbuf *err)\n> >  {\n> > +     int ret = 0;\n> > +\n> >       if (!ref->value) {\n> > -             if (populate_value(ref, err))\n> > -                     return -1;\n> > +             if ((ret = populate_value(ref, err)))\n> > +                     return ret;\n>\n> The new variable only needs to be in this scope, and does not have\n> to be shown to the entire function.\n>\n\nMakes sense.\n\n> > @@ -2573,6 +2575,7 @@ int format_ref_array_item(struct ref_array_item *info,\n> >  {\n> >       const char *cp, *sp, *ep;\n> >       struct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n> > +     int ret = 0;\n>\n> This is dubious...\n>\n> >       state.quote_style = format->quote_style;\n> >       push_stack_element(&state.stack);\n> > @@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n> >               if (cp < sp)\n> >                       append_literal(cp, sp, &state);\n> >               pos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n> > -             if (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n> > +             if (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n>\n> Here, if \"ret\" gets assigned any non-zero value, the condition is\n> satisfied, and ...\n>\n> >                   atomv->handler(atomv, &state, error_buf)) {\n> >                       pop_stack_element(&state.stack);\n> > -                     return -1;\n> > +                     return ret ? ret : -1;\n>\n> ... the control flow will leave this function.  Therefore, ...\n>\n> >               }\n> >       }\n> >       if (*cp) {\n> > @@ -2610,7 +2613,7 @@ int format_ref_array_item(struct ref_array_item *info,\n> >       }\n> >       strbuf_addbuf(final_buf, &state.stack->output);\n> >       pop_stack_element(&state.stack);\n> > -     return 0;\n> > +     return ret;\n>\n> ... at this point, \"ret\" can never be anything other than zero.  Am\n> I misreading the patch?\n>\n> If I am not misreading the patch, then \"ret\" does not have to be\n> globally visible in this function---it can have the same scope as\n> \"pos\".\n>\n\nYou are right, It is correct to only return 0 here at the moment.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"427672","messageId":"87wnqtvtph.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"c208b8a45d66556a3f905063bc7c5026ac4f1e82.1623496458.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/8] [GSOC] ref-filter: teach get_object() return useful value","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:22:39Z","receivedAt":"2021-06-17T07:25:20Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> Let `populate_value()`, `get_ref_atom_value()` and\n> `format_ref_array_item()` get the return value of `get_value()`\n> correctly. This can help us later let `cat-file --batch` get the\n> correct error message and return value of `get_value()`.\n>\n> Mentored-by: Christian Couder <christian.couder@gmail.com>\n> Mentored-by: Hariom Verma <hariom18599@gmail.com>\n> Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n> ---\n>  ref-filter.c | 19 +++++++++++--------\n>  1 file changed, 11 insertions(+), 8 deletions(-)\n>\n> diff --git a/ref-filter.c b/ref-filter.c\n> index 8868cf98f090..420c0bf9384f 100644\n> --- a/ref-filter.c\n> +++ b/ref-filter.c\n> @@ -1808,7 +1808,7 @@ static char *get_worktree_path(const struct used_atom *atom, const struct ref_ar\n>  static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n>  {\n>  \tstruct object *obj;\n> -\tint i;\n> +\tint i, ret = 0;\n\nThe usual style is to not bunch up variables based on type, but only if\nthey're related, i.e. we'd do:\n\n    if i, j; /* proceed to use i and j in two for-loops */\n\nBut:\n\n    int i; /* for the for-loop */\n    int ret = 0; /* for our return value */\n\n(Without the comments)\n\n>  \tstruct object_info empty = OBJECT_INFO_INIT;\n>  \n>  \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n> @@ -1965,8 +1965,8 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n>  \n>  \n>  \toi.oid = ref->objectname;\n> -\tif (get_object(ref, 0, &obj, &oi, err))\n> -\t\treturn -1;\n> +\tif ((ret = get_object(ref, 0, &obj, &oi, err)))\n> +\t\treturn ret;\n\nMaybe more personal style, I'd just write this as:\n\n    ret = x();\n    if (!ret)\n        return ret;\n\nMakes it easier to read and balance parens in your head for the common\ncase...\n\n> @@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n>  \t\tif (cp < sp)\n>  \t\t\tappend_literal(cp, sp, &state);\n>  \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n> -\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n> +\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n>  \t\t    atomv->handler(atomv, &state, error_buf)) {\n>  \t\t\tpop_stack_element(&state.stack);\n\n... and use that mental energy on readist stuff like this, which I'd\njust leave as the inline assignment.\n"},{"id":"427673","messageId":"87tulxvtm4.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 0/9] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T07:26:17Z","receivedAt":"2021-06-17T07:27:19Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jun 15 2021, ZheNing Hu via GitGitGadget wrote:\n\n> This patch series make cat-file reuse ref-filter logic, which based on\n> 5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n>\n> Change from last version:\n>\n>  1. Use free_array_item_internal() to solve the memory leak problem.\n>  2. Change commit message of ([GSOC] ref-filter: teach get_object() return\n>     useful value).\n\nI left some comments, but saw after the fact that I'd replied to the v1\nE-Mails by accident, but anyway, the comments were all on things that\nare also in v2, so it worked out in the end. Sorry about the confusion.\n"},{"id":"427674","messageId":"xmqqtulxkl0i.fsf@gitster.g","threadId":"55909","inReplyTo":"878s39x95m.fsf@evledraar.gmail.com","subject":"Re: [PATCH 1/8] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-17T07:28:29Z","receivedAt":"2021-06-17T07:28:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> -\t\tif (strcmp(name, \"body\") &&\n>> -\t\t    !starts_with(name, \"subject\") &&\n>> -\t\t    !starts_with(name, \"trailers\") &&\n>> -\t\t    !starts_with(name, \"contents\"))\n>> +\n>> +\t\tif ((data->type != OBJ_TAG &&\n>> +\t\t     data->type != OBJ_COMMIT) ||\n>> +\t\t    (strcmp(name, \"body\") &&\n>> +\t\t     !starts_with(name, \"subject\") &&\n>> +\t\t     !starts_with(name, \"trailers\") &&\n>> +\t\t     !starts_with(name, \"contents\")))\n>\n> We have 4 \"real\" object types, commit, tree, blob, tag. Do you really\n> mean \"not tag or commit\" here, don't you mean \"is tree or blob\" instead?\n> I.e. do we really want to pass OBJ_NONE etc. here?\n>\n>>  \t\t\tcontinue;\n\nIf somebody throws OBJ_NONE at us by mistake, we do not want to\nhandle such an object and try to extract the subject member from it\nanyway, no?\n\nThe intent of the code here, before the patch, is that \"what we do\nafter the control flow passes this point is about the body, subject,\ntrailers, and contents request, so everybody else should go to the\nnext iteration\".  The caller used to give us an object compatible\nwith these four types of requests, now the caller may throw others,\nhence \"by the way, we know these four kinds of requests make sense\nonly for tags and commits, so everybody else should go to the next\niteration, too\" would be a natural thing to add.  So in that sense,\nI prefer it over \"we know these four types of requests do not make\nsense for blobs and trees\".\n"},{"id":"427675","messageId":"xmqqpmwlkkqp.fsf@gitster.g","threadId":"55909","inReplyTo":"875yydx8oo.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-17T07:34:22Z","receivedAt":"2021-06-17T07:34:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n>> `--tcl`, `--perl` because if our binary raw data is passed to a variable\n>> in the host language, the host language may not support arbitrary binary\n>> data in the variables of its string type.\n>\n> Perl at least deals with that just fine, and to the extent that it\n> doesn't any new problems here would have nothing to do with \\0 being in\n> the data. Perl doesn't have a notion of \"binary has \\0 in it\", it always\n> supports \\0, it has a notion of \"is it utf-8 or not?\", so any encoding\n> problems wouldn't be new. I'd think that the same would be true of\n> Python, but I'm not sure.\n\nDuring an earlier iteration long time ago, as we knew Perl is\ncapable of handling sequence of bytes including NUL, it was decided\nto start more strict by rejecting binary for all languages, which\ncan later be loosened, to limit the scope of the initial\nimplementation.\n\n>> +\tgit hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n>> +\tgit hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n>> +\tgit hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n>> +\tgit hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n>> +\tgit hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n>> +\tgit hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n>> +\tgit hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n>> +\tgit hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n>\n> Hrm, xargs just to avoid:\n>\n>     git update-ref ... $(git hash-object) ?\n\nThat's horrible.  Thanks for noticing.\n\nWe'd want to catch segfaults from both hash-object and update-ref.\nOne way to do so may be\n\n\tO=$(git hash-object -w blob1) &&\n\tgit update-ref refs/myblobs/blob1 \"$O\"\n\nThanks for a review.\n"},{"id":"427676","messageId":"CAOLTT8T3Ov80w8LkUOw-vwqVB_WptiNZOCTPh7N5NGZwroAnnA@mail.gmail.com","threadId":"55909","inReplyTo":"xmqqv96ep7u6.fsf@gitster.g","subject":"Re: [PATCH v2 6/9] [GSOC] ref-filter: introduce free_array_item_internal() function","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T08:03:17Z","receivedAt":"2021-06-17T08:03:34Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2021年6月16日周三 下午3:49写道：\n>\n> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> > Introduce free_array_item_internal() for freeing ref_array_item value.\n> > It will be called internally by free_array_item(), and it will help\n> > `cat-file --batch` free ref_array_item's memory later.\n>\n> As a file local static function, the horrible name free_array_item()\n> was tolerable.  But before exposing a name like that to the outside\n> world, think twice if that is specific enough, and it is not.  There\n> are 47 different kinds of \"array\"s we use in the system, but this\n> new helper function only works with ref_array_item and not on items\n> in any other kinds of arrays.\n>\n\nWell, this is indeed ignored by me.\n\nSo free_array_item() --> free_ref_array_item() and free_array_item()\n--> free_ref_array_item().\n\n> > -/*  Free memory allocated for a ref_array_item */\n> > -static void free_array_item(struct ref_array_item *item)\n> > +void free_array_item_internal(struct ref_array_item *item)\n> >  {\n>\n> And \"internal\" is a horrible name to have as an external name.  You\n> probably can come up with a more appropriate name when you imagine\n> yourself explaining to somebody who is relatively new to this part\n> of the codebase what the difference between free_array_item() and\n> this new helper is, where the difference comes from, why the symref\n> member (and no other member) is so special, etc.\n>\n\nYeah, \"internal\" is an incorrect naming. The main purpose here is\nto only free a ref_array_item's value.\n\n> I _think_ what is special is not the .symref but is the .value\n> field, IOW, you are trying to come up with an interface to free the\n> value part of ref_array_item without touching other things.  But it\n> is not helpful at all to readers if you do not explain why you want\n> to do so.  Why is the .value member so special?  The ability to\n> clear only the .value member without touching other members is useful\n> because ...?\n>\n\nBecause batch_object_write() use a ref_array_item which is allocated on the\nstack, and it's member symref is not used at all. ref_array_item's symref and\nitself do not need free. The original free_array_item() will free all\ndynamically\nallocated content in ref_array_item. It cannot meet our requirements\nat this time:\nonly free the ref_array_item's value.\n\n> In any case, assuming that you'd establish why the .value member is\n> so special to deserve an externally callable function, when external\n> callers do not have to be able to free the item as a whole (i.e.\n> free_array_item() is still file-scope static), in the proposed log\n> message in an updated patch, I would imagine that\n>\n>     free_ref_array_item_value()\n>\n> would be a more suitable name than the _internal thing.  When it\n> happens, you might want to rename the static one to\n> free_ref_array_item() to match, even if it does not have external\n> callers.\n\nMake sence.\n\n--\nZheNing Hu\n"},{"id":"427677","messageId":"CAOLTT8RQDzE=-8VXRQr2-R1m7G11bTUm-Mz+4WHGXJtNNF1dOg@mail.gmail.com","threadId":"55909","inReplyTo":"8735thx8nn.fsf@evledraar.gmail.com","subject":"Re: [PATCH 7/8] [GSOC] cat-file: reuse err buf in batch_objet_write()","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T08:05:44Z","receivedAt":"2021-06-17T08:05:57Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:17写道：\n>\n>\n>\n>\n> On Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n>\n> Subject: s/objet/object/\n\nOK.\n\n--\nZheNing Hu\n"},{"id":"427678","messageId":"CAOLTT8TMOs-FF+EcTZBbxfGnKQipe_nx_eZon=S=PWRTNT4CjA@mail.gmail.com","threadId":"55909","inReplyTo":"875yydx8oo.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T09:22:19Z","receivedAt":"2021-06-17T09:22:34Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:16写道：\n>\n> > The raw data of blob, tree objects may contain '\\0', but most of\n> > the logic in `ref-filter` depends on the output of the atom being\n> > text (specifically, no embedded NULs in it).\n> >\n> > E.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\n> > add the data to the buffer. The raw data of a tree object is\n> > `100644 one\\0...`, only the `100644 one` will be added to the buffer,\n> > which is incorrect.\n> >\n> > Therefore, add a new member in `struct atom_value`: `s_size`, which\n> > can record raw object size, it can help us add raw object data to\n> > the buffer or compare two buffers which contain raw object data.\n>\n> Most of the functions that deal with this already use a strbuf in some\n> way, before we had a const char *, now there's a size_t to go along with\n> it, why not simply use a strbuf in the struct for the data? You'll then\n> get the size and \\0 handling for free, and any functions to deal with\n> conversion can stick to the strbuf API, there seems to be a lot of back\n> and forth now.\n>\n\nYes, strbuf is a suitable choice when using <str,len> pair.\nBut if replace v->s with strbuf, the possible changes will be larger.\n\n> > Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> > `--tcl`, `--perl` because if our binary raw data is passed to a variable\n> > in the host language, the host language may not support arbitrary binary\n> > data in the variables of its string type.\n>\n> Perl at least deals with that just fine, and to the extent that it\n> doesn't any new problems here would have nothing to do with \\0 being in\n> the data. Perl doesn't have a notion of \"binary has \\0 in it\", it always\n> supports \\0, it has a notion of \"is it utf-8 or not?\", so any encoding\n> problems wouldn't be new. I'd think that the same would be true of\n> Python, but I'm not sure.\n>\n\nNot python safe. See [1].\nRegarding the perl language, I support Junio's point of view: it can be\nre-supported in the future.\n\n>\n> > +test_expect_success 'basic atom: refs/tags/testtag *raw' '\n> > +     git cat-file commit refs/tags/testtag^{} >expected &&\n> > +     git for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n> > +     sanitize_pgp <expected >expected.clean &&\n> > +     sanitize_pgp <actual >actual.clean &&\n> > +     echo \"\" >>expected.clean &&\n>\n> Just \"echo\" will do, ditto for the rest. Also odd to go back and forth\n> between populating expected.clean & actual.clean.\n>\n\nAre you saying that sanitize_pgp is not needed?\n\n>\n> > +test_expect_success 'set up refs pointing to binary blob' '\n> > +     printf \"a\\0b\\0c\" >blob1 &&\n> > +     printf \"a\\0c\\0b\" >blob2 &&\n> > +     printf \"\\0a\\0b\\0c\" >blob3 &&\n> > +     printf \"abc\" >blob4 &&\n> > +     printf \"\\0 \\0 \\0 \" >blob5 &&\n> > +     printf \"\\0 \\0a\\0 \" >blob6 &&\n> > +     printf \"  \" >blob7 &&\n> > +     >blob8 &&\n> > +     git hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n> > +     git hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n> > +     git hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n> > +     git hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n> > +     git hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n> > +     git hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n> > +     git hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n> > +     git hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n>\n> Hrm, xargs just to avoid:\n>\n>     git update-ref ... $(git hash-object) ?\n>\n\nI didn’t think about it, just for convenience.\n\n> > +test_expect_success '%(raw) with --python must failed' '\n> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --python\n> > +'\n> > +\n> > +test_expect_success '%(raw) with --tcl must failed' '\n> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n> > +'\n> > +\n> > +test_expect_success '%(raw) with --perl must failed' '\n> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --perl\n> > +'\n> > +\n> > +test_expect_success '%(raw) with --shell must failed' '\n> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --shell\n> > +'\n> > +\n> > +test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n> > +'\n>\n> s/must failed/must fail/, but see question above about encoding in these\n> languages...\n\n\n[1]: https://lore.kernel.org/git/CAOLTT8QR_GRm4TYk0E_eazQ+unVQODc-3L+b4V5JUN5jtZR8uA@mail.gmail.com/\n\nThanks for a review.\n--\nZheNing Hu\n"},{"id":"427682","messageId":"CAOLTT8QiFiXm7M=Cdk22wqztKKKaTnw3iQpv8DEtZp_iKRUdDw@mail.gmail.com","threadId":"55909","inReplyTo":"87zgvpvtu4.fsf@evledraar.gmail.com","subject":"Re: [PATCH 8/8] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T09:53:20Z","receivedAt":"2021-06-17T09:53:36Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:22写道：\n>\n>\n> On Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n>\n> > From: ZheNing Hu <adlternative@gmail.com>\n>\n> >       opt->format.format = format.buf;\n> > +     if (opt->cmdmode == 'c')\n> > +             opt->format.use_textconv = 1;\n> > +     if (opt->cmdmode == 'w')\n> > +             opt->format.use_filters = 1;\n> > +\n>\n>\n> Style nit: if/if -> if/else if, both can't be true.\n>\n\nOk.\n\n> > +             /* get the type and size */\n> > +             if (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n> > +                                     OBJECT_INFO_LOOKUP_REPLACE))\n> > +                     return strbuf_addf_ret(err, 1, _(\"%s missing\"),\n> > +                                            oid_to_hex(&oi->oid));\n> > +\n> > +             oi->info.sizep = NULL;\n> > +             oi->info.typep = NULL;\n> > +             oi->info.contentp = temp_contentp;\n>\n> Here we have a call function and error if bad, then proceed to populate\n> stuff (good)...\n>\n> > +\n> > +             if (use_textconv) {\n> > +                     act_oi = *oi;\n> > +\n> > +                     if(!ref->rest)\n> > +                             return strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n> > +                                                    oid_to_hex(&act_oi.oid));\n> > +                     if (act_oi.type == OBJ_BLOB) {\n> > +                             if (textconv_object(the_repository,\n> > +                                                 ref->rest, 0100644, &act_oi.oid,\n> > +                                                 1, (char **)(&act_oi.content), &act_oi.size)) {\n> > +                                     actual_oi = &act_oi;\n> > +                                     goto success;\n> > +                             }\n> > +                     }\n> > +             }\n>\n>\n> Maybe change the if (A) { if (B) {} ) to if (A && B) {} ?\n>\n\nMaybe it's something like this:\n\n@@ -1762,19 +1762,17 @@ static int get_object(struct ref_array_item\n*ref, int deref, struct object **obj\n                oi->info.typep = NULL;\n                oi->info.contentp = temp_contentp;\n\n-               if (use_textconv) {\n-                       act_oi = *oi;\n+               if (use_textconv && !ref->rest)\n+                       return strbuf_addf_ret(err, -1, _(\"missing\npath for '%s'\"),\n+                                              oid_to_hex(&act_oi.oid));\n\n-                       if(!ref->rest)\n-                               return strbuf_addf_ret(err, -1,\n_(\"missing path for '%s'\"),\n-                                                      oid_to_hex(&act_oi.oid));\n-                       if (act_oi.type == OBJ_BLOB) {\n-                               if (textconv_object(the_repository,\n-                                                   ref->rest,\n0100644, &act_oi.oid,\n-                                                   1, (char\n**)(&act_oi.content), &act_oi.size)) {\n-                                       actual_oi = &act_oi;\n-                                       goto success;\n-                               }\n+               if (use_textconv && oi->type == OBJ_BLOB) {\n+                       act_oi = *oi;\n+                       if (textconv_object(the_repository,\n+                                           ref->rest, 0100644, &act_oi.oid,\n+                                           1, (char\n**)(&act_oi.content), &act_oi.size)) {\n+                               actual_oi = &act_oi;\n+                               goto success;\n                        }\n                }\n        }\n\n> >       }\n> >       if (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n> >                                    OBJECT_INFO_LOOKUP_REPLACE))\n> > @@ -1748,19 +1786,43 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n> >               BUG(\"Object size is less than zero.\");\n> >\n> >       if (oi->info.contentp) {\n> > -             *obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n> > +             if (use_filters) {\n> > +                     if(!ref->rest)\n> > +                             return strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n> > +                                                    oid_to_hex(&oi->oid));\n> > +                     if (oi->type == OBJ_BLOB) {\n> > +                             struct strbuf strbuf = STRBUF_INIT;\n> > +                             struct checkout_metadata meta;\n> > +                             act_oi = *oi;\n> > +\n> > +                             init_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n> > +                             if (convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta)) {\n> > +                                     act_oi.size = strbuf.len;\n> > +                                     act_oi.content = strbuf_detach(&strbuf, NULL);\n> > +                                     actual_oi = &act_oi;\n> > +                             } else {\n> > +                                     die(\"could not convert '%s' %s\",\n> > +                                         oid_to_hex(&oi->oid), ref->rest);\n> > +                             }\n>\n> ... but here instead of \"if (!x) { bad } do stuff\" we have if (x) {do\n> stuff} else { bad }\". Better to get the die out of the way, and avoid\n> the indentation on the \"do stuff\" IMO.\n>\n\nAnd this part:\n\n@@ -1786,25 +1784,21 @@ static int get_object(struct ref_array_item\n*ref, int deref, struct object **obj\n                BUG(\"Object size is less than zero.\");\n\n        if (oi->info.contentp) {\n-               if (use_filters) {\n-                       if(!ref->rest)\n-                               return strbuf_addf_ret(err, -1,\n_(\"missing path for '%s'\"),\n-                                                      oid_to_hex(&oi->oid));\n-                       if (oi->type == OBJ_BLOB) {\n-                               struct strbuf strbuf = STRBUF_INIT;\n-                               struct checkout_metadata meta;\n-                               act_oi = *oi;\n-\n-                               init_checkout_metadata(&meta, NULL,\nNULL, &act_oi.oid);\n-                               if\n(convert_to_working_tree(&the_index, ref->rest, act_oi.content,\nact_oi.size, &strbuf, &meta)) {\n-                                       act_oi.size = strbuf.len;\n-                                       act_oi.content =\nstrbuf_detach(&strbuf, NULL);\n-                                       actual_oi = &act_oi;\n-                               } else {\n-                                       die(\"could not convert '%s' %s\",\n-                                           oid_to_hex(&oi->oid), ref->rest);\n-                               }\n-                       }\n+               if (use_filters && !ref->rest)\n+                       return strbuf_addf_ret(err, -1, _(\"missing\npath for '%s'\"),\n+                                               oid_to_hex(&oi->oid));\n+               if (use_filters && oi->type == OBJ_BLOB) {\n+                       struct strbuf strbuf = STRBUF_INIT;\n+                       struct checkout_metadata meta;\n+                       act_oi = *oi;\n+\n+                       init_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+                       if (!convert_to_working_tree(&the_index,\nref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n+                               die(\"could not convert '%s' %s\",\n+                                       oid_to_hex(&oi->oid), ref->rest);\n+                       act_oi.size = strbuf.len;\n+                       act_oi.content = strbuf_detach(&strbuf, NULL);\n+                       actual_oi = &act_oi;\n                }\n\n\n>\n> > -#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n> > +#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, 0, 0, -1 }\n>\n> Not a new problem, but an earlier cleanup to simply change this to\n> designated initializers would be welcome, see recent work of mine in\n> fsck.h for an example.\n>\n> I.e. we keep churning on changing this *_INIT just to populate the one\n> -1 field at the end, can also be simply:\n>\n>     #define FOO_INIT { .that_field = 1 }\n>\n\nI agree, this method is better. We don't need to care about which members\nso many \"0\" point to.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"427683","messageId":"CAOLTT8RPMf6YcHT4Qi=NFwh_5ZR7c5xgj6E94GGWMJ+749zS0g@mail.gmail.com","threadId":"55909","inReplyTo":"87wnqtvtph.fsf@evledraar.gmail.com","subject":"Re: [PATCH 5/8] [GSOC] ref-filter: teach get_object() return useful value","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T10:01:05Z","receivedAt":"2021-06-17T10:01:20Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:25写道：\n>\n>\n> On Sat, Jun 12 2021, ZheNing Hu via GitGitGadget wrote:\n>\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> > Let `populate_value()`, `get_ref_atom_value()` and\n> > `format_ref_array_item()` get the return value of `get_value()`\n> > correctly. This can help us later let `cat-file --batch` get the\n> > correct error message and return value of `get_value()`.\n> >\n> > Mentored-by: Christian Couder <christian.couder@gmail.com>\n> > Mentored-by: Hariom Verma <hariom18599@gmail.com>\n> > Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n> > ---\n> >  ref-filter.c | 19 +++++++++++--------\n> >  1 file changed, 11 insertions(+), 8 deletions(-)\n> >\n> > diff --git a/ref-filter.c b/ref-filter.c\n> > index 8868cf98f090..420c0bf9384f 100644\n> > --- a/ref-filter.c\n> > +++ b/ref-filter.c\n> > @@ -1808,7 +1808,7 @@ static char *get_worktree_path(const struct used_atom *atom, const struct ref_ar\n> >  static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n> >  {\n> >       struct object *obj;\n> > -     int i;\n> > +     int i, ret = 0;\n>\n> The usual style is to not bunch up variables based on type, but only if\n> they're related, i.e. we'd do:\n>\n>     if i, j; /* proceed to use i and j in two for-loops */\n>\n> But:\n>\n>     int i; /* for the for-loop */\n>     int ret = 0; /* for our return value */\n>\n> (Without the comments)\n>\n\nI agree.\n\n> >       struct object_info empty = OBJECT_INFO_INIT;\n> >\n> >       CALLOC_ARRAY(ref->value, used_atom_cnt);\n> > @@ -1965,8 +1965,8 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n> >\n> >\n> >       oi.oid = ref->objectname;\n> > -     if (get_object(ref, 0, &obj, &oi, err))\n> > -             return -1;\n> > +     if ((ret = get_object(ref, 0, &obj, &oi, err)))\n> > +             return ret;\n>\n> Maybe more personal style, I'd just write this as:\n>\n>     ret = x();\n>     if (!ret)\n>         return ret;\n>\n> Makes it easier to read and balance parens in your head for the common\n> case...\n>\n\nYeah. This way it will be more readable.\n\n> > @@ -2585,10 +2588,10 @@ int format_ref_array_item(struct ref_array_item *info,\n> >               if (cp < sp)\n> >                       append_literal(cp, sp, &state);\n> >               pos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n> > -             if (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n> > +             if (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n> >                   atomv->handler(atomv, &state, error_buf)) {\n> >                       pop_stack_element(&state.stack);\n>\n> ... and use that mental energy on readist stuff like this, which I'd\n> just leave as the inline assignment.\n\nIndeed, inline assignment must be used in this case.\n\nThanks. I will pay attention to these style issues later.\n--\nZheNing Hu\n"},{"id":"427688","messageId":"CAOLTT8Rm0PonXyn_wFqC1ommUpAd_MwO30gGrxsuEBR179qyEw@mail.gmail.com","threadId":"55909","inReplyTo":"87tulxvtm4.fsf@evledraar.gmail.com","subject":"Re: [PATCH v2 0/9] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T10:02:39Z","receivedAt":"2021-06-17T10:02:53Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:27写道：\n>\n>\n> On Tue, Jun 15 2021, ZheNing Hu via GitGitGadget wrote:\n>\n> > This patch series make cat-file reuse ref-filter logic, which based on\n> > 5a5b5f78 ([GSOC] ref-filter: add %(rest) atom)\n> >\n> > Change from last version:\n> >\n> >  1. Use free_array_item_internal() to solve the memory leak problem.\n> >  2. Change commit message of ([GSOC] ref-filter: teach get_object() return\n> >     useful value).\n>\n> I left some comments, but saw after the fact that I'd replied to the v1\n> E-Mails by accident, but anyway, the comments were all on things that\n> are also in v2, so it worked out in the end. Sorry about the confusion.\n\nIt's okay. Your comments are very useful.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"427757","messageId":"87r1h0wnwg.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"CAOLTT8TMOs-FF+EcTZBbxfGnKQipe_nx_eZon=S=PWRTNT4CjA@mail.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-17T14:37:34Z","receivedAt":"2021-06-17T14:45:25Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Jun 17 2021, ZheNing Hu wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午3:16写道：\n>>\n>> > The raw data of blob, tree objects may contain '\\0', but most of\n>> > the logic in `ref-filter` depends on the output of the atom being\n>> > text (specifically, no embedded NULs in it).\n>> >\n>> > E.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\n>> > add the data to the buffer. The raw data of a tree object is\n>> > `100644 one\\0...`, only the `100644 one` will be added to the buffer,\n>> > which is incorrect.\n>> >\n>> > Therefore, add a new member in `struct atom_value`: `s_size`, which\n>> > can record raw object size, it can help us add raw object data to\n>> > the buffer or compare two buffers which contain raw object data.\n>>\n>> Most of the functions that deal with this already use a strbuf in some\n>> way, before we had a const char *, now there's a size_t to go along with\n>> it, why not simply use a strbuf in the struct for the data? You'll then\n>> get the size and \\0 handling for free, and any functions to deal with\n>> conversion can stick to the strbuf API, there seems to be a lot of back\n>> and forth now.\n>>\n>\n> Yes, strbuf is a suitable choice when using <str,len> pair.\n> But if replace v->s with strbuf, the possible changes will be larger.\n\nI for one would like to see it done that way, those changes are usually\neasy to read. Also it seems a large part of 2/8 is extra new code\nbecause we didn't do that, e.g. getting length differently if something\nis a strbuf or not, passing char*/size_t pairs to new functions etc.\n\n>> > Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n>> > `--tcl`, `--perl` because if our binary raw data is passed to a variable\n>> > in the host language, the host language may not support arbitrary binary\n>> > data in the variables of its string type.\n>>\n>> Perl at least deals with that just fine, and to the extent that it\n>> doesn't any new problems here would have nothing to do with \\0 being in\n>> the data. Perl doesn't have a notion of \"binary has \\0 in it\", it always\n>> supports \\0, it has a notion of \"is it utf-8 or not?\", so any encoding\n>> problems wouldn't be new. I'd think that the same would be true of\n>> Python, but I'm not sure.\n>>\n>\n> Not python safe. See [1].\n> Regarding the perl language, I support Junio's point of view: it can be\n> re-supported in the future.\n\nAh, I'd missed that. Anyway, if it's easy it seems you discovered that\nPerl deals with it correctly, so we could just have it support this.\n\n>>\n>> > +test_expect_success 'basic atom: refs/tags/testtag *raw' '\n>> > +     git cat-file commit refs/tags/testtag^{} >expected &&\n>> > +     git for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n>> > +     sanitize_pgp <expected >expected.clean &&\n>> > +     sanitize_pgp <actual >actual.clean &&\n>> > +     echo \"\" >>expected.clean &&\n>>\n>> Just \"echo\" will do, ditto for the rest. Also odd to go back and forth\n>> between populating expected.clean & actual.clean.\n>>\n>\n> Are you saying that sanitize_pgp is not needed?\n\nNo that instead of:\n\n    echo \"\" >x\n\nYou can do:\n\n    echo >x\n\nAnd also that going back and forth between populating different files is\nconfusing, i.e. this:\n\n\n    echo a >x\n    echo c >y\n    echo b >>x\n\nis better as:\n\n    echo a >x\n    echo b >>x\n    echo c >y\n\n\n>>\n>> > +test_expect_success 'set up refs pointing to binary blob' '\n>> > +     printf \"a\\0b\\0c\" >blob1 &&\n>> > +     printf \"a\\0c\\0b\" >blob2 &&\n>> > +     printf \"\\0a\\0b\\0c\" >blob3 &&\n>> > +     printf \"abc\" >blob4 &&\n>> > +     printf \"\\0 \\0 \\0 \" >blob5 &&\n>> > +     printf \"\\0 \\0a\\0 \" >blob6 &&\n>> > +     printf \"  \" >blob7 &&\n>> > +     >blob8 &&\n>> > +     git hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n>> > +     git hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n>> > +     git hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n>> > +     git hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n>> > +     git hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n>> > +     git hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n>> > +     git hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n>> > +     git hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n>>\n>> Hrm, xargs just to avoid:\n>>\n>>     git update-ref ... $(git hash-object) ?\n>>\n>\n> I didn’t think about it, just for convenience.\n\n*nod*, Junio had a good suggestion.\n\n>> > +test_expect_success '%(raw) with --python must failed' '\n>> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --python\n>> > +'\n>> > +\n>> > +test_expect_success '%(raw) with --tcl must failed' '\n>> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n>> > +'\n>> > +\n>> > +test_expect_success '%(raw) with --perl must failed' '\n>> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --perl\n>> > +'\n>> > +\n>> > +test_expect_success '%(raw) with --shell must failed' '\n>> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --shell\n>> > +'\n>> > +\n>> > +test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n>> > +     test_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n>> > +'\n>>\n>> s/must failed/must fail/, but see question above about encoding in these\n>> languages...\n>\n>\n> [1]: https://lore.kernel.org/git/CAOLTT8QR_GRm4TYk0E_eazQ+unVQODc-3L+b4V5JUN5jtZR8uA@mail.gmail.com/\n>\n> Thanks for a review.\n\n"},{"id":"427766","messageId":"CAOLTT8SGkObKWMZax-k4KXgsN+7ezvOMkRU-zt+c2zon2Ta3pA@mail.gmail.com","threadId":"55909","inReplyTo":"87r1h0wnwg.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-17T16:14:37Z","receivedAt":"2021-06-17T16:14:52Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午10:45写道：\n>\n>\n> On Thu, Jun 17 2021, ZheNing Hu wrote:\n>\n> >\n> > Yes, strbuf is a suitable choice when using <str,len> pair.\n> > But if replace v->s with strbuf, the possible changes will be larger.\n>\n> I for one would like to see it done that way, those changes are usually\n> easy to read. Also it seems a large part of 2/8 is extra new code\n> because we didn't do that, e.g. getting length differently if something\n> is a strbuf or not, passing char*/size_t pairs to new functions etc.\n>\n\nAfter some refactoring, I found that there are two problems:\n1. There are a lot of codes like this in ref-filter to fill v->s:\n\nv->s = show_ref(...)\nv->s = copy_email(...)\n\nIt is very difficult to modify here: We know that show_ref()\nor copy_email() will allocate a block of memory to v->s, but\nif v->s is a strbuf, what should we do? In copy_email(), we\ncan pass the v->s to copy_email() and use strbuf_add()/strbuf_addstr()\ninstead of xstrdup() and xmemdupz(). But show_ref() will call\nexternal functions like shorten_unambiguous_ref(), we don’t know\nwhether it will return us NULL or a dynamically allocated memory.\nIf continue to pass v->s to the inner function, it is not a feasible\nmethod. Or we can use strbuf_attach() + strlen(), I'm not sure\nthis is a good method.\n\n2. See:\n\n-       for (i = 0; i < used_atom_cnt; i++) {\n+       for (i = 0; i < used_atom_cnt; i++) {\n                struct atom_value *v = &ref->value[i];\n-               if (v->s == NULL && used_atom[i].source == SOURCE_NONE)\n+               if (v->s.len == 0 && used_atom[i].source == SOURCE_NONE)\n                        return strbuf_addf_ret(err, -1, _(\"missing\nobject %s for %s\"),\n\noid_to_hex(&ref->objectname), ref->refname);\n        }\n\nIn the case of using strbuf, I don’t know how to distinguish between an empty\nstrbuf and NULL. It can be easily distinguished by using c-style \"const char*\".\n\n> >\n> > Not python safe. See [1].\n> > Regarding the perl language, I support Junio's point of view: it can be\n> > re-supported in the future.\n>\n> Ah, I'd missed that. Anyway, if it's easy it seems you discovered that\n> Perl deals with it correctly, so we could just have it support this.\n>\n\nWell, it's ok, support for perl will be put in a separate commit.\n\n> >>\n> >> > +test_expect_success 'basic atom: refs/tags/testtag *raw' '\n> >> > +     git cat-file commit refs/tags/testtag^{} >expected &&\n> >> > +     git for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n> >> > +     sanitize_pgp <expected >expected.clean &&\n> >> > +     sanitize_pgp <actual >actual.clean &&\n> >> > +     echo \"\" >>expected.clean &&\n> >>\n> >> Just \"echo\" will do, ditto for the rest. Also odd to go back and forth\n> >> between populating expected.clean & actual.clean.\n> >>\n> >\n> > Are you saying that sanitize_pgp is not needed?\n>\n> No that instead of:\n>\n>     echo \"\" >x\n>\n> You can do:\n>\n>     echo >x\n>\n> And also that going back and forth between populating different files is\n> confusing, i.e. this:\n>\n>\n>     echo a >x\n>     echo c >y\n>     echo b >>x\n>\n> is better as:\n>\n>     echo a >x\n>     echo b >>x\n>     echo c >y\n>\n>\n\nThanks, I get what you meant now.\n\n> >>\n> >> > +test_expect_success 'set up refs pointing to binary blob' '\n> >> > +     printf \"a\\0b\\0c\" >blob1 &&\n> >> > +     printf \"a\\0c\\0b\" >blob2 &&\n> >> > +     printf \"\\0a\\0b\\0c\" >blob3 &&\n> >> > +     printf \"abc\" >blob4 &&\n> >> > +     printf \"\\0 \\0 \\0 \" >blob5 &&\n> >> > +     printf \"\\0 \\0a\\0 \" >blob6 &&\n> >> > +     printf \"  \" >blob7 &&\n> >> > +     >blob8 &&\n> >> > +     git hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n> >> > +     git hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n> >> > +     git hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n> >> > +     git hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n> >> > +     git hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n> >> > +     git hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n> >> > +     git hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n> >> > +     git hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n> >>\n> >> Hrm, xargs just to avoid:\n> >>\n> >>     git update-ref ... $(git hash-object) ?\n> >>\n> >\n> > I didn’t think about it, just for convenience.\n>\n> *nod*, Junio had a good suggestion.\n>\n\nok.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"427816","messageId":"87r1gz4faf.fsf@evledraar.gmail.com","threadId":"55909","inReplyTo":"CAOLTT8SGkObKWMZax-k4KXgsN+7ezvOMkRU-zt+c2zon2Ta3pA@mail.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-06-18T10:49:02Z","receivedAt":"2021-06-18T10:51:09Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Jun 18 2021, ZheNing Hu wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2021年6月17日周四 下午10:45写道：\n>>\n>>\n>> On Thu, Jun 17 2021, ZheNing Hu wrote:\n>>\n>> >\n>> > Yes, strbuf is a suitable choice when using <str,len> pair.\n>> > But if replace v->s with strbuf, the possible changes will be larger.\n>>\n>> I for one would like to see it done that way, those changes are usually\n>> easy to read. Also it seems a large part of 2/8 is extra new code\n>> because we didn't do that, e.g. getting length differently if something\n>> is a strbuf or not, passing char*/size_t pairs to new functions etc.\n>>\n>\n> After some refactoring, I found that there are two problems:\n> 1. There are a lot of codes like this in ref-filter to fill v->s:\n>\n> v->s = show_ref(...)\n> v->s = copy_email(...)\n>\n> It is very difficult to modify here: We know that show_ref()\n> or copy_email() will allocate a block of memory to v->s, but\n> if v->s is a strbuf, what should we do? In copy_email(), we\n> can pass the v->s to copy_email() and use strbuf_add()/strbuf_addstr()\n> instead of xstrdup() and xmemdupz(). But show_ref() will call\n> external functions like shorten_unambiguous_ref(), we don’t know\n> whether it will return us NULL or a dynamically allocated memory.\n> If continue to pass v->s to the inner function, it is not a feasible\n> method. Or we can use strbuf_attach() + strlen(), I'm not sure\n> this is a good method.\n>\n> 2. See:\n>\n> -       for (i = 0; i < used_atom_cnt; i++) {\n> +       for (i = 0; i < used_atom_cnt; i++) {\n>                 struct atom_value *v = &ref->value[i];\n> -               if (v->s == NULL && used_atom[i].source == SOURCE_NONE)\n> +               if (v->s.len == 0 && used_atom[i].source == SOURCE_NONE)\n>                         return strbuf_addf_ret(err, -1, _(\"missing\n> object %s for %s\"),\n>\n> oid_to_hex(&ref->objectname), ref->refname);\n>         }\n>\n> In the case of using strbuf, I don’t know how to distinguish between an empty\n> strbuf and NULL. It can be easily distinguished by using c-style \"const char*\".\n\nYes, sometimes it's just too much of a hassle, looking at\nshorten_unambiguous_ref() which returns a xstrdup()'d value that could\nindeed be strbuf_attach'd. I haven't tried the conversion myself,\nperhaps it's too much hassle.\n\nJust a suggestion from reading your patch in isolation.\n\n\n>> >\n>> > Not python safe. See [1].\n>> > Regarding the perl language, I support Junio's point of view: it can be\n>> > re-supported in the future.\n>>\n>> Ah, I'd missed that. Anyway, if it's easy it seems you discovered that\n>> Perl deals with it correctly, so we could just have it support this.\n>>\n>\n> Well, it's ok, support for perl will be put in a separate commit.\n>\n>> >>\n>> >> > +test_expect_success 'basic atom: refs/tags/testtag *raw' '\n>> >> > +     git cat-file commit refs/tags/testtag^{} >expected &&\n>> >> > +     git for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n>> >> > +     sanitize_pgp <expected >expected.clean &&\n>> >> > +     sanitize_pgp <actual >actual.clean &&\n>> >> > +     echo \"\" >>expected.clean &&\n>> >>\n>> >> Just \"echo\" will do, ditto for the rest. Also odd to go back and forth\n>> >> between populating expected.clean & actual.clean.\n>> >>\n>> >\n>> > Are you saying that sanitize_pgp is not needed?\n>>\n>> No that instead of:\n>>\n>>     echo \"\" >x\n>>\n>> You can do:\n>>\n>>     echo >x\n>>\n>> And also that going back and forth between populating different files is\n>> confusing, i.e. this:\n>>\n>>\n>>     echo a >x\n>>     echo c >y\n>>     echo b >>x\n>>\n>> is better as:\n>>\n>>     echo a >x\n>>     echo b >>x\n>>     echo c >y\n>>\n>>\n>\n> Thanks, I get what you meant now.\n>\n>> >>\n>> >> > +test_expect_success 'set up refs pointing to binary blob' '\n>> >> > +     printf \"a\\0b\\0c\" >blob1 &&\n>> >> > +     printf \"a\\0c\\0b\" >blob2 &&\n>> >> > +     printf \"\\0a\\0b\\0c\" >blob3 &&\n>> >> > +     printf \"abc\" >blob4 &&\n>> >> > +     printf \"\\0 \\0 \\0 \" >blob5 &&\n>> >> > +     printf \"\\0 \\0a\\0 \" >blob6 &&\n>> >> > +     printf \"  \" >blob7 &&\n>> >> > +     >blob8 &&\n>> >> > +     git hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n>> >> > +     git hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n>> >> > +     git hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n>> >> > +     git hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n>> >> > +     git hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n>> >> > +     git hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n>> >> > +     git hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n>> >> > +     git hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n>> >>\n>> >> Hrm, xargs just to avoid:\n>> >>\n>> >>     git update-ref ... $(git hash-object) ?\n>> >>\n>> >\n>> > I didn’t think about it, just for convenience.\n>>\n>> *nod*, Junio had a good suggestion.\n>>\n>\n> ok.\n>\n> Thanks.\n\n"},{"id":"427827","messageId":"CAP8UFD1O0PLeThi+_DcSQ3U7Vughode+dRM0b=H5V4J3i_Nn3w@mail.gmail.com","threadId":"55909","inReplyTo":"87r1gz4faf.fsf@evledraar.gmail.com","subject":"Re: [PATCH 2/8] [GSOC] ref-filter: add %(raw) atom","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-06-18T13:47:02Z","receivedAt":"2021-06-18T13:47:19Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Jun 18, 2021 at 12:51 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n\n> On Fri, Jun 18 2021, ZheNing Hu wrote:\n\n> > After some refactoring, I found that there are two problems:\n> > 1. There are a lot of codes like this in ref-filter to fill v->s:\n> >\n> > v->s = show_ref(...)\n> > v->s = copy_email(...)\n> >\n> > It is very difficult to modify here: We know that show_ref()\n> > or copy_email() will allocate a block of memory to v->s, but\n> > if v->s is a strbuf, what should we do? In copy_email(), we\n> > can pass the v->s to copy_email() and use strbuf_add()/strbuf_addstr()\n> > instead of xstrdup() and xmemdupz(). But show_ref() will call\n> > external functions like shorten_unambiguous_ref(), we don’t know\n> > whether it will return us NULL or a dynamically allocated memory.\n> > If continue to pass v->s to the inner function, it is not a feasible\n> > method. Or we can use strbuf_attach() + strlen(), I'm not sure\n> > this is a good method.\n\nIf you resend this patch, it might be a good idea to add a short\nversion of the above explanations into the commit message.\n\n[...]\n\n> > In the case of using strbuf, I don’t know how to distinguish between an empty\n> > strbuf and NULL. It can be easily distinguished by using c-style \"const char*\".\n\nMaybe this could also be part of the explanation.\n\n> Yes, sometimes it's just too much of a hassle, looking at\n> shorten_unambiguous_ref() which returns a xstrdup()'d value that could\n> indeed be strbuf_attach'd. I haven't tried the conversion myself,\n> perhaps it's too much hassle.\n>\n> Just a suggestion from reading your patch in isolation.\n\nYeah, thanks for the review anyway!\n"},{"id":"427946","messageId":"f72ad9cc5e8b1f1139784df9ac7178f1561f70bb.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 01/10] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:51Z","receivedAt":"2021-06-19T07:03:10Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c6..5cee6512fbaf 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"427947","messageId":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v2.git.1623763746.gitgitgadget@gmail.com","subject":"[PATCH v3 00/10] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:50Z","receivedAt":"2021-06-19T07:03:11Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic. As part of the\ncontent related to zh/ref-filter-raw-data has been modified, this patch\ncurrently sets the base on zh/ref-filter-atom-type until\nzh/ref-filter-raw-data is stable And will rebase to it.\n\nf72ad9c~1..39a0d93 of this patch belongs to zh/ref-filter-raw-data. Change\nfrom last version:\n\n 1. Let --format=%(raw) re-support --perl.\n 2. Correct some errors in the test.\n 3. Add the reason why strbuf is not used to replace atom_value's member s\n    in the commit message.\n\n39a0d93..86ac3bc of this patch is another part. Change from last version:\n\n 1. Modify free_array_item_internal() to free_ref_array_item_value().\n 2. Change some errors of the code style.\n 3. Change some commit messages.\n\nZheNing Hu (10):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: --format=%(raw) re-support --perl\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: pass get_object() return value to their callers\n  [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_object_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 267 ++++++-----------------\n builtin/tag.c                      |   2 +-\n quote.c                            |  17 ++\n quote.h                            |   1 +\n ref-filter.c                       | 331 +++++++++++++++++++++++------\n ref-filter.h                       |  14 +-\n t/t1006-cat-file.sh                | 252 ++++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 235 ++++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 14 files changed, 872 insertions(+), 276 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n\nRange-diff vs v2:\n\n  1:  48d256db5c34 =  1:  f72ad9cc5e8b [GSOC] ref-filter: add obj-type check in grab contents\n  2:  abee6a03becb !  2:  ab497d66c116 [GSOC] ref-filter: add %(raw) atom\n     @@ Commit message\n          `100644 one\\0...`, only the `100644 one` will be added to the buffer,\n          which is incorrect.\n      \n     -    Therefore, add a new member in `struct atom_value`: `s_size`, which\n     -    can record raw object size, it can help us add raw object data to\n     -    the buffer or compare two buffers which contain raw object data.\n     +    Therefore, we need to find a way to record the length of the\n     +    atom_value's member `s`. Although strbuf can already record the\n     +    string and its length, if we want to replace the type of atom_value's\n     +    member `s` with strbuf, many places in ref-filter that are filled\n     +    with dynamically allocated mermory in `v->s` are not easy to replace.\n     +    At the same time, we need to check if `v->s == NULL` in\n     +    populate_value(), and strbuf cannot easily distinguish NULL and empty\n     +    strings, but c-style \"const char *\" can do it. So add a new member in\n     +    `struct atom_value`: `s_size`, which can record raw object size, it\n     +    can help us add raw object data to the buffer or compare two buffers\n     +    which contain raw object data.\n      \n          Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n     -    `--tcl`, `--perl` because if our binary raw data is passed to a variable\n     -    in the host language, the host language may not support arbitrary binary\n     -    data in the variables of its string type.\n     +    `--tcl`, `--perl` because if our binary raw data is passed to a\n     +    variable in the host language, the host language may not support\n     +    arbitrary binary data in the variables of its string type.\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n     +    Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n          Helped-by: Felipe Contreras <felipe.contreras@gmail.com>\n          Helped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\n          Helped-by: Junio C Hamano <gitster@pobox.com>\n     @@ t/t6300-for-each-ref.sh: test_atom tag contents 'Tagging at 1151968727\n      +\tgit cat-file commit refs/tags/testtag^{} >expected &&\n      +\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n      +\tsanitize_pgp <expected >expected.clean &&\n     ++\techo >>expected.clean &&\n      +\tsanitize_pgp <actual >actual.clean &&\n     -+\techo \"\" >>expected.clean &&\n      +\ttest_cmp expected.clean actual.clean\n      +'\n      +\n     @@ t/t6300-for-each-ref.sh: test_atom refs/tags/signed-empty contents:body ''\n      +\tgit cat-file tag refs/tags/signed-empty >expected &&\n      +\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n      +\tsanitize_pgp <expected >expected.clean &&\n     ++\techo >>expected.clean &&\n      +\tsanitize_pgp <actual >actual.clean &&\n     -+\techo \"\" >>expected.clean &&\n      +\ttest_cmp expected.clean actual.clean\n      +'\n      +\n     @@ t/t6300-for-each-ref.sh: test_atom refs/tags/signed-short contents:signature \"$s\n      +\tgit cat-file tag refs/tags/signed-short >expected &&\n      +\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n      +\tsanitize_pgp <expected >expected.clean &&\n     ++\techo >>expected.clean &&\n      +\tsanitize_pgp <actual >actual.clean &&\n     -+\techo \"\" >>expected.clean &&\n      +\ttest_cmp expected.clean actual.clean\n      +'\n      +\n     @@ t/t6300-for-each-ref.sh: test_atom refs/tags/signed-long contents \"subject line\n      +\tgit cat-file tag refs/tags/signed-long >expected &&\n      +\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n      +\tsanitize_pgp <expected >expected.clean &&\n     ++\techo >>expected.clean &&\n      +\tsanitize_pgp <actual >actual.clean &&\n     -+\techo \"\" >>expected.clean &&\n      +\ttest_cmp expected.clean actual.clean\n      +'\n      +\n     @@ t/t6300-for-each-ref.sh: test_atom refs/mytrees/first contents:body \"\"\n       \n      +test_expect_success 'basic atom: refs/mytrees/first raw' '\n      +\tgit cat-file tree refs/mytrees/first >expected &&\n     -+\techo \"\" >>expected &&\n     ++\techo >>expected &&\n      +\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n      +\ttest_cmp expected actual &&\n      +\tgit cat-file -s refs/mytrees/first >expected &&\n     @@ t/t6300-for-each-ref.sh: test_atom refs/myblobs/first contents:body \"\"\n       \n      +test_expect_success 'basic atom: refs/myblobs/first raw' '\n      +\tgit cat-file blob refs/myblobs/first >expected &&\n     -+\techo \"\" >>expected &&\n     ++\techo >>expected &&\n      +\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n      +\ttest_cmp expected actual &&\n      +\tgit cat-file -s refs/myblobs/first >expected &&\n     @@ t/t6300-for-each-ref.sh: test_atom refs/myblobs/first contents:body \"\"\n      +\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n      +\tprintf \"  \" >blob7 &&\n      +\t>blob8 &&\n     -+\tgit hash-object blob1 -w | xargs git update-ref refs/myblobs/blob1 &&\n     -+\tgit hash-object blob2 -w | xargs git update-ref refs/myblobs/blob2 &&\n     -+\tgit hash-object blob3 -w | xargs git update-ref refs/myblobs/blob3 &&\n     -+\tgit hash-object blob4 -w | xargs git update-ref refs/myblobs/blob4 &&\n     -+\tgit hash-object blob5 -w | xargs git update-ref refs/myblobs/blob5 &&\n     -+\tgit hash-object blob6 -w | xargs git update-ref refs/myblobs/blob6 &&\n     -+\tgit hash-object blob7 -w | xargs git update-ref refs/myblobs/blob7 &&\n     -+\tgit hash-object blob8 -w | xargs git update-ref refs/myblobs/blob8\n     ++\tobj=$(git hash-object -w blob1) &&\n     ++        git update-ref refs/myblobs/blob1 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob2) &&\n     ++        git update-ref refs/myblobs/blob2 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob3) &&\n     ++        git update-ref refs/myblobs/blob3 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob4) &&\n     ++        git update-ref refs/myblobs/blob4 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob5) &&\n     ++        git update-ref refs/myblobs/blob5 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob6) &&\n     ++        git update-ref refs/myblobs/blob6 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob7) &&\n     ++        git update-ref refs/myblobs/blob7 \"$obj\" &&\n     ++\tobj=$(git hash-object -w blob8) &&\n     ++        git update-ref refs/myblobs/blob8 \"$obj\"\n      +'\n      +\n      +test_expect_success 'Verify sorts with raw' '\n     @@ t/t6300-for-each-ref.sh: test_atom refs/myblobs/first contents:body \"\"\n      +\t\trefs/myblobs/ refs/heads/ >actual &&\n      +\ttest_cmp expected actual\n      +'\n     ++\n      +test_expect_success 'validate raw atom with %(if:notequals)' '\n      +\tcat >expected <<-EOF &&\n      +\trefs/heads/ambiguous\n     @@ t/t6300-for-each-ref.sh: test_atom refs/myblobs/first contents:body \"\"\n      +\ttest_cmp expected actual\n      +'\n      +\n     -+test_expect_success '%(raw) with --python must failed' '\n     ++test_expect_success '%(raw) with --python must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n      +'\n      +\n     -+test_expect_success '%(raw) with --tcl must failed' '\n     ++test_expect_success '%(raw) with --tcl must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n      +'\n      +\n     -+test_expect_success '%(raw) with --perl must failed' '\n     ++test_expect_success '%(raw) with --perl must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n      +'\n      +\n     -+test_expect_success '%(raw) with --shell must failed' '\n     ++test_expect_success '%(raw) with --shell must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n      +'\n      +\n     -+test_expect_success '%(raw) with --shell and --sort=raw must failed' '\n     ++test_expect_success '%(raw) with --shell and --sort=raw must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n      +'\n      +\n  -:  ------------ >  3:  b54dbc431e04 [GSOC] ref-filter: --format=%(raw) re-support --perl\n  3:  c99d1d070a18 =  4:  9fbbb3c492f5 [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  4:  5a5b5f78aeea !  5:  39a0d93c7bc1 [GSOC] ref-filter: add %(rest) atom\n     @@ ref-filter.c: static struct {\n       \t * Please update $__git_ref_fieldlist in git-completion.bash\n       \t * when you add new atoms\n      @@ ref-filter.c: int verify_ref_format(struct ref_format *format)\n     - \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n       \t\tif (at < 0)\n       \t\t\tdie(\"%s\", err.buf);\n     + \n      +\t\tif (used_atom[at].atom_type == ATOM_REST)\n      +\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n      +\n     - \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n     - \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n     - \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n     + \t\tif ((format->quote_style == QUOTE_PYTHON ||\n     + \t\t     format->quote_style == QUOTE_SHELL ||\n     + \t\t     format->quote_style == QUOTE_TCL) &&\n      @@ ref-filter.c: static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n       \t\t\tv->handler = else_atom_handler;\n       \t\t\tv->s = xstrdup(\"\");\n     @@ t/t3203-branch-output.sh: test_expect_success 'git branch --format option' '\n       \ttest_cmp expect actual\n       '\n       \n     -+test_expect_success 'git branch with --format=%(rest) must failed' '\n     ++test_expect_success 'git branch with --format=%(rest) must fail' '\n      +\ttest_must_fail git branch --format=\"%(rest)\" >actual\n      +'\n      +\n     @@ t/t6300-for-each-ref.sh: test_expect_success 'basic atom: head contents:trailers\n       \ttest_cmp expect actual.clean\n       '\n       \n     -+test_expect_success 'basic atom: rest must failed' '\n     ++test_expect_success 'basic atom: rest must fail' '\n      +\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n      +'\n      +\n     @@ t/t7004-tag.sh: test_expect_success '--format should list tags as per format giv\n       \ttest_cmp expect actual\n       '\n       \n     -+test_expect_success 'git tag -l with --format=\"%(rest)\" must failed' '\n     ++test_expect_success 'git tag -l with --format=\"%(rest)\" must fail' '\n      +\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n      +'\n      +\n     @@ t/t7030-verify-tag.sh: test_expect_success GPG 'verifying tag with --format' '\n       \ttest_cmp expect actual\n       '\n       \n     -+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must failed' '\n     ++test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must fail' '\n      +\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n      +'\n      +\n  5:  49063372e003 !  6:  35a376db1fc1 [GSOC] ref-filter: teach get_object() return useful value\n     @@ Metadata\n      Author: ZheNing Hu <adlternative@gmail.com>\n      \n       ## Commit message ##\n     -    [GSOC] ref-filter: teach get_object() return useful value\n     +    [GSOC] ref-filter: pass get_object() return value to their callers\n      \n     -    Let `populate_value()`, `get_ref_atom_value()` and\n     -    `format_ref_array_item()` get the return value of `get_object()`\n     -    correctly. This can help us later let `cat-file --batch` get the\n     -    correct error message and return value of `get_object()`.\n     +    Since in the refactor of `git cat-file --batch` later,\n     +    oid_object_info_extended() in get_object() will be used to obtain\n     +    the info of an object with it's oid. When the object cannot be\n     +    obtained in the git repository, `cat-file --batch` expects to output\n     +    \"<oid> missing\" and continue the next oid query instead of letting\n     +    Git exit. In other error conditions, Git should exit normally. So we\n     +    can achieve this function by passing the return value of get_object().\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n     +    Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n          Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n      \n       ## ref-filter.c ##\n     -@@ ref-filter.c: static char *get_worktree_path(const struct used_atom *atom, const struct ref_ar\n     - static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n     +@@ ref-filter.c: static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n       {\n       \tstruct object *obj;\n     --\tint i;\n     -+\tint i, ret = 0;\n     + \tint i;\n     ++\tint ret = 0;\n       \tstruct object_info empty = OBJECT_INFO_INIT;\n       \n       \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n     @@ ref-filter.c: static int populate_value(struct ref_array_item *ref, struct strbu\n       \toi.oid = ref->objectname;\n      -\tif (get_object(ref, 0, &obj, &oi, err))\n      -\t\treturn -1;\n     -+\tif ((ret = get_object(ref, 0, &obj, &oi, err)))\n     ++\tret = get_object(ref, 0, &obj, &oi, err);\n     ++\tif (ret)\n      +\t\treturn ret;\n       \n       \t/*\n       \t * If there is no atom that wants to know about tagged\n     -@@ ref-filter.c: static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n     - static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n     +@@ ref-filter.c: static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n       \t\t\t      struct atom_value **v, struct strbuf *err)\n       {\n     -+\tint ret = 0;\n     -+\n       \tif (!ref->value) {\n      -\t\tif (populate_value(ref, err))\n      -\t\t\treturn -1;\n     -+\t\tif ((ret = populate_value(ref, err)))\n     ++\t\tint ret = populate_value(ref, err);\n     ++\n     ++\t\tif (ret)\n      +\t\t\treturn ret;\n       \t\tfill_missing_values(ref->value);\n       \t}\n     @@ ref-filter.c: int format_ref_array_item(struct ref_array_item *info,\n       \t\t}\n       \t}\n       \tif (*cp) {\n     -@@ ref-filter.c: int format_ref_array_item(struct ref_array_item *info,\n     - \t}\n     - \tstrbuf_addbuf(final_buf, &state.stack->output);\n     - \tpop_stack_element(&state.stack);\n     --\treturn 0;\n     -+\treturn ret;\n     - }\n     - \n     - void pretty_print_ref(const char *name, const struct object_id *oid,\n  6:  d2f2563eb76a !  7:  8c1d683ec6e9 [GSOC] ref-filter: introduce free_array_item_internal() function\n     @@ Metadata\n      Author: ZheNing Hu <adlternative@gmail.com>\n      \n       ## Commit message ##\n     -    [GSOC] ref-filter: introduce free_array_item_internal() function\n     +    [GSOC] ref-filter: introduce free_ref_array_item_value() function\n      \n     -    Introduce free_array_item_internal() for freeing ref_array_item value.\n     +    When we use ref_array_item which is not dynamically allocated and\n     +    want to free the space of its member \"value\" after the end of use,\n     +    free_array_item() does not meet our needs, because it tries to free\n     +    ref_array_item itself and its member \"symref\".\n     +\n     +    Introduce free_ref_array_item_value() for freeing ref_array_item value.\n          It will be called internally by free_array_item(), and it will help\n     -    `cat-file --batch` free ref_array_item's memory later.\n     +    `cat-file --batch` free ref_array_item's value memory later.\n      \n     +    Helped-by: Junio C Hamano <gitster@pobox.com>\n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n          Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n     @@ ref-filter.c: static int ref_filter_handler(const char *refname, const struct ob\n       \n      -/*  Free memory allocated for a ref_array_item */\n      -static void free_array_item(struct ref_array_item *item)\n     -+void free_array_item_internal(struct ref_array_item *item)\n     ++void free_ref_array_item_value(struct ref_array_item *item)\n       {\n      -\tfree((char *)item->symref);\n       \tif (item->value) {\n     @@ ref-filter.c: static int ref_filter_handler(const char *refname, const struct ob\n      +static void free_array_item(struct ref_array_item *item)\n      +{\n      +\tfree((char *)item->symref);\n     -+\tfree_array_item_internal(item);\n     ++\tfree_ref_array_item_value(item);\n       \tfree(item);\n       }\n       \n     @@ ref-filter.h: struct ref_format {\n       int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n       /*  Clear all memory allocated to ref_array */\n       void ref_array_clear(struct ref_array *array);\n     -+/* Free array item's value */\n     -+void free_array_item_internal(struct ref_array_item *item);\n     ++/* Free ref_array_item's value */\n     ++void free_ref_array_item_value(struct ref_array_item *item);\n       /*  Used to verify if the given format is correct and to parse out the used atoms */\n       int verify_ref_format(struct ref_format *format);\n       /*  Sort the given ref_array as per the ref_sorting provided */\n  7:  765337a46ab0 !  8:  bd534a266a40 [GSOC] cat-file: reuse ref-filter logic\n     @@ Commit message\n          `verify_ref_format()` to check atoms.\n          4. Use `has_object_file()` in `batch_one_object()` to check\n          whether the input object exists.\n     -    5. Use `format_ref_array_item()` in `batch_object_write()` to\n     +    5. Let get_object() return 1 and print \"<oid> missing\" instead\n     +    of returning -1 and printing \"missing object <oid> for <refname>\",\n     +    this can help `format_ref_array_item()` just report that the\n     +    object is missing without letting Git exit.\n     +    6. Use `format_ref_array_item()` in `batch_object_write()` to\n          get the formatted data corresponding to the object. If the\n          return value of `format_ref_array_item()` is equals to zero,\n          use `batch_write()` to print object data; else if the return\n          value less than zero, use `die()` to print the error message\n          and exit; else return value greater than zero, only print the\n          error message, but not exit.\n     -    6. Let get_object() return 1 and print \"<oid> missing\" instead\n     -    of returning -1 and printing \"missing object <oid> for <refname>\",\n     -    this can help `format_ref_array_item()` just report that the\n     -    object is missing without letting Git exit.\n     +    7. Use free_ref_array_item_value() to free ref_array_item's\n     +    value.\n      \n          Most of the atoms in `for-each-ref --format` are now supported,\n          such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      +\t\tprintf(\"%s\\n\", err.buf);\n      +\t\tfflush(stdout);\n       \t}\n     -+\tfree_array_item_internal(&item);\n     ++\tfree_ref_array_item_value(&item);\n      +\tstrbuf_release(&err);\n       }\n       \n     @@ builtin/cat-file.c: int cmd_cat_file(int argc, const char **argv, const char *pr\n      \n       ## ref-filter.c ##\n      @@ ref-filter.c: int verify_ref_format(struct ref_format *format)\n     - \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n       \t\tif (at < 0)\n       \t\t\tdie(\"%s\", err.buf);\n     + \n      -\t\tif (used_atom[at].atom_type == ATOM_REST)\n      -\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n      +\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n     @@ ref-filter.c: int verify_ref_format(struct ref_format *format)\n      +\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n      +\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n       \n     - \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n     - \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n     + \t\tif ((format->quote_style == QUOTE_PYTHON ||\n     + \t\t     format->quote_style == QUOTE_SHELL ||\n      @@ ref-filter.c: static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n       \t}\n       \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n     @@ t/t1006-cat-file.sh: test_expect_success 'cat-file --unordered works' '\n      +batch_test_atom() {\n      +\tif test \"$3\" = \"fail\"\n      +\tthen\n     -+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 mast failed\" \"\n     ++\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 must fail\" \"\n      +\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n      +\t\t\t$1\n      +\t\t\tEOF\n  8:  058b304686fd !  9:  b66ab0f2d569 [GSOC] cat-file: reuse err buf in batch_objet_write()\n     @@ Metadata\n      Author: ZheNing Hu <adlternative@gmail.com>\n      \n       ## Commit message ##\n     -    [GSOC] cat-file: reuse err buf in batch_objet_write()\n     +    [GSOC] cat-file: reuse err buf in batch_object_write()\n      \n          Reuse the `err` buffer in batch_object_write(), as the\n          buffer `scratch` does. This will reduce the overhead\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      +\t\tprintf(\"%s\\n\", err->buf);\n       \t\tfflush(stdout);\n       \t}\n     - \tfree_array_item_internal(&item);\n     + \tfree_ref_array_item_value(&item);\n      -\tstrbuf_release(&err);\n       }\n       \n  9:  cbf7d51933ea ! 10:  86ac3bcaecea [GSOC] cat-file: re-implement --textconv, --filters options\n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const st\n       \topt->format.format = format.buf;\n      +\tif (opt->cmdmode == 'c')\n      +\t\topt->format.use_textconv = 1;\n     -+\tif (opt->cmdmode == 'w')\n     ++\telse if (opt->cmdmode == 'w')\n      +\t\topt->format.use_filters = 1;\n      +\n       \tif (verify_ref_format(&opt->format))\n     @@ ref-filter.c: int verify_ref_format(struct ref_format *format)\n      +\t\tuse_filters = format->use_filters;\n      +\t\tuse_textconv = format->use_textconv;\n      +\n     - \t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n     - \t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n     - \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n     + \t\tif ((format->quote_style == QUOTE_PYTHON ||\n     + \t\t     format->quote_style == QUOTE_SHELL ||\n     + \t\t     format->quote_style == QUOTE_TCL) &&\n      @@ ref-filter.c: static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n       {\n       \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n     @@ ref-filter.c: static int get_object(struct ref_array_item *ref, int deref, struc\n      +\t\toi->info.typep = NULL;\n      +\t\toi->info.contentp = temp_contentp;\n      +\n     -+\t\tif (use_textconv) {\n     ++\t\tif (use_textconv && !ref->rest)\n     ++\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n     ++\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n     ++\t\tif (use_textconv && oi->type == OBJ_BLOB) {\n      +\t\t\tact_oi = *oi;\n     -+\n     -+\t\t\tif(!ref->rest)\n     -+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n     -+\t\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n     -+\t\t\tif (act_oi.type == OBJ_BLOB) {\n     -+\t\t\t\tif (textconv_object(the_repository,\n     -+\t\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n     -+\t\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n     -+\t\t\t\t\tactual_oi = &act_oi;\n     -+\t\t\t\t\tgoto success;\n     -+\t\t\t\t}\n     ++\t\t\tif (textconv_object(the_repository,\n     ++\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n     ++\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n     ++\t\t\t\tactual_oi = &act_oi;\n     ++\t\t\t\tgoto success;\n      +\t\t\t}\n      +\t\t}\n       \t}\n     @@ ref-filter.c: static int get_object(struct ref_array_item *ref, int deref, struc\n       \n       \tif (oi->info.contentp) {\n      -\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n     -+\t\tif (use_filters) {\n     -+\t\t\tif(!ref->rest)\n     -+\t\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n     -+\t\t\t\t\t\t       oid_to_hex(&oi->oid));\n     -+\t\t\tif (oi->type == OBJ_BLOB) {\n     -+\t\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n     -+\t\t\t\tstruct checkout_metadata meta;\n     -+\t\t\t\tact_oi = *oi;\n     ++\t\tif (use_filters && !ref->rest)\n     ++\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n     ++\t\t\t\t\t       oid_to_hex(&oi->oid));\n     ++\t\tif (use_filters && oi->type == OBJ_BLOB) {\n     ++\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n     ++\t\t\tstruct checkout_metadata meta;\n     ++\t\t\tact_oi = *oi;\n      +\n     -+\t\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n     -+\t\t\t\tif (convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta)) {\n     -+\t\t\t\t\tact_oi.size = strbuf.len;\n     -+\t\t\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n     -+\t\t\t\t\tactual_oi = &act_oi;\n     -+\t\t\t\t} else {\n     -+\t\t\t\t\tdie(\"could not convert '%s' %s\",\n     -+\t\t\t\t\t    oid_to_hex(&oi->oid), ref->rest);\n     -+\t\t\t\t}\n     -+\t\t\t}\n     ++\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n     ++\t\t\tif (!convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n     ++\t\t\t\tdie(\"could not convert '%s' %s\",\n     ++\t\t\t\t\toid_to_hex(&oi->oid), ref->rest);\n     ++\t\t\tact_oi.size = strbuf.len;\n     ++\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n     ++\t\t\tactual_oi = &act_oi;\n      +\t\t}\n      +\n      +success:\n     @@ ref-filter.h: struct ref_format {\n       };\n       \n      -#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n     -+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, 0, 0, -1 }\n     ++#define REF_FORMAT_INIT { .use_color = -1 }\n       \n       /*  Macros for checking --merged and --no-merged options */\n       #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n\n-- \ngitgitgadget\n"},{"id":"427948","messageId":"9fbbb3c492f5830d412c17662b086bbfcb76e499.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 04/10] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:54Z","receivedAt":"2021-06-19T07:03:13Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc098242..452558ec9575 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 797b20ffa612..d01a0266fb89 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2553,9 +2553,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2600,7 +2600,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a718965..74fb423fc89f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"427949","messageId":"ab497d66c1167981181173f19cfe7d8857c348a3.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 02/10] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:52Z","receivedAt":"2021-06-19T07:03:14Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, we need to find a way to record the length of the\natom_value's member `s`. Although strbuf can already record the\nstring and its length, if we want to replace the type of atom_value's\nmember `s` with strbuf, many places in ref-filter that are filled\nwith dynamically allocated mermory in `v->s` are not easy to replace.\nAt the same time, we need to check if `v->s == NULL` in\npopulate_value(), and strbuf cannot easily distinguish NULL and empty\nstrings, but c-style \"const char *\" can do it. So add a new member in\n`struct atom_value`: `s_size`, which can record raw object size, it\ncan help us add raw object data to the buffer or compare two buffers\nwhich contain raw object data.\n\nBeyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, `--perl` because if our binary raw data is passed to a\nvariable in the host language, the host language may not support\narbitrary binary data in the variables of its string type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 216 +++++++++++++++++++++++++++++\n 3 files changed, 337 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de706..7f1f0a1ca3b6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the host language may not support arbitrary binary data in the\n+variables of its string type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fbaf..7822be903071 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4d..9c5379e2f56f 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,174 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tobj=$(git hash-object -w blob1) &&\n+        git update-ref refs/myblobs/blob1 \"$obj\" &&\n+\tobj=$(git hash-object -w blob2) &&\n+        git update-ref refs/myblobs/blob2 \"$obj\" &&\n+\tobj=$(git hash-object -w blob3) &&\n+        git update-ref refs/myblobs/blob3 \"$obj\" &&\n+\tobj=$(git hash-object -w blob4) &&\n+        git update-ref refs/myblobs/blob4 \"$obj\" &&\n+\tobj=$(git hash-object -w blob5) &&\n+        git update-ref refs/myblobs/blob5 \"$obj\" &&\n+\tobj=$(git hash-object -w blob6) &&\n+        git update-ref refs/myblobs/blob6 \"$obj\" &&\n+\tobj=$(git hash-object -w blob7) &&\n+        git update-ref refs/myblobs/blob7 \"$obj\" &&\n+\tobj=$(git hash-object -w blob8) &&\n+        git update-ref refs/myblobs/blob8 \"$obj\"\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"427950","messageId":"8c1d683ec6e959b7511a66eb5ea1103cd49c45f9.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 07/10] [GSOC] ref-filter: introduce free_ref_array_item_value() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:57Z","receivedAt":"2021-06-19T07:03:16Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nWhen we use ref_array_item which is not dynamically allocated and\nwant to free the space of its member \"value\" after the end of use,\nfree_array_item() does not meet our needs, because it tries to free\nref_array_item itself and its member \"symref\".\n\nIntroduce free_ref_array_item_value() for freeing ref_array_item value.\nIt will be called internally by free_array_item(), and it will help\n`cat-file --batch` free ref_array_item's value memory later.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 ++++++++---\n ref-filter.h |  2 ++\n 2 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 58def6ccd33a..22315d4809dc 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -2291,16 +2291,21 @@ static int ref_filter_handler(const char *refname, const struct object_id *oid,\n \treturn 0;\n }\n \n-/*  Free memory allocated for a ref_array_item */\n-static void free_array_item(struct ref_array_item *item)\n+void free_ref_array_item_value(struct ref_array_item *item)\n {\n-\tfree((char *)item->symref);\n \tif (item->value) {\n \t\tint i;\n \t\tfor (i = 0; i < used_atom_cnt; i++)\n \t\t\tfree((char *)item->value[i].s);\n \t\tfree(item->value);\n \t}\n+}\n+\n+/*  Free memory allocated for a ref_array_item */\n+static void free_array_item(struct ref_array_item *item)\n+{\n+\tfree((char *)item->symref);\n+\tfree_ref_array_item_value(item);\n \tfree(item);\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 9dc07476a584..76f9af7b4676 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -111,6 +111,8 @@ struct ref_format {\n int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n /*  Clear all memory allocated to ref_array */\n void ref_array_clear(struct ref_array *array);\n+/* Free ref_array_item's value */\n+void free_ref_array_item_value(struct ref_array_item *item);\n /*  Used to verify if the given format is correct and to parse out the used atoms */\n int verify_ref_format(struct ref_format *format);\n /*  Sort the given ref_array as per the ref_sorting provided */\n-- \ngitgitgadget\n\n"},{"id":"427951","messageId":"39a0d93c7bc1d30cc76c2fabef6ff66a9ae3e738.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 05/10] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:55Z","receivedAt":"2021-06-19T07:03:18Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex d01a0266fb89..10c78de9cfa4 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1010,6 +1021,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1927,6 +1941,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2144,6 +2164,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89f..9dc07476a584 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a00..6e94c6db7b5a 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must fail' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 5556063c347d..82c0ad2cb115 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1211,6 +1211,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883e..082be85dffc7 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602bf..10faa645157e 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"427952","messageId":"35a376db1fc144c68f72273363882ded42501cfd.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 06/10] [GSOC] ref-filter: pass get_object() return value to their callers","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:56Z","receivedAt":"2021-06-19T07:03:19Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nSince in the refactor of `git cat-file --batch` later,\noid_object_info_extended() in get_object() will be used to obtain\nthe info of an object with it's oid. When the object cannot be\nobtained in the git repository, `cat-file --batch` expects to output\n\"<oid> missing\" and continue the next oid query instead of letting\nGit exit. In other error conditions, Git should exit normally. So we\ncan achieve this function by passing the return value of get_object().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 10c78de9cfa4..58def6ccd33a 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1816,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n \tint i;\n+\tint ret = 0;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1972,8 +1973,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tret = get_object(ref, 0, &obj, &oi, err);\n+\tif (ret)\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -2005,8 +2007,10 @@ static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tint ret = populate_value(ref, err);\n+\n+\t\tif (ret)\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2580,6 +2584,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret = 0;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2592,10 +2597,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n-- \ngitgitgadget\n\n"},{"id":"427953","messageId":"b66ab0f2d5699a8b48220bebaa9e2cc2ccfad4fa.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 09/10] [GSOC] cat-file: reuse err buf in batch_object_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:59Z","receivedAt":"2021-06-19T07:03:21Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex cbedb8e1c471..b5204493bd56 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,32 +212,33 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret = 0;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (!ret) {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n \t} else if (ret < 0) {\n-\t\tdie(\"%s\\n\", err.buf);\n+\t\tdie(\"%s\\n\", err->buf);\n \t} else {\n \t\t/* when ret > 0 , don't call die and print the err to stdout*/\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n \t}\n \tfree_ref_array_item_value(&item);\n-\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -291,7 +292,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -299,13 +300,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -361,6 +363,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -389,6 +392,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\tcb.opt = opt;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (opt->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -413,6 +417,7 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -441,12 +446,13 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, opt, &data);\n \t}\n \n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"427954","messageId":"bd534a266a401a8edbf8c4d12a2d9e44fcc79d70.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 08/10] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:58Z","receivedAt":"2021-06-19T07:03:31Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, the following\nmethods are used:\n\n1. Add `cat_file_mode` member in struct `ref_format`, this can\nhelp us reject atoms in verify_ref_format() which cat-file\ncannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n2. Change the type of member `format` in struct `batch_options`\nto `ref_format`, We can add format data in it.\n3. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n4. Use `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists.\n5. Let get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\",\nthis can help `format_ref_array_item()` just report that the\nobject is missing without letting Git exit.\n6. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue less than zero, use `die()` to print the error message\nand exit; else return value greater than zero, only print the\nerror message, but not exit.\n7. Use free_ref_array_item_value() to free ref_array_item's\nvalue.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have refname.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 248 +++++++-------------------------\n ref-filter.c                   |  15 +-\n ref-filter.h                   |   3 +-\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n t/t6301-for-each-ref-errors.sh |   2 +-\n 6 files changed, 323 insertions(+), 203 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd9..ef8ab952b2fa 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e83..cbedb8e1c471 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,87 +209,31 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret = 0;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (!ret) {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n+\t} else if (ret < 0) {\n+\t\tdie(\"%s\\n\", err.buf);\n+\t} else {\n+\t\t/* when ret > 0 , don't call die and print the err to stdout*/\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n \t}\n+\tfree_ref_array_item_value(&item);\n+\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n@@ -428,6 +284,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n@@ -488,42 +351,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *opt, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (opt->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (!opt->format.format)\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n+\telse\n+\t\tstrbuf_addstr(&format, opt->format.format);\n \tif (opt->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\topt->format.format = format.buf;\n+\tif (verify_ref_format(&opt->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (opt->cmdmode || opt->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (opt->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -556,6 +411,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -588,18 +444,13 @@ static int batch_objects(struct batch_options *opt)\n \t\tbatch_one_object(input.buf, &output, opt, &data);\n \t}\n \n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -622,7 +473,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -631,7 +482,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -670,6 +523,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -713,7 +567,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 22315d4809dc..181d99c92735 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1021,8 +1021,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n@@ -1742,8 +1749,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 76f9af7b4676..e5fe26e3fdf9 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n@@ -86,7 +87,7 @@ struct ref_format {\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74ad..c028e6548fed 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 must fail\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab534..3553f84a00c1 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"427955","messageId":"b54dbc431e043c3ea2530fa8d1b4ba7e743b3d61.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 03/10] [GSOC] ref-filter: --format=%(raw) re-support --perl","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:02:53Z","receivedAt":"2021-06-19T07:03:33Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause the perl language can handle binary data correctly,\nadd the function perl_quote_buf_with_len(), which can specify\nthe length of the data and prevent the data from being truncated\nat '\\0' to help `--format=\"%(raw)\"` re-support `--perl`.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |  2 +-\n quote.c                            | 17 +++++++++++++++++\n quote.h                            |  1 +\n ref-filter.c                       | 15 +++++++++++----\n t/t6300-for-each-ref.sh            | 19 +++++++++++++++++--\n 5 files changed, 47 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 7f1f0a1ca3b6..ea9b438c16f6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -241,7 +241,7 @@ raw:size::\n \tThe raw data size of the object.\n \n Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n-`--perl` because the host language may not support arbitrary binary data in the\n+because the host language may not support arbitrary binary data in the\n variables of its string type.\n \n The message in a commit or a tag object is `contents`, from which\ndiff --git a/quote.c b/quote.c\nindex 8a3a5e39eb12..26719d21d1e7 100644\n--- a/quote.c\n+++ b/quote.c\n@@ -471,6 +471,23 @@ void perl_quote_buf(struct strbuf *sb, const char *src)\n \tstrbuf_addch(sb, sq);\n }\n \n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len)\n+{\n+\tconst char sq = '\\'';\n+\tconst char bq = '\\\\';\n+\tconst char *c = src;\n+\tconst char *end = src + len;\n+\n+\tstrbuf_addch(sb, sq);\n+\twhile (c != end) {\n+\t\tif (*c == sq || *c == bq)\n+\t\t\tstrbuf_addch(sb, bq);\n+\t\tstrbuf_addch(sb, *c);\n+\t\tc++;\n+\t}\n+\tstrbuf_addch(sb, sq);\n+}\n+\n void python_quote_buf(struct strbuf *sb, const char *src)\n {\n \tconst char sq = '\\'';\ndiff --git a/quote.h b/quote.h\nindex 768cc6338e27..0fe69e264b0e 100644\n--- a/quote.h\n+++ b/quote.h\n@@ -94,6 +94,7 @@ char *quote_path(const char *in, const char *prefix, struct strbuf *out, unsigne\n \n /* quoting as a string literal for other languages */\n void perl_quote_buf(struct strbuf *sb, const char *src);\n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len);\n void python_quote_buf(struct strbuf *sb, const char *src);\n void tcl_quote_buf(struct strbuf *sb, const char *src);\n void basic_regex_quote_buf(struct strbuf *sb, const char *src);\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be903071..797b20ffa612 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -742,7 +742,10 @@ static void quote_formatting(struct strbuf *s, const char *str, size_t len, int\n \t\tsq_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PERL:\n-\t\tperl_quote_buf(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tperl_quote_buf_with_len(s, str, len);\n+\t\telse\n+\t\t\tperl_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PYTHON:\n \t\tpython_quote_buf(s, str);\n@@ -1006,10 +1009,14 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n-\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\n+\t\tif ((format->quote_style == QUOTE_PYTHON ||\n+\t\t     format->quote_style == QUOTE_SHELL ||\n+\t\t     format->quote_style == QUOTE_TCL) &&\n+\t\t     used_atom[at].atom_type == ATOM_RAW &&\n+\t\t     used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n-\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n+\t\t\t      \"--python, --shell, --tcl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9c5379e2f56f..5556063c347d 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -915,8 +915,23 @@ test_expect_success '%(raw) with --tcl must fail' '\n \ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n '\n \n-test_expect_success '%(raw) with --perl must fail' '\n-\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+test_expect_success '%(raw) with --perl' '\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob1 --perl | perl > actual &&\n+\tcmp blob1 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob3 --perl | perl > actual &&\n+\tcmp blob3 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob8 --perl | perl > actual &&\n+\tcmp blob8 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/first --perl | perl > actual &&\n+\tcmp one actual &&\n+\tgit cat-file tree refs/mytrees/first > expected &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/mytrees/first --perl | perl > actual &&\n+\tcmp expected actual\n '\n \n test_expect_success '%(raw) with --shell must fail' '\n-- \ngitgitgadget\n\n"},{"id":"427956","messageId":"86ac3bcaecea9f1b0411173ea5599a9d8127c2dd.1624086181.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v3 10/10] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-19T07:03:00Z","receivedAt":"2021-06-19T07:03:33Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  5 ++++\n ref-filter.c       | 59 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  4 +++-\n 3 files changed, 65 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex b5204493bd56..a78cc076aa59 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -377,6 +377,11 @@ static int batch_objects(struct batch_options *opt, const struct option *options\n \tif (opt->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \topt->format.format = format.buf;\n+\tif (opt->cmdmode == 'c')\n+\t\topt->format.use_textconv = 1;\n+\telse if (opt->cmdmode == 'w')\n+\t\topt->format.use_filters = 1;\n+\n \tif (verify_ref_format(&opt->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 181d99c92735..99b87742b0fd 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1031,6 +1035,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1742,10 +1749,38 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\tif (use_textconv && oi->type == OBJ_BLOB) {\n+\t\t\tact_oi = *oi;\n+\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\tgoto success;\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1755,19 +1790,39 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\tif (use_filters && oi->type == OBJ_BLOB) {\n+\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\tstruct checkout_metadata meta;\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\tif (!convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n+\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\toid_to_hex(&oi->oid), ref->rest);\n+\t\t\tact_oi.size = strbuf.len;\n+\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\tactual_oi = &act_oi;\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex e5fe26e3fdf9..497e3e93632f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,6 +80,8 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n \n@@ -87,7 +89,7 @@ struct ref_format {\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n+#define REF_FORMAT_INIT { .use_color = -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n-- \ngitgitgadget\n"},{"id":"428017","messageId":"CAP8UFD18iumUEgayGxL612MmbK9_2uDpHz3i7aqJ5zYSh7skqg@mail.gmail.com","threadId":"55909","inReplyTo":"bd534a266a401a8edbf8c4d12a2d9e44fcc79d70.1624086181.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 08/10] [GSOC] cat-file: reuse ref-filter logic","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-06-21T05:55:42Z","receivedAt":"2021-06-21T05:55:57Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Sat, Jun 19, 2021 at 9:03 AM ZheNing Hu via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> In order to let cat-file use ref-filter logic, the following\n> methods are used:\n\nMaybe: s/the following methods are used/let's do the following/\n\n> 1. Add `cat_file_mode` member in struct `ref_format`, this can\n> help us reject atoms in verify_ref_format() which cat-file\n> cannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n> 2. Change the type of member `format` in struct `batch_options`\n> to `ref_format`, We can add format data in it.\n\nNot sure what \"We can add format data in it.\" means.\n\n> 3. Let `batch_objects()` add atoms to format, and use\n> `verify_ref_format()` to check atoms.\n> 4. Use `has_object_file()` in `batch_one_object()` to check\n> whether the input object exists.\n> 5. Let get_object() return 1 and print \"<oid> missing\" instead\n> of returning -1 and printing \"missing object <oid> for <refname>\",\n> this can help `format_ref_array_item()` just report that the\n> object is missing without letting Git exit.\n> 6. Use `format_ref_array_item()` in `batch_object_write()` to\n> get the formatted data corresponding to the object. If the\n> return value of `format_ref_array_item()` is equals to zero,\n> use `batch_write()` to print object data; else if the return\n> value less than zero, use `die()` to print the error message\n> and exit; else return value greater than zero, only print the\n\ns/else return value greater/else if the return value is greater/\n\n> error message, but not exit.\n\ns/not exit/don't exit/\n\n> 7. Use free_ref_array_item_value() to free ref_array_item's\n> value.\n\nThat looks like a lot of changes in a single commit. I wonder if this\ncommit could be split.\n\n> Most of the atoms in `for-each-ref --format` are now supported,\n> such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n> `%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n> `%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n> `%(flag)`, `%(HEAD)`, because our objects don't have refname.\n\ns/refname/a refname/\n\nIt might be worth talking a bit about possible performance changes.\n\n[...]\n\n> +       ret = format_ref_array_item(&item, &opt->format, scratch, &err);\n> +       if (!ret) {\n> +               strbuf_addch(scratch, '\\n');\n> +               batch_write(opt, scratch->buf, scratch->len);\n> +       } else if (ret < 0) {\n> +               die(\"%s\\n\", err.buf);\n\nThis if (ret < 0) could be checked first.\n\n> +       } else {\n> +               /* when ret > 0 , don't call die and print the err to stdout*/\n\nI think it would be more helpful to tell what ret > 0 means, rather\nthan what we do below (which can easily be seen).\n\n> +               printf(\"%s\\n\", err.buf);\n> +               fflush(stdout);\n>         }\n\nFor example:\n\n       if (ret < 0) {\n               die(\"%s\\n\", err.buf);\n       if (ret) {\n               /* ret > 0 means ... */\n               printf(\"%s\\n\", err.buf);\n               fflush(stdout);\n       } else {\n               strbuf_addch(scratch, '\\n');\n               batch_write(opt, scratch->buf, scratch->len);\n       }\n\n> +       free_ref_array_item_value(&item);\n> +       strbuf_release(&err);\n>  }\n\n[...]\n\n> +static int batch_objects(struct batch_options *opt, const struct option *options)\n\nIt's unfortunate that one argument is called \"opt\" and the other one\n\"options\". I wonder if the first one could be called \"batch\" as it\nseems to be called this way somewhere else.\n\n> +       if (!opt->format.format)\n> +               strbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n> +       else\n> +               strbuf_addstr(&format, opt->format.format);\n\nIf there is no reason for the condition to be (!X) over just (X), I\nthink the latter is a bit better.\n\n>         if (opt->print_contents)\n> -               data.info.typep = &data.type;\n> +               strbuf_addstr(&format, \"\\n%(raw)\");\n> +       opt->format.format = format.buf;\n\nI wonder if this should be:\n\n      opt->format.format = strbuf_detach(&format, NULL);\n\n> +       if (verify_ref_format(&opt->format))\n> +               usage_with_options(cat_file_usage, options);\n\n[...]\n\n> @@ -86,7 +87,7 @@ struct ref_format {\n>         int need_color_reset_at_eol;\n>  };\n>\n> -#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n> +#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n\nMaybe this can already be changed to a designated initializer, like\nÆvar suggested recently.\n\n[...]\n\n> +test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n\nIf this test is about checking that %(rest) works with both a branch\nand a tag, it might be better to say it more explicitly.\n\n> +       cat >expected <<-EOF &&\n> +       123 commit 123\n> +       456 tag 456\n> +       EOF\n> +       git cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n> +       refs/heads/main 123\n> +       refs/tags/testtag 456\n> +       EOF\n> +       test_cmp expected actual\n> +'\n"},{"id":"428033","messageId":"CAOLTT8TU+E+GaxR+fMvH5Rc-C+BEJDpqjmmrF_BO5kG-i-_gsQ@mail.gmail.com","threadId":"55909","inReplyTo":"CAP8UFD18iumUEgayGxL612MmbK9_2uDpHz3i7aqJ5zYSh7skqg@mail.gmail.com","subject":"Re: [PATCH v3 08/10] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-21T13:05:32Z","receivedAt":"2021-06-21T13:05:46Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Christian Couder <christian.couder@gmail.com> 于2021年6月21日周一 下午1:55写道：\n>\n> > 1. Add `cat_file_mode` member in struct `ref_format`, this can\n> > help us reject atoms in verify_ref_format() which cat-file\n> > cannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n> > 2. Change the type of member `format` in struct `batch_options`\n> > to `ref_format`, We can add format data in it.\n>\n> Not sure what \"We can add format data in it.\" means.\n\nWell, there is something wrong with the expression here. What I want\nto express is that we can fill its member \"format\" with the atoms like\n\"%(objectname) %(refname)\", and then pass it to the ref-filter.\n\n> > 7. Use free_ref_array_item_value() to free ref_array_item's\n> > value.\n>\n> That looks like a lot of changes in a single commit. I wonder if this\n> commit could be split.\n>\n\nYeah, But I don’t know if I should take it apart step by step, If taken apart,\nthose intermediate commits are likely to fail the test.\n\n> > Most of the atoms in `for-each-ref --format` are now supported,\n> > such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n> > `%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n> > `%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n> > `%(flag)`, `%(HEAD)`, because our objects don't have refname.\n>\n> s/refname/a refname/\n>\n> It might be worth talking a bit about possible performance changes.\n>\n\nMakes sense. The performance has indeed deteriorated.\n\n> [...]\n>\n> > +       ret = format_ref_array_item(&item, &opt->format, scratch, &err);\n> > +       if (!ret) {\n> > +               strbuf_addch(scratch, '\\n');\n> > +               batch_write(opt, scratch->buf, scratch->len);\n> > +       } else if (ret < 0) {\n> > +               die(\"%s\\n\", err.buf);\n>\n> This if (ret < 0) could be checked first.\n\nYes, it is better to put error checking first.\n\n>\n> > +       } else {\n> > +               /* when ret > 0 , don't call die and print the err to stdout*/\n>\n> I think it would be more helpful to tell what ret > 0 means, rather\n> than what we do below (which can easily be seen).\n>\n\nAh, There is indeed only one situation for ret > 0 for the time being:\nShow \"<oid> missing\" without exiting Git.\n\n> > +               printf(\"%s\\n\", err.buf);\n> > +               fflush(stdout);\n> >         }\n>\n> For example:\n>\n>        if (ret < 0) {\n>                die(\"%s\\n\", err.buf);\n>        if (ret) {\n>                /* ret > 0 means ... */\n>                printf(\"%s\\n\", err.buf);\n>                fflush(stdout);\n>        } else {\n>                strbuf_addch(scratch, '\\n');\n>                batch_write(opt, scratch->buf, scratch->len);\n>        }\n>\n\nYes, this might be better.\n\n> > +       free_ref_array_item_value(&item);\n> > +       strbuf_release(&err);\n> >  }\n>\n> [...]\n>\n> > +static int batch_objects(struct batch_options *opt, const struct option *options)\n>\n> It's unfortunate that one argument is called \"opt\" and the other one\n> \"options\". I wonder if the first one could be called \"batch\" as it\n> seems to be called this way somewhere else.\n>\n\nOK.\n\n> > +       if (!opt->format.format)\n> > +               strbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n> > +       else\n> > +               strbuf_addstr(&format, opt->format.format);\n>\n> If there is no reason for the condition to be (!X) over just (X), I\n> think the latter is a bit better.\n>\n\nAlthough I think it doesn’t matter which one to use first,\nI will still follow your suggestions.\n\n> >         if (opt->print_contents)\n> > -               data.info.typep = &data.type;\n> > +               strbuf_addstr(&format, \"\\n%(raw)\");\n> > +       opt->format.format = format.buf;\n>\n> I wonder if this should be:\n>\n>       opt->format.format = strbuf_detach(&format, NULL);\n>\n\nNo. here our opt->format.format will not be changed, it would\nbe better for us to use `strbuf_release(&format)` for resource\nrecovery. (strbuf_detach() will forced to let us free opt->format.format)\n\n> > +       if (verify_ref_format(&opt->format))\n> > +               usage_with_options(cat_file_usage, options);\n>\n> [...]\n>\n> > @@ -86,7 +87,7 @@ struct ref_format {\n> >         int need_color_reset_at_eol;\n> >  };\n> >\n> > -#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n> > +#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n>\n> Maybe this can already be changed to a designated initializer, like\n> Ævar suggested recently.\n>\n\nI agree.\n\n> [...]\n>\n> > +test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n>\n> If this test is about checking that %(rest) works with both a branch\n> and a tag, it might be better to say it more explicitly.\n>\n\nOK.\n\nThanks.\n--\nZheNing Hu\n"},{"id":"428158","messageId":"f72ad9cc5e8b1f1139784df9ac7178f1561f70bb.1624332054.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 01/14] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:41Z","receivedAt":"2021-06-22T03:21:04Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c6..5cee6512fbaf 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"428159","messageId":"b54dbc431e043c3ea2530fa8d1b4ba7e743b3d61.1624332054.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 03/14] [GSOC] ref-filter: --format=%(raw) re-support --perl","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:43Z","receivedAt":"2021-06-22T03:21:04Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause the perl language can handle binary data correctly,\nadd the function perl_quote_buf_with_len(), which can specify\nthe length of the data and prevent the data from being truncated\nat '\\0' to help `--format=\"%(raw)\"` re-support `--perl`.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |  2 +-\n quote.c                            | 17 +++++++++++++++++\n quote.h                            |  1 +\n ref-filter.c                       | 15 +++++++++++----\n t/t6300-for-each-ref.sh            | 19 +++++++++++++++++--\n 5 files changed, 47 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 7f1f0a1ca3b6..ea9b438c16f6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -241,7 +241,7 @@ raw:size::\n \tThe raw data size of the object.\n \n Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n-`--perl` because the host language may not support arbitrary binary data in the\n+because the host language may not support arbitrary binary data in the\n variables of its string type.\n \n The message in a commit or a tag object is `contents`, from which\ndiff --git a/quote.c b/quote.c\nindex 8a3a5e39eb12..26719d21d1e7 100644\n--- a/quote.c\n+++ b/quote.c\n@@ -471,6 +471,23 @@ void perl_quote_buf(struct strbuf *sb, const char *src)\n \tstrbuf_addch(sb, sq);\n }\n \n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len)\n+{\n+\tconst char sq = '\\'';\n+\tconst char bq = '\\\\';\n+\tconst char *c = src;\n+\tconst char *end = src + len;\n+\n+\tstrbuf_addch(sb, sq);\n+\twhile (c != end) {\n+\t\tif (*c == sq || *c == bq)\n+\t\t\tstrbuf_addch(sb, bq);\n+\t\tstrbuf_addch(sb, *c);\n+\t\tc++;\n+\t}\n+\tstrbuf_addch(sb, sq);\n+}\n+\n void python_quote_buf(struct strbuf *sb, const char *src)\n {\n \tconst char sq = '\\'';\ndiff --git a/quote.h b/quote.h\nindex 768cc6338e27..0fe69e264b0e 100644\n--- a/quote.h\n+++ b/quote.h\n@@ -94,6 +94,7 @@ char *quote_path(const char *in, const char *prefix, struct strbuf *out, unsigne\n \n /* quoting as a string literal for other languages */\n void perl_quote_buf(struct strbuf *sb, const char *src);\n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len);\n void python_quote_buf(struct strbuf *sb, const char *src);\n void tcl_quote_buf(struct strbuf *sb, const char *src);\n void basic_regex_quote_buf(struct strbuf *sb, const char *src);\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be903071..797b20ffa612 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -742,7 +742,10 @@ static void quote_formatting(struct strbuf *s, const char *str, size_t len, int\n \t\tsq_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PERL:\n-\t\tperl_quote_buf(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tperl_quote_buf_with_len(s, str, len);\n+\t\telse\n+\t\t\tperl_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PYTHON:\n \t\tpython_quote_buf(s, str);\n@@ -1006,10 +1009,14 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n-\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\n+\t\tif ((format->quote_style == QUOTE_PYTHON ||\n+\t\t     format->quote_style == QUOTE_SHELL ||\n+\t\t     format->quote_style == QUOTE_TCL) &&\n+\t\t     used_atom[at].atom_type == ATOM_RAW &&\n+\t\t     used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n-\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n+\t\t\t      \"--python, --shell, --tcl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9c5379e2f56f..5556063c347d 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -915,8 +915,23 @@ test_expect_success '%(raw) with --tcl must fail' '\n \ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n '\n \n-test_expect_success '%(raw) with --perl must fail' '\n-\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+test_expect_success '%(raw) with --perl' '\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob1 --perl | perl > actual &&\n+\tcmp blob1 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob3 --perl | perl > actual &&\n+\tcmp blob3 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob8 --perl | perl > actual &&\n+\tcmp blob8 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/first --perl | perl > actual &&\n+\tcmp one actual &&\n+\tgit cat-file tree refs/mytrees/first > expected &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/mytrees/first --perl | perl > actual &&\n+\tcmp expected actual\n '\n \n test_expect_success '%(raw) with --shell must fail' '\n-- \ngitgitgadget\n\n"},{"id":"428160","messageId":"ab497d66c1167981181173f19cfe7d8857c348a3.1624332054.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 02/14] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:42Z","receivedAt":"2021-06-22T03:21:04Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, we need to find a way to record the length of the\natom_value's member `s`. Although strbuf can already record the\nstring and its length, if we want to replace the type of atom_value's\nmember `s` with strbuf, many places in ref-filter that are filled\nwith dynamically allocated mermory in `v->s` are not easy to replace.\nAt the same time, we need to check if `v->s == NULL` in\npopulate_value(), and strbuf cannot easily distinguish NULL and empty\nstrings, but c-style \"const char *\" can do it. So add a new member in\n`struct atom_value`: `s_size`, which can record raw object size, it\ncan help us add raw object data to the buffer or compare two buffers\nwhich contain raw object data.\n\nBeyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, `--perl` because if our binary raw data is passed to a\nvariable in the host language, the host language may not support\narbitrary binary data in the variables of its string type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 216 +++++++++++++++++++++++++++++\n 3 files changed, 337 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de706..7f1f0a1ca3b6 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the host language may not support arbitrary binary data in the\n+variables of its string type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fbaf..7822be903071 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4d..9c5379e2f56f 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,174 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tobj=$(git hash-object -w blob1) &&\n+        git update-ref refs/myblobs/blob1 \"$obj\" &&\n+\tobj=$(git hash-object -w blob2) &&\n+        git update-ref refs/myblobs/blob2 \"$obj\" &&\n+\tobj=$(git hash-object -w blob3) &&\n+        git update-ref refs/myblobs/blob3 \"$obj\" &&\n+\tobj=$(git hash-object -w blob4) &&\n+        git update-ref refs/myblobs/blob4 \"$obj\" &&\n+\tobj=$(git hash-object -w blob5) &&\n+        git update-ref refs/myblobs/blob5 \"$obj\" &&\n+\tobj=$(git hash-object -w blob6) &&\n+        git update-ref refs/myblobs/blob6 \"$obj\" &&\n+\tobj=$(git hash-object -w blob7) &&\n+        git update-ref refs/myblobs/blob7 \"$obj\" &&\n+\tobj=$(git hash-object -w blob8) &&\n+        git update-ref refs/myblobs/blob8 \"$obj\"\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"428161","messageId":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v3.git.1624086181.gitgitgadget@gmail.com","subject":"[PATCH v4 00/14] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:40Z","receivedAt":"2021-06-22T03:21:04Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic.\n\nChange from last version:\n\n 1. At the suggestion of Christian Couder, split \"[GSOC] cat-file: reuse\n    ref-filter logic\" into multiple small commits: \"[GSOC] cat-file: add\n    has_object_file() check\", \"[GSOC] ref-filter: modify the error message\n    and value in get_object\", \"[GSOC] ref-filter: add cat_file_mode in\n    struct ref_format\".\n 2. Change batch_objects parameter name from \"opt\" to \"batch\".\n 3. Modify test subject.\n 4. Begin to use the explicit initialization of REF_FORMAT_INIT in the\n    design of %(rest).\n 5. Fix some grammatical errors and code style issues.\n\nZheNing Hu (14):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: --format=%(raw) re-support --perl\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: pass get_object() return value to their callers\n  [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  [GSOC] ref-filter: modify the error message and value in get_object\n  [GSOC] cat-file: add has_object_file() check\n  [GSOC] cat-file: change batch_objects parameter name\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_object_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 277 +++++++-----------------\n builtin/tag.c                      |   2 +-\n quote.c                            |  17 ++\n quote.h                            |   1 +\n ref-filter.c                       | 331 +++++++++++++++++++++++------\n ref-filter.h                       |  14 +-\n t/t1006-cat-file.sh                | 252 ++++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 235 ++++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 14 files changed, 879 insertions(+), 279 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n\nRange-diff vs v3:\n\n  1:  f72ad9cc5e8b =  1:  f72ad9cc5e8b [GSOC] ref-filter: add obj-type check in grab contents\n  2:  ab497d66c116 =  2:  ab497d66c116 [GSOC] ref-filter: add %(raw) atom\n  3:  b54dbc431e04 =  3:  b54dbc431e04 [GSOC] ref-filter: --format=%(raw) re-support --perl\n  4:  9fbbb3c492f5 =  4:  9fbbb3c492f5 [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  5:  39a0d93c7bc1 !  5:  08aa44e5e57b [GSOC] ref-filter: add %(rest) atom\n     @@ ref-filter.h: struct ref_format {\n       };\n       \n      -#define REF_FORMAT_INIT { NULL, 0, -1 }\n     -+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n     ++#define REF_FORMAT_INIT { .use_color = -1 }\n       \n       /*  Macros for checking --merged and --no-merged options */\n       #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n  6:  35a376db1fc1 =  6:  05682bccf9f9 [GSOC] ref-filter: pass get_object() return value to their callers\n  7:  8c1d683ec6e9 =  7:  06db6cd6f1f9 [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  -:  ------------ >  8:  b0d9e139935f [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  -:  ------------ >  9:  db7dd8b042c2 [GSOC] ref-filter: modify the error message and value in get_object\n  -:  ------------ > 10:  6b577969734e [GSOC] cat-file: add has_object_file() check\n  -:  ------------ > 11:  069aa203666a [GSOC] cat-file: change batch_objects parameter name\n  8:  bd534a266a40 ! 12:  258ec0a46c56 [GSOC] cat-file: reuse ref-filter logic\n     @@ Metadata\n       ## Commit message ##\n          [GSOC] cat-file: reuse ref-filter logic\n      \n     -    In order to let cat-file use ref-filter logic, the following\n     -    methods are used:\n     +    In order to let cat-file use ref-filter logic, let's do the\n     +    following:\n      \n     -    1. Add `cat_file_mode` member in struct `ref_format`, this can\n     -    help us reject atoms in verify_ref_format() which cat-file\n     -    cannot use, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n     -    2. Change the type of member `format` in struct `batch_options`\n     -    to `ref_format`, We can add format data in it.\n     -    3. Let `batch_objects()` add atoms to format, and use\n     +    1. Change the type of member `format` in struct `batch_options`\n     +    to `ref_format`, we will pass it to ref-filter later.\n     +    2. Let `batch_objects()` add atoms to format, and use\n          `verify_ref_format()` to check atoms.\n     -    4. Use `has_object_file()` in `batch_one_object()` to check\n     -    whether the input object exists.\n     -    5. Let get_object() return 1 and print \"<oid> missing\" instead\n     -    of returning -1 and printing \"missing object <oid> for <refname>\",\n     -    this can help `format_ref_array_item()` just report that the\n     -    object is missing without letting Git exit.\n     -    6. Use `format_ref_array_item()` in `batch_object_write()` to\n     +    3. Use `format_ref_array_item()` in `batch_object_write()` to\n          get the formatted data corresponding to the object. If the\n          return value of `format_ref_array_item()` is equals to zero,\n          use `batch_write()` to print object data; else if the return\n     -    value less than zero, use `die()` to print the error message\n     -    and exit; else return value greater than zero, only print the\n     -    error message, but not exit.\n     -    7. Use free_ref_array_item_value() to free ref_array_item's\n     +    value is less than zero, use `die()` to print the error message\n     +    and exit; else if return value is greater than zero, only print\n     +    the error message, but don't exit.\n     +    4. Use free_ref_array_item_value() to free ref_array_item's\n          value.\n      \n          Most of the atoms in `for-each-ref --format` are now supported,\n          such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n          `%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n          `%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n     -    `%(flag)`, `%(HEAD)`, because our objects don't have refname.\n     +    `%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n     +\n     +    The performance for `git cat-file --batch-all-objects\n     +    --batch-check` on the Git repository itself with performance\n     +    testing tool `hyperfine` changes from 669.4 ms ±  31.1 ms to\n     +    1.134 s ±  0.063 s.\n     +\n     +    The performance for `git cat-file --batch-all-objects --batch\n     +    >/dev/null` on the Git repository itself with performance testing\n     +    tool `time` change from \"27.37s user 0.29s system 98% cpu 28.089\n     +    total\" to \"33.69s user 1.54s system 87% cpu 40.258 total\".\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      -\t\tprint_object_or_die(opt, data);\n      -\t\tbatch_write(opt, \"\\n\", 1);\n      +\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n     -+\tif (!ret) {\n     -+\t\tstrbuf_addch(scratch, '\\n');\n     -+\t\tbatch_write(opt, scratch->buf, scratch->len);\n     -+\t} else if (ret < 0) {\n     ++\tif (ret < 0) {\n      +\t\tdie(\"%s\\n\", err.buf);\n     -+\t} else {\n     -+\t\t/* when ret > 0 , don't call die and print the err to stdout*/\n     ++\t} if (ret) {\n     ++\t\t/* ret > 0 means when the object corresponding to oid\n     ++\t\t * cannot be found in format_ref_array_item(), we only print\n     ++\t\t * the error message.\n     ++\t\t */\n      +\t\tprintf(\"%s\\n\", err.buf);\n      +\t\tfflush(stdout);\n     ++\t} else {\n     ++\t\tstrbuf_addch(scratch, '\\n');\n     ++\t\tbatch_write(opt, scratch->buf, scratch->len);\n       \t}\n      +\tfree_ref_array_item_value(&item);\n      +\tstrbuf_release(&err);\n       }\n       \n       static void batch_one_object(const char *obj_name,\n     -@@ builtin/cat-file.c: static void batch_one_object(const char *obj_name,\n     - \t\treturn;\n     - \t}\n     - \n     -+\tif (!has_object_file(&data->oid)) {\n     -+\t\tprintf(\"%s missing\\n\",\n     -+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n     -+\t\tfflush(stdout);\n     -+\t\treturn;\n     -+\t}\n     -+\n     - \tbatch_object_write(obj_name, scratch, opt, data);\n     - }\n     - \n      @@ builtin/cat-file.c: static int batch_unordered_packed(const struct object_id *oid,\n       \treturn batch_unordered_object(oid, data);\n       }\n       \n     --static int batch_objects(struct batch_options *opt)\n     +-static int batch_objects(struct batch_options *batch)\n      +static const char * const cat_file_usage[] = {\n      +\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n      +\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n      +\tNULL\n      +};\n      +\n     -+static int batch_objects(struct batch_options *opt, const struct option *options)\n     ++static int batch_objects(struct batch_options *batch, const struct option *options)\n       {\n       \tstruct strbuf input = STRBUF_INIT;\n       \tstruct strbuf output = STRBUF_INIT;\n     @@ builtin/cat-file.c: static int batch_unordered_packed(const struct object_id *oi\n       \tint save_warning;\n       \tint retval = 0;\n       \n     --\tif (!opt->format)\n     --\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n     +-\tif (!batch->format)\n     +-\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n      -\n      -\t/*\n      -\t * Expand once with our special mark_query flag, which will prime the\n     @@ builtin/cat-file.c: static int batch_unordered_packed(const struct object_id *oi\n      -\t */\n       \tmemset(&data, 0, sizeof(data));\n      -\tdata.mark_query = 1;\n     --\tstrbuf_expand(&output, opt->format, expand_format, &data);\n     +-\tstrbuf_expand(&output, batch->format, expand_format, &data);\n      -\tdata.mark_query = 0;\n      -\tstrbuf_release(&output);\n     --\tif (opt->cmdmode)\n     +-\tif (batch->cmdmode)\n      -\t\tdata.split_on_whitespace = 1;\n      -\n     --\tif (opt->all_objects) {\n     +-\tif (batch->all_objects) {\n      -\t\tstruct object_info empty = OBJECT_INFO_INIT;\n      -\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n      -\t\t\tdata.skip_object_info = 1;\n     @@ builtin/cat-file.c: static int batch_unordered_packed(const struct object_id *oi\n      -\t * If we are printing out the object, then always fill in the type,\n      -\t * since we will want to decide whether or not to stream.\n      -\t */\n     -+\tif (!opt->format.format)\n     -+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n     ++\tif (batch->format.format)\n     ++\t\tstrbuf_addstr(&format, batch->format.format);\n      +\telse\n     -+\t\tstrbuf_addstr(&format, opt->format.format);\n     - \tif (opt->print_contents)\n     ++\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n     + \tif (batch->print_contents)\n      -\t\tdata.info.typep = &data.type;\n      +\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n     -+\topt->format.format = format.buf;\n     -+\tif (verify_ref_format(&opt->format))\n     ++\tbatch->format.format = format.buf;\n     ++\tif (verify_ref_format(&batch->format))\n      +\t\tusage_with_options(cat_file_usage, options);\n      +\n     -+\tif (opt->cmdmode || opt->format.use_rest)\n     ++\tif (batch->cmdmode || batch->format.use_rest)\n      +\t\tdata.split_on_whitespace = 1;\n       \n     - \tif (opt->all_objects) {\n     + \tif (batch->all_objects) {\n       \t\tstruct object_cb_data cb;\n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt)\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch)\n       \t\t\toid_array_clear(&sa);\n       \t\t}\n       \n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt)\n       \t\tstrbuf_release(&output);\n       \t\treturn 0;\n       \t}\n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt)\n     - \t\tbatch_one_object(input.buf, &output, opt, &data);\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch)\n     + \t\tbatch_one_object(input.buf, &output, batch, &data);\n       \t}\n       \n      +\tstrbuf_release(&format);\n     @@ builtin/cat-file.c: int cmd_cat_file(int argc, const char **argv, const char *pr\n       \tif (unknown_type && opt != 't' && opt != 's')\n       \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\n      \n     - ## ref-filter.c ##\n     -@@ ref-filter.c: int verify_ref_format(struct ref_format *format)\n     - \t\tif (at < 0)\n     - \t\t\tdie(\"%s\", err.buf);\n     - \n     --\t\tif (used_atom[at].atom_type == ATOM_REST)\n     --\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n     -+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n     -+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n     -+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n     -+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n     - \n     - \t\tif ((format->quote_style == QUOTE_PYTHON ||\n     - \t\t     format->quote_style == QUOTE_SHELL ||\n     -@@ ref-filter.c: static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n     - \t}\n     - \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n     - \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n     --\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n     --\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n     -+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n     -+\t\t\t\t       oid_to_hex(&oi->oid));\n     - \tif (oi->info.disk_sizep && oi->disk_size < 0)\n     - \t\tBUG(\"Object size is less than zero.\");\n     - \n     -\n     - ## ref-filter.h ##\n     -@@ ref-filter.h: struct ref_format {\n     - \t */\n     - \tconst char *format;\n     - \tconst char *rest;\n     -+\tint cat_file_mode;\n     - \tint quote_style;\n     - \tint use_rest;\n     - \tint use_color;\n     -@@ ref-filter.h: struct ref_format {\n     - \tint need_color_reset_at_eol;\n     - };\n     - \n     --#define REF_FORMAT_INIT { NULL, NULL, 0, 0, -1 }\n     -+#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n     - \n     - /*  Macros for checking --merged and --no-merged options */\n     - #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n     -\n       ## t/t1006-cat-file.sh ##\n      @@ t/t1006-cat-file.sh: test_expect_success 'cat-file --unordered works' '\n       \ttest_cmp expect actual\n     @@ t/t1006-cat-file.sh: test_expect_success 'cat-file --unordered works' '\n      +batch_test_atom refs/heads/main '%(INVALID)' fail\n      +batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n      +\n     -+test_expect_success 'cat-file refs/heads/main refs/tags/testtag %(rest)' '\n     ++test_expect_success '%(rest) works with both a branch and a tag' '\n      +\tcat >expected <<-EOF &&\n      +\t123 commit 123\n      +\t456 tag 456\n     @@ t/t1006-cat-file.sh: test_expect_success 'cat-file --unordered works' '\n      +'\n      +\n       test_done\n     -\n     - ## t/t6301-for-each-ref-errors.sh ##\n     -@@ t/t6301-for-each-ref-errors.sh: test_expect_success 'Missing objects are reported correctly' '\n     - \tr=refs/heads/missing &&\n     - \techo $MISSING >.git/$r &&\n     - \ttest_when_finished \"rm -f .git/$r\" &&\n     --\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n     -+\techo \"fatal: $MISSING missing\" >missing-err &&\n     - \ttest_must_fail git for-each-ref 2>err &&\n     - \ttest_cmp missing-err err &&\n     - \t(\n  9:  b66ab0f2d569 ! 13:  bda6aae9a6c9 [GSOC] cat-file: reuse err buf in batch_object_write()\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n       \n      -\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n      +\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n     - \tif (!ret) {\n     - \t\tstrbuf_addch(scratch, '\\n');\n     - \t\tbatch_write(opt, scratch->buf, scratch->len);\n     - \t} else if (ret < 0) {\n     + \tif (ret < 0) {\n      -\t\tdie(\"%s\\n\", err.buf);\n      +\t\tdie(\"%s\\n\", err->buf);\n     - \t} else {\n     - \t\t/* when ret > 0 , don't call die and print the err to stdout*/\n     + \t} if (ret) {\n     + \t\t/* ret > 0 means when the object corresponding to oid\n     + \t\t * cannot be found in format_ref_array_item(), we only print\n     + \t\t * the error message.\n     + \t\t */\n      -\t\tprintf(\"%s\\n\", err.buf);\n      +\t\tprintf(\"%s\\n\", err->buf);\n       \t\tfflush(stdout);\n     + \t} else {\n     + \t\tstrbuf_addch(scratch, '\\n');\n     + \t\tbatch_write(opt, scratch->buf, scratch->len);\n       \t}\n       \tfree_ref_array_item_value(&item);\n      -\tstrbuf_release(&err);\n     @@ builtin/cat-file.c: struct object_cb_data {\n       \treturn 0;\n       }\n       \n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const struct option *options\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch, const struct option *optio\n       {\n       \tstruct strbuf input = STRBUF_INIT;\n       \tstruct strbuf output = STRBUF_INIT;\n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const st\n       \tstruct strbuf format = STRBUF_INIT;\n       \tstruct expand_data data;\n       \tint save_warning;\n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const struct option *options\n     - \t\tcb.opt = opt;\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch, const struct option *optio\n     + \t\tcb.opt = batch;\n       \t\tcb.expand = &data;\n       \t\tcb.scratch = &output;\n      +\t\tcb.err = &err;\n       \n     - \t\tif (opt->unordered) {\n     + \t\tif (batch->unordered) {\n       \t\t\tstruct oidset seen = OIDSET_INIT;\n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const struct option *options\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch, const struct option *optio\n       \n       \t\tstrbuf_release(&format);\n       \t\tstrbuf_release(&output);\n     @@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const st\n       \t\treturn 0;\n       \t}\n       \n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const struct option *options\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch, const struct option *optio\n       \t\t\tdata.rest = p;\n       \t\t}\n       \n     --\t\tbatch_one_object(input.buf, &output, opt, &data);\n     -+\t\tbatch_one_object(input.buf, &output, &err, opt, &data);\n     +-\t\tbatch_one_object(input.buf, &output, batch, &data);\n     ++\t\tbatch_one_object(input.buf, &output, &err, batch, &data);\n       \t}\n       \n       \tstrbuf_release(&format);\n 10:  86ac3bcaecea ! 14:  d1114a2bd743 [GSOC] cat-file: re-implement --textconv, --filters options\n     @@ Commit message\n          Signed-off-by: ZheNing Hu <adlternative@gmail.com>\n      \n       ## builtin/cat-file.c ##\n     -@@ builtin/cat-file.c: static int batch_objects(struct batch_options *opt, const struct option *options\n     - \tif (opt->print_contents)\n     +@@ builtin/cat-file.c: static int batch_objects(struct batch_options *batch, const struct option *optio\n     + \tif (batch->print_contents)\n       \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n     - \topt->format.format = format.buf;\n     -+\tif (opt->cmdmode == 'c')\n     -+\t\topt->format.use_textconv = 1;\n     -+\telse if (opt->cmdmode == 'w')\n     -+\t\topt->format.use_filters = 1;\n     + \tbatch->format.format = format.buf;\n      +\n     - \tif (verify_ref_format(&opt->format))\n     ++\tif (batch->cmdmode == 'c')\n     ++\t\tbatch->format.use_textconv = 1;\n     ++\telse if (batch->cmdmode == 'w')\n     ++\t\tbatch->format.use_filters = 1;\n     ++\n     + \tif (verify_ref_format(&batch->format))\n       \t\tusage_with_options(cat_file_usage, options);\n       \n      \n     @@ ref-filter.h: struct ref_format {\n       \tint use_rest;\n       \tint use_color;\n       \n     -@@ ref-filter.h: struct ref_format {\n     - \tint need_color_reset_at_eol;\n     - };\n     - \n     --#define REF_FORMAT_INIT { NULL, NULL, 0, 0, 0, -1 }\n     -+#define REF_FORMAT_INIT { .use_color = -1 }\n     - \n     - /*  Macros for checking --merged and --no-merged options */\n     - #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\n\n-- \ngitgitgadget\n"},{"id":"428162","messageId":"9fbbb3c492f5830d412c17662b086bbfcb76e499.1624332054.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 04/14] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:44Z","receivedAt":"2021-06-22T03:21:08Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc098242..452558ec9575 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 797b20ffa612..d01a0266fb89 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2553,9 +2553,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2600,7 +2600,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a718965..74fb423fc89f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"428163","messageId":"05682bccf9f947ce77ad035c02038286d0661c35.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 06/14] [GSOC] ref-filter: pass get_object() return value to their callers","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:46Z","receivedAt":"2021-06-22T03:21:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nSince in the refactor of `git cat-file --batch` later,\noid_object_info_extended() in get_object() will be used to obtain\nthe info of an object with it's oid. When the object cannot be\nobtained in the git repository, `cat-file --batch` expects to output\n\"<oid> missing\" and continue the next oid query instead of letting\nGit exit. In other error conditions, Git should exit normally. So we\ncan achieve this function by passing the return value of get_object().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 10c78de9cfa4..58def6ccd33a 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1816,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n \tint i;\n+\tint ret = 0;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1972,8 +1973,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tret = get_object(ref, 0, &obj, &oi, err);\n+\tif (ret)\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -2005,8 +2007,10 @@ static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tint ret = populate_value(ref, err);\n+\n+\t\tif (ret)\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2580,6 +2584,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret = 0;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2592,10 +2597,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n-- \ngitgitgadget\n\n"},{"id":"428164","messageId":"08aa44e5e57b8cabd41ea1dc10b0b43ea7a63d7a.1624332054.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 05/14] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:45Z","receivedAt":"2021-06-22T03:21:10Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex d01a0266fb89..10c78de9cfa4 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1010,6 +1021,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1927,6 +1941,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2144,6 +2164,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89f..c15dee8d6b95 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { .use_color = -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a00..6e94c6db7b5a 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must fail' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 5556063c347d..82c0ad2cb115 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1211,6 +1211,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883e..082be85dffc7 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602bf..10faa645157e 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"428165","messageId":"b0d9e139935f362817dd145df137e9ef33335e19.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 08/14] [GSOC] ref-filter: add cat_file_mode in struct ref_format","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:48Z","receivedAt":"2021-06-22T03:21:12Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd `cat_file_mode` member in struct `ref_format`, when\n`cat-file --batch` use ref-filter logic later, it can help\nus reject atoms in verify_ref_format() which cat-file cannot\nuse, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 +++++++++--\n ref-filter.h |  1 +\n 2 files changed, 10 insertions(+), 2 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 22315d4809dc..f21f41df0d88 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1021,8 +1021,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 44e6dc05ac2f..053980a6a426 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n-- \ngitgitgadget\n\n"},{"id":"428166","messageId":"06db6cd6f1f93ddd3be2e8ae3efdebeb82a2ca00.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 07/14] [GSOC] ref-filter: introduce free_ref_array_item_value() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:47Z","receivedAt":"2021-06-22T03:21:13Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nWhen we use ref_array_item which is not dynamically allocated and\nwant to free the space of its member \"value\" after the end of use,\nfree_array_item() does not meet our needs, because it tries to free\nref_array_item itself and its member \"symref\".\n\nIntroduce free_ref_array_item_value() for freeing ref_array_item value.\nIt will be called internally by free_array_item(), and it will help\n`cat-file --batch` free ref_array_item's value memory later.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 ++++++++---\n ref-filter.h |  2 ++\n 2 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 58def6ccd33a..22315d4809dc 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -2291,16 +2291,21 @@ static int ref_filter_handler(const char *refname, const struct object_id *oid,\n \treturn 0;\n }\n \n-/*  Free memory allocated for a ref_array_item */\n-static void free_array_item(struct ref_array_item *item)\n+void free_ref_array_item_value(struct ref_array_item *item)\n {\n-\tfree((char *)item->symref);\n \tif (item->value) {\n \t\tint i;\n \t\tfor (i = 0; i < used_atom_cnt; i++)\n \t\t\tfree((char *)item->value[i].s);\n \t\tfree(item->value);\n \t}\n+}\n+\n+/*  Free memory allocated for a ref_array_item */\n+static void free_array_item(struct ref_array_item *item)\n+{\n+\tfree((char *)item->symref);\n+\tfree_ref_array_item_value(item);\n \tfree(item);\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex c15dee8d6b95..44e6dc05ac2f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -111,6 +111,8 @@ struct ref_format {\n int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n /*  Clear all memory allocated to ref_array */\n void ref_array_clear(struct ref_array *array);\n+/* Free ref_array_item's value */\n+void free_ref_array_item_value(struct ref_array_item *item);\n /*  Used to verify if the given format is correct and to parse out the used atoms */\n int verify_ref_format(struct ref_format *format);\n /*  Sort the given ref_array as per the ref_sorting provided */\n-- \ngitgitgadget\n\n"},{"id":"428167","messageId":"db7dd8b042c281afdccb650f9679e88a4e378b65.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 09/14] [GSOC] ref-filter: modify the error message and value in get_object","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:49Z","receivedAt":"2021-06-22T03:21:14Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nLet get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\"\nif oid_object_info_extended() unable to find the data corresponding\nto oid. When `cat-file --batch` use ref-filter logic later it can\nhelp `format_ref_array_item()` just report that the object is missing\nwithout letting Git exit.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c                   | 4 ++--\n t/t6301-for-each-ref-errors.sh | 2 +-\n 2 files changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex f21f41df0d88..181d99c92735 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1749,8 +1749,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab534..3553f84a00c1 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"428168","messageId":"6b577969734e9d69dbfbc6f0d523f64454dacab8.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 10/14] [GSOC] cat-file: add has_object_file() check","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:50Z","receivedAt":"2021-06-22T03:21:17Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists. This can help us reject\nthe missing oid when we let `cat-file --batch` use ref-filter\nlogic later.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 7 +++++++\n 1 file changed, 7 insertions(+)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e83..9fd3c04ff20b 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -428,6 +428,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"428169","messageId":"069aa203666a4701e22078e9bc35a8c1ff4d5f15.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 11/14] [GSOC] cat-file: change batch_objects parameter name","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:51Z","receivedAt":"2021-06-22T03:21:19Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nSince cat-file reuses ref-filter logic later will add the\nformal parameter \"const struct option *options\" to\nbatch_objects(), the two synonymous parameters of \"opt\"\nand \"options\" may confuse readers, so change batch_options\nparameter of batch_objects() from \"opt\" to \"batch\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 +++++++++++-----------\n 1 file changed, 11 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 9fd3c04ff20b..cd84c39df968 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -495,7 +495,7 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static int batch_objects(struct batch_options *batch)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n@@ -503,8 +503,8 @@ static int batch_objects(struct batch_options *opt)\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n+\tif (!batch->format)\n+\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n \n \t/*\n \t * Expand once with our special mark_query flag, which will prime the\n@@ -513,13 +513,13 @@ static int batch_objects(struct batch_options *opt)\n \t */\n \tmemset(&data, 0, sizeof(data));\n \tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n+\tstrbuf_expand(&output, batch->format, expand_format, &data);\n \tdata.mark_query = 0;\n \tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n+\tif (batch->cmdmode)\n \t\tdata.split_on_whitespace = 1;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_info empty = OBJECT_INFO_INIT;\n \t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n \t\t\tdata.skip_object_info = 1;\n@@ -529,20 +529,20 @@ static int batch_objects(struct batch_options *opt)\n \t * If we are printing out the object, then always fill in the type,\n \t * since we will want to decide whether or not to stream.\n \t */\n-\tif (opt->print_contents)\n+\tif (batch->print_contents)\n \t\tdata.info.typep = &data.type;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n \n \t\tif (has_promisor_remote())\n \t\t\twarning(\"This repository uses promisor remotes. Some objects may not be loaded.\");\n \n-\t\tcb.opt = opt;\n+\t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n \n-\t\tif (opt->unordered) {\n+\t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n \n \t\t\tcb.seen = &seen;\n@@ -592,7 +592,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n \tstrbuf_release(&input);\n-- \ngitgitgadget\n\n"},{"id":"428170","messageId":"258ec0a46c56c7e66710a2913efc4d4cbb1d5944.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 12/14] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:52Z","receivedAt":"2021-06-22T03:21:22Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, let's do the\nfollowing:\n\n1. Change the type of member `format` in struct `batch_options`\nto `ref_format`, we will pass it to ref-filter later.\n2. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n3. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue is less than zero, use `die()` to print the error message\nand exit; else if return value is greater than zero, only print\nthe error message, but don't exit.\n4. Use free_ref_array_item_value() to free ref_array_item's\nvalue.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n\nThe performance for `git cat-file --batch-all-objects\n--batch-check` on the Git repository itself with performance\ntesting tool `hyperfine` changes from 669.4 ms ±  31.1 ms to\n1.134 s ±  0.063 s.\n\nThe performance for `git cat-file --batch-all-objects --batch\n>/dev/null` on the Git repository itself with performance testing\ntool `time` change from \"27.37s user 0.29s system 98% cpu 28.089\ntotal\" to \"33.69s user 1.54s system 87% cpu 40.258 total\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 244 ++++++-------------------------\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n 3 files changed, 305 insertions(+), 197 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd9..ef8ab952b2fa 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex cd84c39df968..0e7ad038e5fb 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,87 +209,34 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret = 0;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (ret < 0) {\n+\t\tdie(\"%s\\n\", err.buf);\n+\t} if (ret) {\n+\t\t/* ret > 0 means when the object corresponding to oid\n+\t\t * cannot be found in format_ref_array_item(), we only print\n+\t\t * the error message.\n+\t\t */\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n+\t} else {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n+\tfree_ref_array_item_value(&item);\n+\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n@@ -495,42 +354,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *batch)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *batch, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!batch->format)\n-\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, batch->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (batch->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (batch->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (batch->format.format)\n+\t\tstrbuf_addstr(&format, batch->format.format);\n+\telse\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n \tif (batch->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\tbatch->format.format = format.buf;\n+\tif (verify_ref_format(&batch->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (batch->cmdmode || batch->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -563,6 +414,7 @@ static int batch_objects(struct batch_options *batch)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -595,18 +447,13 @@ static int batch_objects(struct batch_options *batch)\n \t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -629,7 +476,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -638,7 +485,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -677,6 +526,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -720,7 +570,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74ad..69eb627774d9 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 must fail\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success '%(rest) works with both a branch and a tag' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"428171","messageId":"bda6aae9a6c9dc97ba40fab6ffe30f4af6161f12.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 13/14] [GSOC] cat-file: reuse err buf in batch_object_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:53Z","receivedAt":"2021-06-22T03:21:26Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 0e7ad038e5fb..27403326e7a7 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,35 +212,36 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret = 0;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (ret < 0) {\n-\t\tdie(\"%s\\n\", err.buf);\n+\t\tdie(\"%s\\n\", err->buf);\n \t} if (ret) {\n \t\t/* ret > 0 means when the object corresponding to oid\n \t\t * cannot be found in format_ref_array_item(), we only print\n \t\t * the error message.\n \t\t */\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n \t} else {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n \tfree_ref_array_item_value(&item);\n-\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -294,7 +295,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -302,13 +303,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -364,6 +366,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -392,6 +395,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -416,6 +420,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -444,12 +449,13 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, batch, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, batch, &data);\n \t}\n \n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"428172","messageId":"d1114a2bd743241854cac540d6fb1369f050fbd1.1624332055.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v4 14/14] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-22T03:20:54Z","receivedAt":"2021-06-22T03:21:28Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  6 +++++\n ref-filter.c       | 59 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  2 ++\n 3 files changed, 65 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 27403326e7a7..f3140d927f7d 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -380,6 +380,12 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \tif (batch->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \tbatch->format.format = format.buf;\n+\n+\tif (batch->cmdmode == 'c')\n+\t\tbatch->format.use_textconv = 1;\n+\telse if (batch->cmdmode == 'w')\n+\t\tbatch->format.use_filters = 1;\n+\n \tif (verify_ref_format(&batch->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 181d99c92735..99b87742b0fd 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1031,6 +1035,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1742,10 +1749,38 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\tif (use_textconv && oi->type == OBJ_BLOB) {\n+\t\t\tact_oi = *oi;\n+\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\tgoto success;\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1755,19 +1790,39 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\tif (use_filters && oi->type == OBJ_BLOB) {\n+\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\tstruct checkout_metadata meta;\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\tif (!convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n+\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\toid_to_hex(&oi->oid), ref->rest);\n+\t\t\tact_oi.size = strbuf.len;\n+\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\tactual_oi = &act_oi;\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 053980a6a426..497e3e93632f 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,6 +80,8 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n \n-- \ngitgitgadget\n"},{"id":"428356","messageId":"4bce3cc5-e027-fa85-f400-df529e370a2b@gmail.com","threadId":"55909","inReplyTo":"05682bccf9f947ce77ad035c02038286d0661c35.1624332055.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 06/14] [GSOC] ref-filter: pass get_object() return value to their callers","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-06-24T04:02:36Z","receivedAt":"2021-06-24T04:03:37Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 22/06/21 10.20, ZheNing Hu via GitGitGadget wrote:\n> From: ZheNing Hu <adlternative@gmail.com>\n> \n> Since in the refactor of `git cat-file --batch` later,\n> oid_object_info_extended() in get_object() will be used to obtain\n> the info of an object with it's oid. When the object cannot be\n> obtained in the git repository, `cat-file --batch` expects to output\n> \"<oid> missing\" and continue the next oid query instead of letting\n> Git exit. In other error conditions, Git should exit normally. So we\n> can achieve this function by passing the return value of get_object().\n> \n\ns/Since/Because/\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"428357","messageId":"88fc452c-77be-cfaa-0587-ed752049674b@gmail.com","threadId":"55909","inReplyTo":"069aa203666a4701e22078e9bc35a8c1ff4d5f15.1624332055.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 11/14] [GSOC] cat-file: change batch_objects parameter name","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-06-24T04:07:02Z","receivedAt":"2021-06-24T04:07:09Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 22/06/21 10.20, ZheNing Hu via GitGitGadget wrote:\n> From: ZheNing Hu <adlternative@gmail.com>\n> \n> Since cat-file reuses ref-filter logic later will add the\n> formal parameter \"const struct option *options\" to\n> batch_objects(), the two synonymous parameters of \"opt\"\n> and \"options\" may confuse readers, so change batch_options\n> parameter of batch_objects() from \"opt\" to \"batch\".\n> \n\nBetter say \"Because later cat-file reuses ref-filter logic that will add \nparameter ... \".\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"428358","messageId":"7c43f6d4-7940-de5e-f424-a2ea50739b66@gmail.com","threadId":"55909","inReplyTo":"ab497d66c1167981181173f19cfe7d8857c348a3.1624332054.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 02/14] [GSOC] ref-filter: add %(raw) atom","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-06-24T04:14:07Z","receivedAt":"2021-06-24T04:16:39Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 22/06/21 10.20, ZheNing Hu via GitGitGadget wrote:\n> Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> `--tcl`, `--perl` because if our binary raw data is passed to a\n> variable in the host language, the host language may not support\n> arbitrary binary data in the variables of its string type.\n\nBetter say:\n\n\"Note that `--format=%(raw)` cannot be used with `--python`, `--shell`, \n`--tcl`, and `--perl` because if the binary raw data is passed to a \nvariable in such languages, these may not support arbitrary binary data \nin their string variable type.\"\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"428364","messageId":"CAOLTT8QhT85dAwxUTag=ktf9XVfZunfhSgswJLGeyszSxx7Q=g@mail.gmail.com","threadId":"55909","inReplyTo":"7c43f6d4-7940-de5e-f424-a2ea50739b66@gmail.com","subject":"Re: [PATCH v4 02/14] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-24T08:23:36Z","receivedAt":"2021-06-24T08:23:51Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Bagas Sanjaya <bagasdotme@gmail.com> 于2021年6月24日周四 下午12:14写道：\n>\n> On 22/06/21 10.20, ZheNing Hu via GitGitGadget wrote:\n> > Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> > `--tcl`, `--perl` because if our binary raw data is passed to a\n> > variable in the host language, the host language may not support\n> > arbitrary binary data in the variables of its string type.\n>\n> Better say:\n>\n> \"Note that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> `--tcl`, and `--perl` because if the binary raw data is passed to a\n> variable in such languages, these may not support arbitrary binary data\n> in their string variable type.\"\n>\n\nThanks, Bagas Sanjaya, I will change all of them.\n\n\n> --\n> An old man doll... just what I always wanted! - Clara\n\n--\nZheNing Hu\n"},{"id":"428468","messageId":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v4.git.1624332054.gitgitgadget@gmail.com","subject":"[PATCH v5 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:10Z","receivedAt":"2021-06-25T16:02:34Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic.\n\nChange from last version:\n\n 1. At the suggestion of Bagas Sanjaya, modified the expression of submitted\n    information.\n 2. Remove grab_oid() function in ref-filter to reduce repeated checks.\n\nZheNing Hu (15):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: --format=%(raw) re-support --perl\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: pass get_object() return value to their callers\n  [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  [GSOC] ref-filter: modify the error message and value in get_object\n  [GSOC] cat-file: add has_object_file() check\n  [GSOC] cat-file: change batch_objects parameter name\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_object_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n  [GSOC] ref-filter: remove grab_oid() function\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 277 ++++++----------------\n builtin/tag.c                      |   2 +-\n quote.c                            |  17 ++\n quote.h                            |   1 +\n ref-filter.c                       | 357 ++++++++++++++++++++++-------\n ref-filter.h                       |  14 +-\n t/t1006-cat-file.sh                | 252 ++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 235 +++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 14 files changed, 888 insertions(+), 296 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n\nRange-diff vs v4:\n\n  1:  f72ad9cc5e8 =  1:  f72ad9cc5e8 [GSOC] ref-filter: add obj-type check in grab contents\n  2:  ab497d66c11 !  2:  4e473838b9d [GSOC] ref-filter: add %(raw) atom\n     @@ Commit message\n          can help us add raw object data to the buffer or compare two buffers\n          which contain raw object data.\n      \n     -    Beyond, `--format=%(raw)` cannot be used with `--python`, `--shell`,\n     -    `--tcl`, `--perl` because if our binary raw data is passed to a\n     -    variable in the host language, the host language may not support\n     -    arbitrary binary data in the variables of its string type.\n     +    Note that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n     +    `--tcl`, and `--perl` because if the binary raw data is passed to a\n     +    variable in such languages, these may not support arbitrary binary data\n     +    in their string variable type.\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n  3:  b54dbc431e0 =  3:  765cf08a108 [GSOC] ref-filter: --format=%(raw) re-support --perl\n  4:  9fbbb3c492f =  4:  d2aeafd0ef3 [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  5:  08aa44e5e57 =  5:  1ca3a42f041 [GSOC] ref-filter: add %(rest) atom\n  6:  05682bccf9f !  6:  67f1a3cca9a [GSOC] ref-filter: pass get_object() return value to their callers\n     @@ Metadata\n       ## Commit message ##\n          [GSOC] ref-filter: pass get_object() return value to their callers\n      \n     -    Since in the refactor of `git cat-file --batch` later,\n     +    Because in the refactor of `git cat-file --batch` later,\n          oid_object_info_extended() in get_object() will be used to obtain\n          the info of an object with it's oid. When the object cannot be\n          obtained in the git repository, `cat-file --batch` expects to output\n  7:  06db6cd6f1f =  7:  2a48a48e81c [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  8:  b0d9e139935 =  8:  be55005be75 [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  9:  db7dd8b042c =  9:  937f88b7837 [GSOC] ref-filter: modify the error message and value in get_object\n 10:  6b577969734 = 10:  45657499c55 [GSOC] cat-file: add has_object_file() check\n 11:  069aa203666 ! 11:  bf5c0a017ad [GSOC] cat-file: change batch_objects parameter name\n     @@ Metadata\n       ## Commit message ##\n          [GSOC] cat-file: change batch_objects parameter name\n      \n     -    Since cat-file reuses ref-filter logic later will add the\n     -    formal parameter \"const struct option *options\" to\n     -    batch_objects(), the two synonymous parameters of \"opt\"\n     -    and \"options\" may confuse readers, so change batch_options\n     -    parameter of batch_objects() from \"opt\" to \"batch\".\n     +    Because later cat-file reuses ref-filter logic that will add\n     +    parameter \"const struct option *options\" to batch_objects(),\n     +    the two synonymous parameters of \"opt\" and \"options\" may\n     +    confuse readers, so change batch_options parameter of\n     +    batch_objects() from \"opt\" to \"batch\".\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n 12:  258ec0a46c5 = 12:  370101ba65f [GSOC] cat-file: reuse ref-filter logic\n 13:  bda6aae9a6c = 13:  69eef47065d [GSOC] cat-file: reuse err buf in batch_object_write()\n 14:  d1114a2bd74 = 14:  a7ac037a946 [GSOC] cat-file: re-implement --textconv, --filters options\n  -:  ----------- > 15:  843de8864a9 [GSOC] ref-filter: remove grab_oid() function\n\n-- \ngitgitgadget\n"},{"id":"428469","messageId":"4e473838b9d2651a8e4be27332697c2ba354db5a.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 02/15] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:12Z","receivedAt":"2021-06-25T16:02:36Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, we need to find a way to record the length of the\natom_value's member `s`. Although strbuf can already record the\nstring and its length, if we want to replace the type of atom_value's\nmember `s` with strbuf, many places in ref-filter that are filled\nwith dynamically allocated mermory in `v->s` are not easy to replace.\nAt the same time, we need to check if `v->s == NULL` in\npopulate_value(), and strbuf cannot easily distinguish NULL and empty\nstrings, but c-style \"const char *\" can do it. So add a new member in\n`struct atom_value`: `s_size`, which can record raw object size, it\ncan help us add raw object data to the buffer or compare two buffers\nwhich contain raw object data.\n\nNote that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, and `--perl` because if the binary raw data is passed to a\nvariable in such languages, these may not support arbitrary binary data\nin their string variable type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 216 +++++++++++++++++++++++++++++\n 3 files changed, 337 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de70..7f1f0a1ca3b 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the host language may not support arbitrary binary data in the\n+variables of its string type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fba..7822be90307 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4..9c5379e2f56 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,174 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tobj=$(git hash-object -w blob1) &&\n+        git update-ref refs/myblobs/blob1 \"$obj\" &&\n+\tobj=$(git hash-object -w blob2) &&\n+        git update-ref refs/myblobs/blob2 \"$obj\" &&\n+\tobj=$(git hash-object -w blob3) &&\n+        git update-ref refs/myblobs/blob3 \"$obj\" &&\n+\tobj=$(git hash-object -w blob4) &&\n+        git update-ref refs/myblobs/blob4 \"$obj\" &&\n+\tobj=$(git hash-object -w blob5) &&\n+        git update-ref refs/myblobs/blob5 \"$obj\" &&\n+\tobj=$(git hash-object -w blob6) &&\n+        git update-ref refs/myblobs/blob6 \"$obj\" &&\n+\tobj=$(git hash-object -w blob7) &&\n+        git update-ref refs/myblobs/blob7 \"$obj\" &&\n+\tobj=$(git hash-object -w blob8) &&\n+        git update-ref refs/myblobs/blob8 \"$obj\"\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"428470","messageId":"f72ad9cc5e8b1f1139784df9ac7178f1561f70bb.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 01/15] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:11Z","receivedAt":"2021-06-25T16:02:38Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c..5cee6512fba 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"428471","messageId":"765cf08a108ffca4b096ac46f6be6fb26601f718.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 03/15] [GSOC] ref-filter: --format=%(raw) re-support --perl","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:13Z","receivedAt":"2021-06-25T16:02:39Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause the perl language can handle binary data correctly,\nadd the function perl_quote_buf_with_len(), which can specify\nthe length of the data and prevent the data from being truncated\nat '\\0' to help `--format=\"%(raw)\"` re-support `--perl`.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |  2 +-\n quote.c                            | 17 +++++++++++++++++\n quote.h                            |  1 +\n ref-filter.c                       | 15 +++++++++++----\n t/t6300-for-each-ref.sh            | 19 +++++++++++++++++--\n 5 files changed, 47 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 7f1f0a1ca3b..ea9b438c16f 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -241,7 +241,7 @@ raw:size::\n \tThe raw data size of the object.\n \n Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n-`--perl` because the host language may not support arbitrary binary data in the\n+because the host language may not support arbitrary binary data in the\n variables of its string type.\n \n The message in a commit or a tag object is `contents`, from which\ndiff --git a/quote.c b/quote.c\nindex 8a3a5e39eb1..26719d21d1e 100644\n--- a/quote.c\n+++ b/quote.c\n@@ -471,6 +471,23 @@ void perl_quote_buf(struct strbuf *sb, const char *src)\n \tstrbuf_addch(sb, sq);\n }\n \n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len)\n+{\n+\tconst char sq = '\\'';\n+\tconst char bq = '\\\\';\n+\tconst char *c = src;\n+\tconst char *end = src + len;\n+\n+\tstrbuf_addch(sb, sq);\n+\twhile (c != end) {\n+\t\tif (*c == sq || *c == bq)\n+\t\t\tstrbuf_addch(sb, bq);\n+\t\tstrbuf_addch(sb, *c);\n+\t\tc++;\n+\t}\n+\tstrbuf_addch(sb, sq);\n+}\n+\n void python_quote_buf(struct strbuf *sb, const char *src)\n {\n \tconst char sq = '\\'';\ndiff --git a/quote.h b/quote.h\nindex 768cc6338e2..0fe69e264b0 100644\n--- a/quote.h\n+++ b/quote.h\n@@ -94,6 +94,7 @@ char *quote_path(const char *in, const char *prefix, struct strbuf *out, unsigne\n \n /* quoting as a string literal for other languages */\n void perl_quote_buf(struct strbuf *sb, const char *src);\n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len);\n void python_quote_buf(struct strbuf *sb, const char *src);\n void tcl_quote_buf(struct strbuf *sb, const char *src);\n void basic_regex_quote_buf(struct strbuf *sb, const char *src);\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be90307..797b20ffa61 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -742,7 +742,10 @@ static void quote_formatting(struct strbuf *s, const char *str, size_t len, int\n \t\tsq_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PERL:\n-\t\tperl_quote_buf(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tperl_quote_buf_with_len(s, str, len);\n+\t\telse\n+\t\t\tperl_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PYTHON:\n \t\tpython_quote_buf(s, str);\n@@ -1006,10 +1009,14 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n-\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\n+\t\tif ((format->quote_style == QUOTE_PYTHON ||\n+\t\t     format->quote_style == QUOTE_SHELL ||\n+\t\t     format->quote_style == QUOTE_TCL) &&\n+\t\t     used_atom[at].atom_type == ATOM_RAW &&\n+\t\t     used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n-\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n+\t\t\t      \"--python, --shell, --tcl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9c5379e2f56..5556063c347 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -915,8 +915,23 @@ test_expect_success '%(raw) with --tcl must fail' '\n \ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n '\n \n-test_expect_success '%(raw) with --perl must fail' '\n-\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+test_expect_success '%(raw) with --perl' '\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob1 --perl | perl > actual &&\n+\tcmp blob1 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob3 --perl | perl > actual &&\n+\tcmp blob3 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob8 --perl | perl > actual &&\n+\tcmp blob8 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/first --perl | perl > actual &&\n+\tcmp one actual &&\n+\tgit cat-file tree refs/mytrees/first > expected &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/mytrees/first --perl | perl > actual &&\n+\tcmp expected actual\n '\n \n test_expect_success '%(raw) with --shell must fail' '\n-- \ngitgitgadget\n\n"},{"id":"428472","messageId":"1ca3a42f041087e3640ec7e3a8980798edaf9857.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 05/15] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:15Z","receivedAt":"2021-06-25T16:02:43Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex d01a0266fb8..10c78de9cfa 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1010,6 +1021,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1927,6 +1941,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2144,6 +2164,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89..c15dee8d6b9 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { .use_color = -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a0..6e94c6db7b5 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must fail' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 5556063c347..82c0ad2cb11 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1211,6 +1211,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883..082be85dffc 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602b..10faa645157 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"428473","messageId":"d2aeafd0ef3ed3193ec858af546f21faf9986a32.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 04/15] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:14Z","receivedAt":"2021-06-25T16:02:45Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc09824..452558ec957 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 797b20ffa61..d01a0266fb8 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2553,9 +2553,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2600,7 +2600,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a71896..74fb423fc89 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"428474","messageId":"67f1a3cca9a85fa94cb0158de71b36ce3dd6b444.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 06/15] [GSOC] ref-filter: pass get_object() return value to their callers","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:16Z","receivedAt":"2021-06-25T16:02:48Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause in the refactor of `git cat-file --batch` later,\noid_object_info_extended() in get_object() will be used to obtain\nthe info of an object with it's oid. When the object cannot be\nobtained in the git repository, `cat-file --batch` expects to output\n\"<oid> missing\" and continue the next oid query instead of letting\nGit exit. In other error conditions, Git should exit normally. So we\ncan achieve this function by passing the return value of get_object().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 10c78de9cfa..58def6ccd33 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1816,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n \tint i;\n+\tint ret = 0;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1972,8 +1973,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tret = get_object(ref, 0, &obj, &oi, err);\n+\tif (ret)\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -2005,8 +2007,10 @@ static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tint ret = populate_value(ref, err);\n+\n+\t\tif (ret)\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2580,6 +2584,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret = 0;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2592,10 +2597,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n-- \ngitgitgadget\n\n"},{"id":"428475","messageId":"be55005be75333e22d00ae63c8bd1fc69718b369.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 08/15] [GSOC] ref-filter: add cat_file_mode in struct ref_format","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:18Z","receivedAt":"2021-06-25T16:02:49Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd `cat_file_mode` member in struct `ref_format`, when\n`cat-file --batch` use ref-filter logic later, it can help\nus reject atoms in verify_ref_format() which cat-file cannot\nuse, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 +++++++++--\n ref-filter.h |  1 +\n 2 files changed, 10 insertions(+), 2 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 22315d4809d..f21f41df0d8 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1021,8 +1021,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 44e6dc05ac2..053980a6a42 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n-- \ngitgitgadget\n\n"},{"id":"428476","messageId":"2a48a48e81c6389c1eb8dd943bb8e323ed574bd2.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 07/15] [GSOC] ref-filter: introduce free_ref_array_item_value() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:17Z","receivedAt":"2021-06-25T16:02:51Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nWhen we use ref_array_item which is not dynamically allocated and\nwant to free the space of its member \"value\" after the end of use,\nfree_array_item() does not meet our needs, because it tries to free\nref_array_item itself and its member \"symref\".\n\nIntroduce free_ref_array_item_value() for freeing ref_array_item value.\nIt will be called internally by free_array_item(), and it will help\n`cat-file --batch` free ref_array_item's value memory later.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 ++++++++---\n ref-filter.h |  2 ++\n 2 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 58def6ccd33..22315d4809d 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -2291,16 +2291,21 @@ static int ref_filter_handler(const char *refname, const struct object_id *oid,\n \treturn 0;\n }\n \n-/*  Free memory allocated for a ref_array_item */\n-static void free_array_item(struct ref_array_item *item)\n+void free_ref_array_item_value(struct ref_array_item *item)\n {\n-\tfree((char *)item->symref);\n \tif (item->value) {\n \t\tint i;\n \t\tfor (i = 0; i < used_atom_cnt; i++)\n \t\t\tfree((char *)item->value[i].s);\n \t\tfree(item->value);\n \t}\n+}\n+\n+/*  Free memory allocated for a ref_array_item */\n+static void free_array_item(struct ref_array_item *item)\n+{\n+\tfree((char *)item->symref);\n+\tfree_ref_array_item_value(item);\n \tfree(item);\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex c15dee8d6b9..44e6dc05ac2 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -111,6 +111,8 @@ struct ref_format {\n int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n /*  Clear all memory allocated to ref_array */\n void ref_array_clear(struct ref_array *array);\n+/* Free ref_array_item's value */\n+void free_ref_array_item_value(struct ref_array_item *item);\n /*  Used to verify if the given format is correct and to parse out the used atoms */\n int verify_ref_format(struct ref_format *format);\n /*  Sort the given ref_array as per the ref_sorting provided */\n-- \ngitgitgadget\n\n"},{"id":"428477","messageId":"45657499c55ed91eb78498be4355459b00f76404.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 10/15] [GSOC] cat-file: add has_object_file() check","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:20Z","receivedAt":"2021-06-25T16:02:51Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists. This can help us reject\nthe missing oid when we let `cat-file --batch` use ref-filter\nlogic later.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 7 +++++++\n 1 file changed, 7 insertions(+)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e8..9fd3c04ff20 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -428,6 +428,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"428478","messageId":"937f88b78371d1a4497c8dc389a499ab51446075.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 09/15] [GSOC] ref-filter: modify the error message and value in get_object","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:19Z","receivedAt":"2021-06-25T16:02:54Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nLet get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\"\nif oid_object_info_extended() unable to find the data corresponding\nto oid. When `cat-file --batch` use ref-filter logic later it can\nhelp `format_ref_array_item()` just report that the object is missing\nwithout letting Git exit.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c                   | 4 ++--\n t/t6301-for-each-ref-errors.sh | 2 +-\n 2 files changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex f21f41df0d8..181d99c9273 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1749,8 +1749,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab53..3553f84a00c 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"428479","messageId":"bf5c0a017ad28c587e6d54304202a17d2bc0f1fd.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 11/15] [GSOC] cat-file: change batch_objects parameter name","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:21Z","receivedAt":"2021-06-25T16:02:55Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause later cat-file reuses ref-filter logic that will add\nparameter \"const struct option *options\" to batch_objects(),\nthe two synonymous parameters of \"opt\" and \"options\" may\nconfuse readers, so change batch_options parameter of\nbatch_objects() from \"opt\" to \"batch\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 +++++++++++-----------\n 1 file changed, 11 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 9fd3c04ff20..cd84c39df96 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -495,7 +495,7 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static int batch_objects(struct batch_options *batch)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n@@ -503,8 +503,8 @@ static int batch_objects(struct batch_options *opt)\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n+\tif (!batch->format)\n+\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n \n \t/*\n \t * Expand once with our special mark_query flag, which will prime the\n@@ -513,13 +513,13 @@ static int batch_objects(struct batch_options *opt)\n \t */\n \tmemset(&data, 0, sizeof(data));\n \tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n+\tstrbuf_expand(&output, batch->format, expand_format, &data);\n \tdata.mark_query = 0;\n \tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n+\tif (batch->cmdmode)\n \t\tdata.split_on_whitespace = 1;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_info empty = OBJECT_INFO_INIT;\n \t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n \t\t\tdata.skip_object_info = 1;\n@@ -529,20 +529,20 @@ static int batch_objects(struct batch_options *opt)\n \t * If we are printing out the object, then always fill in the type,\n \t * since we will want to decide whether or not to stream.\n \t */\n-\tif (opt->print_contents)\n+\tif (batch->print_contents)\n \t\tdata.info.typep = &data.type;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n \n \t\tif (has_promisor_remote())\n \t\t\twarning(\"This repository uses promisor remotes. Some objects may not be loaded.\");\n \n-\t\tcb.opt = opt;\n+\t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n \n-\t\tif (opt->unordered) {\n+\t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n \n \t\t\tcb.seen = &seen;\n@@ -592,7 +592,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n \tstrbuf_release(&input);\n-- \ngitgitgadget\n\n"},{"id":"428480","messageId":"370101ba65f0989487360366f8b83144a6641a04.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:22Z","receivedAt":"2021-06-25T16:02:55Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, let's do the\nfollowing:\n\n1. Change the type of member `format` in struct `batch_options`\nto `ref_format`, we will pass it to ref-filter later.\n2. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n3. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue is less than zero, use `die()` to print the error message\nand exit; else if return value is greater than zero, only print\nthe error message, but don't exit.\n4. Use free_ref_array_item_value() to free ref_array_item's\nvalue.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n\nThe performance for `git cat-file --batch-all-objects\n--batch-check` on the Git repository itself with performance\ntesting tool `hyperfine` changes from 669.4 ms ±  31.1 ms to\n1.134 s ±  0.063 s.\n\nThe performance for `git cat-file --batch-all-objects --batch\n>/dev/null` on the Git repository itself with performance testing\ntool `time` change from \"27.37s user 0.29s system 98% cpu 28.089\ntotal\" to \"33.69s user 1.54s system 87% cpu 40.258 total\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 244 ++++++-------------------------\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n 3 files changed, 305 insertions(+), 197 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd..ef8ab952b2f 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex cd84c39df96..0e7ad038e5f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,87 +209,34 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret = 0;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (ret < 0) {\n+\t\tdie(\"%s\\n\", err.buf);\n+\t} if (ret) {\n+\t\t/* ret > 0 means when the object corresponding to oid\n+\t\t * cannot be found in format_ref_array_item(), we only print\n+\t\t * the error message.\n+\t\t */\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n+\t} else {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n+\tfree_ref_array_item_value(&item);\n+\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n@@ -495,42 +354,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *batch)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *batch, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!batch->format)\n-\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, batch->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (batch->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (batch->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (batch->format.format)\n+\t\tstrbuf_addstr(&format, batch->format.format);\n+\telse\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n \tif (batch->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\tbatch->format.format = format.buf;\n+\tif (verify_ref_format(&batch->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (batch->cmdmode || batch->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -563,6 +414,7 @@ static int batch_objects(struct batch_options *batch)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -595,18 +447,13 @@ static int batch_objects(struct batch_options *batch)\n \t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -629,7 +476,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -638,7 +485,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -677,6 +526,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -720,7 +570,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74a..69eb627774d 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 must fail\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success '%(rest) works with both a branch and a tag' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"428481","messageId":"69eef47065d27cc997a26c853691534c5d84df6d.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 13/15] [GSOC] cat-file: reuse err buf in batch_object_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:23Z","receivedAt":"2021-06-25T16:02:56Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 0e7ad038e5f..27403326e7a 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,35 +212,36 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret = 0;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (ret < 0) {\n-\t\tdie(\"%s\\n\", err.buf);\n+\t\tdie(\"%s\\n\", err->buf);\n \t} if (ret) {\n \t\t/* ret > 0 means when the object corresponding to oid\n \t\t * cannot be found in format_ref_array_item(), we only print\n \t\t * the error message.\n \t\t */\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n \t} else {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n \tfree_ref_array_item_value(&item);\n-\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -294,7 +295,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -302,13 +303,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -364,6 +366,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -392,6 +395,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -416,6 +420,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -444,12 +449,13 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, batch, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, batch, &data);\n \t}\n \n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"428482","messageId":"a7ac037a94686585f1f91e74ff1ecc8402dc28a7.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 14/15] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:24Z","receivedAt":"2021-06-25T16:02:58Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  6 +++++\n ref-filter.c       | 59 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  2 ++\n 3 files changed, 65 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 27403326e7a..f3140d927f7 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -380,6 +380,12 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \tif (batch->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \tbatch->format.format = format.buf;\n+\n+\tif (batch->cmdmode == 'c')\n+\t\tbatch->format.use_textconv = 1;\n+\telse if (batch->cmdmode == 'w')\n+\t\tbatch->format.use_filters = 1;\n+\n \tif (verify_ref_format(&batch->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 181d99c9273..99b87742b0f 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1031,6 +1035,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1742,10 +1749,38 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\tif (use_textconv && oi->type == OBJ_BLOB) {\n+\t\t\tact_oi = *oi;\n+\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\tgoto success;\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1755,19 +1790,39 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\tif (use_filters && oi->type == OBJ_BLOB) {\n+\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\tstruct checkout_metadata meta;\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\tif (!convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n+\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\toid_to_hex(&oi->oid), ref->rest);\n+\t\t\tact_oi.size = strbuf.len;\n+\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\tactual_oi = &act_oi;\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 053980a6a42..497e3e93632 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,6 +80,8 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n \n-- \ngitgitgadget\n\n"},{"id":"428483","messageId":"843de8864a9fa0b89b5dede0d24a982959b0ad1a.1624636945.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v5 15/15] [GSOC] ref-filter: remove grab_oid() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-25T16:02:25Z","receivedAt":"2021-06-25T16:02:59Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause \"atom_type == ATOM_OBJECTNAME\" implies the condition\nof `starts_with(name, \"objectname\")`, \"atom_type == ATOM_TREE\"\nimplies the condition of `starts_with(name, \"tree\")`, so the\ncheck for `starts_with(name, field)` in grab_oid() is redundant.\n\nSo Remove the grab_oid() from ref-filter, to reduce repeated check.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 26 +++++++++-----------------\n 1 file changed, 9 insertions(+), 17 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 99b87742b0f..ab53c1fd22f 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1071,16 +1071,6 @@ static const char *do_grab_oid(const char *field, const struct object_id *oid,\n \t}\n }\n \n-static int grab_oid(const char *name, const char *field, const struct object_id *oid,\n-\t\t    struct atom_value *v, struct used_atom *atom)\n-{\n-\tif (starts_with(name, field)) {\n-\t\tv->s = xstrdup(do_grab_oid(field, oid, atom));\n-\t\treturn 1;\n-\t}\n-\treturn 0;\n-}\n-\n /* See grab_values */\n static void grab_common_values(struct atom_value *val, int deref, struct expand_data *oi)\n {\n@@ -1106,8 +1096,9 @@ static void grab_common_values(struct atom_value *val, int deref, struct expand_\n \t\t\t}\n \t\t} else if (atom_type == ATOM_DELTABASE)\n \t\t\tv->s = xstrdup(oid_to_hex(&oi->delta_base_oid));\n-\t\telse if (atom_type == ATOM_OBJECTNAME && deref)\n-\t\t\tgrab_oid(name, \"objectname\", &oi->oid, v, &used_atom[i]);\n+\t\telse if (atom_type == ATOM_OBJECTNAME && deref) {\n+\t\t\tv->s = xstrdup(do_grab_oid(\"objectname\", &oi->oid, &used_atom[i]));\n+\t\t}\n \t}\n }\n \n@@ -1148,9 +1139,10 @@ static void grab_commit_values(struct atom_value *val, int deref, struct object\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (atom_type == ATOM_TREE &&\n-\t\t    grab_oid(name, \"tree\", get_commit_tree_oid(commit), v, &used_atom[i]))\n+\t\tif (atom_type == ATOM_TREE) {\n+\t\t\tv->s = xstrdup(do_grab_oid(\"tree\", get_commit_tree_oid(commit), &used_atom[i]));\n \t\t\tcontinue;\n+\t\t}\n \t\tif (atom_type == ATOM_NUMPARENT) {\n \t\t\tv->value = commit_list_count(commit->parents);\n \t\t\tv->s = xstrfmt(\"%lu\", (unsigned long)v->value);\n@@ -1971,9 +1963,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\t\tv->s = xstrdup(buf + 1);\n \t\t\t}\n \t\t\tcontinue;\n-\t\t} else if (!deref && atom_type == ATOM_OBJECTNAME &&\n-\t\t\t   grab_oid(name, \"objectname\", &ref->objectname, v, atom)) {\n-\t\t\t\tcontinue;\n+\t\t} else if (!deref && atom_type == ATOM_OBJECTNAME) {\n+\t\t\t   v->s = xstrdup(do_grab_oid(\"objectname\", &ref->objectname, atom));\n+\t\t\t   continue;\n \t\t} else if (atom_type == ATOM_HEAD) {\n \t\t\tif (atom->u.head && !strcmp(ref->refname, atom->u.head))\n \t\t\t\tv->s = xstrdup(\"*\");\n-- \ngitgitgadget\n"},{"id":"428510","messageId":"946c95bc-d57a-5b55-0dcf-5c4d6f980396@gmail.com","threadId":"55909","inReplyTo":"4e473838b9d2651a8e4be27332697c2ba354db5a.1624636945.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 02/15] [GSOC] ref-filter: add %(raw) atom","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-06-26T00:42:07Z","receivedAt":"2021-06-26T00:42:13Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 25/06/21 23.02, ZheNing Hu via GitGitGadget wrote:\n> Note that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> `--tcl`, and `--perl` because if the binary raw data is passed to a\n> variable in such languages, these may not support arbitrary binary data\n> in their string variable type.\n> \n\nCommit message looks OK, but...\n\n> +Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n> +`--perl` because the host language may not support arbitrary binary data in the\n> +variables of its string type.\n> +\n\nSeems like out of sync between commit message and the docs change above. \nDid you mean the (unsupported) host languages are Python, BASH script, \nTCL/TK, and Perl respectively? If so, the docs should say:\n\n\"Note that `--format=%(raw) can not be used with `--python`, `--shell`, \n`-tcl`, and `--perl` because such languages may not support arbitrary \nbinary data in their string variable type.\"\n\nThanks.\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"428529","messageId":"CA+CkUQ9jWY8KDJxeAk9kDSCGgQLuBuaLEASrGfbA2xnN7nuBBw@mail.gmail.com","threadId":"55909","inReplyTo":"370101ba65f0989487360366f8b83144a6641a04.1624636945.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"Hariom verma","fromEmail":"hariom18599@gmail.com","sentAt":"2021-06-26T17:26:58Z","receivedAt":"2021-06-26T17:27:13Z","isPatch":true,"sender":{"key":"hariom18599@gmail.com","avatar":"https://avatars.githubusercontent.com/u/37576387?v=4"},"body":"Hi,\n\nOn Fri, Jun 25, 2021 at 9:32 PM ZheNing Hu via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> +       if (ret < 0) {\n> +               die(\"%s\\n\", err.buf);\n> +       } if (ret) {\n> +               /* ret > 0 means when the object corresponding to oid\n> +                * cannot be found in format_ref_array_item(), we only print\n> +                * the error message.\n> +                */\n> +               printf(\"%s\\n\", err.buf);\n> +               fflush(stdout);\n> +       } else {\n> +               strbuf_addch(scratch, '\\n');\n> +               batch_write(opt, scratch->buf, scratch->len);\n>         }\n> +       free_ref_array_item_value(&item);\n> +       strbuf_release(&err);\n>  }\n\nI think you can get rid of braces in condition `ret < 0`:\n\n```\n        if (ret < 0)\n                die(\"%s\\n\", err->buf);\n        if (ret) {\n                /* ret > 0 means when the object corresponding to oid\n                 * cannot be found in format_ref_array_item(), we only print\n                 * the error message.\n                 */\n                printf(\"%s\\n\", err->buf);\n                fflush(stdout);\n        } else {\n                strbuf_addch(scratch, '\\n');\n                batch_write(opt, scratch->buf, scratch->len);\n        }\n```\n\nThanks,\nHariom.\n"},{"id":"428531","messageId":"CA+CkUQ9XR4TjEea0Z4pHBeOdQi7fuTLPtzi01JdTKSS38=4CMg@mail.gmail.com","threadId":"55909","inReplyTo":"370101ba65f0989487360366f8b83144a6641a04.1624636945.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"Hariom verma","fromEmail":"hariom18599@gmail.com","sentAt":"2021-06-26T18:08:17Z","receivedAt":"2021-06-26T18:08:33Z","isPatch":true,"sender":{"key":"hariom18599@gmail.com","avatar":"https://avatars.githubusercontent.com/u/37576387?v=4"},"body":"On Fri, Jun 25, 2021 at 9:32 PM ZheNing Hu via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n>  static void batch_object_write(const char *obj_name,\n>                                struct strbuf *scratch,\n>                                struct batch_options *opt,\n>                                struct expand_data *data)\n>  {\n> -       if (!data->skip_object_info &&\n> -           oid_object_info_extended(the_repository, &data->oid, &data->info,\n> -                                    OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n> -               printf(\"%s missing\\n\",\n> -                      obj_name ? obj_name : oid_to_hex(&data->oid));\n> -               fflush(stdout);\n> -               return;\n> -       }\n> +       int ret = 0;\n\nNo need to initialize `ret` with 0. Later we are going to assign it\nwith the return value of `format_ref_array_item()` anyway.\n\n> +       struct strbuf err = STRBUF_INIT;\n> +       struct ref_array_item item = { data->oid, data->rest };\n>\n>         strbuf_reset(scratch);\n> -       strbuf_expand(scratch, opt->format, expand_format, data);\n> -       strbuf_addch(scratch, '\\n');\n> -       batch_write(opt, scratch->buf, scratch->len);\n>\n> -       if (opt->print_contents) {\n> -               print_object_or_die(opt, data);\n> -               batch_write(opt, \"\\n\", 1);\n> +       ret = format_ref_array_item(&item, &opt->format, scratch, &err);\n\nHere.\n\n-- \nHariom\n"},{"id":"428550","messageId":"CAOLTT8QOi0wYoouqaWn43CKR1bZT6U8v+T+6MMbq1R-4wjqBPg@mail.gmail.com","threadId":"55909","inReplyTo":"CA+CkUQ9jWY8KDJxeAk9kDSCGgQLuBuaLEASrGfbA2xnN7nuBBw@mail.gmail.com","subject":"Re: [PATCH v5 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-27T11:31:28Z","receivedAt":"2021-06-27T11:31:44Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Hariom verma <hariom18599@gmail.com> 于2021年6月27日周日 上午1:27写道：\n>\n> Hi,\n>\n> On Fri, Jun 25, 2021 at 9:32 PM ZheNing Hu via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> > +       if (ret < 0) {\n> > +               die(\"%s\\n\", err.buf);\n> > +       } if (ret) {\n> > +               /* ret > 0 means when the object corresponding to oid\n> > +                * cannot be found in format_ref_array_item(), we only print\n> > +                * the error message.\n> > +                */\n> > +               printf(\"%s\\n\", err.buf);\n> > +               fflush(stdout);\n> > +       } else {\n> > +               strbuf_addch(scratch, '\\n');\n> > +               batch_write(opt, scratch->buf, scratch->len);\n> >         }\n> > +       free_ref_array_item_value(&item);\n> > +       strbuf_release(&err);\n> >  }\n>\n> I think you can get rid of braces in condition `ret < 0`:\n>\n\nMake sences. ;-)\n\n> ```\n>         if (ret < 0)\n>                 die(\"%s\\n\", err->buf);\n>         if (ret) {\n>                 /* ret > 0 means when the object corresponding to oid\n>                  * cannot be found in format_ref_array_item(), we only print\n>                  * the error message.\n>                  */\n>                 printf(\"%s\\n\", err->buf);\n>                 fflush(stdout);\n>         } else {\n>                 strbuf_addch(scratch, '\\n');\n>                 batch_write(opt, scratch->buf, scratch->len);\n>         }\n> ```\n>\n> Thanks,\n> Hariom.\n\nThanks,\nZheNing Hu\n"},{"id":"428551","messageId":"CAOLTT8Qq4sgR9DnCk=+ovHUEqNeSkg-07fOpJQu5povaZFXgiA@mail.gmail.com","threadId":"55909","inReplyTo":"CA+CkUQ9XR4TjEea0Z4pHBeOdQi7fuTLPtzi01JdTKSS38=4CMg@mail.gmail.com","subject":"Re: [PATCH v5 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-27T11:34:14Z","receivedAt":"2021-06-27T11:34:29Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Hariom verma <hariom18599@gmail.com> 于2021年6月27日周日 上午2:08写道：\n>\n> On Fri, Jun 25, 2021 at 9:32 PM ZheNing Hu via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: ZheNing Hu <adlternative@gmail.com>\n> >\n> >  static void batch_object_write(const char *obj_name,\n> >                                struct strbuf *scratch,\n> >                                struct batch_options *opt,\n> >                                struct expand_data *data)\n> >  {\n> > -       if (!data->skip_object_info &&\n> > -           oid_object_info_extended(the_repository, &data->oid, &data->info,\n> > -                                    OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n> > -               printf(\"%s missing\\n\",\n> > -                      obj_name ? obj_name : oid_to_hex(&data->oid));\n> > -               fflush(stdout);\n> > -               return;\n> > -       }\n> > +       int ret = 0;\n>\n> No need to initialize `ret` with 0. Later we are going to assign it\n> with the return value of `format_ref_array_item()` anyway.\n>\n\nI agree. It is worth noting that there are similar `int ret = 0` in\nref-filter.c,\nthey should be changed too.\n\n> > +       struct strbuf err = STRBUF_INIT;\n> > +       struct ref_array_item item = { data->oid, data->rest };\n> >\n> >         strbuf_reset(scratch);\n> > -       strbuf_expand(scratch, opt->format, expand_format, data);\n> > -       strbuf_addch(scratch, '\\n');\n> > -       batch_write(opt, scratch->buf, scratch->len);\n> >\n> > -       if (opt->print_contents) {\n> > -               print_object_or_die(opt, data);\n> > -               batch_write(opt, \"\\n\", 1);\n> > +       ret = format_ref_array_item(&item, &opt->format, scratch, &err);\n>\n> Here.\n>\n> --\n> Hariom\n\n--\nZheNing Hu\n"},{"id":"428552","messageId":"CAOLTT8TK4cC1P-h+=wag8OdPc3C_Acd-dHY55NuA48UuuABAiA@mail.gmail.com","threadId":"55909","inReplyTo":"946c95bc-d57a-5b55-0dcf-5c4d6f980396@gmail.com","subject":"Re: [PATCH v5 02/15] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-27T11:43:02Z","receivedAt":"2021-06-27T11:43:16Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Bagas Sanjaya <bagasdotme@gmail.com> 于2021年6月26日周六 上午8:42写道：\n>\n> On 25/06/21 23.02, ZheNing Hu via GitGitGadget wrote:\n> > Note that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n> > `--tcl`, and `--perl` because if the binary raw data is passed to a\n> > variable in such languages, these may not support arbitrary binary data\n> > in their string variable type.\n> >\n>\n> Commit message looks OK, but...\n>\n> > +Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n> > +`--perl` because the host language may not support arbitrary binary data in the\n> > +variables of its string type.\n> > +\n>\n> Seems like out of sync between commit message and the docs change above.\n> Did you mean the (unsupported) host languages are Python, BASH script,\n> TCL/TK, and Perl respectively? If so, the docs should say:\n>\n\nIndeed so. I will change them too.\n\n> \"Note that `--format=%(raw) can not be used with `--python`, `--shell`,\n> `-tcl`, and `--perl` because such languages may not support arbitrary\n> binary data in their string variable type.\"\n>\n> Thanks.\n>\n> --\n> An old man doll... just what I always wanted! - Clara\n\nThanks.\n--\nZheNing Hu\n"},{"id":"428553","messageId":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v5.git.1624636945.gitgitgadget@gmail.com","subject":"[PATCH v6 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:35Z","receivedAt":"2021-06-27T12:35:59Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"This patch series make cat-file reuse ref-filter logic.\n\nChange from last version:\n\n 1. Amend part of the description of git for-each-ref.txt.\n 2. Modify the code style.\n 3. Do not assign the 0 to the variable ret during it's initialization.\n\nZheNing Hu (15):\n  [GSOC] ref-filter: add obj-type check in grab contents\n  [GSOC] ref-filter: add %(raw) atom\n  [GSOC] ref-filter: --format=%(raw) re-support --perl\n  [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  [GSOC] ref-filter: add %(rest) atom\n  [GSOC] ref-filter: pass get_object() return value to their callers\n  [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  [GSOC] ref-filter: modify the error message and value in get_object\n  [GSOC] cat-file: add has_object_file() check\n  [GSOC] cat-file: change batch_objects parameter name\n  [GSOC] cat-file: reuse ref-filter logic\n  [GSOC] cat-file: reuse err buf in batch_object_write()\n  [GSOC] cat-file: re-implement --textconv, --filters options\n  [GSOC] ref-filter: remove grab_oid() function\n\n Documentation/git-cat-file.txt     |   6 +\n Documentation/git-for-each-ref.txt |   9 +\n builtin/cat-file.c                 | 277 ++++++----------------\n builtin/tag.c                      |   2 +-\n quote.c                            |  17 ++\n quote.h                            |   1 +\n ref-filter.c                       | 357 ++++++++++++++++++++++-------\n ref-filter.h                       |  14 +-\n t/t1006-cat-file.sh                | 252 ++++++++++++++++++++\n t/t3203-branch-output.sh           |   4 +\n t/t6300-for-each-ref.sh            | 235 +++++++++++++++++++\n t/t6301-for-each-ref-errors.sh     |   2 +-\n t/t7004-tag.sh                     |   4 +\n t/t7030-verify-tag.sh              |   4 +\n 14 files changed, 888 insertions(+), 296 deletions(-)\n\n\nbase-commit: 1197f1a46360d3ae96bd9c15908a3a6f8e562207\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-980%2Fadlternative%2Fcat-file-batch-refactor-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-980/adlternative/cat-file-batch-refactor-v6\nPull-Request: https://github.com/gitgitgadget/git/pull/980\n\nRange-diff vs v5:\n\n  1:  f72ad9cc5e8 =  1:  f72ad9cc5e8 [GSOC] ref-filter: add obj-type check in grab contents\n  2:  4e473838b9d !  2:  d9bc50c4ae6 [GSOC] ref-filter: add %(raw) atom\n     @@ Commit message\n      \n          Mentored-by: Christian Couder <christian.couder@gmail.com>\n          Mentored-by: Hariom Verma <hariom18599@gmail.com>\n     +    Helped-by: Bagas Sanjaya <bagasdotme@gmail.com>\n          Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n          Helped-by: Felipe Contreras <felipe.contreras@gmail.com>\n          Helped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\n     @@ Documentation/git-for-each-ref.txt: and `date` to extract the named component.\n      +\tThe raw data size of the object.\n      +\n      +Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n     -+`--perl` because the host language may not support arbitrary binary data in the\n     -+variables of its string type.\n     ++`--perl` because the such language may not support arbitrary binary data in their\n     ++string variable type.\n      +\n       The message in a commit or a tag object is `contents`, from which\n       `contents:<part>` can be used to extract various parts out of:\n     @@ t/t6300-for-each-ref.sh: test_atom refs/myblobs/first contents:body \"\"\n      +\tprintf \"  \" >blob7 &&\n      +\t>blob8 &&\n      +\tobj=$(git hash-object -w blob1) &&\n     -+        git update-ref refs/myblobs/blob1 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob1 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob2) &&\n     -+        git update-ref refs/myblobs/blob2 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob2 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob3) &&\n     -+        git update-ref refs/myblobs/blob3 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob3 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob4) &&\n     -+        git update-ref refs/myblobs/blob4 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob4 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob5) &&\n     -+        git update-ref refs/myblobs/blob5 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob5 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob6) &&\n     -+        git update-ref refs/myblobs/blob6 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob6 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob7) &&\n     -+        git update-ref refs/myblobs/blob7 \"$obj\" &&\n     ++\tgit update-ref refs/myblobs/blob7 \"$obj\" &&\n      +\tobj=$(git hash-object -w blob8) &&\n     -+        git update-ref refs/myblobs/blob8 \"$obj\"\n     ++\tgit update-ref refs/myblobs/blob8 \"$obj\"\n      +'\n      +\n      +test_expect_success 'Verify sorts with raw' '\n  3:  765cf08a108 !  3:  47f868f63d9 [GSOC] ref-filter: --format=%(raw) re-support --perl\n     @@ Documentation/git-for-each-ref.txt: raw:size::\n       \tThe raw data size of the object.\n       \n       Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n     --`--perl` because the host language may not support arbitrary binary data in the\n     -+because the host language may not support arbitrary binary data in the\n     - variables of its string type.\n     +-`--perl` because the such language may not support arbitrary binary data in their\n     ++because the such language may not support arbitrary binary data in their\n     + string variable type.\n       \n       The message in a commit or a tag object is `contents`, from which\n      \n  4:  d2aeafd0ef3 =  4:  debca156470 [GSOC] ref-filter: use non-const ref_format in *_atom_parser()\n  5:  1ca3a42f041 =  5:  cb0df2b8207 [GSOC] ref-filter: add %(rest) atom\n  6:  67f1a3cca9a !  6:  9873354930a [GSOC] ref-filter: pass get_object() return value to their callers\n     @@ ref-filter.c: static int populate_value(struct ref_array_item *ref, struct strbu\n       {\n       \tstruct object *obj;\n       \tint i;\n     -+\tint ret = 0;\n     ++\tint ret;\n       \tstruct object_info empty = OBJECT_INFO_INIT;\n       \n       \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n     @@ ref-filter.c: int format_ref_array_item(struct ref_array_item *info,\n       {\n       \tconst char *cp, *sp, *ep;\n       \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n     -+\tint ret = 0;\n     ++\tint ret;\n       \n       \tstate.quote_style = format->quote_style;\n       \tpush_stack_element(&state.stack);\n  7:  2a48a48e81c =  7:  e592c21ea1d [GSOC] ref-filter: introduce free_ref_array_item_value() function\n  8:  be55005be75 =  8:  b6e7757de4c [GSOC] ref-filter: add cat_file_mode in struct ref_format\n  9:  937f88b7837 =  9:  85686187d49 [GSOC] ref-filter: modify the error message and value in get_object\n 10:  45657499c55 = 10:  6037295ee58 [GSOC] cat-file: add has_object_file() check\n 11:  bf5c0a017ad = 11:  32e1ca56389 [GSOC] cat-file: change batch_objects parameter name\n 12:  370101ba65f ! 12:  9a1f0732940 [GSOC] cat-file: reuse ref-filter logic\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      -\t\tfflush(stdout);\n      -\t\treturn;\n      -\t}\n     -+\tint ret = 0;\n     ++\tint ret;\n      +\tstruct strbuf err = STRBUF_INIT;\n      +\tstruct ref_array_item item = { data->oid, data->rest };\n       \n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n      -\t\tprint_object_or_die(opt, data);\n      -\t\tbatch_write(opt, \"\\n\", 1);\n      +\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n     -+\tif (ret < 0) {\n     ++\tif (ret < 0)\n      +\t\tdie(\"%s\\n\", err.buf);\n     -+\t} if (ret) {\n     ++\tif (ret) {\n      +\t\t/* ret > 0 means when the object corresponding to oid\n      +\t\t * cannot be found in format_ref_array_item(), we only print\n      +\t\t * the error message.\n 13:  69eef47065d ! 13:  3fb47584924 [GSOC] cat-file: reuse err buf in batch_object_write()\n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n       \t\t\t       struct batch_options *opt,\n       \t\t\t       struct expand_data *data)\n       {\n     - \tint ret = 0;\n     + \tint ret;\n      -\tstruct strbuf err = STRBUF_INIT;\n       \tstruct ref_array_item item = { data->oid, data->rest };\n       \n     @@ builtin/cat-file.c: static void batch_write(struct batch_options *opt, const voi\n       \n      -\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n      +\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n     - \tif (ret < 0) {\n     + \tif (ret < 0)\n      -\t\tdie(\"%s\\n\", err.buf);\n      +\t\tdie(\"%s\\n\", err->buf);\n     - \t} if (ret) {\n     + \tif (ret) {\n       \t\t/* ret > 0 means when the object corresponding to oid\n       \t\t * cannot be found in format_ref_array_item(), we only print\n       \t\t * the error message.\n 14:  a7ac037a946 = 14:  e0b1a05e711 [GSOC] cat-file: re-implement --textconv, --filters options\n 15:  843de8864a9 = 15:  891d62fd93f [GSOC] ref-filter: remove grab_oid() function\n\n-- \ngitgitgadget\n"},{"id":"428554","messageId":"d9bc50c4ae699d0516581ac67a2d7b602d2e61a4.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 02/15] [GSOC] ref-filter: add %(raw) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:37Z","receivedAt":"2021-06-27T12:35:59Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd new formatting option `%(raw)`, which will print the raw\nobject data without any changes. It will help further to migrate\nall cat-file formatting logic from cat-file to ref-filter.\n\nThe raw data of blob, tree objects may contain '\\0', but most of\nthe logic in `ref-filter` depends on the output of the atom being\ntext (specifically, no embedded NULs in it).\n\nE.g. `quote_formatting()` use `strbuf_addstr()` or `*._quote_buf()`\nadd the data to the buffer. The raw data of a tree object is\n`100644 one\\0...`, only the `100644 one` will be added to the buffer,\nwhich is incorrect.\n\nTherefore, we need to find a way to record the length of the\natom_value's member `s`. Although strbuf can already record the\nstring and its length, if we want to replace the type of atom_value's\nmember `s` with strbuf, many places in ref-filter that are filled\nwith dynamically allocated mermory in `v->s` are not easy to replace.\nAt the same time, we need to check if `v->s == NULL` in\npopulate_value(), and strbuf cannot easily distinguish NULL and empty\nstrings, but c-style \"const char *\" can do it. So add a new member in\n`struct atom_value`: `s_size`, which can record raw object size, it\ncan help us add raw object data to the buffer or compare two buffers\nwhich contain raw object data.\n\nNote that `--format=%(raw)` cannot be used with `--python`, `--shell`,\n`--tcl`, and `--perl` because if the binary raw data is passed to a\nvariable in such languages, these may not support arbitrary binary data\nin their string variable type.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Bagas Sanjaya <bagasdotme@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Felipe Contreras <felipe.contreras@gmail.com>\nHelped-by: Phillip Wood <phillip.wood@dunelm.org.uk>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nBased-on-patch-by: Olga Telezhnaya <olyatelezhnaya@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |   9 ++\n ref-filter.c                       | 139 +++++++++++++++----\n t/t6300-for-each-ref.sh            | 216 +++++++++++++++++++++++++++++\n 3 files changed, 337 insertions(+), 27 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 2ae2478de70..3727a5ffee7 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -235,6 +235,15 @@ and `date` to extract the named component.  For email fields (`authoremail`,\n without angle brackets, and `:localpart` to get the part before the `@` symbol\n out of the trimmed email.\n \n+The raw data in an object is `raw`.\n+\n+raw:size::\n+\tThe raw data size of the object.\n+\n+Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n+`--perl` because the such language may not support arbitrary binary data in their\n+string variable type.\n+\n The message in a commit or a tag object is `contents`, from which\n `contents:<part>` can be used to extract various parts out of:\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 5cee6512fba..7822be90307 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -144,6 +144,7 @@ enum atom_type {\n \tATOM_BODY,\n \tATOM_TRAILERS,\n \tATOM_CONTENTS,\n+\tATOM_RAW,\n \tATOM_UPSTREAM,\n \tATOM_PUSH,\n \tATOM_SYMREF,\n@@ -189,6 +190,9 @@ static struct used_atom {\n \t\t\tstruct process_trailer_options trailer_opts;\n \t\t\tunsigned int nlines;\n \t\t} contents;\n+\t\tstruct {\n+\t\t\tenum { RAW_BARE, RAW_LENGTH } option;\n+\t\t} raw_data;\n \t\tstruct {\n \t\t\tcmp_status cmp_status;\n \t\t\tconst char *str;\n@@ -426,6 +430,18 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n+static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+\t\t\t\tconst char *arg, struct strbuf *err)\n+{\n+\tif (!arg)\n+\t\tatom->u.raw_data.option = RAW_BARE;\n+\telse if (!strcmp(arg, \"size\"))\n+\t\tatom->u.raw_data.option = RAW_LENGTH;\n+\telse\n+\t\treturn strbuf_addf_ret(err, -1, _(\"unrecognized %%(raw) argument: %s\"), arg);\n+\treturn 0;\n+}\n+\n static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n@@ -586,6 +602,7 @@ static struct {\n \t[ATOM_BODY] = { \"body\", SOURCE_OBJ, FIELD_STR, body_atom_parser },\n \t[ATOM_TRAILERS] = { \"trailers\", SOURCE_OBJ, FIELD_STR, trailers_atom_parser },\n \t[ATOM_CONTENTS] = { \"contents\", SOURCE_OBJ, FIELD_STR, contents_atom_parser },\n+\t[ATOM_RAW] = { \"raw\", SOURCE_OBJ, FIELD_STR, raw_atom_parser },\n \t[ATOM_UPSTREAM] = { \"upstream\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_PUSH] = { \"push\", SOURCE_NONE, FIELD_STR, remote_ref_atom_parser },\n \t[ATOM_SYMREF] = { \"symref\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -620,12 +637,15 @@ struct ref_formatting_state {\n \n struct atom_value {\n \tconst char *s;\n+\tsize_t s_size;\n \tint (*handler)(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t       struct strbuf *err);\n \tuintmax_t value; /* used for sorting when not FIELD_STR */\n \tstruct used_atom *atom;\n };\n \n+#define ATOM_VALUE_S_SIZE_INIT (-1)\n+\n /*\n  * Used to parse format string and sort specifiers\n  */\n@@ -644,13 +664,6 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \t\treturn strbuf_addf_ret(err, -1, _(\"malformed field name: %.*s\"),\n \t\t\t\t       (int)(ep-atom), atom);\n \n-\t/* Do we have the atom already used elsewhere? */\n-\tfor (i = 0; i < used_atom_cnt; i++) {\n-\t\tint len = strlen(used_atom[i].name);\n-\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n-\t\t\treturn i;\n-\t}\n-\n \t/*\n \t * If the atom name has a colon, strip it and everything after\n \t * it off - it specifies the format for this entry, and\n@@ -660,6 +673,13 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \targ = memchr(sp, ':', ep - sp);\n \tatom_len = (arg ? arg : ep) - sp;\n \n+\t/* Do we have the atom already used elsewhere? */\n+\tfor (i = 0; i < used_atom_cnt; i++) {\n+\t\tint len = strlen(used_atom[i].name);\n+\t\tif (len == ep - atom && !memcmp(used_atom[i].name, atom, len))\n+\t\t\treturn i;\n+\t}\n+\n \t/* Is the atom a valid one? */\n \tfor (i = 0; i < ARRAY_SIZE(valid_atom); i++) {\n \t\tint len = strlen(valid_atom[i].name);\n@@ -709,11 +729,14 @@ static int parse_ref_filter_atom(const struct ref_format *format,\n \treturn at;\n }\n \n-static void quote_formatting(struct strbuf *s, const char *str, int quote_style)\n+static void quote_formatting(struct strbuf *s, const char *str, size_t len, int quote_style)\n {\n \tswitch (quote_style) {\n \tcase QUOTE_NONE:\n-\t\tstrbuf_addstr(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(s, str, len);\n+\t\telse\n+\t\t\tstrbuf_addstr(s, str);\n \t\tbreak;\n \tcase QUOTE_SHELL:\n \t\tsq_quote_buf(s, str);\n@@ -740,9 +763,12 @@ static int append_atom(struct atom_value *v, struct ref_formatting_state *state,\n \t * encountered.\n \t */\n \tif (!state->stack->prev)\n-\t\tquote_formatting(&state->stack->output, v->s, state->quote_style);\n+\t\tquote_formatting(&state->stack->output, v->s, v->s_size, state->quote_style);\n \telse\n-\t\tstrbuf_addstr(&state->stack->output, v->s);\n+\t\tif (v->s_size != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tstrbuf_add(&state->stack->output, v->s, v->s_size);\n+\t\telse\n+\t\t\tstrbuf_addstr(&state->stack->output, v->s);\n \treturn 0;\n }\n \n@@ -842,21 +868,23 @@ static int if_atom_handler(struct atom_value *atomv, struct ref_formatting_state\n \treturn 0;\n }\n \n-static int is_empty(const char *s)\n+static int is_empty(struct strbuf *buf)\n {\n-\twhile (*s != '\\0') {\n-\t\tif (!isspace(*s))\n-\t\t\treturn 0;\n-\t\ts++;\n-\t}\n-\treturn 1;\n-}\n+\tconst char *cur = buf->buf;\n+\tconst char *end = buf->buf + buf->len;\n+\n+\twhile (cur != end && (isspace(*cur)))\n+\t\tcur++;\n+\n+\treturn cur == end;\n+ }\n \n static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_state *state,\n \t\t\t     struct strbuf *err)\n {\n \tstruct ref_formatting_stack *cur = state->stack;\n \tstruct if_then_else *if_then_else = NULL;\n+\tsize_t str_len = 0;\n \n \tif (cur->at_end == if_then_else_handler)\n \t\tif_then_else = (struct if_then_else *)cur->at_end_data;\n@@ -867,18 +895,22 @@ static int then_atom_handler(struct atom_value *atomv, struct ref_formatting_sta\n \tif (if_then_else->else_atom_seen)\n \t\treturn strbuf_addf_ret(err, -1, _(\"format: %%(then) atom used after %%(else)\"));\n \tif_then_else->then_atom_seen = 1;\n+\tif (if_then_else->str)\n+\t\tstr_len = strlen(if_then_else->str);\n \t/*\n \t * If the 'equals' or 'notequals' attribute is used then\n \t * perform the required comparison. If not, only non-empty\n \t * strings satisfy the 'if' condition.\n \t */\n \tif (if_then_else->cmp_status == COMPARE_EQUAL) {\n-\t\tif (!strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len == cur->output.len &&\n+\t\t    !memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n \t} else if (if_then_else->cmp_status == COMPARE_UNEQUAL) {\n-\t\tif (strcmp(if_then_else->str, cur->output.buf))\n+\t\tif (str_len != cur->output.len ||\n+\t\t    memcmp(if_then_else->str, cur->output.buf, cur->output.len))\n \t\t\tif_then_else->condition_satisfied = 1;\n-\t} else if (cur->output.len && !is_empty(cur->output.buf))\n+\t} else if (cur->output.len && !is_empty(&cur->output))\n \t\tif_then_else->condition_satisfied = 1;\n \tstrbuf_reset(&cur->output);\n \treturn 0;\n@@ -924,7 +956,7 @@ static int end_atom_handler(struct atom_value *atomv, struct ref_formatting_stat\n \t * only on the topmost supporting atom.\n \t */\n \tif (!current->prev->prev) {\n-\t\tquote_formatting(&s, current->output.buf, state->quote_style);\n+\t\tquote_formatting(&s, current->output.buf, current->output.len, state->quote_style);\n \t\tstrbuf_swap(&current->output, &s);\n \t}\n \tstrbuf_release(&s);\n@@ -974,6 +1006,10 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n+\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\t\t\tdie(_(\"--format=%.*s cannot be used with\"\n+\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\n@@ -1362,17 +1398,29 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, struct exp\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n \tvoid *buf = data->content;\n+\tunsigned long buf_size = data->size;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n \t\tconst char *name = atom->name;\n \t\tstruct atom_value *v = &val[i];\n+\t\tenum atom_type atom_type = atom->atom_type;\n \n \t\tif (!!deref != (*name == '*'))\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n \n+\t\tif (atom_type == ATOM_RAW) {\n+\t\t\tif (atom->u.raw_data.option == RAW_BARE) {\n+\t\t\t\tv->s = xmemdupz(buf, buf_size);\n+\t\t\t\tv->s_size = buf_size;\n+\t\t\t} else if (atom->u.raw_data.option == RAW_LENGTH) {\n+\t\t\t\tv->s = xstrfmt(\"%\"PRIuMAX, (uintmax_t)buf_size);\n+\t\t\t}\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif ((data->type != OBJ_TAG &&\n \t\t     data->type != OBJ_COMMIT) ||\n \t\t    (strcmp(name, \"body\") &&\n@@ -1460,9 +1508,11 @@ static void grab_values(struct atom_value *val, int deref, struct object *obj, s\n \t\tbreak;\n \tcase OBJ_TREE:\n \t\t/* grab_tree_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tcase OBJ_BLOB:\n \t\t/* grab_blob_values(val, deref, obj, buf, sz); */\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tbreak;\n \tdefault:\n \t\tdie(\"Eh?  Object of type %d?\", obj->type);\n@@ -1766,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\tconst char *refname;\n \t\tstruct branch *branch = NULL;\n \n+\t\tv->s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tv->handler = append_atom;\n \t\tv->atom = atom;\n \n@@ -2369,6 +2420,19 @@ static int compare_detached_head(struct ref_array_item *a, struct ref_array_item\n \treturn 0;\n }\n \n+static int memcasecmp(const void *vs1, const void *vs2, size_t n)\n+{\n+\tconst char *s1 = vs1, *s2 = vs2;\n+\tconst char *end = s1 + n;\n+\n+\tfor (; s1 < end; s1++, s2++) {\n+\t\tint diff = tolower(*s1) - tolower(*s2);\n+\t\tif (diff)\n+\t\t\treturn diff;\n+\t}\n+\treturn 0;\n+}\n+\n static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, struct ref_array_item *b)\n {\n \tstruct atom_value *va, *vb;\n@@ -2389,10 +2453,30 @@ static int cmp_ref_sorting(struct ref_sorting *s, struct ref_array_item *a, stru\n \t} else if (s->sort_flags & REF_SORTING_VERSION) {\n \t\tcmp = versioncmp(va->s, vb->s);\n \t} else if (cmp_type == FIELD_STR) {\n-\t\tint (*cmp_fn)(const char *, const char *);\n-\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n-\t\t\t? strcasecmp : strcmp;\n-\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\tif (va->s_size == ATOM_VALUE_S_SIZE_INIT &&\n+\t\t    vb->s_size == ATOM_VALUE_S_SIZE_INIT) {\n+\t\t\tint (*cmp_fn)(const char *, const char *);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? strcasecmp : strcmp;\n+\t\t\tcmp = cmp_fn(va->s, vb->s);\n+\t\t} else {\n+\t\t\tsize_t a_size = va->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(va->s) : va->s_size;\n+\t\t\tsize_t b_size = vb->s_size == ATOM_VALUE_S_SIZE_INIT ?\n+\t\t\t\t\tstrlen(vb->s) : vb->s_size;\n+\t\t\tint (*cmp_fn)(const void *, const void *, size_t);\n+\t\t\tcmp_fn = s->sort_flags & REF_SORTING_ICASE\n+\t\t\t\t? memcasecmp : memcmp;\n+\n+\t\t\tcmp = cmp_fn(va->s, vb->s, b_size > a_size ?\n+\t\t\t\t     a_size : b_size);\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (a_size > b_size)\n+\t\t\t\t\tcmp = 1;\n+\t\t\t\telse if (a_size < b_size)\n+\t\t\t\t\tcmp = -1;\n+\t\t\t}\n+\t\t}\n \t} else {\n \t\tif (va->value < vb->value)\n \t\t\tcmp = -1;\n@@ -2492,6 +2576,7 @@ int format_ref_array_item(struct ref_array_item *info,\n \t}\n \tif (format->need_color_reset_at_eol) {\n \t\tstruct atom_value resetv;\n+\t\tresetv.s_size = ATOM_VALUE_S_SIZE_INIT;\n \t\tresetv.s = GIT_COLOR_RESET;\n \t\tif (append_atom(&resetv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 9e0214076b4..18554f62d94 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -130,6 +130,8 @@ test_atom head parent:short=10 ''\n test_atom head numparent 0\n test_atom head object ''\n test_atom head type ''\n+test_atom head raw \"$(git cat-file commit refs/heads/main)\n+\"\n test_atom head '*objectname' ''\n test_atom head '*objecttype' ''\n test_atom head author 'A U Thor <author@example.com> 1151968724 +0200'\n@@ -221,6 +223,15 @@ test_atom tag contents 'Tagging at 1151968727\n '\n test_atom tag HEAD ' '\n \n+test_expect_success 'basic atom: refs/tags/testtag *raw' '\n+\tgit cat-file commit refs/tags/testtag^{} >expected &&\n+\tgit for-each-ref --format=\"%(*raw)\" refs/tags/testtag >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'Check invalid atoms names are errors' '\n \ttest_must_fail git for-each-ref --format=\"%(INVALID)\" refs/heads\n '\n@@ -686,6 +697,15 @@ test_atom refs/tags/signed-empty contents:body ''\n test_atom refs/tags/signed-empty contents:signature \"$sig\"\n test_atom refs/tags/signed-empty contents \"$sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-empty raw' '\n+\tgit cat-file tag refs/tags/signed-empty >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-empty >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-short subject 'subject line'\n test_atom refs/tags/signed-short subject:sanitize 'subject-line'\n test_atom refs/tags/signed-short contents:subject 'subject line'\n@@ -695,6 +715,15 @@ test_atom refs/tags/signed-short contents:signature \"$sig\"\n test_atom refs/tags/signed-short contents \"subject line\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-short raw' '\n+\tgit cat-file tag refs/tags/signed-short >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-short >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_atom refs/tags/signed-long subject 'subject line'\n test_atom refs/tags/signed-long subject:sanitize 'subject-line'\n test_atom refs/tags/signed-long contents:subject 'subject line'\n@@ -708,6 +737,15 @@ test_atom refs/tags/signed-long contents \"subject line\n body contents\n $sig\"\n \n+test_expect_success GPG 'basic atom: refs/tags/signed-long raw' '\n+\tgit cat-file tag refs/tags/signed-long >expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/tags/signed-long >actual &&\n+\tsanitize_pgp <expected >expected.clean &&\n+\techo >>expected.clean &&\n+\tsanitize_pgp <actual >actual.clean &&\n+\ttest_cmp expected.clean actual.clean\n+'\n+\n test_expect_success 'set up refs pointing to tree and blob' '\n \tgit update-ref refs/mytrees/first refs/heads/main^{tree} &&\n \tgit update-ref refs/myblobs/first refs/heads/main:one\n@@ -720,6 +758,16 @@ test_atom refs/mytrees/first contents:body \"\"\n test_atom refs/mytrees/first contents:signature \"\"\n test_atom refs/mytrees/first contents \"\"\n \n+test_expect_success 'basic atom: refs/mytrees/first raw' '\n+\tgit cat-file tree refs/mytrees/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/mytrees/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_atom refs/myblobs/first subject \"\"\n test_atom refs/myblobs/first contents:subject \"\"\n test_atom refs/myblobs/first body \"\"\n@@ -727,6 +775,174 @@ test_atom refs/myblobs/first contents:body \"\"\n test_atom refs/myblobs/first contents:signature \"\"\n test_atom refs/myblobs/first contents \"\"\n \n+test_expect_success 'basic atom: refs/myblobs/first raw' '\n+\tgit cat-file blob refs/myblobs/first >expected &&\n+\techo >>expected &&\n+\tgit for-each-ref --format=\"%(raw)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual &&\n+\tgit cat-file -s refs/myblobs/first >expected &&\n+\tgit for-each-ref --format=\"%(raw:size)\" refs/myblobs/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'set up refs pointing to binary blob' '\n+\tprintf \"a\\0b\\0c\" >blob1 &&\n+\tprintf \"a\\0c\\0b\" >blob2 &&\n+\tprintf \"\\0a\\0b\\0c\" >blob3 &&\n+\tprintf \"abc\" >blob4 &&\n+\tprintf \"\\0 \\0 \\0 \" >blob5 &&\n+\tprintf \"\\0 \\0a\\0 \" >blob6 &&\n+\tprintf \"  \" >blob7 &&\n+\t>blob8 &&\n+\tobj=$(git hash-object -w blob1) &&\n+\tgit update-ref refs/myblobs/blob1 \"$obj\" &&\n+\tobj=$(git hash-object -w blob2) &&\n+\tgit update-ref refs/myblobs/blob2 \"$obj\" &&\n+\tobj=$(git hash-object -w blob3) &&\n+\tgit update-ref refs/myblobs/blob3 \"$obj\" &&\n+\tobj=$(git hash-object -w blob4) &&\n+\tgit update-ref refs/myblobs/blob4 \"$obj\" &&\n+\tobj=$(git hash-object -w blob5) &&\n+\tgit update-ref refs/myblobs/blob5 \"$obj\" &&\n+\tobj=$(git hash-object -w blob6) &&\n+\tgit update-ref refs/myblobs/blob6 \"$obj\" &&\n+\tobj=$(git hash-object -w blob7) &&\n+\tgit update-ref refs/myblobs/blob7 \"$obj\" &&\n+\tobj=$(git hash-object -w blob8) &&\n+\tgit update-ref refs/myblobs/blob8 \"$obj\"\n+'\n+\n+test_expect_success 'Verify sorts with raw' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob7\n+\trefs/mytrees/first\n+\trefs/myblobs/first\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob4\n+\trefs/heads/main\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'Verify sorts with raw:size' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\trefs/myblobs/blob7\n+\trefs/heads/main\n+\trefs/myblobs/blob4\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/mytrees/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname)\" --sort=raw:size \\\n+\t\trefs/heads/main refs/myblobs/ refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:equals)' '\n+\tcat >expected <<-EOF &&\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\trefs/myblobs/blob4\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tnot equals\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:equals=abc)%(raw)%(then)%(refname)%(else)not equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'validate raw atom with %(if:notequals)' '\n+\tcat >expected <<-EOF &&\n+\trefs/heads/ambiguous\n+\trefs/heads/main\n+\trefs/heads/newtag\n+\trefs/myblobs/blob1\n+\trefs/myblobs/blob2\n+\trefs/myblobs/blob3\n+\tequals\n+\trefs/myblobs/blob5\n+\trefs/myblobs/blob6\n+\trefs/myblobs/blob7\n+\trefs/myblobs/blob8\n+\trefs/myblobs/first\n+\tEOF\n+\tgit for-each-ref --format=\"%(if:notequals=abc)%(raw)%(then)%(refname)%(else)equals%(end)\" \\\n+\t\trefs/myblobs/ refs/heads/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'empty raw refs with %(if)' '\n+\tcat >expected <<-EOF &&\n+\trefs/myblobs/blob1 not empty\n+\trefs/myblobs/blob2 not empty\n+\trefs/myblobs/blob3 not empty\n+\trefs/myblobs/blob4 not empty\n+\trefs/myblobs/blob5 not empty\n+\trefs/myblobs/blob6 not empty\n+\trefs/myblobs/blob7 empty\n+\trefs/myblobs/blob8 empty\n+\trefs/myblobs/first not empty\n+\tEOF\n+\tgit for-each-ref --format=\"%(refname) %(if)%(raw)%(then)not empty%(else)empty%(end)\" \\\n+\t\trefs/myblobs/ >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success '%(raw) with --python must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --python\n+'\n+\n+test_expect_success '%(raw) with --tcl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n+'\n+\n+test_expect_success '%(raw) with --perl must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+'\n+\n+test_expect_success '%(raw) with --shell must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --shell\n+'\n+\n+test_expect_success '%(raw) with --shell and --sort=raw must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(raw)\" --sort=raw --shell\n+'\n+\n+test_expect_success '%(raw:size) with --shell' '\n+\tgit for-each-ref --format=\"%(raw:size)\" | while read line\n+\tdo\n+\t\techo \"'\\''$line'\\''\" >>expect\n+\tdone &&\n+\tgit for-each-ref --format=\"%(raw:size)\" --shell >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'for-each-ref --format compare with cat-file --batch' '\n+\tgit rev-parse refs/mytrees/first | git cat-file --batch >expected &&\n+\tgit for-each-ref --format=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" refs/mytrees/first >actual &&\n+\ttest_cmp expected actual\n+'\n+\n test_expect_success 'set up multiple-sort tags' '\n \tfor when in 100000 200000\n \tdo\n-- \ngitgitgadget\n\n"},{"id":"428555","messageId":"f72ad9cc5e8b1f1139784df9ac7178f1561f70bb.1624797350.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 01/15] [GSOC] ref-filter: add obj-type check in grab contents","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:36Z","receivedAt":"2021-06-27T12:36:02Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nOnly tag and commit objects use `grab_sub_body_contents()` to grab\nobject contents in the current codebase.  We want to teach the\nfunction to also handle blobs and trees to get their raw data,\nwithout parsing a blob (whose contents looks like a commit or a tag)\nincorrectly as a commit or a tag.\n\nSkip the block of code that is specific to handling commits and tags\nearly when the given object is of a wrong type to help later\naddition to handle other types of objects in this function.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 24 +++++++++++++++---------\n 1 file changed, 15 insertions(+), 9 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 4db0e40ff4c..5cee6512fba 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1356,11 +1356,12 @@ static void append_lines(struct strbuf *out, const char *buf, unsigned long size\n }\n \n /* See grab_values */\n-static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n+static void grab_sub_body_contents(struct atom_value *val, int deref, struct expand_data *data)\n {\n \tint i;\n \tconst char *subpos = NULL, *bodypos = NULL, *sigpos = NULL;\n \tsize_t sublen = 0, bodylen = 0, nonsiglen = 0, siglen = 0;\n+\tvoid *buf = data->content;\n \n \tfor (i = 0; i < used_atom_cnt; i++) {\n \t\tstruct used_atom *atom = &used_atom[i];\n@@ -1371,10 +1372,13 @@ static void grab_sub_body_contents(struct atom_value *val, int deref, void *buf)\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (strcmp(name, \"body\") &&\n-\t\t    !starts_with(name, \"subject\") &&\n-\t\t    !starts_with(name, \"trailers\") &&\n-\t\t    !starts_with(name, \"contents\"))\n+\n+\t\tif ((data->type != OBJ_TAG &&\n+\t\t     data->type != OBJ_COMMIT) ||\n+\t\t    (strcmp(name, \"body\") &&\n+\t\t     !starts_with(name, \"subject\") &&\n+\t\t     !starts_with(name, \"trailers\") &&\n+\t\t     !starts_with(name, \"contents\")))\n \t\t\tcontinue;\n \t\tif (!subpos)\n \t\t\tfind_subpos(buf,\n@@ -1438,17 +1442,19 @@ static void fill_missing_values(struct atom_value *val)\n  * pointed at by the ref itself; otherwise it is the object the\n  * ref (which is a tag) refers to.\n  */\n-static void grab_values(struct atom_value *val, int deref, struct object *obj, void *buf)\n+static void grab_values(struct atom_value *val, int deref, struct object *obj, struct expand_data *data)\n {\n+\tvoid *buf = data->content;\n+\n \tswitch (obj->type) {\n \tcase OBJ_TAG:\n \t\tgrab_tag_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"tagger\", val, deref, buf);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n \t\tgrab_commit_values(val, deref, obj);\n-\t\tgrab_sub_body_contents(val, deref, buf);\n+\t\tgrab_sub_body_contents(val, deref, data);\n \t\tgrab_person(\"author\", val, deref, buf);\n \t\tgrab_person(\"committer\", val, deref, buf);\n \t\tbreak;\n@@ -1678,7 +1684,7 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi->content);\n+\t\tgrab_values(ref->value, deref, *obj, oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n-- \ngitgitgadget\n\n"},{"id":"428556","messageId":"47f868f63d92c3dee1e62617fdc9b5d95a411d07.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 03/15] [GSOC] ref-filter: --format=%(raw) re-support --perl","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:38Z","receivedAt":"2021-06-27T12:36:02Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause the perl language can handle binary data correctly,\nadd the function perl_quote_buf_with_len(), which can specify\nthe length of the data and prevent the data from being truncated\nat '\\0' to help `--format=\"%(raw)\"` re-support `--perl`.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-for-each-ref.txt |  2 +-\n quote.c                            | 17 +++++++++++++++++\n quote.h                            |  1 +\n ref-filter.c                       | 15 +++++++++++----\n t/t6300-for-each-ref.sh            | 19 +++++++++++++++++--\n 5 files changed, 47 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-for-each-ref.txt b/Documentation/git-for-each-ref.txt\nindex 3727a5ffee7..6f970088f46 100644\n--- a/Documentation/git-for-each-ref.txt\n+++ b/Documentation/git-for-each-ref.txt\n@@ -241,7 +241,7 @@ raw:size::\n \tThe raw data size of the object.\n \n Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n-`--perl` because the such language may not support arbitrary binary data in their\n+because the such language may not support arbitrary binary data in their\n string variable type.\n \n The message in a commit or a tag object is `contents`, from which\ndiff --git a/quote.c b/quote.c\nindex 8a3a5e39eb1..26719d21d1e 100644\n--- a/quote.c\n+++ b/quote.c\n@@ -471,6 +471,23 @@ void perl_quote_buf(struct strbuf *sb, const char *src)\n \tstrbuf_addch(sb, sq);\n }\n \n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len)\n+{\n+\tconst char sq = '\\'';\n+\tconst char bq = '\\\\';\n+\tconst char *c = src;\n+\tconst char *end = src + len;\n+\n+\tstrbuf_addch(sb, sq);\n+\twhile (c != end) {\n+\t\tif (*c == sq || *c == bq)\n+\t\t\tstrbuf_addch(sb, bq);\n+\t\tstrbuf_addch(sb, *c);\n+\t\tc++;\n+\t}\n+\tstrbuf_addch(sb, sq);\n+}\n+\n void python_quote_buf(struct strbuf *sb, const char *src)\n {\n \tconst char sq = '\\'';\ndiff --git a/quote.h b/quote.h\nindex 768cc6338e2..0fe69e264b0 100644\n--- a/quote.h\n+++ b/quote.h\n@@ -94,6 +94,7 @@ char *quote_path(const char *in, const char *prefix, struct strbuf *out, unsigne\n \n /* quoting as a string literal for other languages */\n void perl_quote_buf(struct strbuf *sb, const char *src);\n+void perl_quote_buf_with_len(struct strbuf *sb, const char *src, size_t len);\n void python_quote_buf(struct strbuf *sb, const char *src);\n void tcl_quote_buf(struct strbuf *sb, const char *src);\n void basic_regex_quote_buf(struct strbuf *sb, const char *src);\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 7822be90307..797b20ffa61 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -742,7 +742,10 @@ static void quote_formatting(struct strbuf *s, const char *str, size_t len, int\n \t\tsq_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PERL:\n-\t\tperl_quote_buf(s, str);\n+\t\tif (len != ATOM_VALUE_S_SIZE_INIT)\n+\t\t\tperl_quote_buf_with_len(s, str, len);\n+\t\telse\n+\t\t\tperl_quote_buf(s, str);\n \t\tbreak;\n \tcase QUOTE_PYTHON:\n \t\tpython_quote_buf(s, str);\n@@ -1006,10 +1009,14 @@ int verify_ref_format(struct ref_format *format)\n \t\tat = parse_ref_filter_atom(format, sp + 2, ep, &err);\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n-\t\tif (format->quote_style && used_atom[at].atom_type == ATOM_RAW &&\n-\t\t    used_atom[at].u.raw_data.option == RAW_BARE)\n+\n+\t\tif ((format->quote_style == QUOTE_PYTHON ||\n+\t\t     format->quote_style == QUOTE_SHELL ||\n+\t\t     format->quote_style == QUOTE_TCL) &&\n+\t\t     used_atom[at].atom_type == ATOM_RAW &&\n+\t\t     used_atom[at].u.raw_data.option == RAW_BARE)\n \t\t\tdie(_(\"--format=%.*s cannot be used with\"\n-\t\t\t      \"--python, --shell, --tcl, --perl\"), (int)(ep - sp - 2), sp + 2);\n+\t\t\t      \"--python, --shell, --tcl\"), (int)(ep - sp - 2), sp + 2);\n \t\tcp = ep + 1;\n \n \t\tif (skip_prefix(used_atom[at].name, \"color:\", &color))\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 18554f62d94..0b66e743c58 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -915,8 +915,23 @@ test_expect_success '%(raw) with --tcl must fail' '\n \ttest_must_fail git for-each-ref --format=\"%(raw)\" --tcl\n '\n \n-test_expect_success '%(raw) with --perl must fail' '\n-\ttest_must_fail git for-each-ref --format=\"%(raw)\" --perl\n+test_expect_success '%(raw) with --perl' '\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob1 --perl | perl > actual &&\n+\tcmp blob1 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob3 --perl | perl > actual &&\n+\tcmp blob3 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/blob8 --perl | perl > actual &&\n+\tcmp blob8 actual &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/myblobs/first --perl | perl > actual &&\n+\tcmp one actual &&\n+\tgit cat-file tree refs/mytrees/first > expected &&\n+\tgit for-each-ref --format=\"\\$name= %(raw);\n+print \\\"\\$name\\\"\" refs/mytrees/first --perl | perl > actual &&\n+\tcmp expected actual\n '\n \n test_expect_success '%(raw) with --shell must fail' '\n-- \ngitgitgadget\n\n"},{"id":"428557","messageId":"e592c21ea1d74f37e9a217424e734863c4683b7d.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 07/15] [GSOC] ref-filter: introduce free_ref_array_item_value() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:42Z","receivedAt":"2021-06-27T12:36:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nWhen we use ref_array_item which is not dynamically allocated and\nwant to free the space of its member \"value\" after the end of use,\nfree_array_item() does not meet our needs, because it tries to free\nref_array_item itself and its member \"symref\".\n\nIntroduce free_ref_array_item_value() for freeing ref_array_item value.\nIt will be called internally by free_array_item(), and it will help\n`cat-file --batch` free ref_array_item's value memory later.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 ++++++++---\n ref-filter.h |  2 ++\n 2 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex e4988aa8a24..731e596eaa6 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -2291,16 +2291,21 @@ static int ref_filter_handler(const char *refname, const struct object_id *oid,\n \treturn 0;\n }\n \n-/*  Free memory allocated for a ref_array_item */\n-static void free_array_item(struct ref_array_item *item)\n+void free_ref_array_item_value(struct ref_array_item *item)\n {\n-\tfree((char *)item->symref);\n \tif (item->value) {\n \t\tint i;\n \t\tfor (i = 0; i < used_atom_cnt; i++)\n \t\t\tfree((char *)item->value[i].s);\n \t\tfree(item->value);\n \t}\n+}\n+\n+/*  Free memory allocated for a ref_array_item */\n+static void free_array_item(struct ref_array_item *item)\n+{\n+\tfree((char *)item->symref);\n+\tfree_ref_array_item_value(item);\n \tfree(item);\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex c15dee8d6b9..44e6dc05ac2 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -111,6 +111,8 @@ struct ref_format {\n int filter_refs(struct ref_array *array, struct ref_filter *filter, unsigned int type);\n /*  Clear all memory allocated to ref_array */\n void ref_array_clear(struct ref_array *array);\n+/* Free ref_array_item's value */\n+void free_ref_array_item_value(struct ref_array_item *item);\n /*  Used to verify if the given format is correct and to parse out the used atoms */\n int verify_ref_format(struct ref_format *format);\n /*  Sort the given ref_array as per the ref_sorting provided */\n-- \ngitgitgadget\n\n"},{"id":"428558","messageId":"debca1564704639c70ffe0e90ef47b8b8068c1dd.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 04/15] [GSOC] ref-filter: use non-const ref_format in *_atom_parser()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:39Z","receivedAt":"2021-06-27T12:36:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse non-const ref_format in *_atom_parser(), which can help us\nmodify the members of ref_format in *_atom_parser().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/tag.c |  2 +-\n ref-filter.c  | 44 ++++++++++++++++++++++----------------------\n ref-filter.h  |  4 ++--\n 3 files changed, 25 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 82fcfc09824..452558ec957 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -146,7 +146,7 @@ static int verify_tag(const char *name, const char *ref,\n \t\t      const struct object_id *oid, void *cb_data)\n {\n \tint flags;\n-\tconst struct ref_format *format = cb_data;\n+\tstruct ref_format *format = cb_data;\n \tflags = GPG_VERIFY_VERBOSE;\n \n \tif (format->format)\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 797b20ffa61..d01a0266fb8 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -226,7 +226,7 @@ static int strbuf_addf_ret(struct strbuf *sb, int ret, const char *fmt, ...)\n \treturn ret;\n }\n \n-static int color_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int color_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *color_value, struct strbuf *err)\n {\n \tif (!color_value)\n@@ -264,7 +264,7 @@ static int refname_atom_parser_internal(struct refname_atom *atom, const char *a\n \treturn 0;\n }\n \n-static int remote_ref_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int remote_ref_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tstruct string_list params = STRING_LIST_INIT_DUP;\n@@ -311,7 +311,7 @@ static int remote_ref_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objecttype_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objecttype_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -323,7 +323,7 @@ static int objecttype_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int objectsize_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int objectsize_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -343,7 +343,7 @@ static int objectsize_atom_parser(const struct ref_format *format, struct used_a\n \treturn 0;\n }\n \n-static int deltabase_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int deltabase_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -355,7 +355,7 @@ static int deltabase_atom_parser(const struct ref_format *format, struct used_at\n \treturn 0;\n }\n \n-static int body_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int body_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (arg)\n@@ -364,7 +364,7 @@ static int body_atom_parser(const struct ref_format *format, struct used_atom *a\n \treturn 0;\n }\n \n-static int subject_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int subject_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -376,7 +376,7 @@ static int subject_atom_parser(const struct ref_format *format, struct used_atom\n \treturn 0;\n }\n \n-static int trailers_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int trailers_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tatom->u.contents.trailer_opts.no_divider = 1;\n@@ -402,7 +402,7 @@ static int trailers_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int contents_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int contents_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -430,7 +430,7 @@ static int contents_atom_parser(const struct ref_format *format, struct used_ato\n \treturn 0;\n }\n \n-static int raw_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int raw_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\tconst char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -442,7 +442,7 @@ static int raw_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int oid_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int oid_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t   const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -461,7 +461,7 @@ static int oid_atom_parser(const struct ref_format *format, struct used_atom *at\n \treturn 0;\n }\n \n-static int person_email_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int person_email_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t\t    const char *arg, struct strbuf *err)\n {\n \tif (!arg)\n@@ -475,7 +475,7 @@ static int person_email_atom_parser(const struct ref_format *format, struct used\n \treturn 0;\n }\n \n-static int refname_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int refname_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t       const char *arg, struct strbuf *err)\n {\n \treturn refname_atom_parser_internal(&atom->u.refname, arg, atom->name, err);\n@@ -492,7 +492,7 @@ static align_type parse_align_position(const char *s)\n \treturn -1;\n }\n \n-static int align_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int align_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t     const char *arg, struct strbuf *err)\n {\n \tstruct align *align = &atom->u.align;\n@@ -544,7 +544,7 @@ static int align_atom_parser(const struct ref_format *format, struct used_atom *\n \treturn 0;\n }\n \n-static int if_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t  const char *arg, struct strbuf *err)\n {\n \tif (!arg) {\n@@ -559,7 +559,7 @@ static int if_atom_parser(const struct ref_format *format, struct used_atom *ato\n \treturn 0;\n }\n \n-static int head_atom_parser(const struct ref_format *format, struct used_atom *atom,\n+static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n \tatom->u.head = resolve_refdup(\"HEAD\", RESOLVE_REF_READING, NULL, NULL);\n@@ -570,7 +570,7 @@ static struct {\n \tconst char *name;\n \tinfo_source source;\n \tcmp_type cmp_type;\n-\tint (*parser)(const struct ref_format *format, struct used_atom *atom,\n+\tint (*parser)(struct ref_format *format, struct used_atom *atom,\n \t\t      const char *arg, struct strbuf *err);\n } valid_atom[] = {\n \t[ATOM_REFNAME] = { \"refname\", SOURCE_NONE, FIELD_STR, refname_atom_parser },\n@@ -649,7 +649,7 @@ struct atom_value {\n /*\n  * Used to parse format string and sort specifiers\n  */\n-static int parse_ref_filter_atom(const struct ref_format *format,\n+static int parse_ref_filter_atom(struct ref_format *format,\n \t\t\t\t const char *atom, const char *ep,\n \t\t\t\t struct strbuf *err)\n {\n@@ -2553,9 +2553,9 @@ static void append_literal(const char *cp, const char *ep, struct ref_formatting\n }\n \n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t   const struct ref_format *format,\n-\t\t\t   struct strbuf *final_buf,\n-\t\t\t   struct strbuf *error_buf)\n+\t\t\t  struct ref_format *format,\n+\t\t\t  struct strbuf *final_buf,\n+\t\t\t  struct strbuf *error_buf)\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n@@ -2600,7 +2600,7 @@ int format_ref_array_item(struct ref_array_item *info,\n }\n \n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format)\n+\t\t      struct ref_format *format)\n {\n \tstruct ref_array_item *ref_item;\n \tstruct strbuf output = STRBUF_INIT;\ndiff --git a/ref-filter.h b/ref-filter.h\nindex baf72a71896..74fb423fc89 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -116,7 +116,7 @@ void ref_array_sort(struct ref_sorting *sort, struct ref_array *array);\n void ref_sorting_set_sort_flags_all(struct ref_sorting *sorting, unsigned int mask, int on);\n /*  Based on the given format and quote_style, fill the strbuf */\n int format_ref_array_item(struct ref_array_item *info,\n-\t\t\t  const struct ref_format *format,\n+\t\t\t  struct ref_format *format,\n \t\t\t  struct strbuf *final_buf,\n \t\t\t  struct strbuf *error_buf);\n /*  Parse a single sort specifier and add it to the list */\n@@ -137,7 +137,7 @@ void setup_ref_filter_porcelain_msg(void);\n  * name must be a fully qualified refname.\n  */\n void pretty_print_ref(const char *name, const struct object_id *oid,\n-\t\t      const struct ref_format *format);\n+\t\t      struct ref_format *format);\n \n /*\n  * Push a single ref onto the array; this can be used to construct your own\n-- \ngitgitgadget\n\n"},{"id":"428559","messageId":"b6e7757de4cf93cf2cfc267b33b72874ce4cada4.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 08/15] [GSOC] ref-filter: add cat_file_mode in struct ref_format","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:43Z","receivedAt":"2021-06-27T12:36:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAdd `cat_file_mode` member in struct `ref_format`, when\n`cat-file --batch` use ref-filter logic later, it can help\nus reject atoms in verify_ref_format() which cat-file cannot\nuse, e.g. `%(refname)`, `%(push)`, `%(upstream)`...\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 11 +++++++++--\n ref-filter.h |  1 +\n 2 files changed, 10 insertions(+), 2 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 731e596eaa6..45122959eef 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1021,8 +1021,15 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n-\t\tif (used_atom[at].atom_type == ATOM_REST)\n-\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\t\tif ((!format->cat_file_mode && used_atom[at].atom_type == ATOM_REST) ||\n+\t\t    (format->cat_file_mode && (used_atom[at].atom_type == ATOM_FLAG ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_HEAD ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_PUSH ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_REFNAME ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_SYMREF ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_UPSTREAM ||\n+\t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n+\t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 44e6dc05ac2..053980a6a42 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -78,6 +78,7 @@ struct ref_format {\n \t */\n \tconst char *format;\n \tconst char *rest;\n+\tint cat_file_mode;\n \tint quote_style;\n \tint use_rest;\n \tint use_color;\n-- \ngitgitgadget\n\n"},{"id":"428560","messageId":"cb0df2b8207d95df17126cc00ad0a080724d6be4.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 05/15] [GSOC] ref-filter: add %(rest) atom","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:40Z","receivedAt":"2021-06-27T12:36:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let \"cat-file --batch=%(rest)\" use the ref-filter\ninterface, add %(rest) atom for ref-filter. \"git for-each-ref\",\n\"git branch\", \"git tag\" and \"git verify-tag\" will reject %(rest)\nby default.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c             | 21 +++++++++++++++++++++\n ref-filter.h             |  5 ++++-\n t/t3203-branch-output.sh |  4 ++++\n t/t6300-for-each-ref.sh  |  4 ++++\n t/t7004-tag.sh           |  4 ++++\n t/t7030-verify-tag.sh    |  4 ++++\n 6 files changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex d01a0266fb8..10c78de9cfa 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -157,6 +157,7 @@ enum atom_type {\n \tATOM_IF,\n \tATOM_THEN,\n \tATOM_ELSE,\n+\tATOM_REST,\n };\n \n /*\n@@ -559,6 +560,15 @@ static int if_atom_parser(struct ref_format *format, struct used_atom *atom,\n \treturn 0;\n }\n \n+static int rest_atom_parser(struct ref_format *format, struct used_atom *atom,\n+\t\t\t    const char *arg, struct strbuf *err)\n+{\n+\tif (arg)\n+\t\treturn strbuf_addf_ret(err, -1, _(\"%%(rest) does not take arguments\"));\n+\tformat->use_rest = 1;\n+\treturn 0;\n+}\n+\n static int head_atom_parser(struct ref_format *format, struct used_atom *atom,\n \t\t\t    const char *arg, struct strbuf *unused_err)\n {\n@@ -615,6 +625,7 @@ static struct {\n \t[ATOM_IF] = { \"if\", SOURCE_NONE, FIELD_STR, if_atom_parser },\n \t[ATOM_THEN] = { \"then\", SOURCE_NONE },\n \t[ATOM_ELSE] = { \"else\", SOURCE_NONE },\n+\t[ATOM_REST] = { \"rest\", SOURCE_NONE, FIELD_STR, rest_atom_parser },\n \t/*\n \t * Please update $__git_ref_fieldlist in git-completion.bash\n \t * when you add new atoms\n@@ -1010,6 +1021,9 @@ int verify_ref_format(struct ref_format *format)\n \t\tif (at < 0)\n \t\t\tdie(\"%s\", err.buf);\n \n+\t\tif (used_atom[at].atom_type == ATOM_REST)\n+\t\t\tdie(\"this command reject atom %%(%.*s)\", (int)(ep - sp - 2), sp + 2);\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1927,6 +1941,12 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\tv->handler = else_atom_handler;\n \t\t\tv->s = xstrdup(\"\");\n \t\t\tcontinue;\n+\t\t} else if (atom_type == ATOM_REST) {\n+\t\t\tif (ref->rest)\n+\t\t\t\tv->s = xstrdup(ref->rest);\n+\t\t\telse\n+\t\t\t\tv->s = xstrdup(\"\");\n+\t\t\tcontinue;\n \t\t} else\n \t\t\tcontinue;\n \n@@ -2144,6 +2164,7 @@ static struct ref_array_item *new_ref_array_item(const char *refname,\n \n \tFLEX_ALLOC_STR(ref, refname, refname);\n \toidcpy(&ref->objectname, oid);\n+\tref->rest = NULL;\n \n \treturn ref;\n }\ndiff --git a/ref-filter.h b/ref-filter.h\nindex 74fb423fc89..c15dee8d6b9 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -38,6 +38,7 @@ struct ref_sorting {\n \n struct ref_array_item {\n \tstruct object_id objectname;\n+\tconst char *rest;\n \tint flag;\n \tunsigned int kind;\n \tconst char *symref;\n@@ -76,14 +77,16 @@ struct ref_format {\n \t * verify_ref_format() afterwards to finalize.\n \t */\n \tconst char *format;\n+\tconst char *rest;\n \tint quote_style;\n+\tint use_rest;\n \tint use_color;\n \n \t/* Internal state to ref-filter */\n \tint need_color_reset_at_eol;\n };\n \n-#define REF_FORMAT_INIT { NULL, 0, -1 }\n+#define REF_FORMAT_INIT { .use_color = -1 }\n \n /*  Macros for checking --merged and --no-merged options */\n #define _OPT_MERGED_NO_MERGED(option, filter, h) \\\ndiff --git a/t/t3203-branch-output.sh b/t/t3203-branch-output.sh\nindex 5325b9f67a0..6e94c6db7b5 100755\n--- a/t/t3203-branch-output.sh\n+++ b/t/t3203-branch-output.sh\n@@ -340,6 +340,10 @@ test_expect_success 'git branch --format option' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git branch with --format=%(rest) must fail' '\n+\ttest_must_fail git branch --format=\"%(rest)\" >actual\n+'\n+\n test_expect_success 'worktree colors correct' '\n \tcat >expect <<-EOF &&\n \t* <GREEN>(HEAD detached from fromtag)<RESET>\ndiff --git a/t/t6300-for-each-ref.sh b/t/t6300-for-each-ref.sh\nindex 0b66e743c58..6ca5c2cc19c 100755\n--- a/t/t6300-for-each-ref.sh\n+++ b/t/t6300-for-each-ref.sh\n@@ -1211,6 +1211,10 @@ test_expect_success 'basic atom: head contents:trailers' '\n \ttest_cmp expect actual.clean\n '\n \n+test_expect_success 'basic atom: rest must fail' '\n+\ttest_must_fail git for-each-ref --format=\"%(rest)\" refs/heads/main\n+'\n+\n test_expect_success 'trailer parsing not fooled by --- line' '\n \tgit commit --allow-empty -F - <<-\\EOF &&\n \tthis is the subject\ndiff --git a/t/t7004-tag.sh b/t/t7004-tag.sh\nindex 2f72c5c6883..082be85dffc 100755\n--- a/t/t7004-tag.sh\n+++ b/t/t7004-tag.sh\n@@ -1998,6 +1998,10 @@ test_expect_success '--format should list tags as per format given' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'git tag -l with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git tag -l --format=\"%(rest)\" \"v1*\"\n+'\n+\n test_expect_success \"set up color tests\" '\n \techo \"<RED>v1.0<RESET>\" >expect.color &&\n \techo \"v1.0\" >expect.bare &&\ndiff --git a/t/t7030-verify-tag.sh b/t/t7030-verify-tag.sh\nindex 3cefde9602b..10faa645157 100755\n--- a/t/t7030-verify-tag.sh\n+++ b/t/t7030-verify-tag.sh\n@@ -194,6 +194,10 @@ test_expect_success GPG 'verifying tag with --format' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success GPG 'verifying tag with --format=\"%(rest)\" must fail' '\n+\ttest_must_fail git verify-tag --format=\"%(rest)\" \"fourth-signed\"\n+'\n+\n test_expect_success GPG 'verifying a forged tag with --format should fail silently' '\n \ttest_must_fail git verify-tag --format=\"tagname : %(tag)\" $(cat forged1.tag) >actual-forged &&\n \ttest_must_be_empty actual-forged\n-- \ngitgitgadget\n\n"},{"id":"428561","messageId":"9873354930a51d6480beaf20a8e096bdae247f39.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 06/15] [GSOC] ref-filter: pass get_object() return value to their callers","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:41Z","receivedAt":"2021-06-27T12:36:09Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause in the refactor of `git cat-file --batch` later,\noid_object_info_extended() in get_object() will be used to obtain\nthe info of an object with it's oid. When the object cannot be\nobtained in the git repository, `cat-file --batch` expects to output\n\"<oid> missing\" and continue the next oid query instead of letting\nGit exit. In other error conditions, Git should exit normally. So we\ncan achieve this function by passing the return value of get_object().\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 10c78de9cfa..e4988aa8a24 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1816,6 +1816,7 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n {\n \tstruct object *obj;\n \tint i;\n+\tint ret;\n \tstruct object_info empty = OBJECT_INFO_INIT;\n \n \tCALLOC_ARRAY(ref->value, used_atom_cnt);\n@@ -1972,8 +1973,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \n \n \toi.oid = ref->objectname;\n-\tif (get_object(ref, 0, &obj, &oi, err))\n-\t\treturn -1;\n+\tret = get_object(ref, 0, &obj, &oi, err);\n+\tif (ret)\n+\t\treturn ret;\n \n \t/*\n \t * If there is no atom that wants to know about tagged\n@@ -2005,8 +2007,10 @@ static int get_ref_atom_value(struct ref_array_item *ref, int atom,\n \t\t\t      struct atom_value **v, struct strbuf *err)\n {\n \tif (!ref->value) {\n-\t\tif (populate_value(ref, err))\n-\t\t\treturn -1;\n+\t\tint ret = populate_value(ref, err);\n+\n+\t\tif (ret)\n+\t\t\treturn ret;\n \t\tfill_missing_values(ref->value);\n \t}\n \t*v = &ref->value[atom];\n@@ -2580,6 +2584,7 @@ int format_ref_array_item(struct ref_array_item *info,\n {\n \tconst char *cp, *sp, *ep;\n \tstruct ref_formatting_state state = REF_FORMATTING_STATE_INIT;\n+\tint ret;\n \n \tstate.quote_style = format->quote_style;\n \tpush_stack_element(&state.stack);\n@@ -2592,10 +2597,10 @@ int format_ref_array_item(struct ref_array_item *info,\n \t\tif (cp < sp)\n \t\t\tappend_literal(cp, sp, &state);\n \t\tpos = parse_ref_filter_atom(format, sp + 2, ep, error_buf);\n-\t\tif (pos < 0 || get_ref_atom_value(info, pos, &atomv, error_buf) ||\n+\t\tif (pos < 0 || (ret = get_ref_atom_value(info, pos, &atomv, error_buf)) ||\n \t\t    atomv->handler(atomv, &state, error_buf)) {\n \t\t\tpop_stack_element(&state.stack);\n-\t\t\treturn -1;\n+\t\t\treturn ret ? ret : -1;\n \t\t}\n \t}\n \tif (*cp) {\n-- \ngitgitgadget\n\n"},{"id":"428562","messageId":"85686187d49500b9ac4afaf526f5031ff5842aff.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 09/15] [GSOC] ref-filter: modify the error message and value in get_object","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:44Z","receivedAt":"2021-06-27T12:36:13Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nLet get_object() return 1 and print \"<oid> missing\" instead\nof returning -1 and printing \"missing object <oid> for <refname>\"\nif oid_object_info_extended() unable to find the data corresponding\nto oid. When `cat-file --batch` use ref-filter logic later it can\nhelp `format_ref_array_item()` just report that the object is missing\nwithout letting Git exit.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c                   | 4 ++--\n t/t6301-for-each-ref-errors.sh | 2 +-\n 2 files changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 45122959eef..9ca3dd5557d 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1749,8 +1749,8 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n-\t\treturn strbuf_addf_ret(err, -1, _(\"missing object %s for %s\"),\n-\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n+\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t       oid_to_hex(&oi->oid));\n \tif (oi->info.disk_sizep && oi->disk_size < 0)\n \t\tBUG(\"Object size is less than zero.\");\n \ndiff --git a/t/t6301-for-each-ref-errors.sh b/t/t6301-for-each-ref-errors.sh\nindex 40edf9dab53..3553f84a00c 100755\n--- a/t/t6301-for-each-ref-errors.sh\n+++ b/t/t6301-for-each-ref-errors.sh\n@@ -41,7 +41,7 @@ test_expect_success 'Missing objects are reported correctly' '\n \tr=refs/heads/missing &&\n \techo $MISSING >.git/$r &&\n \ttest_when_finished \"rm -f .git/$r\" &&\n-\techo \"fatal: missing object $MISSING for $r\" >missing-err &&\n+\techo \"fatal: $MISSING missing\" >missing-err &&\n \ttest_must_fail git for-each-ref 2>err &&\n \ttest_cmp missing-err err &&\n \t(\n-- \ngitgitgadget\n\n"},{"id":"428563","messageId":"9a1f07329401434b5960ef8ae002f11b1133506a.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:47Z","receivedAt":"2021-06-27T12:36:13Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nIn order to let cat-file use ref-filter logic, let's do the\nfollowing:\n\n1. Change the type of member `format` in struct `batch_options`\nto `ref_format`, we will pass it to ref-filter later.\n2. Let `batch_objects()` add atoms to format, and use\n`verify_ref_format()` to check atoms.\n3. Use `format_ref_array_item()` in `batch_object_write()` to\nget the formatted data corresponding to the object. If the\nreturn value of `format_ref_array_item()` is equals to zero,\nuse `batch_write()` to print object data; else if the return\nvalue is less than zero, use `die()` to print the error message\nand exit; else if return value is greater than zero, only print\nthe error message, but don't exit.\n4. Use free_ref_array_item_value() to free ref_array_item's\nvalue.\n\nMost of the atoms in `for-each-ref --format` are now supported,\nsuch as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n`%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n\nThe performance for `git cat-file --batch-all-objects\n--batch-check` on the Git repository itself with performance\ntesting tool `hyperfine` changes from 669.4 ms ±  31.1 ms to\n1.134 s ±  0.063 s.\n\nThe performance for `git cat-file --batch-all-objects --batch\n>/dev/null` on the Git repository itself with performance testing\ntool `time` change from \"27.37s user 0.29s system 98% cpu 28.089\ntotal\" to \"33.69s user 1.54s system 87% cpu 40.258 total\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n Documentation/git-cat-file.txt |   6 +\n builtin/cat-file.c             | 244 ++++++-------------------------\n t/t1006-cat-file.sh            | 252 +++++++++++++++++++++++++++++++++\n 3 files changed, 305 insertions(+), 197 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex 4eb0421b3fd..ef8ab952b2f 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -226,6 +226,12 @@ newline. The available atoms are:\n \tafter that first run of whitespace (i.e., the \"rest\" of the\n \tline) are output in place of the `%(rest)` atom.\n \n+Note that most of the atoms in `for-each-ref --format` are now supported,\n+such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n+`%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n+`%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n+`%(flag)`, `%(HEAD)`. See linkgit:git-for-each-ref[1].\n+\n If no format is specified, the default format is `%(objectname)\n %(objecttype) %(objectsize)`.\n \ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex cd84c39df96..5b163551fc6 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -16,6 +16,7 @@\n #include \"packfile.h\"\n #include \"object-store.h\"\n #include \"promisor-remote.h\"\n+#include \"ref-filter.h\"\n \n struct batch_options {\n \tint enabled;\n@@ -25,7 +26,7 @@ struct batch_options {\n \tint all_objects;\n \tint unordered;\n \tint cmdmode; /* may be 'w' or 'c' for --filters or --textconv */\n-\tconst char *format;\n+\tstruct ref_format format;\n };\n \n static const char *force_path;\n@@ -195,99 +196,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \n struct expand_data {\n \tstruct object_id oid;\n-\tenum object_type type;\n-\tunsigned long size;\n-\toff_t disk_size;\n \tconst char *rest;\n-\tstruct object_id delta_base_oid;\n-\n-\t/*\n-\t * If mark_query is true, we do not expand anything, but rather\n-\t * just mark the object_info with items we wish to query.\n-\t */\n-\tint mark_query;\n-\n-\t/*\n-\t * Whether to split the input on whitespace before feeding it to\n-\t * get_sha1; this is decided during the mark_query phase based on\n-\t * whether we have a %(rest) token in our format.\n-\t */\n \tint split_on_whitespace;\n-\n-\t/*\n-\t * After a mark_query run, this object_info is set up to be\n-\t * passed to oid_object_info_extended. It will point to the data\n-\t * elements above, so you can retrieve the response from there.\n-\t */\n-\tstruct object_info info;\n-\n-\t/*\n-\t * This flag will be true if the requested batch format and options\n-\t * don't require us to call oid_object_info, which can then be\n-\t * optimized out.\n-\t */\n-\tunsigned skip_object_info : 1;\n };\n \n-static int is_atom(const char *atom, const char *s, int slen)\n-{\n-\tint alen = strlen(atom);\n-\treturn alen == slen && !memcmp(atom, s, alen);\n-}\n-\n-static void expand_atom(struct strbuf *sb, const char *atom, int len,\n-\t\t\tvoid *vdata)\n-{\n-\tstruct expand_data *data = vdata;\n-\n-\tif (is_atom(\"objectname\", atom, len)) {\n-\t\tif (!data->mark_query)\n-\t\t\tstrbuf_addstr(sb, oid_to_hex(&data->oid));\n-\t} else if (is_atom(\"objecttype\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.typep = &data->type;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb, type_name(data->type));\n-\t} else if (is_atom(\"objectsize\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.sizep = &data->size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX , (uintmax_t)data->size);\n-\t} else if (is_atom(\"objectsize:disk\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.disk_sizep = &data->disk_size;\n-\t\telse\n-\t\t\tstrbuf_addf(sb, \"%\"PRIuMAX, (uintmax_t)data->disk_size);\n-\t} else if (is_atom(\"rest\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->split_on_whitespace = 1;\n-\t\telse if (data->rest)\n-\t\t\tstrbuf_addstr(sb, data->rest);\n-\t} else if (is_atom(\"deltabase\", atom, len)) {\n-\t\tif (data->mark_query)\n-\t\t\tdata->info.delta_base_oid = &data->delta_base_oid;\n-\t\telse\n-\t\t\tstrbuf_addstr(sb,\n-\t\t\t\t      oid_to_hex(&data->delta_base_oid));\n-\t} else\n-\t\tdie(\"unknown format element: %.*s\", len, atom);\n-}\n-\n-static size_t expand_format(struct strbuf *sb, const char *start, void *data)\n-{\n-\tconst char *end;\n-\n-\tif (*start != '(')\n-\t\treturn 0;\n-\tend = strchr(start + 1, ')');\n-\tif (!end)\n-\t\tdie(\"format element '%s' does not end in ')'\", start);\n-\n-\texpand_atom(sb, start + 1, end - start - 1, data);\n-\n-\treturn end - start + 1;\n-}\n-\n static void batch_write(struct batch_options *opt, const void *data, int len)\n {\n \tif (opt->buffer_output) {\n@@ -297,87 +209,34 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \t\twrite_or_die(1, data, len);\n }\n \n-static void print_object_or_die(struct batch_options *opt, struct expand_data *data)\n-{\n-\tconst struct object_id *oid = &data->oid;\n-\n-\tassert(data->info.typep);\n-\n-\tif (data->type == OBJ_BLOB) {\n-\t\tif (opt->buffer_output)\n-\t\t\tfflush(stdout);\n-\t\tif (opt->cmdmode) {\n-\t\t\tchar *contents;\n-\t\t\tunsigned long size;\n-\n-\t\t\tif (!data->rest)\n-\t\t\t\tdie(\"missing path for '%s'\", oid_to_hex(oid));\n-\n-\t\t\tif (opt->cmdmode == 'w') {\n-\t\t\t\tif (filter_object(data->rest, 0100644, oid,\n-\t\t\t\t\t\t  &contents, &size))\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else if (opt->cmdmode == 'c') {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tif (!textconv_object(the_repository,\n-\t\t\t\t\t\t     data->rest, 0100644, oid,\n-\t\t\t\t\t\t     1, &contents, &size))\n-\t\t\t\t\tcontents = read_object_file(oid,\n-\t\t\t\t\t\t\t\t    &type,\n-\t\t\t\t\t\t\t\t    &size);\n-\t\t\t\tif (!contents)\n-\t\t\t\t\tdie(\"could not convert '%s' %s\",\n-\t\t\t\t\t    oid_to_hex(oid), data->rest);\n-\t\t\t} else\n-\t\t\t\tBUG(\"invalid cmdmode: %c\", opt->cmdmode);\n-\t\t\tbatch_write(opt, contents, size);\n-\t\t\tfree(contents);\n-\t\t} else {\n-\t\t\tstream_blob(oid);\n-\t\t}\n-\t}\n-\telse {\n-\t\tenum object_type type;\n-\t\tunsigned long size;\n-\t\tvoid *contents;\n-\n-\t\tcontents = read_object_file(oid, &type, &size);\n-\t\tif (!contents)\n-\t\t\tdie(\"object %s disappeared\", oid_to_hex(oid));\n-\t\tif (type != data->type)\n-\t\t\tdie(\"object %s changed type!?\", oid_to_hex(oid));\n-\t\tif (data->info.sizep && size != data->size)\n-\t\t\tdie(\"object %s changed size!?\", oid_to_hex(oid));\n-\n-\t\tbatch_write(opt, contents, size);\n-\t\tfree(contents);\n-\t}\n-}\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tif (!data->skip_object_info &&\n-\t    oid_object_info_extended(the_repository, &data->oid, &data->info,\n-\t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE) < 0) {\n-\t\tprintf(\"%s missing\\n\",\n-\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n-\t\tfflush(stdout);\n-\t\treturn;\n-\t}\n+\tint ret;\n+\tstruct strbuf err = STRBUF_INIT;\n+\tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n-\tstrbuf_expand(scratch, opt->format, expand_format, data);\n-\tstrbuf_addch(scratch, '\\n');\n-\tbatch_write(opt, scratch->buf, scratch->len);\n \n-\tif (opt->print_contents) {\n-\t\tprint_object_or_die(opt, data);\n-\t\tbatch_write(opt, \"\\n\", 1);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tif (ret < 0)\n+\t\tdie(\"%s\\n\", err.buf);\n+\tif (ret) {\n+\t\t/* ret > 0 means when the object corresponding to oid\n+\t\t * cannot be found in format_ref_array_item(), we only print\n+\t\t * the error message.\n+\t\t */\n+\t\tprintf(\"%s\\n\", err.buf);\n+\t\tfflush(stdout);\n+\t} else {\n+\t\tstrbuf_addch(scratch, '\\n');\n+\t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n+\tfree_ref_array_item_value(&item);\n+\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n@@ -495,42 +354,34 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *batch)\n+static const char * const cat_file_usage[] = {\n+\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n+\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n+\tNULL\n+};\n+\n+static int batch_objects(struct batch_options *batch, const struct option *options)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!batch->format)\n-\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n-\n-\t/*\n-\t * Expand once with our special mark_query flag, which will prime the\n-\t * object_info to be handed to oid_object_info_extended for each\n-\t * object.\n-\t */\n \tmemset(&data, 0, sizeof(data));\n-\tdata.mark_query = 1;\n-\tstrbuf_expand(&output, batch->format, expand_format, &data);\n-\tdata.mark_query = 0;\n-\tstrbuf_release(&output);\n-\tif (batch->cmdmode)\n-\t\tdata.split_on_whitespace = 1;\n-\n-\tif (batch->all_objects) {\n-\t\tstruct object_info empty = OBJECT_INFO_INIT;\n-\t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n-\t\t\tdata.skip_object_info = 1;\n-\t}\n-\n-\t/*\n-\t * If we are printing out the object, then always fill in the type,\n-\t * since we will want to decide whether or not to stream.\n-\t */\n+\tif (batch->format.format)\n+\t\tstrbuf_addstr(&format, batch->format.format);\n+\telse\n+\t\tstrbuf_addstr(&format, \"%(objectname) %(objecttype) %(objectsize)\");\n \tif (batch->print_contents)\n-\t\tdata.info.typep = &data.type;\n+\t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n+\tbatch->format.format = format.buf;\n+\tif (verify_ref_format(&batch->format))\n+\t\tusage_with_options(cat_file_usage, options);\n+\n+\tif (batch->cmdmode || batch->format.use_rest)\n+\t\tdata.split_on_whitespace = 1;\n \n \tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n@@ -563,6 +414,7 @@ static int batch_objects(struct batch_options *batch)\n \t\t\toid_array_clear(&sa);\n \t\t}\n \n+\t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n \t\treturn 0;\n \t}\n@@ -595,18 +447,13 @@ static int batch_objects(struct batch_options *batch)\n \t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n+\tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n \n-static const char * const cat_file_usage[] = {\n-\tN_(\"git cat-file (-t [--allow-unknown-type] | -s [--allow-unknown-type] | -e | -p | <type> | --textconv | --filters) [--path=<path>] <object>\"),\n-\tN_(\"git cat-file (--batch[=<format>] | --batch-check[=<format>]) [--follow-symlinks] [--textconv | --filters]\"),\n-\tNULL\n-};\n-\n static int git_cat_file_config(const char *var, const char *value, void *cb)\n {\n \tif (userdiff_config(var, value) < 0)\n@@ -629,7 +476,7 @@ static int batch_option_callback(const struct option *opt,\n \n \tbo->enabled = 1;\n \tbo->print_contents = !strcmp(opt->long_name, \"batch\");\n-\tbo->format = arg;\n+\tbo->format.format = arg;\n \n \treturn 0;\n }\n@@ -638,7 +485,9 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n \tint opt = 0;\n \tconst char *exp_type = NULL, *obj_name = NULL;\n-\tstruct batch_options batch = {0};\n+\tstruct batch_options batch = {\n+\t\t.format = REF_FORMAT_INIT\n+\t};\n \tint unknown_type = 0;\n \n \tconst struct option options[] = {\n@@ -677,6 +526,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \tgit_config(git_cat_file_config, NULL);\n \n \tbatch.buffer_output = -1;\n+\tbatch.format.cat_file_mode = 1;\n \targc = parse_options(argc, argv, prefix, options, cat_file_usage, 0);\n \n \tif (opt) {\n@@ -720,7 +570,7 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\tbatch.buffer_output = batch.all_objects;\n \n \tif (batch.enabled)\n-\t\treturn batch_objects(&batch);\n+\t\treturn batch_objects(&batch, options);\n \n \tif (unknown_type && opt != 't' && opt != 's')\n \t\tdie(\"git cat-file --allow-unknown-type: use with -s or -t\");\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 5d2dc99b74a..69eb627774d 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -586,4 +586,256 @@ test_expect_success 'cat-file --unordered works' '\n \ttest_cmp expect actual\n '\n \n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+. \"$TEST_DIRECTORY\"/lib-terminal.sh\n+\n+test_expect_success 'cat-file --batch|--batch-check setup' '\n+\techo 1>blob1 &&\n+\tprintf \"a\\0b\\0\\c\" >blob2 &&\n+\tgit add blob1 blob2 &&\n+\tgit commit -m \"Commit Message\" &&\n+\tgit branch -M main &&\n+\tgit tag -a -m \"v0.0.0\" testtag &&\n+\tgit update-ref refs/myblobs/blob1 HEAD:blob1 &&\n+\tgit update-ref refs/myblobs/blob2 HEAD:blob2 &&\n+\tgit update-ref refs/mytrees/tree1 HEAD^{tree}\n+'\n+\n+batch_test_atom() {\n+\tif test \"$3\" = \"fail\"\n+\tthen\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2 must fail\" \"\n+\t\t\ttest_must_fail git cat-file --batch-check='$2' >bad <<-EOF\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\"\n+\telse\n+\t\ttest_expect_${4:-success} $PREREQ \"basic atom: $1 $2\" \"\n+\t\t\tgit for-each-ref --format='$2' $1 >expected &&\n+\t\t\tgit cat-file --batch-check='$2' >actual <<-EOF &&\n+\t\t\t$1\n+\t\t\tEOF\n+\t\t\tsanitize_pgp <actual >actual.clean &&\n+\t\t\tcmp expected actual.clean\n+\t\t\"\n+\tfi\n+}\n+\n+batch_test_atom refs/heads/main '%(refname)' fail\n+batch_test_atom refs/heads/main '%(refname:)' fail\n+batch_test_atom refs/heads/main '%(refname:short)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=2)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(refname:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream)' fail\n+batch_test_atom refs/heads/main '%(upstream:short)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:lstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:rstrip=-2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=2)' fail\n+batch_test_atom refs/heads/main '%(upstream:strip=-2)' fail\n+batch_test_atom refs/heads/main '%(push)' fail\n+batch_test_atom refs/heads/main '%(push:short)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:lstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=1)' fail\n+batch_test_atom refs/heads/main '%(push:rstrip=-1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=1)' fail\n+batch_test_atom refs/heads/main '%(push:strip=-1)' fail\n+batch_test_atom refs/heads/main '%(objecttype)'\n+batch_test_atom refs/heads/main '%(objectsize)'\n+batch_test_atom refs/heads/main '%(objectsize:disk)'\n+batch_test_atom refs/heads/main '%(deltabase)'\n+batch_test_atom refs/heads/main '%(objectname)'\n+batch_test_atom refs/heads/main '%(objectname:short)'\n+batch_test_atom refs/heads/main '%(objectname:short=1)'\n+batch_test_atom refs/heads/main '%(objectname:short=10)'\n+batch_test_atom refs/heads/main '%(tree)'\n+batch_test_atom refs/heads/main '%(tree:short)'\n+batch_test_atom refs/heads/main '%(tree:short=1)'\n+batch_test_atom refs/heads/main '%(tree:short=10)'\n+batch_test_atom refs/heads/main '%(parent)'\n+batch_test_atom refs/heads/main '%(parent:short)'\n+batch_test_atom refs/heads/main '%(parent:short=1)'\n+batch_test_atom refs/heads/main '%(parent:short=10)'\n+batch_test_atom refs/heads/main '%(numparent)'\n+batch_test_atom refs/heads/main '%(object)'\n+batch_test_atom refs/heads/main '%(type)'\n+batch_test_atom refs/heads/main '%(raw)'\n+batch_test_atom refs/heads/main '%(*objectname)'\n+batch_test_atom refs/heads/main '%(*objecttype)'\n+batch_test_atom refs/heads/main '%(author)'\n+batch_test_atom refs/heads/main '%(authorname)'\n+batch_test_atom refs/heads/main '%(authoremail)'\n+batch_test_atom refs/heads/main '%(authoremail:trim)'\n+batch_test_atom refs/heads/main '%(authoremail:localpart)'\n+batch_test_atom refs/heads/main '%(authordate)'\n+batch_test_atom refs/heads/main '%(committer)'\n+batch_test_atom refs/heads/main '%(committername)'\n+batch_test_atom refs/heads/main '%(committeremail)'\n+batch_test_atom refs/heads/main '%(committeremail:trim)'\n+batch_test_atom refs/heads/main '%(committeremail:localpart)'\n+batch_test_atom refs/heads/main '%(committerdate)'\n+batch_test_atom refs/heads/main '%(tag)'\n+batch_test_atom refs/heads/main '%(tagger)'\n+batch_test_atom refs/heads/main '%(taggername)'\n+batch_test_atom refs/heads/main '%(taggeremail)'\n+batch_test_atom refs/heads/main '%(taggeremail:trim)'\n+batch_test_atom refs/heads/main '%(taggeremail:localpart)'\n+batch_test_atom refs/heads/main '%(taggerdate)'\n+batch_test_atom refs/heads/main '%(creator)'\n+batch_test_atom refs/heads/main '%(creatordate)'\n+batch_test_atom refs/heads/main '%(subject)'\n+batch_test_atom refs/heads/main '%(subject:sanitize)'\n+batch_test_atom refs/heads/main '%(contents:subject)'\n+batch_test_atom refs/heads/main '%(body)'\n+batch_test_atom refs/heads/main '%(contents:body)'\n+batch_test_atom refs/heads/main '%(contents:signature)'\n+batch_test_atom refs/heads/main '%(contents)'\n+batch_test_atom refs/heads/main '%(HEAD)' fail\n+batch_test_atom refs/heads/main '%(upstream:track)' fail\n+batch_test_atom refs/heads/main '%(upstream:trackshort)' fail\n+batch_test_atom refs/heads/main '%(upstream:track,nobracket)' fail\n+batch_test_atom refs/heads/main '%(upstream:nobracket,track)' fail\n+batch_test_atom refs/heads/main '%(push:track)' fail\n+batch_test_atom refs/heads/main '%(push:trackshort)' fail\n+batch_test_atom refs/heads/main '%(worktreepath)' fail\n+batch_test_atom refs/heads/main '%(symref)' fail\n+batch_test_atom refs/heads/main '%(flag)' fail\n+\n+batch_test_atom refs/tags/testtag '%(refname)' fail\n+batch_test_atom refs/tags/testtag '%(refname:short)' fail\n+batch_test_atom refs/tags/testtag '%(upstream)' fail\n+batch_test_atom refs/tags/testtag '%(push)' fail\n+batch_test_atom refs/tags/testtag '%(objecttype)'\n+batch_test_atom refs/tags/testtag '%(objectsize)'\n+batch_test_atom refs/tags/testtag '%(objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(*objectsize:disk)'\n+batch_test_atom refs/tags/testtag '%(deltabase)'\n+batch_test_atom refs/tags/testtag '%(*deltabase)'\n+batch_test_atom refs/tags/testtag '%(objectname)'\n+batch_test_atom refs/tags/testtag '%(objectname:short)'\n+batch_test_atom refs/tags/testtag '%(tree)'\n+batch_test_atom refs/tags/testtag '%(tree:short)'\n+batch_test_atom refs/tags/testtag '%(tree:short=1)'\n+batch_test_atom refs/tags/testtag '%(tree:short=10)'\n+batch_test_atom refs/tags/testtag '%(parent)'\n+batch_test_atom refs/tags/testtag '%(parent:short)'\n+batch_test_atom refs/tags/testtag '%(parent:short=1)'\n+batch_test_atom refs/tags/testtag '%(parent:short=10)'\n+batch_test_atom refs/tags/testtag '%(numparent)'\n+batch_test_atom refs/tags/testtag '%(object)'\n+batch_test_atom refs/tags/testtag '%(type)'\n+batch_test_atom refs/tags/testtag '%(*objectname)'\n+batch_test_atom refs/tags/testtag '%(*objecttype)'\n+batch_test_atom refs/tags/testtag '%(author)'\n+batch_test_atom refs/tags/testtag '%(authorname)'\n+batch_test_atom refs/tags/testtag '%(authoremail)'\n+batch_test_atom refs/tags/testtag '%(authoremail:trim)'\n+batch_test_atom refs/tags/testtag '%(authoremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(authordate)'\n+batch_test_atom refs/tags/testtag '%(committer)'\n+batch_test_atom refs/tags/testtag '%(committername)'\n+batch_test_atom refs/tags/testtag '%(committeremail)'\n+batch_test_atom refs/tags/testtag '%(committeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(committeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(committerdate)'\n+batch_test_atom refs/tags/testtag '%(tag)'\n+batch_test_atom refs/tags/testtag '%(tagger)'\n+batch_test_atom refs/tags/testtag '%(taggername)'\n+batch_test_atom refs/tags/testtag '%(taggeremail)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:trim)'\n+batch_test_atom refs/tags/testtag '%(taggeremail:localpart)'\n+batch_test_atom refs/tags/testtag '%(taggerdate)'\n+batch_test_atom refs/tags/testtag '%(creator)'\n+batch_test_atom refs/tags/testtag '%(creatordate)'\n+batch_test_atom refs/tags/testtag '%(subject)'\n+batch_test_atom refs/tags/testtag '%(subject:sanitize)'\n+batch_test_atom refs/tags/testtag '%(contents:subject)'\n+batch_test_atom refs/tags/testtag '%(body)'\n+batch_test_atom refs/tags/testtag '%(contents:body)'\n+batch_test_atom refs/tags/testtag '%(contents:signature)'\n+batch_test_atom refs/tags/testtag '%(contents)'\n+batch_test_atom refs/tags/testtag '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(refname)' fail\n+batch_test_atom refs/myblobs/blob1 '%(upstream)' fail\n+batch_test_atom refs/myblobs/blob1 '%(push)' fail\n+batch_test_atom refs/myblobs/blob1 '%(HEAD)' fail\n+\n+batch_test_atom refs/myblobs/blob1 '%(objectname)'\n+batch_test_atom refs/myblobs/blob1 '%(objecttype)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize)'\n+batch_test_atom refs/myblobs/blob1 '%(objectsize:disk)'\n+batch_test_atom refs/myblobs/blob1 '%(deltabase)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(contents)'\n+batch_test_atom refs/myblobs/blob2 '%(contents)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(raw)'\n+batch_test_atom refs/mytrees/tree1 '%(raw)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(raw:size)'\n+batch_test_atom refs/myblobs/blob2 '%(raw:size)'\n+batch_test_atom refs/mytrees/tree1 '%(raw:size)'\n+\n+batch_test_atom refs/myblobs/blob1 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/myblobs/blob2 '%(if:equals=blob)%(objecttype)%(then)commit%(else)not commit%(end)'\n+batch_test_atom refs/mytrees/tree1 '%(if:equals=tree)%(objecttype)%(then)tree%(else)not tree%(end)'\n+\n+batch_test_atom refs/heads/main '%(align:60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:left,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:middle,60) objectname is %(objectname)%(end)|%(objectname)'\n+batch_test_atom refs/heads/main '%(align:60,right) objectname is %(objectname)%(end)|%(objectname)'\n+\n+batch_test_atom refs/heads/main 'VALID'\n+batch_test_atom refs/heads/main '%(INVALID)' fail\n+batch_test_atom refs/heads/main '%(authordate:INVALID)' fail\n+\n+test_expect_success '%(rest) works with both a branch and a tag' '\n+\tcat >expected <<-EOF &&\n+\t123 commit 123\n+\t456 tag 456\n+\tEOF\n+\tgit cat-file --batch-check=\"%(rest) %(objecttype) %(rest)\" >actual <<-EOF &&\n+\trefs/heads/main 123\n+\trefs/tags/testtag 456\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+batch_test_atom refs/heads/main '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/tags/testtag '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob1 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+batch_test_atom refs/myblobs/blob2 '%(objectname) %(objecttype) %(objectsize)\n+%(raw)'\n+\n+\n+test_expect_success 'cat-file --batch equals to --batch-check with atoms' '\n+\tgit cat-file --batch-check=\"%(objectname) %(objecttype) %(objectsize)\n+%(raw)\" >expected <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tgit cat-file --batch >actual <<-EOF &&\n+\trefs/heads/main\n+\trefs/tags/testtag\n+\tEOF\n+\tcmp expected actual\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"428565","messageId":"32e1ca5638917eca4855bee5a248dc268168465a.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 11/15] [GSOC] cat-file: change batch_objects parameter name","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:46Z","receivedAt":"2021-06-27T12:36:13Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause later cat-file reuses ref-filter logic that will add\nparameter \"const struct option *options\" to batch_objects(),\nthe two synonymous parameters of \"opt\" and \"options\" may\nconfuse readers, so change batch_options parameter of\nbatch_objects() from \"opt\" to \"batch\".\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 +++++++++++-----------\n 1 file changed, 11 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 9fd3c04ff20..cd84c39df96 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -495,7 +495,7 @@ static int batch_unordered_packed(const struct object_id *oid,\n \treturn batch_unordered_object(oid, data);\n }\n \n-static int batch_objects(struct batch_options *opt)\n+static int batch_objects(struct batch_options *batch)\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n@@ -503,8 +503,8 @@ static int batch_objects(struct batch_options *opt)\n \tint save_warning;\n \tint retval = 0;\n \n-\tif (!opt->format)\n-\t\topt->format = \"%(objectname) %(objecttype) %(objectsize)\";\n+\tif (!batch->format)\n+\t\tbatch->format = \"%(objectname) %(objecttype) %(objectsize)\";\n \n \t/*\n \t * Expand once with our special mark_query flag, which will prime the\n@@ -513,13 +513,13 @@ static int batch_objects(struct batch_options *opt)\n \t */\n \tmemset(&data, 0, sizeof(data));\n \tdata.mark_query = 1;\n-\tstrbuf_expand(&output, opt->format, expand_format, &data);\n+\tstrbuf_expand(&output, batch->format, expand_format, &data);\n \tdata.mark_query = 0;\n \tstrbuf_release(&output);\n-\tif (opt->cmdmode)\n+\tif (batch->cmdmode)\n \t\tdata.split_on_whitespace = 1;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_info empty = OBJECT_INFO_INIT;\n \t\tif (!memcmp(&data.info, &empty, sizeof(empty)))\n \t\t\tdata.skip_object_info = 1;\n@@ -529,20 +529,20 @@ static int batch_objects(struct batch_options *opt)\n \t * If we are printing out the object, then always fill in the type,\n \t * since we will want to decide whether or not to stream.\n \t */\n-\tif (opt->print_contents)\n+\tif (batch->print_contents)\n \t\tdata.info.typep = &data.type;\n \n-\tif (opt->all_objects) {\n+\tif (batch->all_objects) {\n \t\tstruct object_cb_data cb;\n \n \t\tif (has_promisor_remote())\n \t\t\twarning(\"This repository uses promisor remotes. Some objects may not be loaded.\");\n \n-\t\tcb.opt = opt;\n+\t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n \n-\t\tif (opt->unordered) {\n+\t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n \n \t\t\tcb.seen = &seen;\n@@ -592,7 +592,7 @@ static int batch_objects(struct batch_options *opt)\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, opt, &data);\n+\t\tbatch_one_object(input.buf, &output, batch, &data);\n \t}\n \n \tstrbuf_release(&input);\n-- \ngitgitgadget\n\n"},{"id":"428564","messageId":"3fb47584924522e0bb32f667bb215b7d3223015f.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 13/15] [GSOC] cat-file: reuse err buf in batch_object_write()","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:48Z","receivedAt":"2021-06-27T12:36:14Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nReuse the `err` buffer in batch_object_write(), as the\nbuffer `scratch` does. This will reduce the overhead\nof multiple allocations of memory of the err buffer.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5b163551fc6..dc604a9879d 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -212,35 +212,36 @@ static void batch_write(struct batch_options *opt, const void *data, int len)\n \n static void batch_object_write(const char *obj_name,\n \t\t\t       struct strbuf *scratch,\n+\t\t\t       struct strbuf *err,\n \t\t\t       struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n \tint ret;\n-\tstruct strbuf err = STRBUF_INIT;\n \tstruct ref_array_item item = { data->oid, data->rest };\n \n \tstrbuf_reset(scratch);\n+\tstrbuf_reset(err);\n \n-\tret = format_ref_array_item(&item, &opt->format, scratch, &err);\n+\tret = format_ref_array_item(&item, &opt->format, scratch, err);\n \tif (ret < 0)\n-\t\tdie(\"%s\\n\", err.buf);\n+\t\tdie(\"%s\\n\", err->buf);\n \tif (ret) {\n \t\t/* ret > 0 means when the object corresponding to oid\n \t\t * cannot be found in format_ref_array_item(), we only print\n \t\t * the error message.\n \t\t */\n-\t\tprintf(\"%s\\n\", err.buf);\n+\t\tprintf(\"%s\\n\", err->buf);\n \t\tfflush(stdout);\n \t} else {\n \t\tstrbuf_addch(scratch, '\\n');\n \t\tbatch_write(opt, scratch->buf, scratch->len);\n \t}\n \tfree_ref_array_item_value(&item);\n-\tstrbuf_release(&err);\n }\n \n static void batch_one_object(const char *obj_name,\n \t\t\t     struct strbuf *scratch,\n+\t\t\t     struct strbuf *err,\n \t\t\t     struct batch_options *opt,\n \t\t\t     struct expand_data *data)\n {\n@@ -294,7 +295,7 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n-\tbatch_object_write(obj_name, scratch, opt, data);\n+\tbatch_object_write(obj_name, scratch, err, opt, data);\n }\n \n struct object_cb_data {\n@@ -302,13 +303,14 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n \tstruct oidset *seen;\n \tstruct strbuf *scratch;\n+\tstruct strbuf *err;\n };\n \n static int batch_object_cb(const struct object_id *oid, void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \toidcpy(&data->expand->oid, oid);\n-\tbatch_object_write(NULL, data->scratch, data->opt, data->expand);\n+\tbatch_object_write(NULL, data->scratch, data->err, data->opt, data->expand);\n \treturn 0;\n }\n \n@@ -364,6 +366,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n {\n \tstruct strbuf input = STRBUF_INIT;\n \tstruct strbuf output = STRBUF_INIT;\n+\tstruct strbuf err = STRBUF_INIT;\n \tstruct strbuf format = STRBUF_INIT;\n \tstruct expand_data data;\n \tint save_warning;\n@@ -392,6 +395,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\tcb.opt = batch;\n \t\tcb.expand = &data;\n \t\tcb.scratch = &output;\n+\t\tcb.err = &err;\n \n \t\tif (batch->unordered) {\n \t\t\tstruct oidset seen = OIDSET_INIT;\n@@ -416,6 +420,7 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \n \t\tstrbuf_release(&format);\n \t\tstrbuf_release(&output);\n+\t\tstrbuf_release(&err);\n \t\treturn 0;\n \t}\n \n@@ -444,12 +449,13 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \t\t\tdata.rest = p;\n \t\t}\n \n-\t\tbatch_one_object(input.buf, &output, batch, &data);\n+\t\tbatch_one_object(input.buf, &output, &err, batch, &data);\n \t}\n \n \tstrbuf_release(&format);\n \tstrbuf_release(&input);\n \tstrbuf_release(&output);\n+\tstrbuf_release(&err);\n \twarn_on_object_refname_ambiguity = save_warning;\n \treturn retval;\n }\n-- \ngitgitgadget\n\n"},{"id":"428566","messageId":"6037295ee58b9bfb5020f26776728fd12b9fbece.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 10/15] [GSOC] cat-file: add has_object_file() check","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:45Z","receivedAt":"2021-06-27T12:36:15Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nUse `has_object_file()` in `batch_one_object()` to check\nwhether the input object exists. This can help us reject\nthe missing oid when we let `cat-file --batch` use ref-filter\nlogic later.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c | 7 +++++++\n 1 file changed, 7 insertions(+)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 5ebf13359e8..9fd3c04ff20 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -428,6 +428,13 @@ static void batch_one_object(const char *obj_name,\n \t\treturn;\n \t}\n \n+\tif (!has_object_file(&data->oid)) {\n+\t\tprintf(\"%s missing\\n\",\n+\t\t       obj_name ? obj_name : oid_to_hex(&data->oid));\n+\t\tfflush(stdout);\n+\t\treturn;\n+\t}\n+\n \tbatch_object_write(obj_name, scratch, opt, data);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"428567","messageId":"891d62fd93f795219518d85aaf0cd50c84d7c652.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 15/15] [GSOC] ref-filter: remove grab_oid() function","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:50Z","receivedAt":"2021-06-27T12:36:16Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nBecause \"atom_type == ATOM_OBJECTNAME\" implies the condition\nof `starts_with(name, \"objectname\")`, \"atom_type == ATOM_TREE\"\nimplies the condition of `starts_with(name, \"tree\")`, so the\ncheck for `starts_with(name, field)` in grab_oid() is redundant.\n\nSo Remove the grab_oid() from ref-filter, to reduce repeated check.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n ref-filter.c | 26 +++++++++-----------------\n 1 file changed, 9 insertions(+), 17 deletions(-)\n\ndiff --git a/ref-filter.c b/ref-filter.c\nindex 427552b8108..55bff4ed9e7 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1071,16 +1071,6 @@ static const char *do_grab_oid(const char *field, const struct object_id *oid,\n \t}\n }\n \n-static int grab_oid(const char *name, const char *field, const struct object_id *oid,\n-\t\t    struct atom_value *v, struct used_atom *atom)\n-{\n-\tif (starts_with(name, field)) {\n-\t\tv->s = xstrdup(do_grab_oid(field, oid, atom));\n-\t\treturn 1;\n-\t}\n-\treturn 0;\n-}\n-\n /* See grab_values */\n static void grab_common_values(struct atom_value *val, int deref, struct expand_data *oi)\n {\n@@ -1106,8 +1096,9 @@ static void grab_common_values(struct atom_value *val, int deref, struct expand_\n \t\t\t}\n \t\t} else if (atom_type == ATOM_DELTABASE)\n \t\t\tv->s = xstrdup(oid_to_hex(&oi->delta_base_oid));\n-\t\telse if (atom_type == ATOM_OBJECTNAME && deref)\n-\t\t\tgrab_oid(name, \"objectname\", &oi->oid, v, &used_atom[i]);\n+\t\telse if (atom_type == ATOM_OBJECTNAME && deref) {\n+\t\t\tv->s = xstrdup(do_grab_oid(\"objectname\", &oi->oid, &used_atom[i]));\n+\t\t}\n \t}\n }\n \n@@ -1148,9 +1139,10 @@ static void grab_commit_values(struct atom_value *val, int deref, struct object\n \t\t\tcontinue;\n \t\tif (deref)\n \t\t\tname++;\n-\t\tif (atom_type == ATOM_TREE &&\n-\t\t    grab_oid(name, \"tree\", get_commit_tree_oid(commit), v, &used_atom[i]))\n+\t\tif (atom_type == ATOM_TREE) {\n+\t\t\tv->s = xstrdup(do_grab_oid(\"tree\", get_commit_tree_oid(commit), &used_atom[i]));\n \t\t\tcontinue;\n+\t\t}\n \t\tif (atom_type == ATOM_NUMPARENT) {\n \t\t\tv->value = commit_list_count(commit->parents);\n \t\t\tv->s = xstrfmt(\"%lu\", (unsigned long)v->value);\n@@ -1971,9 +1963,9 @@ static int populate_value(struct ref_array_item *ref, struct strbuf *err)\n \t\t\t\tv->s = xstrdup(buf + 1);\n \t\t\t}\n \t\t\tcontinue;\n-\t\t} else if (!deref && atom_type == ATOM_OBJECTNAME &&\n-\t\t\t   grab_oid(name, \"objectname\", &ref->objectname, v, atom)) {\n-\t\t\t\tcontinue;\n+\t\t} else if (!deref && atom_type == ATOM_OBJECTNAME) {\n+\t\t\t   v->s = xstrdup(do_grab_oid(\"objectname\", &ref->objectname, atom));\n+\t\t\t   continue;\n \t\t} else if (atom_type == ATOM_HEAD) {\n \t\t\tif (atom->u.head && !strcmp(ref->refname, atom->u.head))\n \t\t\t\tv->s = xstrdup(\"*\");\n-- \ngitgitgadget\n"},{"id":"428568","messageId":"e0b1a05e71127878c48376295c27a46a6397c216.1624797351.git.gitgitgadget@gmail.com","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"[PATCH v6 14/15] [GSOC] cat-file: re-implement --textconv, --filters options","fromName":"ZheNing Hu via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2021-06-27T12:35:49Z","receivedAt":"2021-06-27T12:36:18Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"From: ZheNing Hu <adlternative@gmail.com>\n\nAfter cat-file reuses the ref-filter logic, we re-implement the\nfunctions of --textconv and --filters options.\n\nAdd members `use_textconv` and `use_filters` in struct `ref_format`,\nand use global variables `use_filters` and `use_textconv` in\n`ref-filter.c`, so that we can filter the content of the object\nin get_object(). Use `actual_oi` to record the real expand_data:\nit may point to the original `oi` or the `act_oi` processed by\n`textconv_object()` or `convert_to_working_tree()`. `grab_values()`\nwill grab the contents of `actual_oi` and `grab_common_values()`\nto grab the contents of origin `oi`, this ensures that `%(objectsize)`\nstill uses the size of the unfiltered data.\n\nIn `get_object()`, we made an optimization: Firstly, get the size and\ntype of the object instead of directly getting the object data.\nIf using --textconv, after successfully obtaining the filtered object\ndata, an extra oid_object_info_extended() will be skipped, which can\nreduce the cost of object data copy; If using --filter, the data of\nthe object first will be getted first, and then convert_to_working_tree()\nwill be used to get the filtered object data.\n\nMentored-by: Christian Couder <christian.couder@gmail.com>\nMentored-by: Hariom Verma <hariom18599@gmail.com>\nSigned-off-by: ZheNing Hu <adlternative@gmail.com>\n---\n builtin/cat-file.c |  6 +++++\n ref-filter.c       | 59 ++++++++++++++++++++++++++++++++++++++++++++--\n ref-filter.h       |  2 ++\n 3 files changed, 65 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex dc604a9879d..710085b7d72 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -380,6 +380,12 @@ static int batch_objects(struct batch_options *batch, const struct option *optio\n \tif (batch->print_contents)\n \t\tstrbuf_addstr(&format, \"\\n%(raw)\");\n \tbatch->format.format = format.buf;\n+\n+\tif (batch->cmdmode == 'c')\n+\t\tbatch->format.use_textconv = 1;\n+\telse if (batch->cmdmode == 'w')\n+\t\tbatch->format.use_filters = 1;\n+\n \tif (verify_ref_format(&batch->format))\n \t\tusage_with_options(cat_file_usage, options);\n \ndiff --git a/ref-filter.c b/ref-filter.c\nindex 9ca3dd5557d..427552b8108 100644\n--- a/ref-filter.c\n+++ b/ref-filter.c\n@@ -1,3 +1,4 @@\n+#define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"parse-options.h\"\n@@ -84,6 +85,9 @@ static struct expand_data {\n \tstruct object_info info;\n } oi, oi_deref;\n \n+int use_filters;\n+int use_textconv;\n+\n struct ref_to_worktree_entry {\n \tstruct hashmap_entry ent;\n \tstruct worktree *wt; /* key is wt->head_ref */\n@@ -1031,6 +1035,9 @@ int verify_ref_format(struct ref_format *format)\n \t\t\t\t\t       used_atom[at].atom_type == ATOM_WORKTREEPATH)))\n \t\t\tdie(_(\"this command reject atom %%(%.*s)\"), (int)(ep - sp - 2), sp + 2);\n \n+\t\tuse_filters = format->use_filters;\n+\t\tuse_textconv = format->use_textconv;\n+\n \t\tif ((format->quote_style == QUOTE_PYTHON ||\n \t\t     format->quote_style == QUOTE_SHELL ||\n \t\t     format->quote_style == QUOTE_TCL) &&\n@@ -1742,10 +1749,38 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n {\n \t/* parse_object_buffer() will set eaten to 0 if free() will be needed */\n \tint eaten = 1;\n+\tstruct expand_data *actual_oi = oi;\n+\tstruct expand_data act_oi = {0};\n+\n \tif (oi->info.contentp) {\n \t\t/* We need to know that to use parse_object_buffer properly */\n+\t\tvoid **temp_contentp = oi->info.contentp;\n+\t\toi->info.contentp = NULL;\n \t\toi->info.sizep = &oi->size;\n \t\toi->info.typep = &oi->type;\n+\n+\t\t/* get the type and size */\n+\t\tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n+\t\t\t\t\tOBJECT_INFO_LOOKUP_REPLACE))\n+\t\t\treturn strbuf_addf_ret(err, 1, _(\"%s missing\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\n+\t\toi->info.sizep = NULL;\n+\t\toi->info.typep = NULL;\n+\t\toi->info.contentp = temp_contentp;\n+\n+\t\tif (use_textconv && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&act_oi.oid));\n+\t\tif (use_textconv && oi->type == OBJ_BLOB) {\n+\t\t\tact_oi = *oi;\n+\t\t\tif (textconv_object(the_repository,\n+\t\t\t\t\t    ref->rest, 0100644, &act_oi.oid,\n+\t\t\t\t\t    1, (char **)(&act_oi.content), &act_oi.size)) {\n+\t\t\t\tactual_oi = &act_oi;\n+\t\t\t\tgoto success;\n+\t\t\t}\n+\t\t}\n \t}\n \tif (oid_object_info_extended(the_repository, &oi->oid, &oi->info,\n \t\t\t\t     OBJECT_INFO_LOOKUP_REPLACE))\n@@ -1755,19 +1790,39 @@ static int get_object(struct ref_array_item *ref, int deref, struct object **obj\n \t\tBUG(\"Object size is less than zero.\");\n \n \tif (oi->info.contentp) {\n-\t\t*obj = parse_object_buffer(the_repository, &oi->oid, oi->type, oi->size, oi->content, &eaten);\n+\t\tif (use_filters && !ref->rest)\n+\t\t\treturn strbuf_addf_ret(err, -1, _(\"missing path for '%s'\"),\n+\t\t\t\t\t       oid_to_hex(&oi->oid));\n+\t\tif (use_filters && oi->type == OBJ_BLOB) {\n+\t\t\tstruct strbuf strbuf = STRBUF_INIT;\n+\t\t\tstruct checkout_metadata meta;\n+\t\t\tact_oi = *oi;\n+\n+\t\t\tinit_checkout_metadata(&meta, NULL, NULL, &act_oi.oid);\n+\t\t\tif (!convert_to_working_tree(&the_index, ref->rest, act_oi.content, act_oi.size, &strbuf, &meta))\n+\t\t\t\tdie(\"could not convert '%s' %s\",\n+\t\t\t\t\toid_to_hex(&oi->oid), ref->rest);\n+\t\t\tact_oi.size = strbuf.len;\n+\t\t\tact_oi.content = strbuf_detach(&strbuf, NULL);\n+\t\t\tactual_oi = &act_oi;\n+\t\t}\n+\n+success:\n+\t\t*obj = parse_object_buffer(the_repository, &actual_oi->oid, actual_oi->type, actual_oi->size, actual_oi->content, &eaten);\n \t\tif (!*obj) {\n \t\t\tif (!eaten)\n \t\t\t\tfree(oi->content);\n \t\t\treturn strbuf_addf_ret(err, -1, _(\"parse_object_buffer failed on %s for %s\"),\n \t\t\t\t\t       oid_to_hex(&oi->oid), ref->refname);\n \t\t}\n-\t\tgrab_values(ref->value, deref, *obj, oi);\n+\t\tgrab_values(ref->value, deref, *obj, actual_oi);\n \t}\n \n \tgrab_common_values(ref->value, deref, oi);\n \tif (!eaten)\n \t\tfree(oi->content);\n+\tif (actual_oi != oi)\n+\t\tfree(actual_oi->content);\n \treturn 0;\n }\n \ndiff --git a/ref-filter.h b/ref-filter.h\nindex 053980a6a42..497e3e93632 100644\n--- a/ref-filter.h\n+++ b/ref-filter.h\n@@ -80,6 +80,8 @@ struct ref_format {\n \tconst char *rest;\n \tint cat_file_mode;\n \tint quote_style;\n+\tint use_textconv;\n+\tint use_filters;\n \tint use_rest;\n \tint use_color;\n \n-- \ngitgitgadget\n\n"},{"id":"428593","messageId":"a1c8f663-ebe5-1145-960d-aad9dc07d86c@gmail.com","threadId":"55909","inReplyTo":"d9bc50c4ae699d0516581ac67a2d7b602d2e61a4.1624797351.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 02/15] [GSOC] ref-filter: add %(raw) atom","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-06-28T06:49:10Z","receivedAt":"2021-06-28T06:49:33Z","isPatch":true,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 27/06/21 19.35, ZheNing Hu via GitGitGadget wrote:\n> +Note that `--format=%(raw)` can not be used with `--python`, `--shell`, `--tcl`,\n> +`--perl` because the such language may not support arbitrary binary data in their\n> +string variable type.\n> +\n\ns/the such language/such languages/\n\nThanks.\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"428595","messageId":"CA+CkUQ-gHT=g3KwcBXkp8eBX7wfY_HkT1=hhRiYLap-Av2w5MA@mail.gmail.com","threadId":"55909","inReplyTo":"9a1f07329401434b5960ef8ae002f11b1133506a.1624797351.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"Hariom verma","fromEmail":"hariom18599@gmail.com","sentAt":"2021-06-28T07:46:53Z","receivedAt":"2021-06-28T07:47:07Z","isPatch":true,"sender":{"key":"hariom18599@gmail.com","avatar":"https://avatars.githubusercontent.com/u/37576387?v=4"},"body":"Hi,\n\nOn Sun, Jun 27, 2021 at 6:06 PM ZheNing Hu via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: ZheNing Hu <adlternative@gmail.com>\n>\n> In order to let cat-file use ref-filter logic, let's do the\n> following:\n>\n> 1. Change the type of member `format` in struct `batch_options`\n> to `ref_format`, we will pass it to ref-filter later.\n> 2. Let `batch_objects()` add atoms to format, and use\n> `verify_ref_format()` to check atoms.\n> 3. Use `format_ref_array_item()` in `batch_object_write()` to\n> get the formatted data corresponding to the object. If the\n> return value of `format_ref_array_item()` is equals to zero,\n> use `batch_write()` to print object data; else if the return\n> value is less than zero, use `die()` to print the error message\n> and exit; else if return value is greater than zero, only print\n> the error message, but don't exit.\n> 4. Use free_ref_array_item_value() to free ref_array_item's\n> value.\n>\n> Most of the atoms in `for-each-ref --format` are now supported,\n> such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n> `%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n> `%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n> `%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n\nIt's not clear why some atoms are rejected?\n\nAre we going to support them in later commits? (or sometime in the future)\nOR\nWe are never going to support them. Because they make no sense to\ncat-file? (or whatever the reason)\n\nWhatever is the reason, I think it's a good idea to include it in the\ncommit message.\n\n-- \nHariom\n"},{"id":"428608","messageId":"CAOLTT8RLX2Zn7C3hd5JhiAYK9TptXbraf=jobL7Cj+p=uLitvA@mail.gmail.com","threadId":"55909","inReplyTo":"CA+CkUQ-gHT=g3KwcBXkp8eBX7wfY_HkT1=hhRiYLap-Av2w5MA@mail.gmail.com","subject":"Re: [PATCH v6 12/15] [GSOC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-06-28T13:51:54Z","receivedAt":"2021-06-28T13:52:07Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Hi,\n\nHariom verma <hariom18599@gmail.com> 于2021年6月28日周一 下午3:47写道：\n>\n> Hi,\n>\n> >\n> > Most of the atoms in `for-each-ref --format` are now supported,\n> > such as `%(tree)`, `%(parent)`, `%(author)`, `%(tagger)`, `%(if)`,\n> > `%(then)`, `%(else)`, `%(end)`. But these atoms will be rejected:\n> > `%(refname)`, `%(symref)`, `%(upstream)`, `%(push)`, `%(worktreepath)`,\n> > `%(flag)`, `%(HEAD)`, because our objects don't have a refname.\n>\n> It's not clear why some atoms are rejected?\n>\n> Are we going to support them in later commits? (or sometime in the future)\n> OR\n> We are never going to support them. Because they make no sense to\n> cat-file? (or whatever the reason)\n>\n\nBecause in \"git for-each-ref\"'s \"family\", ref_array_item is generated\nby filter_refs(),\nwhich uses ref_filter_handler() to fill ref_array_item with ref's\ndata. In \"git cat-file\",\nwe care about the object, not the ref. Therefore, ref_array_item is\nonly filled with\n{oid, rest} in batch_object_write() in cat-file.c. We cannot represent\nsome specific\nref-related data in \"git cat-file\", so we cannot have some atoms in ref_filter.\nYes, we probably won't support them in the future.\n\nFrom an object-oriented point of view, the atom supported by\n\"cat-file\" should be a\nparent class, \"for-each-ref\"',\"branch\",\"tag\"... they have more\nspecific object details (ref),\ntheir supported atom should be a derived class, they can support more atoms.\n\n> Whatever is the reason, I think it's a good idea to include it in the\n> commit message.\n>\n\nYeah.  The sentence \"because our objects don't have a refname.\" may not\ncorrectly express the reason for rejecting these atoms.  I will add\nmore descriptions.\n\n> --\n> Hariom\n\nThanks.\n--\nZheNing Hu\n"},{"id":"428856","messageId":"xmqqbl7nt3fh.fsf@gitster.g","threadId":"55909","inReplyTo":"pull.980.v6.git.1624797350.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-06-30T22:04:18Z","receivedAt":"2021-06-30T22:04:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> This patch series make cat-file reuse ref-filter logic.\n\nUnfortunately this seems to interact with your own\nzh/cat-file-batch-fix rather badly.\n\n"},{"id":"428893","messageId":"CAOLTT8S=qeg8QBhMU0fCc600n-qzJBT9P9LvhSVZaXpvsr0__Q@mail.gmail.com","threadId":"55909","inReplyTo":"xmqqbl7nt3fh.fsf@gitster.g","subject":"Re: [PATCH v6 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-07-01T12:39:18Z","receivedAt":"2021-07-01T12:39:32Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2021年7月1日周四 上午6:04写道：\n>\n> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > This patch series make cat-file reuse ref-filter logic.\n>\n> Unfortunately this seems to interact with your own\n> zh/cat-file-batch-fix rather badly.\n>\n\nWell, it's because I didn't base this patch on it.\nThat should be easy to achieve.\n\nBy the way,  I think patches before \"[GSOC] ref-filter: add %(rest) atom\"\nshould belong to \"zh/ref-filter-raw-data\", and patches after that should belong\nto \"zh/cat-file-batch-refactor\".\n\nThanks.\n--\nZheNing Hu\n"},{"id":"428899","messageId":"xmqqtulerudh.fsf@gitster.g","threadId":"55909","inReplyTo":"CAOLTT8S=qeg8QBhMU0fCc600n-qzJBT9P9LvhSVZaXpvsr0__Q@mail.gmail.com","subject":"Re: [PATCH v6 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-07-01T14:17:30Z","receivedAt":"2021-07-01T14:17:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ZheNing Hu <adlternative@gmail.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> 于2021年7月1日周四 上午6:04写道：\n>>\n>> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>> > This patch series make cat-file reuse ref-filter logic.\n>>\n>> Unfortunately this seems to interact with your own\n>> zh/cat-file-batch-fix rather badly.\n>>\n>\n> Well, it's because I didn't base this patch on it.\n> That should be easy to achieve.\n\nIt is preferrable for contributors try merging their individual\ntopics with the rest of 'seen' to see if there are potential\nconflicts (either textual or semantic) before sending their topics\nout.  Not all topics need to build on top of other topics (in fact,\nthe fewer inter-dependencies they have, the better), but in this\ncase, I think it makes sense to build one on top of the other.\n\n"},{"id":"429529","messageId":"CAOLTT8TipBZ+rQmKN_hMpXfPJFKWbfeYash6G50H23e9ez6J2A@mail.gmail.com","threadId":"55909","inReplyTo":"xmqqtulerudh.fsf@gitster.g","subject":"Re: [PATCH v6 00/15] [GSOC][RFC] cat-file: reuse ref-filter logic","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2021-07-09T10:04:11Z","receivedAt":"2021-07-09T10:04:07Z","isPatch":true,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Hi, Junio,\n\nJunio C Hamano <gitster@pobox.com> 于2021年7月1日周四 下午10:17写道：\n>\n> ZheNing Hu <adlternative@gmail.com> writes:\n>\n> > Junio C Hamano <gitster@pobox.com> 于2021年7月1日周四 上午6:04写道：\n> >>\n> >> \"ZheNing Hu via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> >>\n> >> > This patch series make cat-file reuse ref-filter logic.\n> >>\n> >> Unfortunately this seems to interact with your own\n> >> zh/cat-file-batch-fix rather badly.\n> >>\n> >\n> > Well, it's because I didn't base this patch on it.\n> > That should be easy to achieve.\n>\n> It is preferrable for contributors try merging their individual\n> topics with the rest of 'seen' to see if there are potential\n> conflicts (either textual or semantic) before sending their topics\n> out.  Not all topics need to build on top of other topics (in fact,\n> the fewer inter-dependencies they have, the better), but in this\n> case, I think it makes sense to build one on top of the other.\n>\n\nI have a \"rebase\" trouble:\n\nMy new feature branch \"cat-file-batch-refactor-rebase-version\" should base on\nzh/cat-file-batch-fix and  zh/ref-filter-atom-type, so last time I choice\n(bb9a3a8f77 Merge branch 'zh/cat-file-batch-fix' into jch) as the patch base.\n\nBut github only allow me base the patch on a branch, so I choice\n\"gitgitgadget:seen\"\nas my github PR base. It causes that some merge commit include in it. [1]\n\nSo In order to prevent these \"merge\" commits from being sent, the GGG mechanism\nis modified to reject their merge commits.\n\nNow I can't choice a good branch as my patch base... Have any ideas?\n\nThanks.\n\n[1]: https://github.com/gitgitgadget/git/pull/989\n\n--\nZheNing Hu\n"}]}