{"thread":{"id":"51372","subject":"[PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","startedAt":"2019-06-24T13:02:48Z","lastAt":"2019-07-08T17:58:57Z","messageCount":43,"participants":["Nguyễn Thái Ngọc Duy","Johannes Schindelin","Jeff Hostetler","Junio C Hamano","Thomas Gummerer","Duy Nguyen","Derrick Stolee","Ramsay Jones","SZEDER Gábor"],"isPatch":true,"patchVersion":2,"patchTotal":10},"messages":[{"id":"377874","messageId":"20190624130226.17293-1-pclouds@gmail.com","threadId":"51372","inReplyTo":null,"subject":"[PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:16Z","receivedAt":"2019-06-24T13:02:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"v2 main changes:\n\n- --json is renamed --debug-json\n- json field names now use '_' instead of '.' to be friendlier to some\n  languages. I stick to underscore_name instead of camelCase because\n  the former is closer to what we use\n- extension location is printed, in case you need to decode the\n  extension by yourself (previously only the size is printed)\n- all extensions are printed in the same order they appear in the file\n  (previously eoie and ieot are printed first because that's how we\n  parse)\n- resolve undo extension is reorganized a bit to be easier to read\n- tests added. Example json files are in t/t3011\n\nInterdiff\n\ndiff --git a/Documentation/git-ls-files.txt b/Documentation/git-ls-files.txt\nindex 54011c8f65..fec5cb7170 100644\n--- a/Documentation/git-ls-files.txt\n+++ b/Documentation/git-ls-files.txt\n@@ -60,11 +60,6 @@ OPTIONS\n --stage::\n \tShow staged contents' mode bits, object name and stage number in the output.\n \n---json::\n-\tDump the entire index content in JSON format. This is for\n-\tdebugging purposes and the JSON structure may change from time\n-\tto time.\n-\n --directory::\n \tIf a whole directory is classified as \"other\", show just its\n \tname (with a trailing slash) and not its whole contents.\n@@ -167,6 +162,11 @@ a space) at the start of each line:\n \tpossible for manual inspection; the exact format may change at\n \tany time.\n \n+--debug-json::\n+\tDump the entire index content in JSON format. This is for\n+\tdebugging purposes. The JSON structure is subject to change.\n+\tNote that the strings are not always valid UTF-8.\n+\n --eol::\n \tShow <eolinfo> and <eolattr> of files.\n \t<eolinfo> is the file content identification used by Git when\ndiff --git a/builtin/ls-files.c b/builtin/ls-files.c\nindex d00f6d3074..b60cd9ab28 100644\n--- a/builtin/ls-files.c\n+++ b/builtin/ls-files.c\n@@ -545,8 +545,6 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t\tN_(\"show staged contents' object name in the output\")),\n \t\tOPT_BOOL('k', \"killed\", &show_killed,\n \t\t\tN_(\"show files on the filesystem that need to be removed\")),\n-\t\tOPT_BOOL(0, \"json\", &show_json,\n-\t\t\tN_(\"dump index content in json format\")),\n \t\tOPT_BIT(0, \"directory\", &dir.flags,\n \t\t\tN_(\"show 'other' directories' names only\"),\n \t\t\tDIR_SHOW_OTHER_DIRECTORIES),\n@@ -581,6 +579,8 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t\tN_(\"pretend that paths removed since <tree-ish> are still present\")),\n \t\tOPT__ABBREV(&abbrev),\n \t\tOPT_BOOL(0, \"debug\", &debug_mode, N_(\"show debugging data\")),\n+\t\tOPT_BOOL(0, \"debug-json\", &show_json,\n+\t\t\tN_(\"dump index content in JSON format\")),\n \t\tOPT_END()\n \t};\n \n@@ -636,7 +636,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t    \"--error-unmatch\");\n \n \tparse_pathspec(&pathspec, 0,\n-\t\t       PATHSPEC_PREFER_CWD,\n+\t\t       show_json ? PATHSPEC_PREFER_FULL : PATHSPEC_PREFER_CWD,\n \t\t       prefix, argv);\n \n \t/*\n@@ -668,8 +668,14 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\tshow_cached = 1;\n \tif (show_json && (show_stage || show_deleted || show_others ||\n \t\t\t  show_unmerged || show_killed || show_modified ||\n-\t\t\t  show_cached || show_resolve_undo || with_tree))\n-\t\tdie(_(\"--show-json cannot be used with other --show- options, or --with-tree\"));\n+\t\t\t  show_cached || pathspec.nr))\n+\t\tdie(_(\"--debug-json cannot be used with other file selection options\"));\n+\tif (show_json && show_resolve_undo)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--resolve-undo\");\n+\tif (show_json && with_tree)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--with-tree\");\n+\tif (show_json && debug_mode)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--debug\");\n \n \tif (with_tree) {\n \t\t/*\ndiff --git a/cache-tree.c b/cache-tree.c\nindex fc44016fe8..b6a233307e 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -599,18 +599,13 @@ struct cache_tree *cache_tree_read(const char *buffer, unsigned long size,\n \tstruct cache_tree *ret;\n \n \tif (jw) {\n-\t\tjw_object_inline_begin_object(jw, \"cache-tree\");\n-\t\tjw_object_intmax(jw, \"ext-size\", size);\n \t\tjw_object_inline_begin_object(jw, \"root\");\n \t}\n \tif (buffer[0])\n \t\tret = NULL; /* not the whole tree */\n \telse\n \t\tret = read_one(&buffer, &size, jw);\n-\tif (jw) {\n-\t\tjw_end(jw);\t/* root */\n-\t\tjw_end(jw);\t/* cache-tree */\n-\t}\n+\tjw_end_gently(jw);\n \treturn ret;\n }\n \ndiff --git a/dir.c b/dir.c\nindex f389eee24a..8808577ea3 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -2902,13 +2902,15 @@ struct untracked_cache *read_untracked_extension(const void *data,\n \tuc->exclude_per_dir = xstrdup(exclude_per_dir);\n \n \tif (jw) {\n-\t\tjw_object_inline_begin_object(jw, \"untracked-cache\");\n-\t\tjw_object_intmax(jw, \"ext-size\", sz);\n \t\tjw_object_string(jw, \"ident\", ident);\n-\t\tjw_object_oid_stat(jw, \"info/exclude\", &uc->ss_info_exclude);\n-\t\tjw_object_oid_stat(jw, \"excludes-file\", &uc->ss_excludes_file);\n+\t\tjw_object_oid_stat(jw, \"info_exclude\", &uc->ss_info_exclude);\n+\t\tjw_object_oid_stat(jw, \"excludes_file\", &uc->ss_excludes_file);\n \t\tjw_object_intmax(jw, \"flags\", uc->dir_flags);\n-\t\tjw_object_string(jw, \"excludes-per-dir\", uc->exclude_per_dir);\n+\t\tif (uc->dir_flags & DIR_SHOW_OTHER_DIRECTORIES)\n+\t\t\tjw_object_bool(jw, \"show_other_directories\", 1);\n+\t\tif (uc->dir_flags & DIR_HIDE_EMPTY_DIRECTORIES)\n+\t\t\tjw_object_bool(jw, \"hide_empty_directories\", 1);\n+\t\tjw_object_string(jw, \"excludes_per_dir\", uc->exclude_per_dir);\n \t}\n \n \t/* NUL after exclude_per_dir is covered by sizeof(*ouc) */\n@@ -2968,7 +2970,6 @@ struct untracked_cache *read_untracked_extension(const void *data,\n \t\tfree_untracked_cache(uc);\n \t\tuc = NULL;\n \t}\n-\tjw_end_gently(jw);\n \treturn uc;\n }\n \ndiff --git a/fsmonitor.c b/fsmonitor.c\nindex f6ba437255..5ed55ad176 100644\n--- a/fsmonitor.c\n+++ b/fsmonitor.c\n@@ -52,12 +52,9 @@ int read_fsmonitor_extension(struct index_state *istate, const void *data,\n \tistate->fsmonitor_dirty = fsmonitor_dirty;\n \n \tif (istate->jw) {\n-\t\tjw_object_inline_begin_object(istate->jw, \"fsmonitor\");\n \t\tjw_object_intmax(istate->jw, \"version\", hdr_version);\n-\t\tjw_object_intmax(istate->jw, \"last-update\", istate->fsmonitor_last_update);\n+\t\tjw_object_intmax(istate->jw, \"last_update\", istate->fsmonitor_last_update);\n \t\tjw_object_ewah(istate->jw, \"dirty\", fsmonitor_dirty);\n-\t\tjw_object_intmax(istate->jw, \"ext-size\", sz);\n-\t\tjw_end(istate->jw);\n \t}\n \ttrace_printf_key(&trace_fsmonitor, \"read fsmonitor extension successful\");\n \treturn 0;\ndiff --git a/json-writer.c b/json-writer.c\nindex 70403580ca..c0bd302e4e 100644\n--- a/json-writer.c\n+++ b/json-writer.c\n@@ -203,19 +203,25 @@ void jw_object_null(struct json_writer *jw, const char *key)\n \tstrbuf_addstr(&jw->json, \"null\");\n }\n \n+void jw_object_filemode(struct json_writer *jw, const char *key, mode_t mode)\n+{\n+\tobject_common(jw, key);\n+\tstrbuf_addf(&jw->json, \"\\\"%06o\\\"\", mode);\n+}\n+\n void jw_object_stat_data(struct json_writer *jw, const char *name,\n \t\t\t const struct stat_data *sd)\n {\n \tjw_object_inline_begin_object(jw, name);\n-\tjw_object_intmax(jw, \"st_ctime.sec\", sd->sd_ctime.sec);\n-\tjw_object_intmax(jw, \"st_ctime.nsec\", sd->sd_ctime.nsec);\n-\tjw_object_intmax(jw, \"st_mtime.sec\", sd->sd_mtime.sec);\n-\tjw_object_intmax(jw, \"st_mtime.nsec\", sd->sd_mtime.nsec);\n-\tjw_object_intmax(jw, \"st_dev\", sd->sd_dev);\n-\tjw_object_intmax(jw, \"st_ino\", sd->sd_ino);\n-\tjw_object_intmax(jw, \"st_uid\", sd->sd_uid);\n-\tjw_object_intmax(jw, \"st_gid\", sd->sd_gid);\n-\tjw_object_intmax(jw, \"st_size\", sd->sd_size);\n+\tjw_object_intmax(jw, \"ctime_sec\", sd->sd_ctime.sec);\n+\tjw_object_intmax(jw, \"ctime_nsec\", sd->sd_ctime.nsec);\n+\tjw_object_intmax(jw, \"mtime_sec\", sd->sd_mtime.sec);\n+\tjw_object_intmax(jw, \"mtime_nsec\", sd->sd_mtime.nsec);\n+\tjw_object_intmax(jw, \"device\", sd->sd_dev);\n+\tjw_object_intmax(jw, \"inode\", sd->sd_ino);\n+\tjw_object_intmax(jw, \"uid\", sd->sd_uid);\n+\tjw_object_intmax(jw, \"gid\", sd->sd_gid);\n+\tjw_object_intmax(jw, \"size\", sd->sd_size);\n \tjw_end(jw);\n }\n \ndiff --git a/json-writer.h b/json-writer.h\nindex f778e019a2..07d841d52a 100644\n--- a/json-writer.h\n+++ b/json-writer.h\n@@ -42,8 +42,10 @@\n  * of the given strings.\n  */\n \n+#include \"git-compat-util.h\"\n #include \"strbuf.h\"\n \n+struct ewah_bitmap;\n struct stat_data;\n \n struct json_writer\n@@ -83,6 +85,7 @@ void jw_object_true(struct json_writer *jw, const char *key);\n void jw_object_false(struct json_writer *jw, const char *key);\n void jw_object_bool(struct json_writer *jw, const char *key, int value);\n void jw_object_null(struct json_writer *jw, const char *key);\n+void jw_object_filemode(struct json_writer *jw, const char *key, mode_t value);\n void jw_object_stat_data(struct json_writer *jw, const char *key,\n \t\t\t const struct stat_data *sd);\n void jw_object_ewah(struct json_writer *jw, const char *key,\ndiff --git a/read-cache.c b/read-cache.c\nindex d7d9ce7260..c26edcc9d9 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1693,9 +1693,28 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n \treturn 0;\n }\n \n+static struct index_entry_offset_table *do_read_ieot_extension(struct index_state *, const char *, uint32_t);\n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n+\t\t\t\tconst char *map,\n+\t\t\t\tunsigned long *offset)\n {\n+\tint ret = 0;\n+\tconst char *ext = map + *offset;\n+\tuint32_t sz = get_be32(ext + 4);\n+\tconst char *data = ext + 8;\n+\n+\tif (istate->jw) {\n+\t\tchar buf[5];\n+\n+\t\tmemcpy(buf, ext, 4);\n+\t\tbuf[4] = '\\0';\n+\t\tjw_object_inline_begin_object(istate->jw, buf);\n+\n+\t\tjw_object_intmax(istate->jw, \"file_offset\", *offset);\n+\t\tjw_object_intmax(istate->jw, \"ext_size\", sz);\n+\t}\n+\t*offset += sz + 8;\n+\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n \t\tistate->cache_tree = cache_tree_read(data, sz, istate->jw);\n@@ -1704,8 +1723,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tistate->resolve_undo = resolve_undo_read(data, sz, istate->jw);\n \t\tbreak;\n \tcase CACHE_EXT_LINK:\n-\t\tif (read_link_extension(istate, data, sz))\n-\t\t\treturn -1;\n+\t\tret = read_link_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_UNTRACKED:\n \t\tistate->untracked = read_untracked_extension(data, sz, istate->jw);\n@@ -1714,17 +1732,31 @@ static int read_index_extension(struct index_state *istate,\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_ENDOFINDEXENTRIES:\n-\tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n+\t\tif (istate->jw) {\n+\t\t\t/* must be synchronized with read_eoie_extension() */\n+\t\t\tjw_object_intmax(istate->jw, \"offset\", get_be32(data));\n+\t\t\tjw_object_string(istate->jw, \"oid\",\n+\t\t\t\t\t hash_to_hex((const unsigned char*)data + sizeof(uint32_t)));\n+\t\t}\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\n+\tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n+\t\tif (istate->jw) {\n+\t\t\tfree(do_read_ieot_extension(istate, data, sz));\n+\t\t} else {\n+\t\t\t/* already handled in do_read_index() */\n+\t\t}\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n-\t\t\treturn error(_(\"index uses %.4s extension, which we do not understand\"),\n+\t\t\tret = error(_(\"index uses %.4s extension, which we do not understand\"),\n \t\t\t\t     ext);\n-\t\tfprintf_ln(stderr, _(\"ignoring %.4s extension\"), ext);\n+\t\telse\n+\t\t\t  fprintf_ln(stderr, _(\"ignoring %.4s extension\"), ext);\n \t\tbreak;\n \t}\n-\treturn 0;\n+\tjw_end_gently(istate->jw);\n+\treturn ret;\n }\n \n static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n@@ -1911,10 +1943,10 @@ struct index_entry_offset_table\n \tstruct index_entry_offset entries[FLEX_ARRAY];\n };\n \n-static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset, struct json_writer *jw);\n+static struct index_entry_offset_table *read_ieot_extension(struct index_state *istate, const char *mmap, size_t mmap_size, size_t offset);\n static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot);\n \n-static size_t read_eoie_extension(const char *mmap, size_t mmap_size, struct json_writer *jw);\n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n struct load_index_extensions\n@@ -1930,25 +1962,27 @@ static void *load_index_extensions(void *_data)\n {\n \tstruct load_index_extensions *p = _data;\n \tunsigned long src_offset = p->src_offset;\n+\tint dump_json = 0;\n \n \twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n+\t\tif (p->istate->jw && !dump_json) {\n+\t\t\tjw_object_inline_begin_object(p->istate->jw, \"extensions\");\n+\t\t\tdump_json = 1;\n+\t\t}\n+\t\t/*\n+\t\t * After an array of active_nr index entries,\n \t\t * there can be arbitrary number of extended\n \t\t * sections, each of which is prefixed with\n \t\t * extension name (4-byte) and section length\n \t\t * in 4-byte network byte order.\n \t\t */\n-\t\tuint32_t extsize = get_be32(p->mmap + src_offset + 4);\n-\t\tif (read_index_extension(p->istate,\n-\t\t\t\t\t p->mmap + src_offset,\n-\t\t\t\t\t p->mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0) {\n+\t\tif (read_index_extension(p->istate, p->mmap, &src_offset) < 0) {\n \t\t\tmunmap((void *)p->mmap, p->mmap_size);\n \t\t\tdie(_(\"index file corrupt\"));\n \t\t}\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n \t}\n+\tif (dump_json)\n+\t\tjw_end(p->istate->jw);\n \n \treturn NULL;\n }\n@@ -1958,7 +1992,6 @@ static void dump_cache_entry(struct index_state *istate,\n \t\t\t     unsigned long offset,\n \t\t\t     const struct cache_entry *ce)\n {\n-\tstruct strbuf sb = STRBUF_INIT;\n \tstruct json_writer *jw = istate->jw;\n \n \tjw_array_inline_begin_object(jw);\n@@ -1971,28 +2004,28 @@ static void dump_cache_entry(struct index_state *istate,\n \n \tjw_object_string(jw, \"name\", ce->name);\n \n-\tstrbuf_addf(&sb, \"%06o\", ce->ce_mode);\n-\tjw_object_string(jw, \"mode\", sb.buf);\n-\tstrbuf_release(&sb);\n+\tjw_object_filemode(jw, \"mode\", ce->ce_mode);\n \n \tjw_object_intmax(jw, \"flags\", ce->ce_flags);\n \t/*\n \t * again redundant info, just so you don't have to decode\n \t * flags values manually\n \t */\n+\tif (ce->ce_flags & CE_EXTENDED)\n+\t\tjw_object_true(jw, \"extended_flags\");\n \tif (ce->ce_flags & CE_VALID)\n-\t\tjw_object_true(jw, \"assume-unchanged\");\n+\t\tjw_object_true(jw, \"assume_unchanged\");\n \tif (ce->ce_flags & CE_INTENT_TO_ADD)\n-\t\tjw_object_true(jw, \"intent-to-add\");\n+\t\tjw_object_true(jw, \"intent_to_add\");\n \tif (ce->ce_flags & CE_SKIP_WORKTREE)\n-\t\tjw_object_true(jw, \"skip-worktree\");\n+\t\tjw_object_true(jw, \"skip_worktree\");\n \tif (ce_stage(ce))\n \t\tjw_object_intmax(jw, \"stage\", ce_stage(ce));\n \n \tjw_object_string(jw, \"oid\", oid_to_hex(&ce->oid));\n \n \tjw_object_stat_data(jw, \"stat\", &ce->ce_stat_data);\n-\tjw_object_intmax(jw, \"file-offset\", offset);\n+\tjw_object_intmax(jw, \"file_offset\", offset);\n \n \tjw_end(jw);\n }\n@@ -2232,28 +2265,21 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tnr_threads = 1;\n \n \tif (istate->jw) {\n-\t\tsize_t off;\n-\n \t\tjw_object_begin(istate->jw, jw_pretty);\n \t\tjw_object_intmax(istate->jw, \"version\", istate->version);\n \t\tjw_object_string(istate->jw, \"oid\", oid_to_hex(&istate->oid));\n-\t\tjw_object_intmax(istate->jw, \"st_mtime.sec\", istate->timestamp.sec);\n-\t\tjw_object_intmax(istate->jw, \"st_mtime.nsec\", istate->timestamp.nsec);\n+\t\tjw_object_intmax(istate->jw, \"mtime_sec\", istate->timestamp.sec);\n+\t\tjw_object_intmax(istate->jw, \"mtime_nsec\", istate->timestamp.nsec);\n \n \t\t/*\n \t\t * Threading may mess up json writing. This is for\n \t\t * debugging only, so performance is not a concern.\n \t\t */\n \t\tnr_threads = 1;\n-\t\t/* and dump EOIE/IOET extensions even with threading off */\n-\t\toff = read_eoie_extension(mmap, mmap_size, istate->jw);\n-\t\tif (off)\n-\t\t\tfree(read_ieot_extension(mmap, mmap_size,\n-\t\t\t\t\t\t off, istate->jw));\n \t}\n \n \tif (nr_threads > 1) {\n-\t\textension_offset = read_eoie_extension(mmap, mmap_size, NULL);\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n \t\tif (extension_offset) {\n \t\t\tint err;\n \n@@ -2271,7 +2297,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t * to multi-thread the reading of the cache entries.\n \t */\n \tif (extension_offset && nr_threads > 1)\n-\t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset, NULL);\n+\t\tieot = read_ieot_extension(istate, mmap, mmap_size, extension_offset);\n \n \tif (ieot) {\n \t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, nr_threads, ieot);\n@@ -3511,8 +3537,7 @@ int should_validate_cache_entries(void)\n #define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n #define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n \n-static size_t read_eoie_extension(const char *mmap, size_t mmap_size,\n-\t\t\t\t  struct json_writer *jw)\n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n {\n \t/*\n \t * The end of index entries (EOIE) extension is guaranteed to be last\n@@ -3556,12 +3581,6 @@ static size_t read_eoie_extension(const char *mmap, size_t mmap_size,\n \t\treturn 0;\n \tindex += sizeof(uint32_t);\n \n-\tif (jw) {\n-\t\tjw_object_inline_begin_object(jw, \"end-of-index\");\n-\t\tjw_object_intmax(jw, \"offset\", offset);\n-\t\tjw_object_intmax(jw, \"ext-size\", extsize);\n-\t\tjw_object_inline_begin_array(jw, \"extensions\");\n-\t}\n \t/*\n \t * The hash is computed over extension types and their sizes (but not\n \t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n@@ -3590,24 +3609,9 @@ static size_t read_eoie_extension(const char *mmap, size_t mmap_size,\n \n \t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n \n-\t\tif (jw) {\n-\t\t\tchar name[5];\n-\n-\t\t\tjw_array_inline_begin_object(jw);\n-\t\t\tmemcpy(name, mmap + src_offset, 4);\n-\t\t\tname[4] = '\\0';\n-\t\t\tjw_object_string(jw, \"name\",  name);\n-\t\t\tjw_object_intmax(jw, \"size\", extsize);\n-\t\t\tjw_end(jw);\n-\t\t}\n-\n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n-\tif (jw) {\n-\t\tjw_end(jw);\t/* extensions */\n-\t\tjw_end(jw);\t/* end-of-index */\n-\t}\n \tthe_hash_algo->final_fn(hash, &c);\n \tif (!hasheq(hash, (const unsigned char *)index))\n \t\treturn 0;\n@@ -3636,13 +3640,12 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n #define IEOT_VERSION\t(1)\n \n static struct index_entry_offset_table *read_ieot_extension(\n-\tconst char *mmap, size_t mmap_size,\n-\tsize_t offset, struct json_writer *jw)\n+\t\tstruct index_state *istate,\n+\t\tconst char *mmap, size_t mmap_size,\n+\t\tsize_t offset)\n {\n \tconst char *index = NULL;\n-\tuint32_t extsize, ext_version;\n-\tstruct index_entry_offset_table *ieot;\n-\tint i, nr;\n+\tuint32_t extsize;\n \n \t/* find the IEOT extension */\n \tif (!offset)\n@@ -3658,6 +3661,17 @@ static struct index_entry_offset_table *read_ieot_extension(\n \t}\n \tif (!index)\n \t\treturn NULL;\n+\treturn do_read_ieot_extension(istate, index, extsize);\n+}\n+\n+static struct index_entry_offset_table *do_read_ieot_extension(\n+\t\tstruct index_state *istate,\n+\t\tconst char *index,\n+\t\tuint32_t extsize)\n+{\n+\tuint32_t ext_version;\n+\tstruct index_entry_offset_table *ieot;\n+\tint i, nr;\n \n \t/* validate the version is IEOT_VERSION */\n \text_version = get_be32(index);\n@@ -3673,11 +3687,9 @@ static struct index_entry_offset_table *read_ieot_extension(\n \t\terror(\"invalid number of IEOT entries %d\", nr);\n \t\treturn NULL;\n \t}\n-\tif (jw) {\n-\t\tjw_object_inline_begin_object(jw, \"index-entry-offsets\");\n-\t\tjw_object_intmax(jw, \"version\", ext_version);\n-\t\tjw_object_intmax(jw, \"ext-size\", extsize);\n-\t\tjw_object_inline_begin_array(jw, \"entries\");\n+\tif (istate->jw) {\n+\t\tjw_object_intmax(istate->jw, \"version\", ext_version);\n+\t\tjw_object_inline_begin_array(istate->jw, \"entries\");\n \t}\n \tieot = xmalloc(sizeof(struct index_entry_offset_table)\n \t\t       + (nr * sizeof(struct index_entry_offset)));\n@@ -3688,17 +3700,14 @@ static struct index_entry_offset_table *read_ieot_extension(\n \t\tieot->entries[i].nr = get_be32(index);\n \t\tindex += sizeof(uint32_t);\n \n-\t\tif (jw) {\n-\t\t\tjw_array_inline_begin_object(jw);\n-\t\t\tjw_object_intmax(jw, \"offset\", ieot->entries[i].offset);\n-\t\t\tjw_object_intmax(jw, \"count\", ieot->entries[i].nr);\n-\t\t\tjw_end(jw);\n+\t\tif (istate->jw) {\n+\t\t\tjw_array_inline_begin_object(istate->jw);\n+\t\t\tjw_object_intmax(istate->jw, \"offset\", ieot->entries[i].offset);\n+\t\t\tjw_object_intmax(istate->jw, \"count\", ieot->entries[i].nr);\n+\t\t\tjw_end(istate->jw);\n \t\t}\n \t}\n-\tif (jw) {\n-\t\tjw_end(jw);\t/* entries */\n-\t\tjw_end(jw);\t/* index-entry-offsets */\n-\t}\n+\tjw_end_gently(istate->jw);\n \n \treturn ieot;\n }\ndiff --git a/resolve-undo.c b/resolve-undo.c\nindex 999020bc40..68921e3dfe 100644\n--- a/resolve-undo.c\n+++ b/resolve-undo.c\n@@ -50,6 +50,28 @@ void resolve_undo_write(struct strbuf *sb, struct string_list *resolve_undo)\n \t}\n }\n \n+static void dump_resolve_undo(struct json_writer *jw,\n+\t\t\t      const char *path,\n+\t\t\t      const struct resolve_undo_info *ui)\n+{\n+\tint i;\n+\n+\tif (!jw)\n+\t\treturn;\n+\n+\tjw_array_inline_begin_object(jw);\n+\tjw_object_string(jw, \"path\", path);\n+\n+\tjw_object_inline_begin_array(jw, \"stages\");\n+\tfor (i = 0; i < 3; i++) {\n+\t\tjw_array_inline_begin_object(jw);\n+\t\tjw_object_filemode(jw, \"mode\", ui->mode[i]);\n+\t\tjw_object_string(jw, \"oid\", oid_to_hex(&ui->oid[i]));\n+\t\tjw_end(jw);\n+\t}\n+\tjw_end(jw);\n+}\n+\n struct string_list *resolve_undo_read(const char *data, unsigned long size,\n \t\t\t\t      struct json_writer *jw)\n {\n@@ -61,11 +83,7 @@ struct string_list *resolve_undo_read(const char *data, unsigned long size,\n \n \tresolve_undo = xcalloc(1, sizeof(*resolve_undo));\n \tresolve_undo->strdup_strings = 1;\n-\tif (jw) {\n-\t\tjw_object_inline_begin_object(jw, \"resolve-undo\");\n-\t\tjw_object_intmax(jw, \"ext-size\", size);\n-\t\tjw_object_inline_begin_array(jw, \"entries\");\n-\t}\n+\tjw_object_inline_begin_array_gently(jw, \"entries\");\n \n \twhile (size) {\n \t\tstruct string_list_item *lost;\n@@ -102,33 +120,9 @@ struct string_list *resolve_undo_read(const char *data, unsigned long size,\n \t\t\tdata += rawsz;\n \t\t}\n \n-\t\tif (jw) {\n-\t\t\tstruct strbuf sb = STRBUF_INIT;\n-\n-\t\t\tjw_array_inline_begin_object(jw);\n-\t\t\tjw_object_string(jw, \"path\", lost->string);\n-\n-\t\t\tjw_object_inline_begin_array(jw, \"mode\");\n-\t\t\tfor (i = 0; i < 3; i++) {\n-\t\t\t\tstrbuf_addf(&sb, \"%06o\", ui->mode[i]);\n-\t\t\t\tjw_array_string(jw, sb.buf);\n-\t\t\t\tstrbuf_reset(&sb);\n-\t\t\t}\n-\t\t\tjw_end(jw);\n-\n-\t\t\tjw_object_inline_begin_array(jw, \"oid\");\n-\t\t\tfor (i = 0; i < 3; i++)\n-\t\t\t\tjw_array_string(jw, oid_to_hex(&ui->oid[i]));\n-\t\t\tjw_end(jw);\n-\n-\t\t\tjw_end(jw);\n-\t\t\tstrbuf_release(&sb);\n-\t\t}\n-\t}\n-\tif (jw) {\n-\t\tjw_end(jw);\t/* entries */\n-\t\tjw_end(jw);\t/* resolve-undo */\n+\t\tdump_resolve_undo(jw, lost->string, ui);\n \t}\n+\tjw_end_gently(jw);\n \treturn resolve_undo;\n \n error:\ndiff --git a/split-index.c b/split-index.c\nindex d7b4420c92..41552bf771 100644\n--- a/split-index.c\n+++ b/split-index.c\n@@ -17,7 +17,6 @@ int read_link_extension(struct index_state *istate,\n {\n \tconst unsigned char *data = data_;\n \tstruct split_index *si;\n-\tunsigned long original_sz = sz;\n \tint ret;\n \n \tif (sz < the_hash_algo->rawsz)\n@@ -42,12 +41,9 @@ int read_link_extension(struct index_state *istate,\n \t\treturn error(\"garbage at the end of link extension\");\n done:\n \tif (istate->jw) {\n-\t\tjw_object_inline_begin_object(istate->jw, \"split-index\");\n \t\tjw_object_string(istate->jw, \"oid\", oid_to_hex(&si->base_oid));\n-\t\tjw_object_ewah(istate->jw, \"delete-bitmap\", si->delete_bitmap);\n-\t\tjw_object_ewah(istate->jw, \"replace-bitmap\", si->replace_bitmap);\n-\t\tjw_object_intmax(istate->jw, \"ext-size\", original_sz);\n-\t\tjw_end(istate->jw);\n+\t\tjw_object_ewah(istate->jw, \"delete_bitmap\", si->delete_bitmap);\n+\t\tjw_object_ewah(istate->jw, \"replace_bitmap\", si->replace_bitmap);\n \t}\n \treturn 0;\n }\ndiff --git a/t/t3008-ls-files-lazy-init-name-hash.sh b/t/t3008-ls-files-lazy-init-name-hash.sh\nindex 64f047332b..7f918c05f6 100755\n--- a/t/t3008-ls-files-lazy-init-name-hash.sh\n+++ b/t/t3008-ls-files-lazy-init-name-hash.sh\n@@ -4,15 +4,9 @@ test_description='Test the lazy init name hash with various folder structures'\n \n . ./test-lib.sh\n \n-if test 1 -eq $($GIT_BUILD_DIR/t/helper/test-tool online-cpus)\n-then\n-\tskip_all='skipping lazy-init tests, single cpu'\n-\ttest_done\n-fi\n-\n LAZY_THREAD_COST=2000\n \n-test_expect_success 'no buffer overflow in lazy_init_name_hash' '\n+test_expect_success !SINGLE_CPU 'no buffer overflow in lazy_init_name_hash' '\n \t(\n \t    test_seq $LAZY_THREAD_COST | sed \"s/^/a_/\" &&\n \t    echo b/b/b &&\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nnew file mode 100755\nindex 0000000000..9f4ad4c9cf\n--- /dev/null\n+++ b/t/t3011-ls-files-json.sh\n@@ -0,0 +1,106 @@\n+#!/bin/sh\n+\n+test_description='ls-files dumping json'\n+\n+. ./test-lib.sh\n+\n+strip_number() {\n+\tfor name; do\n+\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n+\tdone\n+}\n+\n+strip_string() {\n+\tfor name; do\n+\t\techo 's/\\(\"'$name'\":\\) \".*\"/\\1 <string>/' >>filter.sed\n+\tdone\n+}\n+\n+compare_json() {\n+\tgit ls-files --debug-json >json &&\n+\tsed -f filter.sed json >filtered &&\n+\ttest_cmp \"$TEST_DIRECTORY\"/t3011/\"$1\" filtered\n+}\n+\n+test_expect_success 'setup' '\n+\tmkdir sub &&\n+\techo one >one &&\n+\tgit add one &&\n+\techo 2 >sub/two &&\n+\tgit add sub/two &&\n+\n+\tgit commit -m first &&\n+\tgit update-index --untracked-cache &&\n+\n+\techo intent-to-add >ita &&\n+\tgit add -N ita &&\n+\n+\tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n+\tstrip_number device inode uid gid file_offset ext_size last_update &&\n+\tstrip_string oid ident\n+'\n+\n+test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n+\tcompare_json basic\n+'\n+\n+test_expect_success 'ls-files --json, split index' '\n+\tgit init split &&\n+\t(\n+\t\tcd split &&\n+\t\techo one >one &&\n+\t\tgit add one &&\n+\t\tgit update-index --split-index &&\n+\t\techo updated >>one &&\n+\t\ttest_must_fail git -c splitIndex.maxPercentChange=100 update-index --refresh &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json split-index\n+\t)\n+'\n+\n+test_expect_success 'ls-files --json, fsmonitor extension ' '\n+\tgit init fsmonitor &&\n+\t(\n+\t\tcd fsmonitor &&\n+\t\techo one >one &&\n+\t\tgit add one &&\n+\t\tgit update-index --fsmonitor &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json fsmonitor\n+\t)\n+'\n+\n+test_expect_success 'ls-files --json, rerere extension' '\n+\tgit init rerere &&\n+\t(\n+\t\tcd rerere &&\n+\t\tmkdir fi &&\n+\t\ttest_commit initial fi/le first &&\n+\t\tgit branch side &&\n+\t\ttest_commit second fi/le second &&\n+\t\tgit checkout side &&\n+\t\ttest_commit third fi/le third &&\n+\t\tgit checkout master &&\n+\t\tgit config rerere.enabled true &&\n+\t\ttest_must_fail git merge side &&\n+\t\techo resolved >fi/le &&\n+\t\tgit add fi/le &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json rerere\n+\t)\n+'\n+\n+test_expect_success !SINGLE_CPU 'ls-files --json and multicore extensions' '\n+\tgit init eoie &&\n+\t(\n+\t\tcd eoie &&\n+\t\tgit config index.threads 2 &&\n+\t\ttouch one two three four &&\n+\t\tgit add . &&\n+\t\tcp ../filter.sed . &&\n+\t\tstrip_number offset &&\n+\t\tcompare_json eoie\n+\t)\n+'\n+\n+test_done\ndiff --git a/t/t3011/basic b/t/t3011/basic\nnew file mode 100644\nindex 0000000000..8e049f5350\n--- /dev/null\n+++ b/t/t3011/basic\n@@ -0,0 +1,124 @@\n+{\n+  \"version\": 3,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"ita\",\n+      \"mode\": \"100644\",\n+      \"flags\": 536887296,\n+      \"extended_flags\": true,\n+      \"intent_to_add\": true,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 1,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 2,\n+      \"name\": \"sub/two\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 2\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"TREE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"root\": {\n+        \"oid\": null,\n+        \"subdirs\": [\n+          {\n+            \"name\": \"sub\",\n+            \"oid\": <string>,\n+            \"entry_count\": 1,\n+            \"subdirs\": [\n+            ]\n+          }\n+        ]\n+      }\n+    },\n+    \"UNTR\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"ident\": <string>,\n+      \"info_exclude\": {\n+        \"valid\": true,\n+        \"oid\": <string>,\n+        \"stat\": {\n+          \"ctime_sec\": <number>,\n+          \"ctime_nsec\": <number>,\n+          \"mtime_sec\": <number>,\n+          \"mtime_nsec\": <number>,\n+          \"device\": <number>,\n+          \"inode\": <number>,\n+          \"uid\": <number>,\n+          \"gid\": <number>,\n+          \"size\": 0\n+        }\n+      },\n+      \"excludes_file\": {\n+        \"valid\": true,\n+        \"oid\": <string>,\n+        \"stat\": {\n+          \"ctime_sec\": <number>,\n+          \"ctime_nsec\": <number>,\n+          \"mtime_sec\": <number>,\n+          \"mtime_nsec\": <number>,\n+          \"device\": <number>,\n+          \"inode\": <number>,\n+          \"uid\": <number>,\n+          \"gid\": <number>,\n+          \"size\": 0\n+        }\n+      },\n+      \"flags\": 6,\n+      \"show_other_directories\": true,\n+      \"hide_empty_directories\": true,\n+      \"excludes_per_dir\": \".gitignore\"\n+    }\n+  }\n+}\ndiff --git a/t/t3011/eoie b/t/t3011/eoie\nnew file mode 100644\nindex 0000000000..66a0feb3b6\n--- /dev/null\n+++ b/t/t3011/eoie\n@@ -0,0 +1,107 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"four\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 1,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 2,\n+      \"name\": \"three\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 3,\n+      \"name\": \"two\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"IEOT\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"version\": 1,\n+      \"entries\": [\n+        {\n+          \"offset\": <number>,\n+          \"count\": 2\n+        },\n+        {\n+          \"offset\": <number>,\n+          \"count\": 2\n+        }\n+      ]\n+    },\n+    \"EOIE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"offset\": <number>,\n+      \"oid\": <string>\n+    }\n+  }\n+}\ndiff --git a/t/t3011/fsmonitor b/t/t3011/fsmonitor\nnew file mode 100644\nindex 0000000000..17f2d4a664\n--- /dev/null\n+++ b/t/t3011/fsmonitor\n@@ -0,0 +1,38 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"FSMN\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"version\": 1,\n+      \"last_update\": <number>,\n+      \"dirty\": [\n+        0\n+      ]\n+    }\n+  }\n+}\ndiff --git a/t/t3011/rerere b/t/t3011/rerere\nnew file mode 100644\nindex 0000000000..a8ec4b16ee\n--- /dev/null\n+++ b/t/t3011/rerere\n@@ -0,0 +1,66 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"fi/le\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 9\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"TREE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"root\": {\n+        \"oid\": null,\n+        \"subdirs\": [\n+          {\n+            \"name\": \"fi\",\n+            \"oid\": null,\n+            \"subdirs\": [\n+            ]\n+          }\n+        ]\n+      }\n+    },\n+    \"REUC\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"entries\": [\n+        {\n+          \"path\": \"fi/le\",\n+          \"stages\": [\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            },\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            },\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            }\n+          ]\n+        }\n+      ]\n+    }\n+  }\ndiff --git a/t/t3011/split-index b/t/t3011/split-index\nnew file mode 100644\nindex 0000000000..cdcc4ddded\n--- /dev/null\n+++ b/t/t3011/split-index\n@@ -0,0 +1,39 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"link\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"oid\": <string>,\n+      \"delete_bitmap\": [\n+      ],\n+      \"replace_bitmap\": [\n+        0\n+      ]\n+    }\n+  }\n+}\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex 4b346467df..9d5b273b40 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1611,3 +1611,7 @@ test_lazy_prereq REBASE_P '\n test_lazy_prereq FAIL_PREREQS '\n \ttest -n \"$GIT_TEST_FAIL_PREREQS\"\n '\n+\n+test_lazy_prereq SINGLE_CPU '\n+\ttest \"$(test-tool online-cpus)\" -eq 1\n+'\n\nNguyễn Thái Ngọc Duy (10):\n  ls-files: add --json to dump the index\n  read-cache.c: dump common extension info in json\n  cache-tree.c: dump \"TREE\" extension as json\n  dir.c: dump \"UNTR\" extension as json\n  split-index.c: dump \"link\" extension as json\n  fsmonitor.c: dump \"FSMN\" extension as json\n  resolve-undo.c: dump \"REUC\" extension as json\n  read-cache.c: dump \"EOIE\" extension as json\n  read-cache.c: dump \"IEOT\" extension as json\n  t3008: use the new SINGLE_CPU prereq\n\n Documentation/git-ls-files.txt          |   5 +\n builtin/ls-files.c                      |  38 ++++-\n cache-tree.c                            |  36 ++++-\n cache-tree.h                            |   5 +-\n cache.h                                 |   2 +\n dir.c                                   |  57 +++++++-\n dir.h                                   |   4 +-\n fsmonitor.c                             |   6 +\n json-writer.c                           |  36 +++++\n json-writer.h                           |  32 +++++\n read-cache.c                            | 178 ++++++++++++++++++++----\n resolve-undo.c                          |  30 +++-\n resolve-undo.h                          |   4 +-\n split-index.c                           |   9 +-\n t/t3008-ls-files-lazy-init-name-hash.sh |   8 +-\n t/t3011-ls-files-json.sh (new +x)       | 106 ++++++++++++++\n t/t3011/basic (new)                     | 124 +++++++++++++++++\n t/t3011/eoie (new)                      | 107 ++++++++++++++\n t/t3011/fsmonitor (new)                 |  38 +++++\n t/t3011/rerere (new)                    |  66 +++++++++\n t/t3011/split-index (new)               |  39 ++++++\n t/test-lib.sh                           |   4 +\n 22 files changed, 884 insertions(+), 50 deletions(-)\n create mode 100755 t/t3011-ls-files-json.sh\n create mode 100644 t/t3011/basic\n create mode 100644 t/t3011/eoie\n create mode 100644 t/t3011/fsmonitor\n create mode 100644 t/t3011/rerere\n create mode 100644 t/t3011/split-index\n\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377875","messageId":"20190624130226.17293-2-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:17Z","receivedAt":"2019-06-24T13:02:52Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"So far we don't have a command to basically dump the index file out,\nwith all its glory details. Checking some info, for example, stat\ntime, usually involves either writing new code or firing up \"xxd\" and\ndecoding values by yourself.\n\nThis --json is supposed to help that. It dumps the index in a human\nreadable format but also easy to be processed with tools. And it will\nprint almost enough info to reconstruct the index later.\n\nIn this patch we only dump the main part, not extensions. But at the\nend of the series, the entire index is dumped. The end result could be\nvery verbose even on a small repository such as git.git.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/git-ls-files.txt    |  5 +++\n builtin/ls-files.c                | 38 +++++++++++++---\n cache.h                           |  2 +\n json-writer.c                     | 22 ++++++++++\n json-writer.h                     | 23 ++++++++++\n read-cache.c                      | 72 ++++++++++++++++++++++++++++++-\n t/t3011-ls-files-json.sh (new +x) | 44 +++++++++++++++++++\n t/t3011/basic (new)               | 67 ++++++++++++++++++++++++++++\n 8 files changed, 265 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-ls-files.txt b/Documentation/git-ls-files.txt\nindex 8461c0e83e..fec5cb7170 100644\n--- a/Documentation/git-ls-files.txt\n+++ b/Documentation/git-ls-files.txt\n@@ -162,6 +162,11 @@ a space) at the start of each line:\n \tpossible for manual inspection; the exact format may change at\n \tany time.\n \n+--debug-json::\n+\tDump the entire index content in JSON format. This is for\n+\tdebugging purposes. The JSON structure is subject to change.\n+\tNote that the strings are not always valid UTF-8.\n+\n --eol::\n \tShow <eolinfo> and <eolattr> of files.\n \t<eolinfo> is the file content identification used by Git when\ndiff --git a/builtin/ls-files.c b/builtin/ls-files.c\nindex 7f83c9a6f2..b60cd9ab28 100644\n--- a/builtin/ls-files.c\n+++ b/builtin/ls-files.c\n@@ -8,6 +8,7 @@\n #include \"cache.h\"\n #include \"repository.h\"\n #include \"config.h\"\n+#include \"json-writer.h\"\n #include \"quote.h\"\n #include \"dir.h\"\n #include \"builtin.h\"\n@@ -31,6 +32,7 @@ static int show_modified;\n static int show_killed;\n static int show_valid_bit;\n static int show_fsmonitor_bit;\n+static int show_json;\n static int line_terminator = '\\n';\n static int debug_mode;\n static int show_eol;\n@@ -577,6 +579,8 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t\tN_(\"pretend that paths removed since <tree-ish> are still present\")),\n \t\tOPT__ABBREV(&abbrev),\n \t\tOPT_BOOL(0, \"debug\", &debug_mode, N_(\"show debugging data\")),\n+\t\tOPT_BOOL(0, \"debug-json\", &show_json,\n+\t\t\tN_(\"dump index content in JSON format\")),\n \t\tOPT_END()\n \t};\n \n@@ -632,7 +636,7 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\t    \"--error-unmatch\");\n \n \tparse_pathspec(&pathspec, 0,\n-\t\t       PATHSPEC_PREFER_CWD,\n+\t\t       show_json ? PATHSPEC_PREFER_FULL : PATHSPEC_PREFER_CWD,\n \t\t       prefix, argv);\n \n \t/*\n@@ -660,8 +664,18 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \n \t/* With no flags, we default to showing the cached files */\n \tif (!(show_stage || show_deleted || show_others || show_unmerged ||\n-\t      show_killed || show_modified || show_resolve_undo))\n+\t      show_killed || show_modified || show_resolve_undo || show_json))\n \t\tshow_cached = 1;\n+\tif (show_json && (show_stage || show_deleted || show_others ||\n+\t\t\t  show_unmerged || show_killed || show_modified ||\n+\t\t\t  show_cached || pathspec.nr))\n+\t\tdie(_(\"--debug-json cannot be used with other file selection options\"));\n+\tif (show_json && show_resolve_undo)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--resolve-undo\");\n+\tif (show_json && with_tree)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--with-tree\");\n+\tif (show_json && debug_mode)\n+\t\tdie(_(\"--debug-json cannot be used with %s\"), \"--debug\");\n \n \tif (with_tree) {\n \t\t/*\n@@ -673,10 +687,22 @@ int cmd_ls_files(int argc, const char **argv, const char *cmd_prefix)\n \t\toverlay_tree_on_index(the_repository->index, with_tree, max_prefix);\n \t}\n \n-\tshow_files(the_repository, &dir);\n-\n-\tif (show_resolve_undo)\n-\t\tshow_ru_info(the_repository->index);\n+\tif (!show_json) {\n+\t\tshow_files(the_repository, &dir);\n+\n+\t\tif (show_resolve_undo)\n+\t\t\tshow_ru_info(the_repository->index);\n+\t} else {\n+\t\tstruct json_writer jw = JSON_WRITER_INIT;\n+\n+\t\tdiscard_index(the_repository->index);\n+\t\tthe_repository->index->jw = &jw;\n+\t\tif (repo_read_index(the_repository) < 0)\n+\t\t\tdie(\"index file corrupt\");\n+\t\tputs(jw.json.buf);\n+\t\tthe_repository->index->jw = NULL;\n+\t\tjw_release(&jw);\n+\t}\n \n \tif (ps_matched) {\n \t\tint bad;\ndiff --git a/cache.h b/cache.h\nindex bf20337ef4..84d0aeed20 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -326,6 +326,7 @@ static inline unsigned int canon_mode(unsigned int mode)\n #define UNTRACKED_CHANGED\t(1 << 7)\n #define FSMONITOR_CHANGED\t(1 << 8)\n \n+struct json_writer;\n struct split_index;\n struct untracked_cache;\n \n@@ -350,6 +351,7 @@ struct index_state {\n \tuint64_t fsmonitor_last_update;\n \tstruct ewah_bitmap *fsmonitor_dirty;\n \tstruct mem_pool *ce_mem_pool;\n+\tstruct json_writer *jw;\n };\n \n /* Name hashing */\ndiff --git a/json-writer.c b/json-writer.c\nindex aadb9dbddc..0608726512 100644\n--- a/json-writer.c\n+++ b/json-writer.c\n@@ -202,6 +202,28 @@ void jw_object_null(struct json_writer *jw, const char *key)\n \tstrbuf_addstr(&jw->json, \"null\");\n }\n \n+void jw_object_filemode(struct json_writer *jw, const char *key, mode_t mode)\n+{\n+\tobject_common(jw, key);\n+\tstrbuf_addf(&jw->json, \"\\\"%06o\\\"\", mode);\n+}\n+\n+void jw_object_stat_data(struct json_writer *jw, const char *name,\n+\t\t\t const struct stat_data *sd)\n+{\n+\tjw_object_inline_begin_object(jw, name);\n+\tjw_object_intmax(jw, \"ctime_sec\", sd->sd_ctime.sec);\n+\tjw_object_intmax(jw, \"ctime_nsec\", sd->sd_ctime.nsec);\n+\tjw_object_intmax(jw, \"mtime_sec\", sd->sd_mtime.sec);\n+\tjw_object_intmax(jw, \"mtime_nsec\", sd->sd_mtime.nsec);\n+\tjw_object_intmax(jw, \"device\", sd->sd_dev);\n+\tjw_object_intmax(jw, \"inode\", sd->sd_ino);\n+\tjw_object_intmax(jw, \"uid\", sd->sd_uid);\n+\tjw_object_intmax(jw, \"gid\", sd->sd_gid);\n+\tjw_object_intmax(jw, \"size\", sd->sd_size);\n+\tjw_end(jw);\n+}\n+\n static void increase_indent(struct strbuf *sb,\n \t\t\t    const struct json_writer *jw,\n \t\t\t    int indent)\ndiff --git a/json-writer.h b/json-writer.h\nindex 83906b09c1..c48c4cbf33 100644\n--- a/json-writer.h\n+++ b/json-writer.h\n@@ -42,8 +42,11 @@\n  * of the given strings.\n  */\n \n+#include \"git-compat-util.h\"\n #include \"strbuf.h\"\n \n+struct stat_data;\n+\n struct json_writer\n {\n \t/*\n@@ -81,6 +84,9 @@ void jw_object_true(struct json_writer *jw, const char *key);\n void jw_object_false(struct json_writer *jw, const char *key);\n void jw_object_bool(struct json_writer *jw, const char *key, int value);\n void jw_object_null(struct json_writer *jw, const char *key);\n+void jw_object_filemode(struct json_writer *jw, const char *key, mode_t value);\n+void jw_object_stat_data(struct json_writer *jw, const char *key,\n+\t\t\t const struct stat_data *sd);\n void jw_object_sub_jw(struct json_writer *jw, const char *key,\n \t\t      const struct json_writer *value);\n \n@@ -104,4 +110,21 @@ void jw_array_inline_begin_array(struct json_writer *jw);\n int jw_is_terminated(const struct json_writer *jw);\n void jw_end(struct json_writer *jw);\n \n+/*\n+ * These _gently versions accept NULL json_writer to reduce too much\n+ * branching at the call site.\n+ */\n+static inline void jw_object_inline_begin_array_gently(struct json_writer *jw,\n+\t\t\t\t\t\t       const char *name)\n+{\n+\tif (jw)\n+\t\tjw_object_inline_begin_array(jw, name);\n+}\n+\n+static inline void jw_end_gently(struct json_writer *jw)\n+{\n+\tif (jw)\n+\t\tjw_end(jw);\n+}\n+\n #endif /* JSON_WRITER_H */\ndiff --git a/read-cache.c b/read-cache.c\nindex 4dd22f4f6e..db5147d088 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -25,6 +25,7 @@\n #include \"fsmonitor.h\"\n #include \"thread-utils.h\"\n #include \"progress.h\"\n+#include \"json-writer.h\"\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1952,6 +1953,49 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+static void dump_cache_entry(struct index_state *istate,\n+\t\t\t     int index,\n+\t\t\t     unsigned long offset,\n+\t\t\t     const struct cache_entry *ce)\n+{\n+\tstruct json_writer *jw = istate->jw;\n+\n+\tjw_array_inline_begin_object(jw);\n+\n+\t/*\n+\t * this is technically redundant, but it's for easier\n+\t * navigation when there hundreds of entries\n+\t */\n+\tjw_object_intmax(jw, \"id\", index);\n+\n+\tjw_object_string(jw, \"name\", ce->name);\n+\n+\tjw_object_filemode(jw, \"mode\", ce->ce_mode);\n+\n+\tjw_object_intmax(jw, \"flags\", ce->ce_flags);\n+\t/*\n+\t * again redundant info, just so you don't have to decode\n+\t * flags values manually\n+\t */\n+\tif (ce->ce_flags & CE_EXTENDED)\n+\t\tjw_object_true(jw, \"extended_flags\");\n+\tif (ce->ce_flags & CE_VALID)\n+\t\tjw_object_true(jw, \"assume_unchanged\");\n+\tif (ce->ce_flags & CE_INTENT_TO_ADD)\n+\t\tjw_object_true(jw, \"intent_to_add\");\n+\tif (ce->ce_flags & CE_SKIP_WORKTREE)\n+\t\tjw_object_true(jw, \"skip_worktree\");\n+\tif (ce_stage(ce))\n+\t\tjw_object_intmax(jw, \"stage\", ce_stage(ce));\n+\n+\tjw_object_string(jw, \"oid\", oid_to_hex(&ce->oid));\n+\n+\tjw_object_stat_data(jw, \"stat\", &ce->ce_stat_data);\n+\tjw_object_intmax(jw, \"file_offset\", offset);\n+\n+\tjw_end(jw);\n+}\n+\n /*\n  * A helper function that will load the specified range of cache entries\n  * from the memory mapped file and add them to the given index.\n@@ -1972,6 +2016,9 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n+\t\tif (istate->jw)\n+\t\t\tdump_cache_entry(istate, i, src_offset, ce);\n+\n \t\tsrc_offset += consumed;\n \t\tprevious_ce = ce;\n \t}\n@@ -1983,6 +2030,8 @@ static unsigned long load_all_cache_entries(struct index_state *istate,\n {\n \tunsigned long consumed;\n \n+\tjw_object_inline_begin_array_gently(istate->jw, \"entries\");\n+\n \tif (istate->version == 4) {\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n@@ -1993,6 +2042,8 @@ static unsigned long load_all_cache_entries(struct index_state *istate,\n \n \tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n \t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n+\n+\tjw_end_gently(istate->jw);\n \treturn consumed;\n }\n \n@@ -2120,6 +2171,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tsize_t extension_offset = 0;\n \tint nr_threads, cpus;\n \tstruct index_entry_offset_table *ieot = NULL;\n+\tint jw_pretty = 1;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -2154,6 +2206,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache_nr = ntohl(hdr->hdr_entries);\n \tistate->cache_alloc = alloc_nr(istate->cache_nr);\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n+\tistate->timestamp.sec = st.st_mtime;\n+\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \tistate->initialized = 1;\n \n \tp.istate = istate;\n@@ -2176,6 +2230,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tif (!HAVE_THREADS)\n \t\tnr_threads = 1;\n \n+\tif (istate->jw) {\n+\t\tjw_object_begin(istate->jw, jw_pretty);\n+\t\tjw_object_intmax(istate->jw, \"version\", istate->version);\n+\t\tjw_object_string(istate->jw, \"oid\", oid_to_hex(&istate->oid));\n+\t\tjw_object_intmax(istate->jw, \"mtime_sec\", istate->timestamp.sec);\n+\t\tjw_object_intmax(istate->jw, \"mtime_nsec\", istate->timestamp.nsec);\n+\n+\t\t/*\n+\t\t * Threading may mess up json writing. This is for\n+\t\t * debugging only, so performance is not a concern.\n+\t\t */\n+\t\tnr_threads = 1;\n+\t}\n+\n \tif (nr_threads > 1) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n \t\tif (extension_offset) {\n@@ -2204,8 +2272,6 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n \t}\n \n-\tistate->timestamp.sec = st.st_mtime;\n-\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n \t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n \tif (extension_offset) {\n@@ -2216,6 +2282,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tp.src_offset = src_offset;\n \t\tload_index_extensions(&p);\n \t}\n+\tjw_end_gently(istate->jw);\n+\n \tmunmap((void *)mmap, mmap_size);\n \n \t/*\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nnew file mode 100755\nindex 0000000000..97bcd814be\n--- /dev/null\n+++ b/t/t3011-ls-files-json.sh\n@@ -0,0 +1,44 @@\n+#!/bin/sh\n+\n+test_description='ls-files dumping json'\n+\n+. ./test-lib.sh\n+\n+strip_number() {\n+\tfor name; do\n+\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n+\tdone\n+}\n+\n+strip_string() {\n+\tfor name; do\n+\t\techo 's/\\(\"'$name'\":\\) \".*\"/\\1 <string>/' >>filter.sed\n+\tdone\n+}\n+\n+compare_json() {\n+\tgit ls-files --debug-json >json &&\n+\tsed -f filter.sed json >filtered &&\n+\ttest_cmp \"$TEST_DIRECTORY\"/t3011/\"$1\" filtered\n+}\n+\n+test_expect_success 'setup' '\n+\tmkdir sub &&\n+\techo one >one &&\n+\tgit add one &&\n+\techo 2 >sub/two &&\n+\tgit add sub/two &&\n+\n+\techo intent-to-add >ita &&\n+\tgit add -N ita &&\n+\n+\tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n+\tstrip_number device inode uid gid file_offset ext_size &&\n+\tstrip_string oid ident\n+'\n+\n+test_expect_success 'ls-files --json, main entries' '\n+\tcompare_json basic\n+'\n+\n+test_done\ndiff --git a/t/t3011/basic b/t/t3011/basic\nnew file mode 100644\nindex 0000000000..9436445d90\n--- /dev/null\n+++ b/t/t3011/basic\n@@ -0,0 +1,67 @@\n+{\n+  \"version\": 3,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"ita\",\n+      \"mode\": \"100644\",\n+      \"flags\": 536887296,\n+      \"extended_flags\": true,\n+      \"intent_to_add\": true,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 1,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 2,\n+      \"name\": \"sub/two\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 2\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ]\n+}\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377876","messageId":"20190624130226.17293-3-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 02/10] read-cache.c: dump common extension info in json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:18Z","receivedAt":"2019-06-24T13:02:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 49 +++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 35 insertions(+), 14 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex db5147d088..4accd8bb08 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1694,8 +1694,26 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n+\t\t\t\tconst char *map,\n+\t\t\t\tunsigned long *offset)\n {\n+\tint ret = 0;\n+\tconst char *ext = map + *offset;\n+\tuint32_t sz = get_be32(ext + 4);\n+\tconst char *data = ext + 8;\n+\n+\tif (istate->jw) {\n+\t\tchar buf[5];\n+\n+\t\tmemcpy(buf, ext, 4);\n+\t\tbuf[4] = '\\0';\n+\t\tjw_object_inline_begin_object(istate->jw, buf);\n+\n+\t\tjw_object_intmax(istate->jw, \"file_offset\", *offset);\n+\t\tjw_object_intmax(istate->jw, \"ext_size\", sz);\n+\t}\n+\t*offset += sz + 8;\n+\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n \t\tistate->cache_tree = cache_tree_read(data, sz);\n@@ -1704,8 +1722,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tistate->resolve_undo = resolve_undo_read(data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_LINK:\n-\t\tif (read_link_extension(istate, data, sz))\n-\t\t\treturn -1;\n+\t\tret = read_link_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_UNTRACKED:\n \t\tistate->untracked = read_untracked_extension(data, sz);\n@@ -1719,12 +1736,14 @@ static int read_index_extension(struct index_state *istate,\n \t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n-\t\t\treturn error(_(\"index uses %.4s extension, which we do not understand\"),\n+\t\t\tret = error(_(\"index uses %.4s extension, which we do not understand\"),\n \t\t\t\t     ext);\n-\t\tfprintf_ln(stderr, _(\"ignoring %.4s extension\"), ext);\n+\t\telse\n+\t\t\t  fprintf_ln(stderr, _(\"ignoring %.4s extension\"), ext);\n \t\tbreak;\n \t}\n-\treturn 0;\n+\tjw_end_gently(istate->jw);\n+\treturn ret;\n }\n \n static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n@@ -1930,25 +1949,27 @@ static void *load_index_extensions(void *_data)\n {\n \tstruct load_index_extensions *p = _data;\n \tunsigned long src_offset = p->src_offset;\n+\tint dump_json = 0;\n \n \twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n+\t\tif (p->istate->jw && !dump_json) {\n+\t\t\tjw_object_inline_begin_object(p->istate->jw, \"extensions\");\n+\t\t\tdump_json = 1;\n+\t\t}\n+\t\t/*\n+\t\t * After an array of active_nr index entries,\n \t\t * there can be arbitrary number of extended\n \t\t * sections, each of which is prefixed with\n \t\t * extension name (4-byte) and section length\n \t\t * in 4-byte network byte order.\n \t\t */\n-\t\tuint32_t extsize = get_be32(p->mmap + src_offset + 4);\n-\t\tif (read_index_extension(p->istate,\n-\t\t\t\t\t p->mmap + src_offset,\n-\t\t\t\t\t p->mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0) {\n+\t\tif (read_index_extension(p->istate, p->mmap, &src_offset) < 0) {\n \t\t\tmunmap((void *)p->mmap, p->mmap_size);\n \t\t\tdie(_(\"index file corrupt\"));\n \t\t}\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n \t}\n+\tif (dump_json)\n+\t\tjw_end(p->istate->jw);\n \n \treturn NULL;\n }\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377877","messageId":"20190624130226.17293-4-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 03/10] cache-tree.c: dump \"TREE\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:19Z","receivedAt":"2019-06-24T13:03:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache-tree.c             | 36 +++++++++++++++++++++++++++++++-----\n cache-tree.h             |  5 ++++-\n read-cache.c             |  2 +-\n t/t3011-ls-files-json.sh |  4 +++-\n t/t3011/basic            | 20 +++++++++++++++++++-\n 5 files changed, 58 insertions(+), 9 deletions(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex b13bfaf71e..b6a233307e 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -3,6 +3,7 @@\n #include \"tree.h\"\n #include \"tree-walk.h\"\n #include \"cache-tree.h\"\n+#include \"json-writer.h\"\n #include \"object-store.h\"\n #include \"replace-object.h\"\n \n@@ -492,7 +493,8 @@ void cache_tree_write(struct strbuf *sb, struct cache_tree *root)\n \twrite_one(sb, root, \"\", 0);\n }\n \n-static struct cache_tree *read_one(const char **buffer, unsigned long *size_p)\n+static struct cache_tree *read_one(const char **buffer, unsigned long *size_p,\n+\t\t\t\t   struct json_writer *jw)\n {\n \tconst char *buf = *buffer;\n \tunsigned long size = *size_p;\n@@ -546,6 +548,15 @@ static struct cache_tree *read_one(const char **buffer, unsigned long *size_p)\n \t\t\t*buffer, subtree_nr);\n #endif\n \n+\tif (jw) {\n+\t\tif (it->entry_count >= 0) {\n+\t\t\tjw_object_string(jw, \"oid\", oid_to_hex(&it->oid));\n+\t\t\tjw_object_intmax(jw, \"entry_count\", it->entry_count);\n+\t\t} else {\n+\t\t\tjw_object_null(jw, \"oid\");\n+\t\t}\n+\t\tjw_object_inline_begin_array(jw, \"subdirs\");\n+\t}\n \t/*\n \t * Just a heuristic -- we do not add directories that often but\n \t * we do not want to have to extend it immediately when we do,\n@@ -559,12 +570,18 @@ static struct cache_tree *read_one(const char **buffer, unsigned long *size_p)\n \t\tstruct cache_tree_sub *subtree;\n \t\tconst char *name = buf;\n \n-\t\tsub = read_one(&buf, &size);\n+\t\tif (jw) {\n+\t\t\tjw_array_inline_begin_object(jw);\n+\t\t\tjw_object_string(jw, \"name\", name);\n+\t\t}\n+\t\tsub = read_one(&buf, &size, jw);\n+\t\tjw_end_gently(jw);\n \t\tif (!sub)\n \t\t\tgoto free_return;\n \t\tsubtree = cache_tree_sub(it, name);\n \t\tsubtree->cache_tree = sub;\n \t}\n+\tjw_end_gently(jw);\n \tif (subtree_nr != it->subtree_nr)\n \t\tdie(\"cache-tree: internal error\");\n \t*buffer = buf;\n@@ -576,11 +593,20 @@ static struct cache_tree *read_one(const char **buffer, unsigned long *size_p)\n \treturn NULL;\n }\n \n-struct cache_tree *cache_tree_read(const char *buffer, unsigned long size)\n+struct cache_tree *cache_tree_read(const char *buffer, unsigned long size,\n+\t\t\t\t   struct json_writer *jw)\n {\n+\tstruct cache_tree *ret;\n+\n+\tif (jw) {\n+\t\tjw_object_inline_begin_object(jw, \"root\");\n+\t}\n \tif (buffer[0])\n-\t\treturn NULL; /* not the whole tree */\n-\treturn read_one(&buffer, &size);\n+\t\tret = NULL; /* not the whole tree */\n+\telse\n+\t\tret = read_one(&buffer, &size, jw);\n+\tjw_end_gently(jw);\n+\treturn ret;\n }\n \n static struct cache_tree *cache_tree_find(struct cache_tree *it, const char *path)\ndiff --git a/cache-tree.h b/cache-tree.h\nindex 757bbc48bc..fc3c73284b 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -6,6 +6,8 @@\n #include \"tree-walk.h\"\n \n struct cache_tree;\n+struct json_writer;\n+\n struct cache_tree_sub {\n \tstruct cache_tree *cache_tree;\n \tint count;\t\t/* internally used by update_one() */\n@@ -28,7 +30,8 @@ void cache_tree_invalidate_path(struct index_state *, const char *);\n struct cache_tree_sub *cache_tree_sub(struct cache_tree *, const char *);\n \n void cache_tree_write(struct strbuf *, struct cache_tree *root);\n-struct cache_tree *cache_tree_read(const char *buffer, unsigned long size);\n+struct cache_tree *cache_tree_read(const char *buffer, unsigned long size,\n+\t\t\t\t   struct json_writer *jw);\n \n int cache_tree_fully_valid(struct cache_tree *);\n int cache_tree_update(struct index_state *, int);\ndiff --git a/read-cache.c b/read-cache.c\nindex 4accd8bb08..d09ce42b9a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1716,7 +1716,7 @@ static int read_index_extension(struct index_state *istate,\n \n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n-\t\tistate->cache_tree = cache_tree_read(data, sz);\n+\t\tistate->cache_tree = cache_tree_read(data, sz, istate->jw);\n \t\tbreak;\n \tcase CACHE_EXT_RESOLVE_UNDO:\n \t\tistate->resolve_undo = resolve_undo_read(data, sz);\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex 97bcd814be..fc313f2c9a 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -29,6 +29,8 @@ test_expect_success 'setup' '\n \techo 2 >sub/two &&\n \tgit add sub/two &&\n \n+\tgit commit -m first &&\n+\n \techo intent-to-add >ita &&\n \tgit add -N ita &&\n \n@@ -37,7 +39,7 @@ test_expect_success 'setup' '\n \tstrip_string oid ident\n '\n \n-test_expect_success 'ls-files --json, main entries' '\n+test_expect_success 'ls-files --json, main entries and TREE' '\n \tcompare_json basic\n '\n \ndiff --git a/t/t3011/basic b/t/t3011/basic\nindex 9436445d90..e27f5be5ff 100644\n--- a/t/t3011/basic\n+++ b/t/t3011/basic\n@@ -63,5 +63,23 @@\n       },\n       \"file_offset\": <number>\n     }\n-  ]\n+  ],\n+  \"extensions\": {\n+    \"TREE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"root\": {\n+        \"oid\": null,\n+        \"subdirs\": [\n+          {\n+            \"name\": \"sub\",\n+            \"oid\": <string>,\n+            \"entry_count\": 1,\n+            \"subdirs\": [\n+            ]\n+          }\n+        ]\n+      }\n+    }\n+  }\n }\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377878","messageId":"20190624130226.17293-5-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 04/10] dir.c: dump \"UNTR\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:20Z","receivedAt":"2019-06-24T13:03:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"The big part of UNTR extension is dumped at the end instead of dumping\nas soon as we read it, because we actually \"patch\" some fields in\nuntracked_cache_dir with EWAH bitmaps at the end.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n dir.c                    | 57 +++++++++++++++++++++++++++++++++++++++-\n dir.h                    |  4 ++-\n json-writer.h            |  6 +++++\n read-cache.c             |  2 +-\n t/t3011-ls-files-json.sh |  3 ++-\n t/t3011/basic            | 39 +++++++++++++++++++++++++++\n 6 files changed, 107 insertions(+), 4 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex ba4a51c296..8808577ea3 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -19,6 +19,7 @@\n #include \"varint.h\"\n #include \"ewah/ewok.h\"\n #include \"fsmonitor.h\"\n+#include \"json-writer.h\"\n #include \"submodule-config.h\"\n \n /*\n@@ -2826,7 +2827,42 @@ static void load_oid_stat(struct oid_stat *oid_stat, const unsigned char *data,\n \toid_stat->valid = 1;\n }\n \n-struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz)\n+static void jw_object_oid_stat(struct json_writer *jw, const char *key,\n+\t\t\t       const struct oid_stat *oid_stat)\n+{\n+\tjw_object_inline_begin_object(jw, key);\n+\tjw_object_bool(jw, \"valid\", oid_stat->valid);\n+\tjw_object_string(jw, \"oid\", oid_to_hex(&oid_stat->oid));\n+\tjw_object_stat_data(jw, \"stat\", &oid_stat->stat);\n+\tjw_end(jw);\n+}\n+\n+static void jw_object_untracked_cache_dir(struct json_writer *jw,\n+\t\t\t\t\t  const struct untracked_cache_dir *ucd)\n+{\n+\tint i;\n+\n+\tjw_object_bool(jw, \"valid\", ucd->valid);\n+\tjw_object_bool(jw, \"check-only\", ucd->check_only);\n+\tjw_object_stat_data(jw, \"stat\", &ucd->stat_data);\n+\tjw_object_string(jw, \"exclude-oid\", oid_to_hex(&ucd->exclude_oid));\n+\tjw_object_inline_begin_array(jw, \"untracked\");\n+\tfor (i = 0; i < ucd->untracked_nr; i++)\n+\t\tjw_array_string(jw, ucd->untracked[i]);\n+\tjw_end(jw);\n+\n+\tjw_object_inline_begin_object(jw, \"dirs\");\n+\tfor (i = 0; i < ucd->dirs_nr; i++) {\n+\t\tjw_object_inline_begin_object(jw, ucd->dirs[i]->name);\n+\t\tjw_object_untracked_cache_dir(jw, ucd->dirs[i]);\n+\t\tjw_end(jw);\n+\t}\n+\tjw_end(jw);\n+}\n+\n+struct untracked_cache *read_untracked_extension(const void *data,\n+\t\t\t\t\t\t unsigned long sz,\n+\t\t\t\t\t\t struct json_writer *jw)\n {\n \tstruct untracked_cache *uc;\n \tstruct read_data rd;\n@@ -2864,6 +2900,19 @@ struct untracked_cache *read_untracked_extension(const void *data, unsigned long\n \tuc->dir_flags = get_be32(next + ouc_offset(dir_flags));\n \texclude_per_dir = (const char *)next + exclude_per_dir_offset;\n \tuc->exclude_per_dir = xstrdup(exclude_per_dir);\n+\n+\tif (jw) {\n+\t\tjw_object_string(jw, \"ident\", ident);\n+\t\tjw_object_oid_stat(jw, \"info_exclude\", &uc->ss_info_exclude);\n+\t\tjw_object_oid_stat(jw, \"excludes_file\", &uc->ss_excludes_file);\n+\t\tjw_object_intmax(jw, \"flags\", uc->dir_flags);\n+\t\tif (uc->dir_flags & DIR_SHOW_OTHER_DIRECTORIES)\n+\t\t\tjw_object_bool(jw, \"show_other_directories\", 1);\n+\t\tif (uc->dir_flags & DIR_HIDE_EMPTY_DIRECTORIES)\n+\t\t\tjw_object_bool(jw, \"hide_empty_directories\", 1);\n+\t\tjw_object_string(jw, \"excludes_per_dir\", uc->exclude_per_dir);\n+\t}\n+\n \t/* NUL after exclude_per_dir is covered by sizeof(*ouc) */\n \tnext += exclude_per_dir_offset + strlen(exclude_per_dir) + 1;\n \tif (next >= end)\n@@ -2905,6 +2954,12 @@ struct untracked_cache *read_untracked_extension(const void *data, unsigned long\n \tewah_each_bit(rd.sha1_valid, read_oid, &rd);\n \tnext = rd.data;\n \n+\tif (jw) {\n+\t\tjw_object_inline_begin_object(jw, \"root\");\n+\t\tjw_object_untracked_cache_dir(jw, uc->root);\n+\t\tjw_end(jw);\n+\t}\n+\n done:\n \tfree(rd.ucd);\n \tewah_free(rd.valid);\ndiff --git a/dir.h b/dir.h\nindex 680079bbe3..80efdd05c4 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -6,6 +6,8 @@\n #include \"cache.h\"\n #include \"strbuf.h\"\n \n+struct json_writer;\n+\n struct dir_entry {\n \tunsigned int len;\n \tchar name[FLEX_ARRAY]; /* more */\n@@ -362,7 +364,7 @@ void untracked_cache_remove_from_index(struct index_state *, const char *);\n void untracked_cache_add_to_index(struct index_state *, const char *);\n \n void free_untracked_cache(struct untracked_cache *);\n-struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz);\n+struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz, struct json_writer *jw);\n void write_untracked_extension(struct strbuf *out, struct untracked_cache *untracked);\n void add_untracked_cache(struct index_state *istate);\n void remove_untracked_cache(struct index_state *istate);\ndiff --git a/json-writer.h b/json-writer.h\nindex c48c4cbf33..c3d0fbd1ef 100644\n--- a/json-writer.h\n+++ b/json-writer.h\n@@ -121,6 +121,12 @@ static inline void jw_object_inline_begin_array_gently(struct json_writer *jw,\n \t\tjw_object_inline_begin_array(jw, name);\n }\n \n+static inline void jw_array_inline_begin_object_gently(struct json_writer *jw)\n+{\n+\tif (jw)\n+\t\tjw_array_inline_begin_object(jw);\n+}\n+\n static inline void jw_end_gently(struct json_writer *jw)\n {\n \tif (jw)\ndiff --git a/read-cache.c b/read-cache.c\nindex d09ce42b9a..a70df4b0a5 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1725,7 +1725,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tret = read_link_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_UNTRACKED:\n-\t\tistate->untracked = read_untracked_extension(data, sz);\n+\t\tistate->untracked = read_untracked_extension(data, sz, istate->jw);\n \t\tbreak;\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex fc313f2c9a..082fe8e966 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -30,6 +30,7 @@ test_expect_success 'setup' '\n \tgit add sub/two &&\n \n \tgit commit -m first &&\n+\tgit update-index --untracked-cache &&\n \n \techo intent-to-add >ita &&\n \tgit add -N ita &&\n@@ -39,7 +40,7 @@ test_expect_success 'setup' '\n \tstrip_string oid ident\n '\n \n-test_expect_success 'ls-files --json, main entries and TREE' '\n+test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n \tcompare_json basic\n '\n \ndiff --git a/t/t3011/basic b/t/t3011/basic\nindex e27f5be5ff..8e049f5350 100644\n--- a/t/t3011/basic\n+++ b/t/t3011/basic\n@@ -80,6 +80,45 @@\n           }\n         ]\n       }\n+    },\n+    \"UNTR\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"ident\": <string>,\n+      \"info_exclude\": {\n+        \"valid\": true,\n+        \"oid\": <string>,\n+        \"stat\": {\n+          \"ctime_sec\": <number>,\n+          \"ctime_nsec\": <number>,\n+          \"mtime_sec\": <number>,\n+          \"mtime_nsec\": <number>,\n+          \"device\": <number>,\n+          \"inode\": <number>,\n+          \"uid\": <number>,\n+          \"gid\": <number>,\n+          \"size\": 0\n+        }\n+      },\n+      \"excludes_file\": {\n+        \"valid\": true,\n+        \"oid\": <string>,\n+        \"stat\": {\n+          \"ctime_sec\": <number>,\n+          \"ctime_nsec\": <number>,\n+          \"mtime_sec\": <number>,\n+          \"mtime_nsec\": <number>,\n+          \"device\": <number>,\n+          \"inode\": <number>,\n+          \"uid\": <number>,\n+          \"gid\": <number>,\n+          \"size\": 0\n+        }\n+      },\n+      \"flags\": 6,\n+      \"show_other_directories\": true,\n+      \"hide_empty_directories\": true,\n+      \"excludes_per_dir\": \".gitignore\"\n     }\n   }\n }\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377879","messageId":"20190624130226.17293-6-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:21Z","receivedAt":"2019-06-24T13:03:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n json-writer.c             | 14 ++++++++++++++\n json-writer.h             |  3 +++\n split-index.c             |  9 ++++++++-\n t/t3011-ls-files-json.sh  | 14 ++++++++++++++\n t/t3011/split-index (new) | 39 +++++++++++++++++++++++++++++++++++++++\n 5 files changed, 78 insertions(+), 1 deletion(-)\n\ndiff --git a/json-writer.c b/json-writer.c\nindex 0608726512..c0bd302e4e 100644\n--- a/json-writer.c\n+++ b/json-writer.c\n@@ -1,4 +1,5 @@\n #include \"cache.h\"\n+#include \"ewah/ewok.h\"\n #include \"json-writer.h\"\n \n void jw_init(struct json_writer *jw)\n@@ -224,6 +225,19 @@ void jw_object_stat_data(struct json_writer *jw, const char *name,\n \tjw_end(jw);\n }\n \n+static void dump_ewah_one(size_t pos, void *jw)\n+{\n+\tjw_array_intmax(jw, pos);\n+}\n+\n+void jw_object_ewah(struct json_writer *jw, const char *key,\n+\t\t    struct ewah_bitmap *ewah)\n+{\n+\tjw_object_inline_begin_array(jw, key);\n+\tewah_each_bit(ewah, dump_ewah_one, jw);\n+\tjw_end(jw);\n+}\n+\n static void increase_indent(struct strbuf *sb,\n \t\t\t    const struct json_writer *jw,\n \t\t\t    int indent)\ndiff --git a/json-writer.h b/json-writer.h\nindex c3d0fbd1ef..07d841d52a 100644\n--- a/json-writer.h\n+++ b/json-writer.h\n@@ -45,6 +45,7 @@\n #include \"git-compat-util.h\"\n #include \"strbuf.h\"\n \n+struct ewah_bitmap;\n struct stat_data;\n \n struct json_writer\n@@ -87,6 +88,8 @@ void jw_object_null(struct json_writer *jw, const char *key);\n void jw_object_filemode(struct json_writer *jw, const char *key, mode_t value);\n void jw_object_stat_data(struct json_writer *jw, const char *key,\n \t\t\t const struct stat_data *sd);\n+void jw_object_ewah(struct json_writer *jw, const char *key,\n+\t\t    struct ewah_bitmap *ewah);\n void jw_object_sub_jw(struct json_writer *jw, const char *key,\n \t\t      const struct json_writer *value);\n \ndiff --git a/split-index.c b/split-index.c\nindex e6154e4ea9..41552bf771 100644\n--- a/split-index.c\n+++ b/split-index.c\n@@ -1,4 +1,5 @@\n #include \"cache.h\"\n+#include \"json-writer.h\"\n #include \"split-index.h\"\n #include \"ewah/ewok.h\"\n \n@@ -25,7 +26,7 @@ int read_link_extension(struct index_state *istate,\n \tdata += the_hash_algo->rawsz;\n \tsz -= the_hash_algo->rawsz;\n \tif (!sz)\n-\t\treturn 0;\n+\t\tgoto done;\n \tsi->delete_bitmap = ewah_new();\n \tret = ewah_read_mmap(si->delete_bitmap, data, sz);\n \tif (ret < 0)\n@@ -38,6 +39,12 @@ int read_link_extension(struct index_state *istate,\n \t\treturn error(\"corrupt replace bitmap in link extension\");\n \tif (ret != sz)\n \t\treturn error(\"garbage at the end of link extension\");\n+done:\n+\tif (istate->jw) {\n+\t\tjw_object_string(istate->jw, \"oid\", oid_to_hex(&si->base_oid));\n+\t\tjw_object_ewah(istate->jw, \"delete_bitmap\", si->delete_bitmap);\n+\t\tjw_object_ewah(istate->jw, \"replace_bitmap\", si->replace_bitmap);\n+\t}\n \treturn 0;\n }\n \ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex 082fe8e966..dbb572ce9d 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -44,4 +44,18 @@ test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n \tcompare_json basic\n '\n \n+test_expect_success 'ls-files --json, split index' '\n+\tgit init split &&\n+\t(\n+\t\tcd split &&\n+\t\techo one >one &&\n+\t\tgit add one &&\n+\t\tgit update-index --split-index &&\n+\t\techo updated >>one &&\n+\t\ttest_must_fail git -c splitIndex.maxPercentChange=100 update-index --refresh &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json split-index\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t3011/split-index b/t/t3011/split-index\nnew file mode 100644\nindex 0000000000..cdcc4ddded\n--- /dev/null\n+++ b/t/t3011/split-index\n@@ -0,0 +1,39 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"link\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"oid\": <string>,\n+      \"delete_bitmap\": [\n+      ],\n+      \"replace_bitmap\": [\n+        0\n+      ]\n+    }\n+  }\n+}\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377880","messageId":"20190624130226.17293-7-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 06/10] fsmonitor.c: dump \"FSMN\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:22Z","receivedAt":"2019-06-24T13:03:20Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n fsmonitor.c              |  6 ++++++\n t/t3011-ls-files-json.sh | 14 +++++++++++++-\n t/t3011/fsmonitor (new)  | 38 ++++++++++++++++++++++++++++++++++++++\n 3 files changed, 57 insertions(+), 1 deletion(-)\n\ndiff --git a/fsmonitor.c b/fsmonitor.c\nindex 1dee0aded1..5ed55ad176 100644\n--- a/fsmonitor.c\n+++ b/fsmonitor.c\n@@ -3,6 +3,7 @@\n #include \"dir.h\"\n #include \"ewah/ewok.h\"\n #include \"fsmonitor.h\"\n+#include \"json-writer.h\"\n #include \"run-command.h\"\n #include \"strbuf.h\"\n \n@@ -50,6 +51,11 @@ int read_fsmonitor_extension(struct index_state *istate, const void *data,\n \t}\n \tistate->fsmonitor_dirty = fsmonitor_dirty;\n \n+\tif (istate->jw) {\n+\t\tjw_object_intmax(istate->jw, \"version\", hdr_version);\n+\t\tjw_object_intmax(istate->jw, \"last_update\", istate->fsmonitor_last_update);\n+\t\tjw_object_ewah(istate->jw, \"dirty\", fsmonitor_dirty);\n+\t}\n \ttrace_printf_key(&trace_fsmonitor, \"read fsmonitor extension successful\");\n \treturn 0;\n }\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex dbb572ce9d..25215f83ae 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -36,7 +36,7 @@ test_expect_success 'setup' '\n \tgit add -N ita &&\n \n \tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n-\tstrip_number device inode uid gid file_offset ext_size &&\n+\tstrip_number device inode uid gid file_offset ext_size last_update &&\n \tstrip_string oid ident\n '\n \n@@ -58,4 +58,16 @@ test_expect_success 'ls-files --json, split index' '\n \t)\n '\n \n+test_expect_success 'ls-files --json, fsmonitor extension ' '\n+\tgit init fsmonitor &&\n+\t(\n+\t\tcd fsmonitor &&\n+\t\techo one >one &&\n+\t\tgit add one &&\n+\t\tgit update-index --fsmonitor &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json fsmonitor\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t3011/fsmonitor b/t/t3011/fsmonitor\nnew file mode 100644\nindex 0000000000..17f2d4a664\n--- /dev/null\n+++ b/t/t3011/fsmonitor\n@@ -0,0 +1,38 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 4\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"FSMN\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"version\": 1,\n+      \"last_update\": <number>,\n+      \"dirty\": [\n+        0\n+      ]\n+    }\n+  }\n+}\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377881","messageId":"20190624130226.17293-8-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 07/10] resolve-undo.c: dump \"REUC\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:23Z","receivedAt":"2019-06-24T13:03:26Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c             |  2 +-\n resolve-undo.c           | 30 +++++++++++++++++-\n resolve-undo.h           |  4 ++-\n t/t3011-ls-files-json.sh | 20 ++++++++++++\n t/t3011/rerere (new)     | 66 ++++++++++++++++++++++++++++++++++++++++\n 5 files changed, 119 insertions(+), 3 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex a70df4b0a5..e5183636fc 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1719,7 +1719,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tistate->cache_tree = cache_tree_read(data, sz, istate->jw);\n \t\tbreak;\n \tcase CACHE_EXT_RESOLVE_UNDO:\n-\t\tistate->resolve_undo = resolve_undo_read(data, sz);\n+\t\tistate->resolve_undo = resolve_undo_read(data, sz, istate->jw);\n \t\tbreak;\n \tcase CACHE_EXT_LINK:\n \t\tret = read_link_extension(istate, data, sz);\ndiff --git a/resolve-undo.c b/resolve-undo.c\nindex 236320f179..68921e3dfe 100644\n--- a/resolve-undo.c\n+++ b/resolve-undo.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"dir.h\"\n+#include \"json-writer.h\"\n #include \"resolve-undo.h\"\n #include \"string-list.h\"\n \n@@ -49,7 +50,30 @@ void resolve_undo_write(struct strbuf *sb, struct string_list *resolve_undo)\n \t}\n }\n \n-struct string_list *resolve_undo_read(const char *data, unsigned long size)\n+static void dump_resolve_undo(struct json_writer *jw,\n+\t\t\t      const char *path,\n+\t\t\t      const struct resolve_undo_info *ui)\n+{\n+\tint i;\n+\n+\tif (!jw)\n+\t\treturn;\n+\n+\tjw_array_inline_begin_object(jw);\n+\tjw_object_string(jw, \"path\", path);\n+\n+\tjw_object_inline_begin_array(jw, \"stages\");\n+\tfor (i = 0; i < 3; i++) {\n+\t\tjw_array_inline_begin_object(jw);\n+\t\tjw_object_filemode(jw, \"mode\", ui->mode[i]);\n+\t\tjw_object_string(jw, \"oid\", oid_to_hex(&ui->oid[i]));\n+\t\tjw_end(jw);\n+\t}\n+\tjw_end(jw);\n+}\n+\n+struct string_list *resolve_undo_read(const char *data, unsigned long size,\n+\t\t\t\t      struct json_writer *jw)\n {\n \tstruct string_list *resolve_undo;\n \tsize_t len;\n@@ -59,6 +83,7 @@ struct string_list *resolve_undo_read(const char *data, unsigned long size)\n \n \tresolve_undo = xcalloc(1, sizeof(*resolve_undo));\n \tresolve_undo->strdup_strings = 1;\n+\tjw_object_inline_begin_array_gently(jw, \"entries\");\n \n \twhile (size) {\n \t\tstruct string_list_item *lost;\n@@ -94,7 +119,10 @@ struct string_list *resolve_undo_read(const char *data, unsigned long size)\n \t\t\tsize -= rawsz;\n \t\t\tdata += rawsz;\n \t\t}\n+\n+\t\tdump_resolve_undo(jw, lost->string, ui);\n \t}\n+\tjw_end_gently(jw);\n \treturn resolve_undo;\n \n error:\ndiff --git a/resolve-undo.h b/resolve-undo.h\nindex 2b3f0f901e..46b4e93a7e 100644\n--- a/resolve-undo.h\n+++ b/resolve-undo.h\n@@ -3,6 +3,8 @@\n \n #include \"cache.h\"\n \n+struct json_writer;\n+\n struct resolve_undo_info {\n \tunsigned int mode[3];\n \tstruct object_id oid[3];\n@@ -10,7 +12,7 @@ struct resolve_undo_info {\n \n void record_resolve_undo(struct index_state *, struct cache_entry *);\n void resolve_undo_write(struct strbuf *, struct string_list *);\n-struct string_list *resolve_undo_read(const char *, unsigned long);\n+struct string_list *resolve_undo_read(const char *, unsigned long, struct json_writer *);\n void resolve_undo_clear_index(struct index_state *);\n int unmerge_index_entry_at(struct index_state *, int);\n void unmerge_index(struct index_state *, const struct pathspec *);\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex 25215f83ae..dc57138f5b 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -70,4 +70,24 @@ test_expect_success 'ls-files --json, fsmonitor extension ' '\n \t)\n '\n \n+test_expect_success 'ls-files --json, rerere extension' '\n+\tgit init rerere &&\n+\t(\n+\t\tcd rerere &&\n+\t\tmkdir fi &&\n+\t\ttest_commit initial fi/le first &&\n+\t\tgit branch side &&\n+\t\ttest_commit second fi/le second &&\n+\t\tgit checkout side &&\n+\t\ttest_commit third fi/le third &&\n+\t\tgit checkout master &&\n+\t\tgit config rerere.enabled true &&\n+\t\ttest_must_fail git merge side &&\n+\t\techo resolved >fi/le &&\n+\t\tgit add fi/le &&\n+\t\tcp ../filter.sed . &&\n+\t\tcompare_json rerere\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t3011/rerere b/t/t3011/rerere\nnew file mode 100644\nindex 0000000000..a8ec4b16ee\n--- /dev/null\n+++ b/t/t3011/rerere\n@@ -0,0 +1,66 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"fi/le\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 9\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"TREE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"root\": {\n+        \"oid\": null,\n+        \"subdirs\": [\n+          {\n+            \"name\": \"fi\",\n+            \"oid\": null,\n+            \"subdirs\": [\n+            ]\n+          }\n+        ]\n+      }\n+    },\n+    \"REUC\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"entries\": [\n+        {\n+          \"path\": \"fi/le\",\n+          \"stages\": [\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            },\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            },\n+            {\n+              \"mode\": \"100644\",\n+              \"oid\": <string>\n+            }\n+          ]\n+        }\n+      ]\n+    }\n+  }\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377883","messageId":"20190624130226.17293-10-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 09/10] read-cache.c: dump \"IEOT\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:25Z","receivedAt":"2019-06-24T13:03:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 43 ++++++++++++++++++++++++++++++++++++-------\n t/t3011/eoie | 13 ++++++++++++-\n 2 files changed, 48 insertions(+), 8 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 37491dd03d..c26edcc9d9 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1693,6 +1693,7 @@ static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n \treturn 0;\n }\n \n+static struct index_entry_offset_table *do_read_ieot_extension(struct index_state *, const char *, uint32_t);\n static int read_index_extension(struct index_state *istate,\n \t\t\t\tconst char *map,\n \t\t\t\tunsigned long *offset)\n@@ -1740,7 +1741,11 @@ static int read_index_extension(struct index_state *istate,\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\n \tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n-\t\t/* already handled in do_read_index() */\n+\t\tif (istate->jw) {\n+\t\t\tfree(do_read_ieot_extension(istate, data, sz));\n+\t\t} else {\n+\t\t\t/* already handled in do_read_index() */\n+\t\t}\n \t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n@@ -1938,7 +1943,7 @@ struct index_entry_offset_table\n \tstruct index_entry_offset entries[FLEX_ARRAY];\n };\n \n-static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset);\n+static struct index_entry_offset_table *read_ieot_extension(struct index_state *istate, const char *mmap, size_t mmap_size, size_t offset);\n static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot);\n \n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n@@ -2292,7 +2297,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t * to multi-thread the reading of the cache entries.\n \t */\n \tif (extension_offset && nr_threads > 1)\n-\t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n+\t\tieot = read_ieot_extension(istate, mmap, mmap_size, extension_offset);\n \n \tif (ieot) {\n \t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, nr_threads, ieot);\n@@ -3634,12 +3639,13 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n \n #define IEOT_VERSION\t(1)\n \n-static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n+static struct index_entry_offset_table *read_ieot_extension(\n+\t\tstruct index_state *istate,\n+\t\tconst char *mmap, size_t mmap_size,\n+\t\tsize_t offset)\n {\n \tconst char *index = NULL;\n-\tuint32_t extsize, ext_version;\n-\tstruct index_entry_offset_table *ieot;\n-\tint i, nr;\n+\tuint32_t extsize;\n \n \t/* find the IEOT extension */\n \tif (!offset)\n@@ -3655,6 +3661,17 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n \t}\n \tif (!index)\n \t\treturn NULL;\n+\treturn do_read_ieot_extension(istate, index, extsize);\n+}\n+\n+static struct index_entry_offset_table *do_read_ieot_extension(\n+\t\tstruct index_state *istate,\n+\t\tconst char *index,\n+\t\tuint32_t extsize)\n+{\n+\tuint32_t ext_version;\n+\tstruct index_entry_offset_table *ieot;\n+\tint i, nr;\n \n \t/* validate the version is IEOT_VERSION */\n \text_version = get_be32(index);\n@@ -3670,6 +3687,10 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n \t\terror(\"invalid number of IEOT entries %d\", nr);\n \t\treturn NULL;\n \t}\n+\tif (istate->jw) {\n+\t\tjw_object_intmax(istate->jw, \"version\", ext_version);\n+\t\tjw_object_inline_begin_array(istate->jw, \"entries\");\n+\t}\n \tieot = xmalloc(sizeof(struct index_entry_offset_table)\n \t\t       + (nr * sizeof(struct index_entry_offset)));\n \tieot->nr = nr;\n@@ -3678,7 +3699,15 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n \t\tindex += sizeof(uint32_t);\n \t\tieot->entries[i].nr = get_be32(index);\n \t\tindex += sizeof(uint32_t);\n+\n+\t\tif (istate->jw) {\n+\t\t\tjw_array_inline_begin_object(istate->jw);\n+\t\t\tjw_object_intmax(istate->jw, \"offset\", ieot->entries[i].offset);\n+\t\t\tjw_object_intmax(istate->jw, \"count\", ieot->entries[i].nr);\n+\t\t\tjw_end(istate->jw);\n+\t\t}\n \t}\n+\tjw_end_gently(istate->jw);\n \n \treturn ieot;\n }\ndiff --git a/t/t3011/eoie b/t/t3011/eoie\nindex 85ec61517b..66a0feb3b6 100644\n--- a/t/t3011/eoie\n+++ b/t/t3011/eoie\n@@ -84,7 +84,18 @@\n   \"extensions\": {\n     \"IEOT\": {\n       \"file_offset\": <number>,\n-      \"ext_size\": <number>\n+      \"ext_size\": <number>,\n+      \"version\": 1,\n+      \"entries\": [\n+        {\n+          \"offset\": <number>,\n+          \"count\": 2\n+        },\n+        {\n+          \"offset\": <number>,\n+          \"count\": 2\n+        }\n+      ]\n     },\n     \"EOIE\": {\n       \"file_offset\": <number>,\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377884","messageId":"20190624130226.17293-9-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 08/10] read-cache.c: dump \"EOIE\" extension as json","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:24Z","receivedAt":"2019-06-24T13:03:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c             |  8 ++++\n t/t3011-ls-files-json.sh | 13 ++++++\n t/t3011/eoie (new)       | 96 ++++++++++++++++++++++++++++++++++++++++\n t/test-lib.sh            |  4 ++\n 4 files changed, 121 insertions(+)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex e5183636fc..37491dd03d 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1731,6 +1731,14 @@ static int read_index_extension(struct index_state *istate,\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\tif (istate->jw) {\n+\t\t\t/* must be synchronized with read_eoie_extension() */\n+\t\t\tjw_object_intmax(istate->jw, \"offset\", get_be32(data));\n+\t\t\tjw_object_string(istate->jw, \"oid\",\n+\t\t\t\t\t hash_to_hex((const unsigned char*)data + sizeof(uint32_t)));\n+\t\t}\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex dc57138f5b..9f4ad4c9cf 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -90,4 +90,17 @@ test_expect_success 'ls-files --json, rerere extension' '\n \t)\n '\n \n+test_expect_success !SINGLE_CPU 'ls-files --json and multicore extensions' '\n+\tgit init eoie &&\n+\t(\n+\t\tcd eoie &&\n+\t\tgit config index.threads 2 &&\n+\t\ttouch one two three four &&\n+\t\tgit add . &&\n+\t\tcp ../filter.sed . &&\n+\t\tstrip_number offset &&\n+\t\tcompare_json eoie\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t3011/eoie b/t/t3011/eoie\nnew file mode 100644\nindex 0000000000..85ec61517b\n--- /dev/null\n+++ b/t/t3011/eoie\n@@ -0,0 +1,96 @@\n+{\n+  \"version\": 2,\n+  \"oid\": <string>,\n+  \"mtime_sec\": <number>,\n+  \"mtime_nsec\": <number>,\n+  \"entries\": [\n+    {\n+      \"id\": 0,\n+      \"name\": \"four\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 1,\n+      \"name\": \"one\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 2,\n+      \"name\": \"three\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    },\n+    {\n+      \"id\": 3,\n+      \"name\": \"two\",\n+      \"mode\": \"100644\",\n+      \"flags\": 0,\n+      \"oid\": <string>,\n+      \"stat\": {\n+        \"ctime_sec\": <number>,\n+        \"ctime_nsec\": <number>,\n+        \"mtime_sec\": <number>,\n+        \"mtime_nsec\": <number>,\n+        \"device\": <number>,\n+        \"inode\": <number>,\n+        \"uid\": <number>,\n+        \"gid\": <number>,\n+        \"size\": 0\n+      },\n+      \"file_offset\": <number>\n+    }\n+  ],\n+  \"extensions\": {\n+    \"IEOT\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>\n+    },\n+    \"EOIE\": {\n+      \"file_offset\": <number>,\n+      \"ext_size\": <number>,\n+      \"offset\": <number>,\n+      \"oid\": <string>\n+    }\n+  }\n+}\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex 4b346467df..9d5b273b40 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1611,3 +1611,7 @@ test_lazy_prereq REBASE_P '\n test_lazy_prereq FAIL_PREREQS '\n \ttest -n \"$GIT_TEST_FAIL_PREREQS\"\n '\n+\n+test_lazy_prereq SINGLE_CPU '\n+\ttest \"$(test-tool online-cpus)\" -eq 1\n+'\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377885","messageId":"20190624130226.17293-11-pclouds@gmail.com","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"[PATCH v2 10/10] t3008: use the new SINGLE_CPU prereq","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-24T13:02:26Z","receivedAt":"2019-06-24T13:03:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3008-ls-files-lazy-init-name-hash.sh | 8 +-------\n 1 file changed, 1 insertion(+), 7 deletions(-)\n\ndiff --git a/t/t3008-ls-files-lazy-init-name-hash.sh b/t/t3008-ls-files-lazy-init-name-hash.sh\nindex 64f047332b..7f918c05f6 100755\n--- a/t/t3008-ls-files-lazy-init-name-hash.sh\n+++ b/t/t3008-ls-files-lazy-init-name-hash.sh\n@@ -4,15 +4,9 @@ test_description='Test the lazy init name hash with various folder structures'\n \n . ./test-lib.sh\n \n-if test 1 -eq $($GIT_BUILD_DIR/t/helper/test-tool online-cpus)\n-then\n-\tskip_all='skipping lazy-init tests, single cpu'\n-\ttest_done\n-fi\n-\n LAZY_THREAD_COST=2000\n \n-test_expect_success 'no buffer overflow in lazy_init_name_hash' '\n+test_expect_success !SINGLE_CPU 'no buffer overflow in lazy_init_name_hash' '\n \t(\n \t    test_seq $LAZY_THREAD_COST | sed \"s/^/a_/\" &&\n \t    echo b/b/b &&\n-- \n2.22.0.rc0.322.g2b0371e29a\n\n"},{"id":"377907","messageId":"nycvar.QRO.7.76.6.1906241954290.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"20190624130226.17293-1-pclouds@gmail.com","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-24T18:00:20Z","receivedAt":"2019-06-24T18:00:12Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Mon, 24 Jun 2019, Nguyễn Thái Ngọc Duy wrote:\n\n> - json field names now use '_' instead of '.' to be friendlier to some\n>   languages. I stick to underscore_name instead of camelCase because\n>   the former is closer to what we use\n\nThis is not a good reason. People who are used to read JSON will stumble\nover this all the time because it is so uncommon.\n\n> - extension location is printed, in case you need to decode the\n>   extension by yourself (previously only the size is printed)\n> - all extensions are printed in the same order they appear in the file\n>   (previously eoie and ieot are printed first because that's how we\n>   parse)\n> - resolve undo extension is reorganized a bit to be easier to read\n> - tests added. Example json files are in t/t3011\n\nIt might actually make sense to optionally disable showing extensions.\n\nYou also forgot to mention that you explicitly disable handling\n`<pathspec>`, which I find a bit odd, personally, as that would probably\ncome in real handy at times, especially when we offer this as a better way\nfor 3rd-party applications to interact with Git (which I think will be the\nuse case for this feature that will be _far_ more common than using it for\ndebugging).\n\nCiao,\nDscho\n"},{"id":"377917","messageId":"0367673b-aa5a-49de-87f6-d52beb1af4c4@jeffhostetler.com","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906241954290.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-24T18:39:19Z","receivedAt":"2019-06-24T18:39:23Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/24/2019 2:00 PM, Johannes Schindelin wrote:\n> Hi Duy,\n> \n> On Mon, 24 Jun 2019, Nguyễn Thái Ngọc Duy wrote:\n> \n>> - json field names now use '_' instead of '.' to be friendlier to some\n>>    languages. I stick to underscore_name instead of camelCase because\n>>    the former is closer to what we use\n> \n> This is not a good reason. People who are used to read JSON will stumble\n> over this all the time because it is so uncommon.\n\nGetting rid of \".\" and \"-\" in field names is the important part.\nThese confuse some languages and make us drop into object[\"<field>\"]\nsyntax in my experience.\n\nAs for \"_\" or camelCase (or PascalCase), I'm not sure it matters one\nway or the other.  Personally, I'd vote for underscores.\n\nJeff\n"},{"id":"377924","messageId":"755a4cfe-fd6b-044b-dca2-05eebfa518b1@jeffhostetler.com","threadId":"51372","inReplyTo":"20190624130226.17293-2-pclouds@gmail.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-24T19:15:54Z","receivedAt":"2019-06-24T19:15:58Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/24/2019 9:02 AM, Nguyễn Thái Ngọc Duy wrote:\n> So far we don't have a command to basically dump the index file out,\n> with all its glory details. Checking some info, for example, stat\n> time, usually involves either writing new code or firing up \"xxd\" and\n> decoding values by yourself.\n> \n> This --json is supposed to help that. It dumps the index in a human\n> readable format but also easy to be processed with tools. And it will\n> print almost enough info to reconstruct the index later.\n> \n> In this patch we only dump the main part, not extensions. But at the\n> end of the series, the entire index is dumped. The end result could be\n> very verbose even on a small repository such as git.git.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n\n...\n\n> diff --git a/json-writer.c b/json-writer.c\n> index aadb9dbddc..0608726512 100644\n> --- a/json-writer.c\n> +++ b/json-writer.c\n> @@ -202,6 +202,28 @@ void jw_object_null(struct json_writer *jw, const char *key)\n>   \tstrbuf_addstr(&jw->json, \"null\");\n>   }\n>   \n> +void jw_object_filemode(struct json_writer *jw, const char *key, mode_t mode)\n> +{\n> +\tobject_common(jw, key);\n> +\tstrbuf_addf(&jw->json, \"\\\"%06o\\\"\", mode);\n> +}\n> +\n> +void jw_object_stat_data(struct json_writer *jw, const char *name,\n> +\t\t\t const struct stat_data *sd)\n\nShould this be in json_writer.c or in read-cache.c ?\nCurrently, json_writer.c is concerned with formatting\nJSON on basic/scalar types.  Do we want to start\nextending it to handle arbitrary structures?  Or would\nit be better for the code that defines/manipulates the\nstructure to define a \"stat_data_dump_json()\" function.\n\nI'm torn on the jw_object_filemode() function, JSON format\nlimits us to decimal integers and there are places where\nI'd like to have hex, or in this case octal values.\n\nI'm thinking it'd be better to have a helper function in\nread-cache.c that formats a local strbuf and calls\njs_object_string(&jw, key, buf);\n\n\n> +{\n> +\tjw_object_inline_begin_object(jw, name);\n> +\tjw_object_intmax(jw, \"ctime_sec\", sd->sd_ctime.sec);\n> +\tjw_object_intmax(jw, \"ctime_nsec\", sd->sd_ctime.nsec);\n> +\tjw_object_intmax(jw, \"mtime_sec\", sd->sd_mtime.sec);\n> +\tjw_object_intmax(jw, \"mtime_nsec\", sd->sd_mtime.nsec);\n\nIt'd be nice if we could also have a formatted date\nfor the mtime and ctime in addition to the integer\nvalues.  (I'm not sure whether you'd always want them\nor make it a verbose option.)\n\n> +\tjw_object_intmax(jw, \"device\", sd->sd_dev);\n> +\tjw_object_intmax(jw, \"inode\", sd->sd_ino);\n> +\tjw_object_intmax(jw, \"uid\", sd->sd_uid);\n> +\tjw_object_intmax(jw, \"gid\", sd->sd_gid);\n> +\tjw_object_intmax(jw, \"size\", sd->sd_size);\n> +\tjw_end(jw);\n> +}\n> +\n>   static void increase_indent(struct strbuf *sb,\n>   \t\t\t    const struct json_writer *jw,\n>   \t\t\t    int indent)\n> diff --git a/json-writer.h b/json-writer.h\n> index 83906b09c1..c48c4cbf33 100644\n> --- a/json-writer.h\n> +++ b/json-writer.h\n> @@ -42,8 +42,11 @@\n>    * of the given strings.\n>    */\n>   \n> +#include \"git-compat-util.h\"\n>   #include \"strbuf.h\"\n>   \n> +struct stat_data;\n> +\n>   struct json_writer\n>   {\n>   \t/*\n> @@ -81,6 +84,9 @@ void jw_object_true(struct json_writer *jw, const char *key);\n>   void jw_object_false(struct json_writer *jw, const char *key);\n>   void jw_object_bool(struct json_writer *jw, const char *key, int value);\n>   void jw_object_null(struct json_writer *jw, const char *key);\n> +void jw_object_filemode(struct json_writer *jw, const char *key, mode_t value);\n> +void jw_object_stat_data(struct json_writer *jw, const char *key,\n> +\t\t\t const struct stat_data *sd);\n>   void jw_object_sub_jw(struct json_writer *jw, const char *key,\n>   \t\t      const struct json_writer *value);\n>   \n> @@ -104,4 +110,21 @@ void jw_array_inline_begin_array(struct json_writer *jw);\n>   int jw_is_terminated(const struct json_writer *jw);\n>   void jw_end(struct json_writer *jw);\n>   \n> +/*\n> + * These _gently versions accept NULL json_writer to reduce too much\n> + * branching at the call site.\n> + */\n> +static inline void jw_object_inline_begin_array_gently(struct json_writer *jw,\n> +\t\t\t\t\t\t       const char *name)\n> +{\n> +\tif (jw)\n> +\t\tjw_object_inline_begin_array(jw, name);\n> +}\n> +\n> +static inline void jw_end_gently(struct json_writer *jw)\n> +{\n> +\tif (jw)\n> +\t\tjw_end(jw);\n> +}\n> +\n>   #endif /* JSON_WRITER_H */\n> diff --git a/read-cache.c b/read-cache.c\n> index 4dd22f4f6e..db5147d088 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -25,6 +25,7 @@\n>   #include \"fsmonitor.h\"\n>   #include \"thread-utils.h\"\n>   #include \"progress.h\"\n> +#include \"json-writer.h\"\n>   \n>   /* Mask for the name length in ce_flags in the on-disk index */\n>   \n> @@ -1952,6 +1953,49 @@ static void *load_index_extensions(void *_data)\n>   \treturn NULL;\n>   }\n>   \n> +static void dump_cache_entry(struct index_state *istate,\n> +\t\t\t     int index,\n> +\t\t\t     unsigned long offset,\n> +\t\t\t     const struct cache_entry *ce)\n> +{\n> +\tstruct json_writer *jw = istate->jw;\n> +\n> +\tjw_array_inline_begin_object(jw);\n> +\n> +\t/*\n> +\t * this is technically redundant, but it's for easier\n> +\t * navigation when there hundreds of entries\n> +\t */\n> +\tjw_object_intmax(jw, \"id\", index);\n> +\n> +\tjw_object_string(jw, \"name\", ce->name);\n> +\n> +\tjw_object_filemode(jw, \"mode\", ce->ce_mode);\n> +\n> +\tjw_object_intmax(jw, \"flags\", ce->ce_flags);\n\nIt would be nice to have the flags as a hex-formatted string\nin addition to (or instead of) the decimal integer value.\n\n> +\t/*\n> +\t * again redundant info, just so you don't have to decode\n> +\t * flags values manually\n> +\t */\n> +\tif (ce->ce_flags & CE_EXTENDED)\n> +\t\tjw_object_true(jw, \"extended_flags\");\n> +\tif (ce->ce_flags & CE_VALID)\n> +\t\tjw_object_true(jw, \"assume_unchanged\");\n> +\tif (ce->ce_flags & CE_INTENT_TO_ADD)\n> +\t\tjw_object_true(jw, \"intent_to_add\");\n> +\tif (ce->ce_flags & CE_SKIP_WORKTREE)\n> +\t\tjw_object_true(jw, \"skip_worktree\");\n> +\tif (ce_stage(ce))\n> +\t\tjw_object_intmax(jw, \"stage\", ce_stage(ce));\n> +\n> +\tjw_object_string(jw, \"oid\", oid_to_hex(&ce->oid));\n> +\n> +\tjw_object_stat_data(jw, \"stat\", &ce->ce_stat_data);\n> +\tjw_object_intmax(jw, \"file_offset\", offset);\n> +\n> +\tjw_end(jw);\n> +}\n> +\n>   /*\n>    * A helper function that will load the specified range of cache entries\n>    * from the memory mapped file and add them to the given index.\n> @@ -1972,6 +2016,9 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n>   \t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n>   \t\tset_index_entry(istate, i, ce);\n>   \n> +\t\tif (istate->jw)\n> +\t\t\tdump_cache_entry(istate, i, src_offset, ce);\n> +\n>   \t\tsrc_offset += consumed;\n>   \t\tprevious_ce = ce;\n>   \t}\n> @@ -1983,6 +2030,8 @@ static unsigned long load_all_cache_entries(struct index_state *istate,\n>   {\n>   \tunsigned long consumed;\n>   \n> +\tjw_object_inline_begin_array_gently(istate->jw, \"entries\");\n> +\n>   \tif (istate->version == 4) {\n>   \t\tmem_pool_init(&istate->ce_mem_pool,\n>   \t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n> @@ -1993,6 +2042,8 @@ static unsigned long load_all_cache_entries(struct index_state *istate,\n>   \n>   \tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n>   \t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n> +\n> +\tjw_end_gently(istate->jw);\n>   \treturn consumed;\n>   }\n>   \n> @@ -2120,6 +2171,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>   \tsize_t extension_offset = 0;\n>   \tint nr_threads, cpus;\n>   \tstruct index_entry_offset_table *ieot = NULL;\n> +\tint jw_pretty = 1;\n>   \n>   \tif (istate->initialized)\n>   \t\treturn istate->cache_nr;\n> @@ -2154,6 +2206,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>   \tistate->cache_nr = ntohl(hdr->hdr_entries);\n>   \tistate->cache_alloc = alloc_nr(istate->cache_nr);\n>   \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n> +\tistate->timestamp.sec = st.st_mtime;\n> +\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n>   \tistate->initialized = 1;\n>   \n>   \tp.istate = istate;\n> @@ -2176,6 +2230,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>   \tif (!HAVE_THREADS)\n>   \t\tnr_threads = 1;\n>   \n> +\tif (istate->jw) {\n> +\t\tjw_object_begin(istate->jw, jw_pretty);\n> +\t\tjw_object_intmax(istate->jw, \"version\", istate->version);\n> +\t\tjw_object_string(istate->jw, \"oid\", oid_to_hex(&istate->oid));\n> +\t\tjw_object_intmax(istate->jw, \"mtime_sec\", istate->timestamp.sec);\n> +\t\tjw_object_intmax(istate->jw, \"mtime_nsec\", istate->timestamp.nsec);\n\nagain, it would be nice to also have a formated version of the mtime.\n\n> +\n> +\t\t/*\n> +\t\t * Threading may mess up json writing. This is for\n> +\t\t * debugging only, so performance is not a concern.\n> +\t\t */\n> +\t\tnr_threads = 1;\n\nyes. we should turn off threading when dumping to json.\n\n> +\t}\n> +\n>   \tif (nr_threads > 1) {\n>   \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n>   \t\tif (extension_offset) {\n> @@ -2204,8 +2272,6 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>   \t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n>   \t}\n>   \n> -\tistate->timestamp.sec = st.st_mtime;\n> -\tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n>   \n>   \t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n>   \tif (extension_offset) {\n> @@ -2216,6 +2282,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>   \t\tp.src_offset = src_offset;\n>   \t\tload_index_extensions(&p);\n>   \t}\n> +\tjw_end_gently(istate->jw);\n> +\n>   \tmunmap((void *)mmap, mmap_size);\n>   \n>   \t/*\n> \n...\n\nThanks\nJeff\n"},{"id":"377926","messageId":"ec6bdf73-0c05-80cd-ecab-9a6ebf6aad6e@jeffhostetler.com","threadId":"51372","inReplyTo":"20190624130226.17293-5-pclouds@gmail.com","subject":"Re: [PATCH v2 04/10] dir.c: dump \"UNTR\" extension as json","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-24T19:32:54Z","receivedAt":"2019-06-24T19:32:58Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/24/2019 9:02 AM, Nguyễn Thái Ngọc Duy wrote:\n> The big part of UNTR extension is dumped at the end instead of dumping\n> as soon as we read it, because we actually \"patch\" some fields in\n> untracked_cache_dir with EWAH bitmaps at the end.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   dir.c                    | 57 +++++++++++++++++++++++++++++++++++++++-\n>   dir.h                    |  4 ++-\n>   json-writer.h            |  6 +++++\n>   read-cache.c             |  2 +-\n>   t/t3011-ls-files-json.sh |  3 ++-\n>   t/t3011/basic            | 39 +++++++++++++++++++++++++++\n>   6 files changed, 107 insertions(+), 4 deletions(-)\n> \n> diff --git a/dir.c b/dir.c\n> index ba4a51c296..8808577ea3 100644\n> --- a/dir.c\n> +++ b/dir.c\n> @@ -19,6 +19,7 @@\n>   #include \"varint.h\"\n>   #include \"ewah/ewok.h\"\n>   #include \"fsmonitor.h\"\n> +#include \"json-writer.h\"\n>   #include \"submodule-config.h\"\n>   \n>   /*\n> @@ -2826,7 +2827,42 @@ static void load_oid_stat(struct oid_stat *oid_stat, const unsigned char *data,\n>   \toid_stat->valid = 1;\n>   }\n>   \n> -struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz)\n> +static void jw_object_oid_stat(struct json_writer *jw, const char *key,\n> +\t\t\t       const struct oid_stat *oid_stat)\n\nI was hoping to reserve the jw_, jw_object_, and jw_array_ prefixes\nto refer to public functions in the json-writer.[ch] files, rather\nthan private functions elsewhere in the code.  I think it is less\nconfusing that way.  IMHO.\n\nIt's fine to have such helper functions for local structure types,\nbut I think it'd be better for them to not be named like other\npublic functions.\n\n> +{\n> +\tjw_object_inline_begin_object(jw, key);\n> +\tjw_object_bool(jw, \"valid\", oid_stat->valid);\n> +\tjw_object_string(jw, \"oid\", oid_to_hex(&oid_stat->oid));\n> +\tjw_object_stat_data(jw, \"stat\", &oid_stat->stat);\n> +\tjw_end(jw);\n> +}\n> +\n> +static void jw_object_untracked_cache_dir(struct json_writer *jw,\n> +\t\t\t\t\t  const struct untracked_cache_dir *ucd)\n> +{\n> +\tint i;\n> +\n> +\tjw_object_bool(jw, \"valid\", ucd->valid);\n> +\tjw_object_bool(jw, \"check-only\", ucd->check_only);\n> +\tjw_object_stat_data(jw, \"stat\", &ucd->stat_data);\n> +\tjw_object_string(jw, \"exclude-oid\", oid_to_hex(&ucd->exclude_oid));\n> +\tjw_object_inline_begin_array(jw, \"untracked\");\n> +\tfor (i = 0; i < ucd->untracked_nr; i++)\n> +\t\tjw_array_string(jw, ucd->untracked[i]);\n> +\tjw_end(jw);\n> +\n> +\tjw_object_inline_begin_object(jw, \"dirs\");\n> +\tfor (i = 0; i < ucd->dirs_nr; i++) {\n> +\t\tjw_object_inline_begin_object(jw, ucd->dirs[i]->name);\n> +\t\tjw_object_untracked_cache_dir(jw, ucd->dirs[i]);\n> +\t\tjw_end(jw);\n> +\t}\n> +\tjw_end(jw);\n> +}\n> +\n> +struct untracked_cache *read_untracked_extension(const void *data,\n> +\t\t\t\t\t\t unsigned long sz,\n> +\t\t\t\t\t\t struct json_writer *jw)\n>   {\n>   \tstruct untracked_cache *uc;\n>   \tstruct read_data rd;\n> @@ -2864,6 +2900,19 @@ struct untracked_cache *read_untracked_extension(const void *data, unsigned long\n>   \tuc->dir_flags = get_be32(next + ouc_offset(dir_flags));\n>   \texclude_per_dir = (const char *)next + exclude_per_dir_offset;\n>   \tuc->exclude_per_dir = xstrdup(exclude_per_dir);\n> +\n> +\tif (jw) {\n> +\t\tjw_object_string(jw, \"ident\", ident);\n> +\t\tjw_object_oid_stat(jw, \"info_exclude\", &uc->ss_info_exclude);\n> +\t\tjw_object_oid_stat(jw, \"excludes_file\", &uc->ss_excludes_file);\n> +\t\tjw_object_intmax(jw, \"flags\", uc->dir_flags);\n> +\t\tif (uc->dir_flags & DIR_SHOW_OTHER_DIRECTORIES)\n> +\t\t\tjw_object_bool(jw, \"show_other_directories\", 1);\n> +\t\tif (uc->dir_flags & DIR_HIDE_EMPTY_DIRECTORIES)\n> +\t\t\tjw_object_bool(jw, \"hide_empty_directories\", 1);\n> +\t\tjw_object_string(jw, \"excludes_per_dir\", uc->exclude_per_dir);\n> +\t}\n> +\n>   \t/* NUL after exclude_per_dir is covered by sizeof(*ouc) */\n>   \tnext += exclude_per_dir_offset + strlen(exclude_per_dir) + 1;\n>   \tif (next >= end)\n> @@ -2905,6 +2954,12 @@ struct untracked_cache *read_untracked_extension(const void *data, unsigned long\n>   \tewah_each_bit(rd.sha1_valid, read_oid, &rd);\n>   \tnext = rd.data;\n>   \n> +\tif (jw) {\n> +\t\tjw_object_inline_begin_object(jw, \"root\");\n> +\t\tjw_object_untracked_cache_dir(jw, uc->root);\n> +\t\tjw_end(jw);\n> +\t}\n> +\n>   done:\n>   \tfree(rd.ucd);\n>   \tewah_free(rd.valid);\n> diff --git a/dir.h b/dir.h\n> index 680079bbe3..80efdd05c4 100644\n> --- a/dir.h\n> +++ b/dir.h\n> @@ -6,6 +6,8 @@\n>   #include \"cache.h\"\n>   #include \"strbuf.h\"\n>   \n> +struct json_writer;\n> +\n>   struct dir_entry {\n>   \tunsigned int len;\n>   \tchar name[FLEX_ARRAY]; /* more */\n> @@ -362,7 +364,7 @@ void untracked_cache_remove_from_index(struct index_state *, const char *);\n>   void untracked_cache_add_to_index(struct index_state *, const char *);\n>   \n>   void free_untracked_cache(struct untracked_cache *);\n> -struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz);\n> +struct untracked_cache *read_untracked_extension(const void *data, unsigned long sz, struct json_writer *jw);\n>   void write_untracked_extension(struct strbuf *out, struct untracked_cache *untracked);\n>   void add_untracked_cache(struct index_state *istate);\n>   void remove_untracked_cache(struct index_state *istate);\n> diff --git a/json-writer.h b/json-writer.h\n> index c48c4cbf33..c3d0fbd1ef 100644\n> --- a/json-writer.h\n> +++ b/json-writer.h\n> @@ -121,6 +121,12 @@ static inline void jw_object_inline_begin_array_gently(struct json_writer *jw,\n>   \t\tjw_object_inline_begin_array(jw, name);\n>   }\n>   \n> +static inline void jw_array_inline_begin_object_gently(struct json_writer *jw)\n> +{\n> +\tif (jw)\n> +\t\tjw_array_inline_begin_object(jw);\n> +}\n> +\n>   static inline void jw_end_gently(struct json_writer *jw)\n>   {\n>   \tif (jw)\n\nI'm not sure about the need for these _gently versions, but\nmaybe make them macros in json-writer.h\n\nJeff\n"},{"id":"377927","messageId":"xmqq5zouhhrw.fsf@gitster-ct.c.googlers.com","threadId":"51372","inReplyTo":"755a4cfe-fd6b-044b-dca2-05eebfa518b1@jeffhostetler.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-06-24T20:04:19Z","receivedAt":"2019-06-24T20:04:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff Hostetler <git@jeffhostetler.com> writes:\n\n> On 6/24/2019 9:02 AM, Nguyễn Thái Ngọc Duy wrote:\n> ...\n>> +void jw_object_stat_data(struct json_writer *jw, const char *name,\n>> +\t\t\t const struct stat_data *sd)\n>\n> Should this be in json_writer.c or in read-cache.c ?\n> Currently, json_writer.c is concerned with formatting\n> JSON on basic/scalar types.\n\nThat's an interesting question.\n\nIn the longer run, is it reasonable to assume that the current\ncallers of these functions will stay to be the only places that\nwould want to say \"For this path, various attributes like timestamps\nand other stuff are these values\"?  The answer is probably no.\n\nWould it be reasonable to assume that the future callers are all\nlikely to be closely related to what functions that are currently in\nread-cache.c do and would appear only in that file?  I think the\nanswer would also be no.  Functions in entry.c may want to report\n\"I've created/updated these filesystem entities and their stat info\nnow look like so\", for example.  So (assuming that we would want to\nrepresent \"struct stat\" the same way no matter who wants to write it\nout), I think it is reasonable to have this function on the json\nside, not on Git side.\n\nI do not think it is a bad idea to have a layer that sits above\njson-writer.c and below read-cache.c that give us standardized\nmapping from in-core objects (like \"struct stat\") to objects in json\n(they are set of key-value pairs after all).  Call it json-schema.c\nor something?\n\nWhat other \"higher than scaler\" types do we foresee that we'd need\nto serialize in a standardised way?  If we do not have that many\nyet, perhaps a split like that might be a bit premature and we can\nstart by having the function in json-writer.c side, not in the API\nconsumer side.\n\nThanks.\n"},{"id":"377928","messageId":"55f81571-ba45-edcf-49bd-05418cc309c5@jeffhostetler.com","threadId":"51372","inReplyTo":"20190624130226.17293-6-pclouds@gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-24T20:06:24Z","receivedAt":"2019-06-24T20:06:27Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/24/2019 9:02 AM, Nguyễn Thái Ngọc Duy wrote:\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   json-writer.c             | 14 ++++++++++++++\n>   json-writer.h             |  3 +++\n>   split-index.c             |  9 ++++++++-\n>   t/t3011-ls-files-json.sh  | 14 ++++++++++++++\n>   t/t3011/split-index (new) | 39 +++++++++++++++++++++++++++++++++++++++\n>   5 files changed, 78 insertions(+), 1 deletion(-)\n> \n> diff --git a/json-writer.c b/json-writer.c\n> index 0608726512..c0bd302e4e 100644\n> --- a/json-writer.c\n> +++ b/json-writer.c\n> @@ -1,4 +1,5 @@\n>   #include \"cache.h\"\n> +#include \"ewah/ewok.h\"\n>   #include \"json-writer.h\"\n>   \n>   void jw_init(struct json_writer *jw)\n> @@ -224,6 +225,19 @@ void jw_object_stat_data(struct json_writer *jw, const char *name,\n>   \tjw_end(jw);\n>   }\n>   \n> +static void dump_ewah_one(size_t pos, void *jw)\n> +{\n> +\tjw_array_intmax(jw, pos);\n> +}\n> +\n> +void jw_object_ewah(struct json_writer *jw, const char *key,\n> +\t\t    struct ewah_bitmap *ewah)\n> +{\n> +\tjw_object_inline_begin_array(jw, key);\n> +\tewah_each_bit(ewah, dump_ewah_one, jw);\n> +\tjw_end(jw);\n> +}\n> +\n\nAs I said in an earlier commit in this series, I'd prefer\nthat we keep such data-structure-specific helper functions\nin the source file that defines the data structure rather\nthan adding them to json-writer.[ch]\n\nI'm curious how big these EWAHs will be in practice and\nhow useful an array of integers will be (especially as the\npretty format will be one integer per line).  Perhaps it\nwould helpful to have an extended example in one of the\ntests.\n\nWould it be better to have the caller of ewah_each_bit()\nbuild a hex or bit string in a strbuf and then write it\nas a single string?\n\nJeff\n"},{"id":"377946","messageId":"20190625090554.GA2423@hank.intra.tgummerer.com","threadId":"51372","inReplyTo":"20190624130226.17293-2-pclouds@gmail.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2019-06-25T09:05:54Z","receivedAt":"2019-06-25T09:06:00Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 06/24, Nguyễn Thái Ngọc Duy wrote:\n> So far we don't have a command to basically dump the index file out,\n> with all its glory details. Checking some info, for example, stat\n> time, usually involves either writing new code or firing up \"xxd\" and\n> decoding values by yourself.\n> \n> This --json is supposed to help that. It dumps the index in a human\n> readable format but also easy to be processed with tools. And it will\n> print almost enough info to reconstruct the index later.\n> \n> In this patch we only dump the main part, not extensions. But at the\n> end of the series, the entire index is dumped. The end result could be\n> very verbose even on a small repository such as git.git.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  Documentation/git-ls-files.txt    |  5 +++\n>  builtin/ls-files.c                | 38 +++++++++++++---\n>  cache.h                           |  2 +\n>  json-writer.c                     | 22 ++++++++++\n>  json-writer.h                     | 23 ++++++++++\n>  read-cache.c                      | 72 ++++++++++++++++++++++++++++++-\n>  t/t3011-ls-files-json.sh (new +x) | 44 +++++++++++++++++++\n>  t/t3011/basic (new)               | 67 ++++++++++++++++++++++++++++\n>  8 files changed, 265 insertions(+), 8 deletions(-)\n>\n> [...]\n>\n> diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> new file mode 100755\n> index 0000000000..97bcd814be\n> --- /dev/null\n> +++ b/t/t3011-ls-files-json.sh\n> @@ -0,0 +1,44 @@\n> +#!/bin/sh\n> +\n> +test_description='ls-files dumping json'\n> +\n> +. ./test-lib.sh\n> +\n> +strip_number() {\n> +\tfor name; do\n> +\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n> +\tdone\n> +}\n> +\n> +strip_string() {\n> +\tfor name; do\n> +\t\techo 's/\\(\"'$name'\":\\) \".*\"/\\1 <string>/' >>filter.sed\n> +\tdone\n> +}\n> +\n> +compare_json() {\n> +\tgit ls-files --debug-json >json &&\n> +\tsed -f filter.sed json >filtered &&\n> +\ttest_cmp \"$TEST_DIRECTORY\"/t3011/\"$1\" filtered\n> +}\n> +\n> +test_expect_success 'setup' '\n> +\tmkdir sub &&\n> +\techo one >one &&\n> +\tgit add one &&\n> +\techo 2 >sub/two &&\n> +\tgit add sub/two &&\n> +\n> +\techo intent-to-add >ita &&\n> +\tgit add -N ita &&\n> +\n> +\tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n> +\tstrip_number device inode uid gid file_offset ext_size &&\n> +\tstrip_string oid ident\n> +'\n> +\n> +test_expect_success 'ls-files --json, main entries' '\n> +\tcompare_json basic\n> +'\n> +\n> +test_done\n> diff --git a/t/t3011/basic b/t/t3011/basic\n> new file mode 100644\n> index 0000000000..9436445d90\n> --- /dev/null\n> +++ b/t/t3011/basic\n> @@ -0,0 +1,67 @@\n> +{\n> +  \"version\": 3,\n\nThis will break the test suite when 'GIT_TEST_INDEX_VERSION' is set to\n4 for example.  I think this applies to a few other tests in later\npatches as well.\n\n> +  \"oid\": <string>,\n> +  \"mtime_sec\": <number>,\n> +  \"mtime_nsec\": <number>,\n> +  \"entries\": [\n> +    {\n> +      \"id\": 0,\n> +      \"name\": \"ita\",\n> +      \"mode\": \"100644\",\n> +      \"flags\": 536887296,\n> +      \"extended_flags\": true,\n> +      \"intent_to_add\": true,\n> +      \"oid\": <string>,\n> +      \"stat\": {\n> +        \"ctime_sec\": <number>,\n> +        \"ctime_nsec\": <number>,\n> +        \"mtime_sec\": <number>,\n> +        \"mtime_nsec\": <number>,\n> +        \"device\": <number>,\n> +        \"inode\": <number>,\n> +        \"uid\": <number>,\n> +        \"gid\": <number>,\n> +        \"size\": 0\n> +      },\n> +      \"file_offset\": <number>\n> +    },\n> +    {\n> +      \"id\": 1,\n> +      \"name\": \"one\",\n> +      \"mode\": \"100644\",\n> +      \"flags\": 0,\n> +      \"oid\": <string>,\n> +      \"stat\": {\n> +        \"ctime_sec\": <number>,\n> +        \"ctime_nsec\": <number>,\n> +        \"mtime_sec\": <number>,\n> +        \"mtime_nsec\": <number>,\n> +        \"device\": <number>,\n> +        \"inode\": <number>,\n> +        \"uid\": <number>,\n> +        \"gid\": <number>,\n> +        \"size\": 4\n> +      },\n> +      \"file_offset\": <number>\n> +    },\n> +    {\n> +      \"id\": 2,\n> +      \"name\": \"sub/two\",\n> +      \"mode\": \"100644\",\n> +      \"flags\": 0,\n> +      \"oid\": <string>,\n> +      \"stat\": {\n> +        \"ctime_sec\": <number>,\n> +        \"ctime_nsec\": <number>,\n> +        \"mtime_sec\": <number>,\n> +        \"mtime_nsec\": <number>,\n> +        \"device\": <number>,\n> +        \"inode\": <number>,\n> +        \"uid\": <number>,\n> +        \"gid\": <number>,\n> +        \"size\": 2\n> +      },\n> +      \"file_offset\": <number>\n> +    }\n> +  ]\n> +}\n> -- \n> 2.22.0.rc0.322.g2b0371e29a\n> \n"},{"id":"377947","messageId":"CACsJy8BsT-GaVvEmqfk5n1jGmkcLG_bRjqcU0M3yefBmNSxmnA@mail.gmail.com","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906241954290.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-25T09:05:51Z","receivedAt":"2019-06-25T09:06:19Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jun 25, 2019 at 1:00 AM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> > - extension location is printed, in case you need to decode the\n> >   extension by yourself (previously only the size is printed)\n> > - all extensions are printed in the same order they appear in the file\n> >   (previously eoie and ieot are printed first because that's how we\n> >   parse)\n> > - resolve undo extension is reorganized a bit to be easier to read\n> > - tests added. Example json files are in t/t3011\n>\n> It might actually make sense to optionally disable showing extensions.\n>\n> You also forgot to mention that you explicitly disable handling\n> `<pathspec>`, which I find a bit odd, personally, as that would probably\n> come in real handy at times,\n\nNo. I mentioned the land of high level languages before. Filtering in\nany Python, Ruby, Scheme, JavaScript, Java is a piece of cake and much\nmore flexible than pathspec. Even with shell scripts, jq could do a\nmuch better job than pathspec. If you filter by pathspec, good luck\ntrying that on extensions.\n\nIt's the same reason why I will not provide a flexible way to disable\nextensions. I'm not starting a JSON API for Git. I provide an index\nfile in JSON format. You do what you want with it. You have a format\neasy enough to import to native data structures of your favorite\nlanguage.\n\n> especially when we offer this as a better way\n> for 3rd-party applications to interact with Git (which I think will be the\n> use case for this feature that will be _far_ more common than using it for\n> debugging).\n\nWe may have conflicting goals. For me, first priority is the debug\ntool for Git developers. 3rd-party support is a stretch. I could move\nall this back to test-tool, then you can provide a 3rd-party API if\nyou want. Or I'll withdraw this series and go back to my original\nplan.\n-- \nDuy\n"},{"id":"377951","messageId":"nycvar.QRO.7.76.6.1906251116020.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"755a4cfe-fd6b-044b-dca2-05eebfa518b1@jeffhostetler.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T09:21:48Z","receivedAt":"2019-06-25T09:21:39Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jeff,\n\nOn Mon, 24 Jun 2019, Jeff Hostetler wrote:\n\n> On 6/24/2019 9:02 AM, Nguyễn Thái Ngọc Duy wrote:\n> [...]\n> > +{\n> > +\tjw_object_inline_begin_object(jw, name);\n> > +\tjw_object_intmax(jw, \"ctime_sec\", sd->sd_ctime.sec);\n> > +\tjw_object_intmax(jw, \"ctime_nsec\", sd->sd_ctime.nsec);\n> > +\tjw_object_intmax(jw, \"mtime_sec\", sd->sd_mtime.sec);\n> > +\tjw_object_intmax(jw, \"mtime_nsec\", sd->sd_mtime.nsec);\n>\n> It'd be nice if we could also have a formatted date\n> for the mtime and ctime in addition to the integer\n> values.  (I'm not sure whether you'd always want them\n> or make it a verbose option.)\n\nIt would be more consistent with JSON conventions to have only the\nformatted date, preferably in ISO-8601 format [*1*], I would think.\n\nAnd for debugging, it would also make more sense, in my opinion.\n\n> [...]\n> > diff --git a/read-cache.c b/read-cache.c\n> > index 4dd22f4f6e..db5147d088 100644\n> > --- a/read-cache.c\n> > +++ b/read-cache.c\n> > @@ -25,6 +25,7 @@\n> >   #include \"fsmonitor.h\"\n> >   #include \"thread-utils.h\"\n> >   #include \"progress.h\"\n> > +#include \"json-writer.h\"\n> >\n> >   /* Mask for the name length in ce_flags in the on-disk index */\n> >\n> > @@ -1952,6 +1953,49 @@ static void *load_index_extensions(void *_data)\n> >   \treturn NULL;\n> >   }\n> >\n> > +static void dump_cache_entry(struct index_state *istate,\n> > +\t\t\t     int index,\n> > +\t\t\t     unsigned long offset,\n> > +\t\t\t     const struct cache_entry *ce)\n> > +{\n> > +\tstruct json_writer *jw = istate->jw;\n> > +\n> > +\tjw_array_inline_begin_object(jw);\n> > +\n> > +\t/*\n> > +\t * this is technically redundant, but it's for easier\n> > +\t * navigation when there hundreds of entries\n> > +\t */\n> > +\tjw_object_intmax(jw, \"id\", index);\n> > +\n> > +\tjw_object_string(jw, \"name\", ce->name);\n> > +\n> > +\tjw_object_filemode(jw, \"mode\", ce->ce_mode);\n> > +\n> > +\tjw_object_intmax(jw, \"flags\", ce->ce_flags);\n>\n> It would be nice to have the flags as a hex-formatted string\n> in addition to (or instead of) the decimal integer value.\n\nInstead of, please, prefixed with \"0x\" to make it clear that this is hex.\n\nAnd \"mode\" in octal, please, prefixed with \"0\", as that is the convention\nto display file modes.\n\nThanks,\nDscho\n\nFootnote *1*: It is too bad that JSON leaves the exact date format\nunspecified, leading to a lot of inconsistency and putting the burden on\nparsers:\nhttps://stackoverflow.com/questions/10286204/the-right-json-date-format\n"},{"id":"377952","messageId":"20190625093837.GC21407@hank.intra.tgummerer.com","threadId":"51372","inReplyTo":"CACsJy8BsT-GaVvEmqfk5n1jGmkcLG_bRjqcU0M3yefBmNSxmnA@mail.gmail.com","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2019-06-25T09:38:37Z","receivedAt":"2019-06-25T09:38:42Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 06/25, Duy Nguyen wrote:\n> On Tue, Jun 25, 2019 at 1:00 AM Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> > especially when we offer this as a better way\n> > for 3rd-party applications to interact with Git (which I think will be the\n> > use case for this feature that will be _far_ more common than using it for\n> > debugging).\n> \n> We may have conflicting goals. For me, first priority is the debug\n> tool for Git developers. 3rd-party support is a stretch. I could move\n> all this back to test-tool, then you can provide a 3rd-party API if\n> you want. Or I'll withdraw this series and go back to my original\n> plan.\n\nFWIW, I am very much in favor of this series, and would have found\nsomething like this very useful many times in the past when I was\ndigging into the index code.  So I'd be more than happy to just have\nthis as debug tool, rather than as 3rd-party API.\n\nI'd also be fine with this living in test-tool, as long as it's\nsomewhere in git.git where it's easily usable, I'd find it helpful.\n\nThanks for working on this!\n"},{"id":"377954","messageId":"nycvar.QRO.7.76.6.1906251142580.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"20190624130226.17293-2-pclouds@gmail.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T09:44:53Z","receivedAt":"2019-06-25T09:44:45Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Mon, 24 Jun 2019, Nguyễn Thái Ngọc Duy wrote:\n\n> diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> new file mode 100755\n> index 0000000000..97bcd814be\n> --- /dev/null\n> +++ b/t/t3011-ls-files-json.sh\n> @@ -0,0 +1,44 @@\n> +#!/bin/sh\n> +\n> +test_description='ls-files dumping json'\n> +\n> +. ./test-lib.sh\n> +\n> +strip_number() {\n> +\tfor name; do\n> +\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n\nThis does not do what you think it does, in Ubuntu Xenial and on macOS:\n\nhttps://dev.azure.com/gitgitgadget/git/_build/results?buildId=11408&view=ms.vss-test-web.build-test-results-tab&runId=27736&paneView=debug&resultId=105613\n\nThe `\\1` is expanded to the ASCII character 001. Therefore your test cases\nfail on almost all platforms.\n\nFunnily enough, they pass on Windows...\n\nCiao,\nJohannes\n"},{"id":"377955","messageId":"CACsJy8CEaT7QGrOsoQw6k9H2A5DYW5ZJR1=Qs45TiJv+9sMBdQ@mail.gmail.com","threadId":"51372","inReplyTo":"755a4cfe-fd6b-044b-dca2-05eebfa518b1@jeffhostetler.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-25T09:52:27Z","receivedAt":"2019-06-25T09:52:56Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jun 25, 2019 at 2:15 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> > @@ -202,6 +202,28 @@ void jw_object_null(struct json_writer *jw, const char *key)\n> >       strbuf_addstr(&jw->json, \"null\");\n> >   }\n> >\n> > +void jw_object_filemode(struct json_writer *jw, const char *key, mode_t mode)\n> > +{\n> > +     object_common(jw, key);\n> > +     strbuf_addf(&jw->json, \"\\\"%06o\\\"\", mode);\n> > +}\n> > +\n> > +void jw_object_stat_data(struct json_writer *jw, const char *name,\n> > +                      const struct stat_data *sd)\n>\n> Should this be in json_writer.c or in read-cache.c ?\n> Currently, json_writer.c is concerned with formatting\n> JSON on basic/scalar types.  Do we want to start\n> extending it to handle arbitrary structures?  Or would\n> it be better for the code that defines/manipulates the\n> structure to define a \"stat_data_dump_json()\" function.\n>\n> I'm torn on the jw_object_filemode() function, JSON format\n> limits us to decimal integers and there are places where\n> I'd like to have hex, or in this case octal values.\n>\n> I'm thinking it'd be better to have a helper function in\n> read-cache.c that formats a local strbuf and calls\n> js_object_string(&jw, key, buf);\n\nI can move these back to read-cache.c. Though if we have a lot more jw\nhelpers like this (hard to tell at the moment) then perhaps we can\nhave json-writer-utils.c or something to group them together. That\nkeep the \"boring\" code out of main logic code in read-cache.c and\nother call sites.\n\n> > @@ -1952,6 +1953,49 @@ static void *load_index_extensions(void *_data)\n> >       return NULL;\n> >   }\n> >\n> > +static void dump_cache_entry(struct index_state *istate,\n> > +                          int index,\n> > +                          unsigned long offset,\n> > +                          const struct cache_entry *ce)\n> > +{\n> > +     struct json_writer *jw = istate->jw;\n> > +\n> > +     jw_array_inline_begin_object(jw);\n> > +\n> > +     /*\n> > +      * this is technically redundant, but it's for easier\n> > +      * navigation when there hundreds of entries\n> > +      */\n> > +     jw_object_intmax(jw, \"id\", index);\n> > +\n> > +     jw_object_string(jw, \"name\", ce->name);\n> > +\n> > +     jw_object_filemode(jw, \"mode\", ce->ce_mode);\n> > +\n> > +     jw_object_intmax(jw, \"flags\", ce->ce_flags);\n>\n> It would be nice to have the flags as a hex-formatted string\n> in addition to (or instead of) the decimal integer value.\n\nI'm not against reformatting it in hex string, but is there a value in\nit? ce_flags is expanded in the code below so that you don't have to\ndecode it yourself when you read each entry. The \"flags\" field here is\nfor further processing in tools. I'm trying to see if looking at hex\nvalues helps, but I'm still not seeing it...\n\n> > +     /*\n> > +      * again redundant info, just so you don't have to decode\n> > +      * flags values manually\n> > +      */\n> > +     if (ce->ce_flags & CE_EXTENDED)\n> > +             jw_object_true(jw, \"extended_flags\");\n> > +     if (ce->ce_flags & CE_VALID)\n> > +             jw_object_true(jw, \"assume_unchanged\");\n> > +     if (ce->ce_flags & CE_INTENT_TO_ADD)\n> > +             jw_object_true(jw, \"intent_to_add\");\n> > +     if (ce->ce_flags & CE_SKIP_WORKTREE)\n> > +             jw_object_true(jw, \"skip_worktree\");\n> > +     if (ce_stage(ce))\n> > +             jw_object_intmax(jw, \"stage\", ce_stage(ce));\n-- \nDuy\n"},{"id":"377962","messageId":"CACsJy8BjhQD-g69dr-yDCycgfrHZ8xJLgjD=LanRUBxAN6=Zrg@mail.gmail.com","threadId":"51372","inReplyTo":"55f81571-ba45-edcf-49bd-05418cc309c5@jeffhostetler.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-25T10:29:49Z","receivedAt":"2019-06-25T10:30:20Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> I'm curious how big these EWAHs will be in practice and\n> how useful an array of integers will be (especially as the\n> pretty format will be one integer per line).  Perhaps it\n> would helpful to have an extended example in one of the\n> tests.\n\nIt's one integer per updated entry. So if you have a giant index and\nupdated every single one of them, the EWAH bitmap contains that many\nintegers.\n\nIf it was easy to just merge these bitmaps back to the entry (e.g. in\nthis example, add \"replaced\": true to entry zero) I would have done\nit. But we dump as we stream and it's already too late to do it.\n\n> Would it be better to have the caller of ewah_each_bit()\n> build a hex or bit string in a strbuf and then write it\n> as a single string?\n\nI don't think the current EWAH representation is easy to read in the\nfirst place. You'll probably have to run through some script to update\nthe main entries part and will have a much better view, but that's\npretty quick. If it's for scripts, then it's probably best to keep as\nan array of integers, not a string. Less post processing.\n\nAnother reason for not merging to one string (might not be a very good\nargument though) is to help diff between two indexes.\nOne-number-per-line works well with \"git diff --no-index\" while one\nlong string is a bit harder. I did this kind of comparison when I made\nchanges in read-cache.c and wanted to check if the new index file is\ncompletely broken, or just slighly broken.\n-- \nDuy\n"},{"id":"377968","messageId":"nycvar.QRO.7.76.6.1906251311280.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"CACsJy8BsT-GaVvEmqfk5n1jGmkcLG_bRjqcU0M3yefBmNSxmnA@mail.gmail.com","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T11:27:24Z","receivedAt":"2019-06-25T11:27:20Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Tue, 25 Jun 2019, Duy Nguyen wrote:\n\n> On Tue, Jun 25, 2019 at 1:00 AM Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> > > - extension location is printed, in case you need to decode the\n> > >   extension by yourself (previously only the size is printed)\n> > > - all extensions are printed in the same order they appear in the file\n> > >   (previously eoie and ieot are printed first because that's how we\n> > >   parse)\n> > > - resolve undo extension is reorganized a bit to be easier to read\n> > > - tests added. Example json files are in t/t3011\n> >\n> > It might actually make sense to optionally disable showing extensions.\n> >\n> > You also forgot to mention that you explicitly disable handling\n> > `<pathspec>`, which I find a bit odd, personally, as that would probably\n> > come in real handy at times,\n>\n> No. I mentioned the land of high level languages before. Filtering in\n> any Python, Ruby, Scheme, JavaScript, Java is a piece of cake and much\n> more flexible than pathspec.\n\nI heard that type of argument before. I was working on the initial Windows\nport of Git, uh, of course I was working on a big refactoring of a big C++\napplication backed by a database. A colleague suggested that filtering\ncould be done much better in C++, on the desktop, than in SQL. And so they\nchanged the paradigm to \"simplify\" the SQL query, and instead dropped the\nunwanted data after it had hit the RAM of the client machine.\n\nTurns out it was a bad idea. A _really_ bad idea. Because it required\ndownloading 30MB of data for about several dozens computers in parallel,\nat the start of every shift.\n\nThis change was reverted in one big hurry, and the colleague was tasked to\nlearn them some SQL.\n\nWhy am I telling you this story? Because you fall into the exact same trap\nas my colleague.\n\nIn this instance, it may not be so much network bandwidth, but it is still\nquite a suboptimal idea to render JSON for possibly tens of thousands of\nfiles, then parse the same JSON on the receiving side, the spend even more\ntime to drop all but a dozen files.\n\nAnd this is _even more_ relevant when you want to debug things.\n\nIn short: I am quite puzzled why this is even debated here. There is a\nreason, a good reason, why `git ls-files` accepts pathspecs. I would not\nwant to ignore the lessons of history as willfully here.\n\n> Even with shell scripts, jq could do a much better job than pathspec. If\n> you filter by pathspec, good luck trying that on extensions.\n\nYou keep harping on extensions, but the reality of the matter is that they\nare rarely interesting. I would even wager a bet that we will end up\nexcluding them from the JSON output by default.\n\nMost of the times when I had to decode the index file manually in the\npast, it was about the regular file entries.\n\nThere was *one* week in which I had to decode the untracked cache a bit,\nto the point where I patched the test helper locally to help me with that.\n\nIf my experience in debugging these things is any indicator, extensions do\nnot matter even 10% of the non-extension data.\n\nAnd that's not even taking into account the third-party software that\ncould definitely benefit from having this JSON format as query result.\n\nIn my work as Git for Windows maintainer, I do hear about the needs of\nthird-party software developers quite a bit, so I would claim that I know\na bit about what they need, why the NUL-terminated format is not a good\nmatch, and how much a JSON-based API would help.\n\nSo while that is not your pet, it will be the most useful part of the\noutcome of your work.\n\n> It's the same reason why I will not provide a flexible way to disable\n> extensions. I'm not starting a JSON API for Git. I provide an index\n> file in JSON format. You do what you want with it. You have a format\n> easy enough to import to native data structures of your favorite\n> language.\n\nI understand that you don't care.\n\nYour patch series is just too good a start on something truly useful to\npass up on the opportunity.\n\n> > especially when we offer this as a better way for 3rd-party\n> > applications to interact with Git (which I think will be the use case\n> > for this feature that will be _far_ more common than using it for\n> > debugging).\n>\n> We may have conflicting goals. For me, first priority is the debug\n> tool for Git developers. 3rd-party support is a stretch. I could move\n> all this back to test-tool, then you can provide a 3rd-party API if\n> you want. Or I'll withdraw this series and go back to my original\n> plan.\n\nYou don't need JSON if you want to debug things. That would be a lot of\nlove lost, if debugging was your goal.\n\nI guess I'll wait until your patch series hits `next`, and then try to\nfind some time to work on that feature.\n\nCiao,\nJohannes\n"},{"id":"377969","messageId":"nycvar.QRO.7.76.6.1906251328320.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906251142580.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T11:31:07Z","receivedAt":"2019-06-25T11:30:56Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Tue, 25 Jun 2019, Johannes Schindelin wrote:\n\n> On Mon, 24 Jun 2019, Nguyễn Thái Ngọc Duy wrote:\n>\n> > diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> > new file mode 100755\n> > index 0000000000..97bcd814be\n> > --- /dev/null\n> > +++ b/t/t3011-ls-files-json.sh\n> > @@ -0,0 +1,44 @@\n> > +#!/bin/sh\n> > +\n> > +test_description='ls-files dumping json'\n> > +\n> > +. ./test-lib.sh\n> > +\n> > +strip_number() {\n> > +\tfor name; do\n> > +\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n>\n> This does not do what you think it does, in Ubuntu Xenial and on macOS:\n>\n> https://dev.azure.com/gitgitgadget/git/_build/results?buildId=11408&view=ms.vss-test-web.build-test-results-tab&runId=27736&paneView=debug&resultId=105613\n>\n> The `\\1` is expanded to the ASCII character 001. Therefore your test cases\n> fail on almost all platforms.\n\nThe `strip_number()`/`strip_string()` approach might look elegant from a\ndesign perspective, but from a readability perspective (and obviously,\nwhen one wants to make those tests more robust and cross-platform), it\nwould be a lot better to do it more explicitly.\n\nThis patch on top of your patch series makes the test run correctly in my\nLinux and Windows setup, and much easier to understand:\n\n-- snipsnap --\ndiff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\nindex 9f4ad4c9cf..8b782c48e0 100755\n--- a/t/t3011-ls-files-json.sh\n+++ b/t/t3011-ls-files-json.sh\n@@ -4,18 +4,6 @@ test_description='ls-files dumping json'\n\n . ./test-lib.sh\n\n-strip_number() {\n-\tfor name; do\n-\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n-\tdone\n-}\n-\n-strip_string() {\n-\tfor name; do\n-\t\techo 's/\\(\"'$name'\":\\) \".*\"/\\1 <string>/' >>filter.sed\n-\tdone\n-}\n-\n compare_json() {\n \tgit ls-files --debug-json >json &&\n \tsed -f filter.sed json >filtered &&\n@@ -35,9 +23,21 @@ test_expect_success 'setup' '\n \techo intent-to-add >ita &&\n \tgit add -N ita &&\n\n-\tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n-\tstrip_number device inode uid gid file_offset ext_size last_update &&\n-\tstrip_string oid ident\n+\tcat >filter.sed <<-\\EOF\n+\ts/\\(\"ctime_sec\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"ctime_nsec\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"mtime_sec\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"mtime_nsec\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"device\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"inode\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"uid\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"gid\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"file_offset\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"ext_size\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"last_update\":\\) [0-9]\\+/\\1 <number>/\n+\ts/\\(\"oid\":\\) \".*\"/\\1 <string>/\n+\ts/\\(\"ident\":\\) \".*\"/\\1 <string>/\n+\tEOF\n '\n\n test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n@@ -98,7 +98,9 @@ test_expect_success !SINGLE_CPU 'ls-files --json and multicore extensions' '\n \t\ttouch one two three four &&\n \t\tgit add . &&\n \t\tcp ../filter.sed . &&\n-\t\tstrip_number offset &&\n+\t\tcat >>filter.sed <<-\\EOF &&\n+\t\ts/\\(\"offset\":\\) [0-9]\\+/\\1 <number>/\n+\t\tEOF\n \t\tcompare_json eoie\n \t)\n '\n"},{"id":"377976","messageId":"CACsJy8B9vd9YaP_FHN-EDEPc_OHgD=MtFu8WymM66PURWX=25Q@mail.gmail.com","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906251311280.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-25T12:06:29Z","receivedAt":"2019-06-25T12:06:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jun 25, 2019 at 6:27 PM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n>\n> Hi Duy,\n>\n> On Tue, 25 Jun 2019, Duy Nguyen wrote:\n>\n> > On Tue, Jun 25, 2019 at 1:00 AM Johannes Schindelin\n> > <Johannes.Schindelin@gmx.de> wrote:\n> > > > - extension location is printed, in case you need to decode the\n> > > >   extension by yourself (previously only the size is printed)\n> > > > - all extensions are printed in the same order they appear in the file\n> > > >   (previously eoie and ieot are printed first because that's how we\n> > > >   parse)\n> > > > - resolve undo extension is reorganized a bit to be easier to read\n> > > > - tests added. Example json files are in t/t3011\n> > >\n> > > It might actually make sense to optionally disable showing extensions.\n> > >\n> > > You also forgot to mention that you explicitly disable handling\n> > > `<pathspec>`, which I find a bit odd, personally, as that would probably\n> > > come in real handy at times,\n> >\n> > No. I mentioned the land of high level languages before. Filtering in\n> > any Python, Ruby, Scheme, JavaScript, Java is a piece of cake and much\n> > more flexible than pathspec.\n>\n> I heard that type of argument before. I was working on the initial Windows\n> port of Git, uh, of course I was working on a big refactoring of a big C++\n> application backed by a database. A colleague suggested that filtering\n> could be done much better in C++, on the desktop, than in SQL. And so they\n> changed the paradigm to \"simplify\" the SQL query, and instead dropped the\n> unwanted data after it had hit the RAM of the client machine.\n>\n> Turns out it was a bad idea. A _really_ bad idea. Because it required\n> downloading 30MB of data for about several dozens computers in parallel,\n> at the start of every shift.\n>\n> This change was reverted in one big hurry, and the colleague was tasked to\n> learn them some SQL.\n>\n> Why am I telling you this story? Because you fall into the exact same trap\n> as my colleague.\n>\n> In this instance, it may not be so much network bandwidth, but it is still\n> quite a suboptimal idea to render JSON for possibly tens of thousands of\n> files, then parse the same JSON on the receiving side, the spend even more\n> time to drop all but a dozen files.\n\nThis was mentioned before [1]. Of course I don't work on giant index\nfiles, but I would assume the cost of parsing JSON (at least with a\nstream-based one, not loading the whole thing in core) is still\ncheaper. And you could still do it in iteration, saving every step\nuntil you have the reasonable small dataset to work on. The other side\nof the story is, are we sure parsing and executing pathspec is cheap?\nI'm not so sure, especially when pathspec code is not exactly\noptimized.\n\nConsider the amount of code to support something like that. I'd rather\nwait until a real example come up and no good solution found without\nmodify git.git, before actually supporting it.\n\n[1] https://public-inbox.org/git/45e49624-be8e-deff-bf9d-aee052991189@gmail.com/\n\n> And this is _even more_ relevant when you want to debug things.\n>\n> In short: I am quite puzzled why this is even debated here. There is a\n> reason, a good reason, why `git ls-files` accepts pathspecs. I would not\n> want to ignore the lessons of history as willfully here.\n\nI guess you and I have different ways of debugging things.\n\n> > Even with shell scripts, jq could do a much better job than pathspec. If\n> > you filter by pathspec, good luck trying that on extensions.\n>\n> You keep harping on extensions, but the reality of the matter is that they\n> are rarely interesting. I would even wager a bet that we will end up\n> excluding them from the JSON output by default.\n>\n> Most of the times when I had to decode the index file manually in the\n> past, it was about the regular file entries.\n>\n> There was *one* week in which I had to decode the untracked cache a bit,\n> to the point where I patched the test helper locally to help me with that.\n>\n> If my experience in debugging these things is any indicator, extensions do\n> not matter even 10% of the non-extension data.\n\nAgain our experiences differ. Mine is mostly about extensions,\nprobably because I had to work on them more often. For normal entries\n\"ls-files --debug\" gives you 99% what's in the index file already.\n\n> > > especially when we offer this as a better way for 3rd-party\n> > > applications to interact with Git (which I think will be the use case\n> > > for this feature that will be _far_ more common than using it for\n> > > debugging).\n> >\n> > We may have conflicting goals. For me, first priority is the debug\n> > tool for Git developers. 3rd-party support is a stretch. I could move\n> > all this back to test-tool, then you can provide a 3rd-party API if\n> > you want. Or I'll withdraw this series and go back to my original\n> > plan.\n>\n> You don't need JSON if you want to debug things. That would be a lot of\n> love lost, if debugging was your goal.\n\nNo, I did think of some other line-based format before I ended up with\nJSON. I did not want to use it in the beginning.\n\nThe thing is, a giant table to cover all fields and entries in the\nmain index is not as easy to navigate, or search even in 'less'. And\nthe hierarchical structure of some extensions is hard to represent in\ngood way (at least without writing lots of code). On top of that I\nstill need some easy way to parse and post-process, like how much\nsaving I could gain if I compressed stat data. And the final nail is\njson-writer.c is already there, much less work.\n\nSo JSON was the best option I found to meet all those points.\n-- \nDuy\n"},{"id":"377977","messageId":"98afb501-ef57-9b64-7ffb-f13cea6fd58a@gmail.com","threadId":"51372","inReplyTo":"CACsJy8BjhQD-g69dr-yDCycgfrHZ8xJLgjD=LanRUBxAN6=Zrg@mail.gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-06-25T12:40:50Z","receivedAt":"2019-06-25T12:40:53Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 6/25/2019 6:29 AM, Duy Nguyen wrote:\n> On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n>> I'm curious how big these EWAHs will be in practice and\n>> how useful an array of integers will be (especially as the\n>> pretty format will be one integer per line).  Perhaps it\n>> would helpful to have an extended example in one of the\n>> tests.\n> \n> It's one integer per updated entry. So if you have a giant index and\n> updated every single one of them, the EWAH bitmap contains that many\n> integers.\n> \n> If it was easy to just merge these bitmaps back to the entry (e.g. in\n> this example, add \"replaced\": true to entry zero) I would have done\n> it. But we dump as we stream and it's already too late to do it.\n> \n>> Would it be better to have the caller of ewah_each_bit()\n>> build a hex or bit string in a strbuf and then write it\n>> as a single string?\n> \n> I don't think the current EWAH representation is easy to read in the\n> first place. You'll probably have to run through some script to update\n> the main entries part and will have a much better view, but that's\n> pretty quick. If it's for scripts, then it's probably best to keep as\n> an array of integers, not a string. Less post processing.\n\nI don't think the intent is to dump the EWAH directly, but instead to\ndump a string of the uncompressed bitmap. Something like:\n\n\t\"delete_bitmap\" : \"01101101101\"\n\ninstead of\n\n\t\"delete_bitmap\" : [ 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1 ]\n\n> Another reason for not merging to one string (might not be a very good\n> argument though) is to help diff between two indexes.\n> One-number-per-line works well with \"git diff --no-index\" while one\n> long string is a bit harder. I did this kind of comparison when I made\n> changes in read-cache.c and wanted to check if the new index file is\n> completely broken, or just slighly broken.\n\nYou're right that the diff of the json output is an interesting\nuse, and the \"single string\" output is not helpful. What about\nbatches of 64-bit strings? For example:\n\n\t\"delete_bitmap\" : [\n\t\t\"0101010101010101010101010101010101010101010101010101010101010101\",\n\t\t\"0101010101010101010101010101010101010101010101010101010101010101\",\n\t\t\"0101010101010101010101010101010101010101010101010101010101010101\",\n\t\t\"01010101010101\"\n\t]\n\nThis could be a happy medium between the two options, but does require\nsome extra work in the formatter.\n\nThanks,\n-Stolee\n"},{"id":"377999","messageId":"nycvar.QRO.7.76.6.1906251557001.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906251328320.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T13:57:50Z","receivedAt":"2019-06-25T13:57:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Tue, 25 Jun 2019, Johannes Schindelin wrote:\n\n> diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> index 9f4ad4c9cf..8b782c48e0 100755\n> --- a/t/t3011-ls-files-json.sh\n> +++ b/t/t3011-ls-files-json.sh\n> @@ -4,18 +4,6 @@ test_description='ls-files dumping json'\n>\n>  . ./test-lib.sh\n>\n> -strip_number() {\n> -\tfor name; do\n> -\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n> -\tdone\n> -}\n> -\n> -strip_string() {\n> -\tfor name; do\n> -\t\techo 's/\\(\"'$name'\":\\) \".*\"/\\1 <string>/' >>filter.sed\n> -\tdone\n> -}\n> -\n>  compare_json() {\n>  \tgit ls-files --debug-json >json &&\n>  \tsed -f filter.sed json >filtered &&\n> @@ -35,9 +23,21 @@ test_expect_success 'setup' '\n>  \techo intent-to-add >ita &&\n>  \tgit add -N ita &&\n>\n> -\tstrip_number ctime_sec ctime_nsec mtime_sec mtime_nsec &&\n> -\tstrip_number device inode uid gid file_offset ext_size last_update &&\n> -\tstrip_string oid ident\n> +\tcat >filter.sed <<-\\EOF\n> +\ts/\\(\"ctime_sec\":\\) [0-9]\\+/\\1 <number>/\n\nAnd of course, \\+ still isn't POSIX, so you have to write [0-9][1-9]*\ninstead.\n\nCiao,\nJohannes\n\n> +\ts/\\(\"ctime_nsec\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"mtime_sec\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"mtime_nsec\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"device\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"inode\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"uid\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"gid\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"file_offset\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"ext_size\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"last_update\":\\) [0-9]\\+/\\1 <number>/\n> +\ts/\\(\"oid\":\\) \".*\"/\\1 <string>/\n> +\ts/\\(\"ident\":\\) \".*\"/\\1 <string>/\n> +\tEOF\n>  '\n>\n>  test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n> @@ -98,7 +98,9 @@ test_expect_success !SINGLE_CPU 'ls-files --json and multicore extensions' '\n>  \t\ttouch one two three four &&\n>  \t\tgit add . &&\n>  \t\tcp ../filter.sed . &&\n> -\t\tstrip_number offset &&\n> +\t\tcat >>filter.sed <<-\\EOF &&\n> +\t\ts/\\(\"offset\":\\) [0-9]\\+/\\1 <number>/\n> +\t\tEOF\n>  \t\tcompare_json eoie\n>  \t)\n>  '\n"},{"id":"378000","messageId":"nycvar.QRO.7.76.6.1906251601240.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"CACsJy8B9vd9YaP_FHN-EDEPc_OHgD=MtFu8WymM66PURWX=25Q@mail.gmail.com","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-25T14:10:38Z","receivedAt":"2019-06-25T14:10:35Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Tue, 25 Jun 2019, Duy Nguyen wrote:\n\n> On Tue, Jun 25, 2019 at 6:27 PM Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> >\n> > On Tue, 25 Jun 2019, Duy Nguyen wrote:\n> >\n> > > On Tue, Jun 25, 2019 at 1:00 AM Johannes Schindelin\n> > > <Johannes.Schindelin@gmx.de> wrote:\n> > > > > - extension location is printed, in case you need to decode the\n> > > > >   extension by yourself (previously only the size is printed)\n> > > > > - all extensions are printed in the same order they appear in the file\n> > > > >   (previously eoie and ieot are printed first because that's how we\n> > > > >   parse)\n> > > > > - resolve undo extension is reorganized a bit to be easier to read\n> > > > > - tests added. Example json files are in t/t3011\n> > > >\n> > > > It might actually make sense to optionally disable showing extensions.\n> > > >\n> > > > You also forgot to mention that you explicitly disable handling\n> > > > `<pathspec>`, which I find a bit odd, personally, as that would probably\n> > > > come in real handy at times,\n> > >\n> > > No. I mentioned the land of high level languages before. Filtering in\n> > > any Python, Ruby, Scheme, JavaScript, Java is a piece of cake and much\n> > > more flexible than pathspec.\n> >\n> > I heard that type of argument before. I was working on the initial Windows\n> > port of Git, uh, of course I was working on a big refactoring of a big C++\n> > application backed by a database. A colleague suggested that filtering\n> > could be done much better in C++, on the desktop, than in SQL. And so they\n> > changed the paradigm to \"simplify\" the SQL query, and instead dropped the\n> > unwanted data after it had hit the RAM of the client machine.\n> >\n> > Turns out it was a bad idea. A _really_ bad idea. Because it required\n> > downloading 30MB of data for about several dozens computers in parallel,\n> > at the start of every shift.\n> >\n> > This change was reverted in one big hurry, and the colleague was tasked to\n> > learn them some SQL.\n> >\n> > Why am I telling you this story? Because you fall into the exact same trap\n> > as my colleague.\n> >\n> > In this instance, it may not be so much network bandwidth, but it is still\n> > quite a suboptimal idea to render JSON for possibly tens of thousands of\n> > files, then parse the same JSON on the receiving side, the spend even more\n> > time to drop all but a dozen files.\n>\n> This was mentioned before [1]. Of course I don't work on giant index\n> files, but I would assume the cost of parsing JSON (at least with a\n> stream-based one, not loading the whole thing in core) is still\n> cheaper.\n\nYou may have heard that a few thousand of my colleagues are working on\nwhat they call the largest repository on this planet.\n\nNo, the cost of parsing JSON only to throw away the majority of the parsed\ninformation is not cheap. It is a clear sign of a design in want of being\nimproved.\n\n> And you could still do it in iteration, saving every step until you have\n> the reasonable small dataset to work on. The other side of the story is,\n> are we sure parsing and executing pathspec is cheap? I'm not so sure,\n> especially when pathspec code is not exactly optimized.\n\nLet's not try to slap on workaround over workaround. Let's fix the root\ncause. (Being: don't filter at the wrong end.)\n\n> Consider the amount of code to support something like that.\n\nGiven that I am pretty familiar with the pathspec machinery due to working\nwith it in the `git stash` and `git add -p` built-ins, I have a very easy\ntime considering the amount of code. It makes me smile how little code\nwill be needed.\n\n> I'd rather wait until a real example come up and no good solution found\n> without modify git.git, before actually supporting it.\n\nOh hey, there you go: Team Explorer. Visual Studio Code. Literally every\nsingle 3rd-party application that needs to deal with real-world loads.\nEvery single one.\n\n> > And this is _even more_ relevant when you want to debug things.\n> >\n> > In short: I am quite puzzled why this is even debated here. There is a\n> > reason, a good reason, why `git ls-files` accepts pathspecs. I would not\n> > want to ignore the lessons of history as willfully here.\n>\n> I guess you and I have different ways of debugging things.\n\nYep, I'm with Lincoln here: Give me six hours to debug a problem and I\nwill spend the first four optimizing the edit-build-test cycle.\n\n> > > Even with shell scripts, jq could do a much better job than pathspec. If\n> > > you filter by pathspec, good luck trying that on extensions.\n> >\n> > You keep harping on extensions, but the reality of the matter is that they\n> > are rarely interesting. I would even wager a bet that we will end up\n> > excluding them from the JSON output by default.\n> >\n> > Most of the times when I had to decode the index file manually in the\n> > past, it was about the regular file entries.\n> >\n> > There was *one* week in which I had to decode the untracked cache a bit,\n> > to the point where I patched the test helper locally to help me with that.\n> >\n> > If my experience in debugging these things is any indicator, extensions do\n> > not matter even 10% of the non-extension data.\n>\n> Again our experiences differ. Mine is mostly about extensions,\n> probably because I had to work on them more often. For normal entries\n> \"ls-files --debug\" gives you 99% what's in the index file already.\n\nLike the device. And the ctime. And the file size. And the uid/gid. Is\nthat what you mean?\n\nI don't know whether I missed a joke or not.\n\n> > You don't need JSON if you want to debug things. That would be a lot of\n> > love lost, if debugging was your goal.\n>\n> No, I did think of some other line-based format before I ended up with\n> JSON. I did not want to use it in the beginning.\n\nThen why bother.\n\n> The thing is, a giant table to cover all fields and entries in the\n> main index is not as easy to navigate, or search even in 'less'. And\n> the hierarchical structure of some extensions is hard to represent in\n> good way (at least without writing lots of code). On top of that I\n> still need some easy way to parse and post-process, like how much\n> saving I could gain if I compressed stat data. And the final nail is\n> json-writer.c is already there, much less work.\n>\n> So JSON was the best option I found to meet all those points.\n\nWell, as I said: you're obviously dead-set to optimize this for debugging\nyour own problems. The beauty of open source is that it can be turned into\nsomething of wider use.\n\nCiao,\nJohannes\n"},{"id":"378024","messageId":"9a95bfcf-9fc7-dedb-d7b5-ebb4855c9ef3@jeffhostetler.com","threadId":"51372","inReplyTo":"CACsJy8CEaT7QGrOsoQw6k9H2A5DYW5ZJR1=Qs45TiJv+9sMBdQ@mail.gmail.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-25T15:37:28Z","receivedAt":"2019-06-25T15:37:32Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/25/2019 5:52 AM, Duy Nguyen wrote:\n> On Tue, Jun 25, 2019 at 2:15 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n>>> @@ -202,6 +202,28 @@ void jw_object_null(struct json_writer *jw, const char *key)\n>>>        strbuf_addstr(&jw->json, \"null\");\n>>>    }\n>>>\n>>> +void jw_object_filemode(struct json_writer *jw, const char *key, mode_t mode)\n>>> +{\n>>> +     object_common(jw, key);\n>>> +     strbuf_addf(&jw->json, \"\\\"%06o\\\"\", mode);\n>>> +}\n>>> +\n>>> +void jw_object_stat_data(struct json_writer *jw, const char *name,\n>>> +                      const struct stat_data *sd)\n>>\n>> Should this be in json_writer.c or in read-cache.c ?\n>> Currently, json_writer.c is concerned with formatting\n>> JSON on basic/scalar types.  Do we want to start\n>> extending it to handle arbitrary structures?  Or would\n>> it be better for the code that defines/manipulates the\n>> structure to define a \"stat_data_dump_json()\" function.\n>>\n>> I'm torn on the jw_object_filemode() function, JSON format\n>> limits us to decimal integers and there are places where\n>> I'd like to have hex, or in this case octal values.\n>>\n>> I'm thinking it'd be better to have a helper function in\n>> read-cache.c that formats a local strbuf and calls\n>> js_object_string(&jw, key, buf);\n> \n> I can move these back to read-cache.c. Though if we have a lot more jw\n> helpers like this (hard to tell at the moment) then perhaps we can\n> have json-writer-utils.c or something to group them together. That\n> keep the \"boring\" code out of main logic code in read-cache.c and\n> other call sites.\n\nyeah, in an utils file or close to the \"constructors\" of the\nstructure types.  either one works.\n\n> \n>>> @@ -1952,6 +1953,49 @@ static void *load_index_extensions(void *_data)\n>>>        return NULL;\n>>>    }\n>>>\n>>> +static void dump_cache_entry(struct index_state *istate,\n>>> +                          int index,\n>>> +                          unsigned long offset,\n>>> +                          const struct cache_entry *ce)\n>>> +{\n>>> +     struct json_writer *jw = istate->jw;\n>>> +\n>>> +     jw_array_inline_begin_object(jw);\n>>> +\n>>> +     /*\n>>> +      * this is technically redundant, but it's for easier\n>>> +      * navigation when there hundreds of entries\n>>> +      */\n>>> +     jw_object_intmax(jw, \"id\", index);\n>>> +\n>>> +     jw_object_string(jw, \"name\", ce->name);\n>>> +\n>>> +     jw_object_filemode(jw, \"mode\", ce->ce_mode);\n>>> +\n>>> +     jw_object_intmax(jw, \"flags\", ce->ce_flags);\n>>\n>> It would be nice to have the flags as a hex-formatted string\n>> in addition to (or instead of) the decimal integer value.\n> \n> I'm not against reformatting it in hex string, but is there a value in\n> it? ce_flags is expanded in the code below so that you don't have to\n> decode it yourself when you read each entry. The \"flags\" field here is\n> for further processing in tools. I'm trying to see if looking at hex\n> values helps, but I'm still not seeing it...\n> \n\nI guess I was thinking of the in-memory bits and thinking\nit'd be useful to be able to dump the index immediately\nafter reading it and then later after some operation or\ntraversal and see the intermediate states.  But I realize\nnow that that's not what you're building.  This is a dump\nit while you're reading it feature (and that's fine).\n\nSo, as long as you have all of the on-disk bits, we should\nbe fine as you suggest.\n\nJeff\n\n\n>>> +     /*\n>>> +      * again redundant info, just so you don't have to decode\n>>> +      * flags values manually\n>>> +      */\n>>> +     if (ce->ce_flags & CE_EXTENDED)\n>>> +             jw_object_true(jw, \"extended_flags\");\n>>> +     if (ce->ce_flags & CE_VALID)\n>>> +             jw_object_true(jw, \"assume_unchanged\");\n>>> +     if (ce->ce_flags & CE_INTENT_TO_ADD)\n>>> +             jw_object_true(jw, \"intent_to_add\");\n>>> +     if (ce->ce_flags & CE_SKIP_WORKTREE)\n>>> +             jw_object_true(jw, \"skip_worktree\");\n>>> +     if (ce_stage(ce))\n>>> +             jw_object_intmax(jw, \"stage\", ce_stage(ce));\n"},{"id":"378027","messageId":"27211d51-c77f-84f8-49c0-4bc104baa266@ramsayjones.plus.com","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906251601240.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2019-06-25T17:08:02Z","receivedAt":"2019-06-25T17:08:12Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 25/06/2019 15:10, Johannes Schindelin wrote:\n> Hi Duy,\n[snip]\n\n>> Again our experiences differ. Mine is mostly about extensions,\n>> probably because I had to work on them more often. For normal entries\n>> \"ls-files --debug\" gives you 99% what's in the index file already.\n> \n> Like the device. And the ctime. And the file size. And the uid/gid. Is\n> that what you mean?\n\nHmm, well I think so:\n\n  $ git ls-files --debug git.c git-compat-util.h \n  git-compat-util.h\n    ctime: 1561457278:502638001\n    mtime: 1561457278:502638001\n    dev: 2049\tino: 262663\n    uid: 1000\tgid: 1000\n    size: 35440\tflags: 0\n  git.c\n    ctime: 1561457278:518646000\n    mtime: 1561457278:518646000\n    dev: 2049\tino: 263083\n    uid: 1000\tgid: 1000\n    size: 26837\tflags: 0\n  $ \n\nI have occasionally added stuff to the '--debug' output\nwhile debugging something, but the above is usually\nsufficient for my uses. (Having said that, I have not\nhad the need to debug extensions [yet!]).\n\nATB,\nRamsay Jones\n"},{"id":"378043","messageId":"xmqqk1d9e1vb.fsf@gitster-ct.c.googlers.com","threadId":"51372","inReplyTo":"nycvar.QRO.7.76.6.1906251142580.44@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-06-25T22:28:24Z","receivedAt":"2019-06-25T22:28:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> +\t\techo 's/\\(\"'$name'\":\\) [0-9]\\+/\\1 <number>/' >>filter.sed\n>\n> This does not do what you think it does, in Ubuntu Xenial and on macOS:\n>\n> https://dev.azure.com/gitgitgadget/git/_build/results?buildId=11408&view=ms.vss-test-web.build-test-results-tab&runId=27736&paneView=debug&resultId=105613\n>\n> The `\\1` is expanded to the ASCII character 001. Therefore your test cases\n> fail on almost all platforms.\n>\n> Funnily enough, they pass on Windows...\n\nbash, dash and /bin/echo behave differently given \n\n    $ echo 'foo \\1 bar'\n\nsome 'echo' suffer from the \"\\<n>\" interpolation.  Some don't.\n\nI think your spelled-out version downthread (except for stepping out\nof BRE which would break your sed script, as you realized) would be\na much readable alternative.\n\nThanks.\n"},{"id":"378066","messageId":"nycvar.QRO.7.76.6.1906261704330.44@tvgsbejvaqbjf.bet","threadId":"51372","inReplyTo":"27211d51-c77f-84f8-49c0-4bc104baa266@ramsayjones.plus.com","subject":"Re: [PATCH v2 00/10] Add 'ls-files --debug-json' to dump the index in json","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-06-26T15:05:06Z","receivedAt":"2019-06-26T15:05:03Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Ramsay,\n\nOn Tue, 25 Jun 2019, Ramsay Jones wrote:\n\n> On 25/06/2019 15:10, Johannes Schindelin wrote:\n> > Hi Duy,\n> [snip]\n>\n> >> Again our experiences differ. Mine is mostly about extensions,\n> >> probably because I had to work on them more often. For normal entries\n> >> \"ls-files --debug\" gives you 99% what's in the index file already.\n> >\n> > Like the device. And the ctime. And the file size. And the uid/gid. Is\n> > that what you mean?\n>\n> Hmm, well I think so:\n>\n>   $ git ls-files --debug git.c git-compat-util.h\n>   git-compat-util.h\n>     ctime: 1561457278:502638001\n>     mtime: 1561457278:502638001\n>     dev: 2049\tino: 262663\n>     uid: 1000\tgid: 1000\n>     size: 35440\tflags: 0\n>   git.c\n>     ctime: 1561457278:518646000\n>     mtime: 1561457278:518646000\n>     dev: 2049\tino: 263083\n>     uid: 1000\tgid: 1000\n>     size: 26837\tflags: 0\n>   $\n>\n> I have occasionally added stuff to the '--debug' output\n> while debugging something, but the above is usually\n> sufficient for my uses. (Having said that, I have not\n> had the need to debug extensions [yet!]).\n\nWell, live and learn. So for debugging purposes, we already have a good\nfacility and would not need JSON at all.\n\nGood to know,\nDscho\n"},{"id":"378077","messageId":"xmqqd0j0cegk.fsf@gitster-ct.c.googlers.com","threadId":"51372","inReplyTo":"20190624130226.17293-2-pclouds@gmail.com","subject":"Re: [PATCH v2 01/10] ls-files: add --json to dump the index","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-06-26T19:51:39Z","receivedAt":"2019-06-26T19:51:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> This --json is supposed to help that. It dumps the index in a human\n\nThis and another one on the title need to become \"--debug-json\".\n\nAre we expecting a reroll to introduce another layer above the most\nprimitive json writer that knows the schema used to represent both\nsystem standard and our application-specific structures, or is the\ncurrent arrangement to have them in json-writer.c until there are\nenough of them to warrant such a split good enough?\n"},{"id":"378150","messageId":"CACsJy8CwWvKNbYvDqWc-zCwEPc_rz-P4y-SvXV-9jL8_XCFjZQ@mail.gmail.com","threadId":"51372","inReplyTo":"98afb501-ef57-9b64-7ffb-f13cea6fd58a@gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-27T10:48:59Z","receivedAt":"2019-06-27T10:49:29Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jun 25, 2019 at 7:40 PM Derrick Stolee <stolee@gmail.com> wrote:\n>\n> On 6/25/2019 6:29 AM, Duy Nguyen wrote:\n> > On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> >> I'm curious how big these EWAHs will be in practice and\n> >> how useful an array of integers will be (especially as the\n> >> pretty format will be one integer per line).  Perhaps it\n> >> would helpful to have an extended example in one of the\n> >> tests.\n> >\n> > It's one integer per updated entry. So if you have a giant index and\n> > updated every single one of them, the EWAH bitmap contains that many\n> > integers.\n> >\n> > If it was easy to just merge these bitmaps back to the entry (e.g. in\n> > this example, add \"replaced\": true to entry zero) I would have done\n> > it. But we dump as we stream and it's already too late to do it.\n> >\n> >> Would it be better to have the caller of ewah_each_bit()\n> >> build a hex or bit string in a strbuf and then write it\n> >> as a single string?\n> >\n> > I don't think the current EWAH representation is easy to read in the\n> > first place. You'll probably have to run through some script to update\n> > the main entries part and will have a much better view, but that's\n> > pretty quick. If it's for scripts, then it's probably best to keep as\n> > an array of integers, not a string. Less post processing.\n>\n> I don't think the intent is to dump the EWAH directly, but instead to\n> dump a string of the uncompressed bitmap. Something like:\n>\n>         \"delete_bitmap\" : \"01101101101\"\n>\n> instead of\n>\n>         \"delete_bitmap\" : [ 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1 ]\n\nI get this part. But the numbers in the array were the position of the\nset bits. It's not showing just the actual bit map.\n\nThe same bitmap would be currently displayed as\n\n \"delete_bitmap\": [ 1, 2, 4, 5, 7, 8, 9, 11 ]\n\nAnd that maps back to the entry[1], entry[2], entry[4]... in the index\nbeing deleted from the base index. So displaying as a real bit map\nactually adds more work for both the reader and the tool because you\nhave to calculate the position either way. And it gets harder if the\nbit you're intereted in is on the far right.\n\n> > Another reason for not merging to one string (might not be a very good\n> > argument though) is to help diff between two indexes.\n> > One-number-per-line works well with \"git diff --no-index\" while one\n> > long string is a bit harder. I did this kind of comparison when I made\n> > changes in read-cache.c and wanted to check if the new index file is\n> > completely broken, or just slighly broken.\n>\n> You're right that the diff of the json output is an interesting\n> use, and the \"single string\" output is not helpful. What about\n> batches of 64-bit strings? For example:\n>\n>         \"delete_bitmap\" : [\n>                 \"0101010101010101010101010101010101010101010101010101010101010101\",\n>                 \"0101010101010101010101010101010101010101010101010101010101010101\",\n>                 \"0101010101010101010101010101010101010101010101010101010101010101\",\n>                 \"01010101010101\"\n>         ]\n>\n> This could be a happy medium between the two options, but does require\n> some extra work in the formatter.\n\nAnd the reader/parser too since you have to join that array back in\none string first.\n--\nDuy\n"},{"id":"378158","messageId":"93562f66-07a7-d074-e225-65afd7ced1d4@jeffhostetler.com","threadId":"51372","inReplyTo":"CACsJy8CwWvKNbYvDqWc-zCwEPc_rz-P4y-SvXV-9jL8_XCFjZQ@mail.gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2019-06-27T13:24:46Z","receivedAt":"2019-06-27T13:24:49Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 6/27/2019 6:48 AM, Duy Nguyen wrote:\n> On Tue, Jun 25, 2019 at 7:40 PM Derrick Stolee <stolee@gmail.com> wrote:\n>>\n>> On 6/25/2019 6:29 AM, Duy Nguyen wrote:\n>>> On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n>>>> I'm curious how big these EWAHs will be in practice and\n>>>> how useful an array of integers will be (especially as the\n>>>> pretty format will be one integer per line).  Perhaps it\n>>>> would helpful to have an extended example in one of the\n>>>> tests.\n>>>\n>>> It's one integer per updated entry. So if you have a giant index and\n>>> updated every single one of them, the EWAH bitmap contains that many\n>>> integers.\n>>>\n>>> If it was easy to just merge these bitmaps back to the entry (e.g. in\n>>> this example, add \"replaced\": true to entry zero) I would have done\n>>> it. But we dump as we stream and it's already too late to do it.\n>>>\n>>>> Would it be better to have the caller of ewah_each_bit()\n>>>> build a hex or bit string in a strbuf and then write it\n>>>> as a single string?\n>>>\n>>> I don't think the current EWAH representation is easy to read in the\n>>> first place. You'll probably have to run through some script to update\n>>> the main entries part and will have a much better view, but that's\n>>> pretty quick. If it's for scripts, then it's probably best to keep as\n>>> an array of integers, not a string. Less post processing.\n>>\n>> I don't think the intent is to dump the EWAH directly, but instead to\n>> dump a string of the uncompressed bitmap. Something like:\n>>\n>>          \"delete_bitmap\" : \"01101101101\"\n>>\n>> instead of\n>>\n>>          \"delete_bitmap\" : [ 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1 ]\n> \n> I get this part. But the numbers in the array were the position of the\n> set bits. It's not showing just the actual bit map.\n> \n> The same bitmap would be currently displayed as\n> \n>   \"delete_bitmap\": [ 1, 2, 4, 5, 7, 8, 9, 11 ]\n> \n> And that maps back to the entry[1], entry[2], entry[4]... in the index\n> being deleted from the base index. So displaying as a real bit map\n> actually adds more work for both the reader and the tool because you\n> have to calculate the position either way. And it gets harder if the\n> bit you're intereted in is on the far right.\n\n\nThanks for the clarification.  That helps.\n\nJeff\n"},{"id":"378159","messageId":"f4f82ab4-2846-34f9-45ee-a2149fb15d17@gmail.com","threadId":"51372","inReplyTo":"93562f66-07a7-d074-e225-65afd7ced1d4@jeffhostetler.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-06-27T13:42:09Z","receivedAt":"2019-06-27T13:42:15Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 6/27/2019 9:24 AM, Jeff Hostetler wrote:\n> On 6/27/2019 6:48 AM, Duy Nguyen wrote:\n>> On Tue, Jun 25, 2019 at 7:40 PM Derrick Stolee <stolee@gmail.com> wrote:\n>>>\n>>> On 6/25/2019 6:29 AM, Duy Nguyen wrote:\n>>>> On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n>>>>> I'm curious how big these EWAHs will be in practice and\n>>>>> how useful an array of integers will be (especially as the\n>>>>> pretty format will be one integer per line).  Perhaps it\n>>>>> would helpful to have an extended example in one of the\n>>>>> tests.\n>>>>\n>>>> It's one integer per updated entry. So if you have a giant index and\n>>>> updated every single one of them, the EWAH bitmap contains that many\n>>>> integers.\n>>>>\n>>>> If it was easy to just merge these bitmaps back to the entry (e.g. in\n>>>> this example, add \"replaced\": true to entry zero) I would have done\n>>>> it. But we dump as we stream and it's already too late to do it.\n>>>>\n>>>>> Would it be better to have the caller of ewah_each_bit()\n>>>>> build a hex or bit string in a strbuf and then write it\n>>>>> as a single string?\n>>>>\n>>>> I don't think the current EWAH representation is easy to read in the\n>>>> first place. You'll probably have to run through some script to update\n>>>> the main entries part and will have a much better view, but that's\n>>>> pretty quick. If it's for scripts, then it's probably best to keep as\n>>>> an array of integers, not a string. Less post processing.\n>>>\n>>> I don't think the intent is to dump the EWAH directly, but instead to\n>>> dump a string of the uncompressed bitmap. Something like:\n>>>\n>>>          \"delete_bitmap\" : \"01101101101\"\n>>>\n>>> instead of\n>>>\n>>>          \"delete_bitmap\" : [ 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1 ]\n>>\n>> I get this part. But the numbers in the array were the position of the\n>> set bits. It's not showing just the actual bit map.\n>>\n>> The same bitmap would be currently displayed as\n>>\n>>   \"delete_bitmap\": [ 1, 2, 4, 5, 7, 8, 9, 11 ]\n>>\n>> And that maps back to the entry[1], entry[2], entry[4]... in the index\n>> being deleted from the base index. So displaying as a real bit map\n>> actually adds more work for both the reader and the tool because you\n>> have to calculate the position either way. And it gets harder if the\n>> bit you're intereted in is on the far right.\n> \n> \n> Thanks for the clarification.  That helps.\n\nSame here! We expect these to be much smaller than the full set, correct?\n\nThanks,\n-Stolee\n\n"},{"id":"378161","messageId":"CACsJy8BeUOv+He5c58iGO47XDkoKkAGQZqH0NM1fZcb2ESFscQ@mail.gmail.com","threadId":"51372","inReplyTo":"f4f82ab4-2846-34f9-45ee-a2149fb15d17@gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-06-27T13:47:54Z","receivedAt":"2019-06-27T13:48:22Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Jun 27, 2019 at 8:42 PM Derrick Stolee <stolee@gmail.com> wrote:\n>\n> On 6/27/2019 9:24 AM, Jeff Hostetler wrote:\n> > On 6/27/2019 6:48 AM, Duy Nguyen wrote:\n> >> On Tue, Jun 25, 2019 at 7:40 PM Derrick Stolee <stolee@gmail.com> wrote:\n> >>>\n> >>> On 6/25/2019 6:29 AM, Duy Nguyen wrote:\n> >>>> On Tue, Jun 25, 2019 at 3:06 AM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> >>>>> I'm curious how big these EWAHs will be in practice and\n> >>>>> how useful an array of integers will be (especially as the\n> >>>>> pretty format will be one integer per line).  Perhaps it\n> >>>>> would helpful to have an extended example in one of the\n> >>>>> tests.\n> >>>>\n> >>>> It's one integer per updated entry. So if you have a giant index and\n> >>>> updated every single one of them, the EWAH bitmap contains that many\n> >>>> integers.\n> >>>>\n> >>>> If it was easy to just merge these bitmaps back to the entry (e.g. in\n> >>>> this example, add \"replaced\": true to entry zero) I would have done\n> >>>> it. But we dump as we stream and it's already too late to do it.\n> >>>>\n> >>>>> Would it be better to have the caller of ewah_each_bit()\n> >>>>> build a hex or bit string in a strbuf and then write it\n> >>>>> as a single string?\n> >>>>\n> >>>> I don't think the current EWAH representation is easy to read in the\n> >>>> first place. You'll probably have to run through some script to update\n> >>>> the main entries part and will have a much better view, but that's\n> >>>> pretty quick. If it's for scripts, then it's probably best to keep as\n> >>>> an array of integers, not a string. Less post processing.\n> >>>\n> >>> I don't think the intent is to dump the EWAH directly, but instead to\n> >>> dump a string of the uncompressed bitmap. Something like:\n> >>>\n> >>>          \"delete_bitmap\" : \"01101101101\"\n> >>>\n> >>> instead of\n> >>>\n> >>>          \"delete_bitmap\" : [ 0, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1 ]\n> >>\n> >> I get this part. But the numbers in the array were the position of the\n> >> set bits. It's not showing just the actual bit map.\n> >>\n> >> The same bitmap would be currently displayed as\n> >>\n> >>   \"delete_bitmap\": [ 1, 2, 4, 5, 7, 8, 9, 11 ]\n> >>\n> >> And that maps back to the entry[1], entry[2], entry[4]... in the index\n> >> being deleted from the base index. So displaying as a real bit map\n> >> actually adds more work for both the reader and the tool because you\n> >> have to calculate the position either way. And it gets harder if the\n> >> bit you're intereted in is on the far right.\n> >\n> >\n> > Thanks for the clarification.  That helps.\n>\n> Same here! We expect these to be much smaller than the full set, correct?\n\nFor split-index, the number of 1 bits should be about the size of your\nworking set, not the index size. In the normal case, then yes it\nshould be much smaller. After a big merge or branch switch, it could\nget as big as the index. But I would hope the logic to re-split the\nindex kicks in, which essentially empties these bitmaps.\n\nEWAH bitmap is also used in UNTR extension if I remember correctly.\nThose bitmaps may have as many bits as the directories you have in the\nindex.\n\n> Thanks,\n> -Stolee\n>\n\n\n-- \nDuy\n"},{"id":"378530","messageId":"20190703090844.GO21574@szeder.dev","threadId":"51372","inReplyTo":"20190624130226.17293-6-pclouds@gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-07-03T09:08:44Z","receivedAt":"2019-07-03T09:08:50Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Jun 24, 2019 at 08:02:21PM +0700, Nguyễn Thái Ngọc Duy wrote:\n> diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> index 082fe8e966..dbb572ce9d 100755\n> --- a/t/t3011-ls-files-json.sh\n> +++ b/t/t3011-ls-files-json.sh\n> @@ -44,4 +44,18 @@ test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n>  \tcompare_json basic\n>  '\n>  \n> +test_expect_success 'ls-files --json, split index' '\n> +\tgit init split &&\n> +\t(\n> +\t\tcd split &&\n> +\t\techo one >one &&\n> +\t\tgit add one &&\n> +\t\tgit update-index --split-index &&\n> +\t\techo updated >>one &&\n> +\t\ttest_must_fail git -c splitIndex.maxPercentChange=100 update-index --refresh &&\n> +\t\tcp ../filter.sed . &&\n> +\t\tcompare_json split-index\n> +\t)\n> +'\n\nI think this test should 'sane_unset GIT_TEST_SPLIT_INDEX'.  Maybe\nit's not absolutely necessary, because the explicit '--split-index'\nand '-c splitIndex.maxPercentChange=100' would already fully control\nwhen index splitting is performed, eliminating any indeterminism\ninherent to GIT_TEST_SPLIT_INDEX...  but unsetting it would reduce the\ncognitive load on future readers.\n\nThe same might apply to GIT_TEST_FSMONITOR in the following patch, and\nperhaps even to GIT_TEST_INDEX_THREADS.\n\n"},{"id":"378604","messageId":"20190704200133.GD20404@szeder.dev","threadId":"51372","inReplyTo":"20190624130226.17293-6-pclouds@gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-07-04T20:01:33Z","receivedAt":"2019-07-04T20:01:40Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Jun 24, 2019 at 08:02:21PM +0700, Nguyễn Thái Ngọc Duy wrote:\n> diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> index 082fe8e966..dbb572ce9d 100755\n> --- a/t/t3011-ls-files-json.sh\n> +++ b/t/t3011-ls-files-json.sh\n> @@ -44,4 +44,18 @@ test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n>  \tcompare_json basic\n>  '\n>  \n> +test_expect_success 'ls-files --json, split index' '\n> +\tgit init split &&\n> +\t(\n> +\t\tcd split &&\n> +\t\techo one >one &&\n> +\t\tgit add one &&\n> +\t\tgit update-index --split-index &&\n> +\t\techo updated >>one &&\n> +\t\ttest_must_fail git -c splitIndex.maxPercentChange=100 update-index --refresh &&\n> +\t\tcp ../filter.sed . &&\n> +\t\tcompare_json split-index\n> +\t)\n> +'\n> +\n>  test_done\n> diff --git a/t/t3011/split-index b/t/t3011/split-index\n> new file mode 100644\n> index 0000000000..cdcc4ddded\n> --- /dev/null\n> +++ b/t/t3011/split-index\n> @@ -0,0 +1,39 @@\n> +{\n> +  \"version\": 2,\n> +  \"oid\": <string>,\n> +  \"mtime_sec\": <number>,\n> +  \"mtime_nsec\": <number>,\n> +  \"entries\": [\n> +    {\n> +      \"id\": 0,\n> +      \"name\": \"\",\n> +      \"mode\": \"100644\",\n> +      \"flags\": 0,\n> +      \"oid\": <string>,\n> +      \"stat\": {\n> +        \"ctime_sec\": <number>,\n> +        \"ctime_nsec\": <number>,\n> +        \"mtime_sec\": <number>,\n> +        \"mtime_nsec\": <number>,\n> +        \"device\": <number>,\n> +        \"inode\": <number>,\n> +        \"uid\": <number>,\n> +        \"gid\": <number>,\n> +        \"size\": 4\n> +      },\n> +      \"file_offset\": <number>\n> +    }\n> +  ],\n> +  \"extensions\": {\n> +    \"link\": {\n> +      \"file_offset\": <number>,\n> +      \"ext_size\": <number>,\n> +      \"oid\": <string>,\n> +      \"delete_bitmap\": [\n> +      ],\n> +      \"replace_bitmap\": [\n> +        0\n> +      ]\n> +    }\n> +  }\n> +}\n\nThis test is flaky, as reported in:\n\n  https://public-inbox.org/git/xmqqftno2mku.fsf@gitster-ct.c.googlers.com/\n\nThis is because it relies on racy behaviour, namely that the following\nthree commands\n\n    echo one >one &&\n    git add one &&\n    git update-index --split-index &&\n\nare executed within the same second, leaving 'one' racily clean.  To\ndeal with the racily clean file, 5581a019ba (split-index: smudge and\nadd racily clean cache entries to split index, 2018-10-11) kicks in,\nand 'one's smudged index entry is stored both in the shared index and\nin the split index.  That's why this test expects the offset 0 in the\n\"replace_bitmap\" array.\n\nHowever, it's possible that a second boundary is crossed between\nwriting to 'one' and splitting the index, and then 'one' is not racily\nclean, and its index entry is only stored in the shared index.\nConsequently, there are no index entries in the split index, so the\n\"replace_bitmap\" array ends up being empty, ultimately failing the\ntest.\n\nA 'test-tool chmtime' invocation or two could make the test\ndeterministic (i.e it could make sure that 'one' is either always\nracily clean or it never is, whichever is preferred).\n\nWhat I still don't understand, however, is that when the test fails\nthis way, then the \"entries\" array ends up being empty as well.  It\nlooks as if the JSON dump included only index entries that were\nactually stored in '.git/index', but omitted entries that were only\npresent in the shared index.  I think this is wrong, and it should\ndump the unified view of the split and shared indexes.  Or include all\nentries from the shared index as well.  Or perhaps I'm completely\nmissing something...\n\n\n"},{"id":"378613","messageId":"CACsJy8CZZAkcuN_hqp6YmMkhKs0ON6b-+Cyo+Q+Jk4zFh0Ve7w@mail.gmail.com","threadId":"51372","inReplyTo":"20190704200133.GD20404@szeder.dev","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2019-07-04T23:54:49Z","receivedAt":"2019-07-04T23:55:18Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Jul 5, 2019 at 3:01 AM SZEDER Gábor <szeder.dev@gmail.com> wrote:\n>\n> On Mon, Jun 24, 2019 at 08:02:21PM +0700, Nguyễn Thái Ngọc Duy wrote:\n> > diff --git a/t/t3011-ls-files-json.sh b/t/t3011-ls-files-json.sh\n> > index 082fe8e966..dbb572ce9d 100755\n> > --- a/t/t3011-ls-files-json.sh\n> > +++ b/t/t3011-ls-files-json.sh\n> > @@ -44,4 +44,18 @@ test_expect_success 'ls-files --json, main entries, UNTR and TREE' '\n> >       compare_json basic\n> >  '\n> >\n> > +test_expect_success 'ls-files --json, split index' '\n> > +     git init split &&\n> > +     (\n> > +             cd split &&\n> > +             echo one >one &&\n> > +             git add one &&\n> > +             git update-index --split-index &&\n> > +             echo updated >>one &&\n> > +             test_must_fail git -c splitIndex.maxPercentChange=100 update-index --refresh &&\n> > +             cp ../filter.sed . &&\n> > +             compare_json split-index\n> > +     )\n> > +'\n> > +\n> >  test_done\n> > diff --git a/t/t3011/split-index b/t/t3011/split-index\n> > new file mode 100644\n> > index 0000000000..cdcc4ddded\n> > --- /dev/null\n> > +++ b/t/t3011/split-index\n> > @@ -0,0 +1,39 @@\n> > +{\n> > +  \"version\": 2,\n> > +  \"oid\": <string>,\n> > +  \"mtime_sec\": <number>,\n> > +  \"mtime_nsec\": <number>,\n> > +  \"entries\": [\n> > +    {\n> > +      \"id\": 0,\n> > +      \"name\": \"\",\n> > +      \"mode\": \"100644\",\n> > +      \"flags\": 0,\n> > +      \"oid\": <string>,\n> > +      \"stat\": {\n> > +        \"ctime_sec\": <number>,\n> > +        \"ctime_nsec\": <number>,\n> > +        \"mtime_sec\": <number>,\n> > +        \"mtime_nsec\": <number>,\n> > +        \"device\": <number>,\n> > +        \"inode\": <number>,\n> > +        \"uid\": <number>,\n> > +        \"gid\": <number>,\n> > +        \"size\": 4\n> > +      },\n> > +      \"file_offset\": <number>\n> > +    }\n> > +  ],\n> > +  \"extensions\": {\n> > +    \"link\": {\n> > +      \"file_offset\": <number>,\n> > +      \"ext_size\": <number>,\n> > +      \"oid\": <string>,\n> > +      \"delete_bitmap\": [\n> > +      ],\n> > +      \"replace_bitmap\": [\n> > +        0\n> > +      ]\n> > +    }\n> > +  }\n> > +}\n>\n> This test is flaky, as reported in:\n>\n>   https://public-inbox.org/git/xmqqftno2mku.fsf@gitster-ct.c.googlers.com/\n>\n> This is because it relies on racy behaviour, namely that the following\n> three commands\n>\n>     echo one >one &&\n>     git add one &&\n>     git update-index --split-index &&\n>\n> are executed within the same second, leaving 'one' racily clean.  To\n> deal with the racily clean file, 5581a019ba (split-index: smudge and\n> add racily clean cache entries to split index, 2018-10-11) kicks in,\n> and 'one's smudged index entry is stored both in the shared index and\n> in the split index.  That's why this test expects the offset 0 in the\n> \"replace_bitmap\" array.\n>\n> However, it's possible that a second boundary is crossed between\n> writing to 'one' and splitting the index, and then 'one' is not racily\n> clean, and its index entry is only stored in the shared index.\n> Consequently, there are no index entries in the split index, so the\n> \"replace_bitmap\" array ends up being empty, ultimately failing the\n> test.\n\nYep. I came up with the same conclusion. But I still have a couple\nother things to update before resending.\n\n> A 'test-tool chmtime' invocation or two could make the test\n> deterministic (i.e it could make sure that 'one' is either always\n> racily clean or it never is, whichever is preferred).\n>\n> What I still don't understand, however, is that when the test fails\n> this way, then the \"entries\" array ends up being empty as well.  It\n> looks as if the JSON dump included only index entries that were\n> actually stored in '.git/index', but omitted entries that were only\n> present in the shared index.  I think this is wrong, and it should\n> dump the unified view of the split and shared indexes.  Or include all\n> entries from the shared index as well.  Or perhaps I'm completely\n> missing something...\n\nThe command is to dump .git/index, not the shared one. And since this\nis not a split index test, rather a (quite low-level) json dump test,\nI did not bother to also dump the shared index, which should look like\na regular one. Producing a unified view in json might not be easy with\nthe current code because it's tied to file reading code, nearly stream\nout json as we read the file, and split-index requires a post\nprocessing step. I could contribute a python script or something to\ncombine shared/main index together. That way you can still see the\ncombined one, but we don't have to add/maintain more C code.\n-- \nDuy\n"},{"id":"378696","messageId":"xmqqd0ikz9ut.fsf@gitster-ct.c.googlers.com","threadId":"51372","inReplyTo":"CACsJy8CZZAkcuN_hqp6YmMkhKs0ON6b-+Cyo+Q+Jk4zFh0Ve7w@mail.gmail.com","subject":"Re: [PATCH v2 05/10] split-index.c: dump \"link\" extension as json","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-07-08T17:58:50Z","receivedAt":"2019-07-08T17:58:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> The command is to dump .git/index, not the shared one. And since this\n> is not a split index test, rather a (quite low-level) json dump test,\n> I did not bother to also dump the shared index, which should look like\n> a regular one. Producing a unified view in json might not be easy with\n> the current code because it's tied to file reading code, nearly stream\n> out json as we read the file, and split-index requires a post\n> processing step. I could contribute a python script or something to\n> combine shared/main index together. That way you can still see the\n> combined one, but we don't have to add/maintain more C code.\n\nWell, such a post-processing is something external scripts shine at\nand exporting the internal data in json format is exactly to support\nthese scripts, so it may make a good first test case ;-)\n"}]}