{"thread":{"id":"57313","subject":"[PATCH 0/5] scalar: implement the subcommand \"diagnose\"","startedAt":"2022-01-26T08:41:53Z","lastAt":"2022-06-20T09:42:30Z","messageCount":140,"participants":["Johannes Schindelin via GitGitGadget","Matthew John Cheetham via GitGitGadget","René Scharfe","Taylor Blau","Derrick Stolee","Elijah Newren","Johannes Schindelin","Junio C Hamano","rsbecker@nexbridge.com","Ævar Arnfjörð Bjarmason","Adam Dinwoodie"],"isPatch":true,"patchVersion":1,"patchTotal":5},"messages":[{"id":"446920","messageId":"pull.1128.git.1643186507.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":null,"subject":"[PATCH 0/5] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:42Z","receivedAt":"2022-01-26T08:41:53Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nJohannes Schindelin (3):\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n  scalar diagnose: show a spinner while staging content\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n contrib/scalar/scalar.c          | 336 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  17 ++\n 3 files changed, 365 insertions(+)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n-- \ngitgitgadget\n"},{"id":"446921","messageId":"ce85506e7a4313a4ae21ef712b84d8396ac45cdc.1643186507.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH 1/5] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:43Z","receivedAt":"2022-01-26T08:41:53Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` command: Instead of writing\nthe `.zip` file directly, we stage the file contents in a Git index of a\ntemporary, bare repository, only to let `git archive` have at it, and\nfinally removing the temporary repository.\n\nAlso note: Due to the frequently-spawned `git hash-object` processes,\nthis command is quite a bit slow on Windows. Should it turn out to be a\nbig problem, the lack of a batch mode of the `hash-object` command could\npotentially be worked around via using `git fast-import` with a crafted\n`stdin`.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 170 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  13 +++\n 3 files changed, 195 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..13f2b0f4d5a 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -259,6 +259,108 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int stage(const char *git_dir, struct strbuf *buf, const char *path)\n+{\n+\tstruct strbuf cacheinfo = STRBUF_INIT;\n+\tstruct child_process cp = CHILD_PROCESS_INIT;\n+\tint res;\n+\n+\tstrbuf_addstr(&cacheinfo, \"100644,\");\n+\n+\tcp.git_cmd = 1;\n+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n+\t\t     \"hash-object\", \"-w\", \"--stdin\", NULL);\n+\tres = pipe_command(&cp, buf->buf, buf->len, &cacheinfo, 256, NULL, 0);\n+\tif (!res) {\n+\t\tstrbuf_rtrim(&cacheinfo);\n+\t\tstrbuf_addch(&cacheinfo, ',');\n+\t\t/* We cannot stage `.git`, use `_git` instead. */\n+\t\tif (starts_with(path, \".git/\"))\n+\t\t\tstrbuf_addf(&cacheinfo, \"_%s\", path + 1);\n+\t\telse\n+\t\t\tstrbuf_addstr(&cacheinfo, path);\n+\n+\t\tchild_process_init(&cp);\n+\t\tcp.git_cmd = 1;\n+\t\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n+\t\t\t     \"update-index\", \"--add\", \"--cacheinfo\",\n+\t\t\t     cacheinfo.buf, NULL);\n+\t\tres = run_command(&cp);\n+\t}\n+\n+\tstrbuf_release(&cacheinfo);\n+\treturn res;\n+}\n+\n+static int stage_file(const char *git_dir, const char *path)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tint res;\n+\n+\tif (strbuf_read_file(&buf, path, 0) < 0)\n+\t\treturn error(_(\"could not read '%s'\"), path);\n+\n+\tres = stage(git_dir, &buf, path);\n+\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n+static int stage_directory(const char *git_dir, const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif ((e->d_type == DT_REG && stage_file(git_dir, buf.buf)) ||\n+\t\t    (e->d_type == DT_DIR && recurse &&\n+\t\t     stage_directory(git_dir, buf.buf, recurse)))\n+\t\t\tres = -1;\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n+static int index_to_zip(const char *git_dir)\n+{\n+\tstruct child_process cp = CHILD_PROCESS_INIT;\n+\tstruct strbuf oid = STRBUF_INIT;\n+\n+\tcp.git_cmd = 1;\n+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"write-tree\", NULL);\n+\tif (pipe_command(&cp, NULL, 0, &oid, the_hash_algo->hexsz + 1,\n+\t\t\t NULL, 0))\n+\t\treturn error(_(\"could not write temporary tree object\"));\n+\n+\tstrbuf_rtrim(&oid);\n+\tchild_process_init(&cp);\n+\tcp.git_cmd = 1;\n+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"archive\", \"-o\", NULL);\n+\tstrvec_pushf(&cp.args, \"%s.zip\", git_dir);\n+\tstrvec_pushl(&cp.args, oid.buf, \"--\", NULL);\n+\tstrbuf_release(&oid);\n+\treturn run_command(&cp);\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -499,6 +601,73 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf tmp_dir = STRBUF_INIT;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &buf);\n+\n+\tstrbuf_addstr(&buf, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&buf, \"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tif (run_git(\"init\", \"-q\", \"-b\", \"dummy\", \"--bare\", buf.buf, NULL)) {\n+\t\tres = error(_(\"could not initialize temporary repository: %s\"),\n+\t\t\t    buf.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstrbuf_realpath(&tmp_dir, buf.buf, 1);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf, \"Collecting diagnostic info into temp folder %s\\n\\n\",\n+\t\t    tmp_dir.buf);\n+\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tfwrite(buf.buf, buf.len, 1, stdout);\n+\n+\tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n+\t\tgoto diagnose_cleanup;\n+\n+\tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n+\t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n+\t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n+\t    (res = stage_directory(tmp_dir.buf, \".git/logs\", 1)) ||\n+\t    (res = stage_directory(tmp_dir.buf, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tres = index_to_zip(tmp_dir.buf);\n+\n+\tif (!res)\n+\t\tres = remove_dir_recursively(&tmp_dir, 0);\n+\n+\tif (!res)\n+\t\tprintf(\"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s.zip'\\n\",\n+\t\t       tmp_dir.buf);\n+\n+diagnose_cleanup:\n+\tstrbuf_release(&tmp_dir);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -800,6 +969,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..ecd06e207c2 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -65,6 +65,19 @@ test_expect_success 'scalar clone' '\n \t)\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar diagnose cloned >out &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_expect_success 'scalar reconfigure' '\n \tgit init one/src &&\n \tscalar register one &&\n-- \ngitgitgadget\n\n"},{"id":"446922","messageId":"f8885b27502408984f687e28b0a6fc9531287276.1643186507.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH 2/5] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:44Z","receivedAt":"2022-01-26T08:41:57Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c | 53 +++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 53 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 13f2b0f4d5a..e26fb2fc018 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -361,6 +361,58 @@ static int index_to_zip(const char *git_dir)\n \treturn run_command(&cp);\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -637,6 +689,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \tfwrite(buf.buf, buf.len, 1, stdout);\n \n \tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n-- \ngitgitgadget\n\n"},{"id":"446923","messageId":"330b36de799f82425c22bec50e6e42f0e495cab8.1643186507.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH 3/5] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:45Z","receivedAt":"2022-01-26T08:41:59Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nTeach the `scalar diagnose` command to gather file size information\nabout pack files.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n---\n contrib/scalar/scalar.c          | 39 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  2 ++\n 2 files changed, 41 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex e26fb2fc018..690933ffdf3 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -653,6 +653,39 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tstruct stat e_stat;\n+\tstruct strbuf file_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Contents of \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&file_path, path);\n+\tstrbuf_addch(&file_path, '/');\n+\tbase_path_len = file_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG) {\n+\t\t\tstrbuf_setlen(&file_path, base_path_len);\n+\t\t\tstrbuf_addstr(&file_path, e->d_name);\n+\t\t\tif (!stat(file_path.buf, &e_stat))\n+\t\t\t\tstrbuf_addf(buf, \"%-70s %16\"PRIuMAX\"\\n\",\n+\t\t\t\t\t    e->d_name,\n+\t\t\t\t\t    (uintmax_t)e_stat.st_size);\n+\t\t}\n+\n+\tstrbuf_release(&file_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -695,6 +728,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n \t\tgoto diagnose_cleanup;\n \n+\tstrbuf_reset(&buf);\n+\tdir_file_stats(&buf, \".git/objects/pack\");\n+\n+\tif ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n+\t\tgoto diagnose_cleanup;\n+\n \tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n \t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n \t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex ecd06e207c2..b1745851e31 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -75,6 +75,8 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n \ttest_file_not_empty out\n '\n \n-- \ngitgitgadget\n\n"},{"id":"446924","messageId":"213f2c94b73f90fc758c2e3872804cf640cb2005.1643186507.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH 4/5] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:46Z","receivedAt":"2022-01-26T08:42:03Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers or large sizes of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n---\n contrib/scalar/scalar.c          | 60 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  2 ++\n 2 files changed, 62 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 690933ffdf3..c0ad4948215 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -686,6 +686,60 @@ static void dir_file_stats(struct strbuf *buf, const char *path)\n \tclosedir(dir);\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -734,6 +788,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \tif ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n \t\tgoto diagnose_cleanup;\n \n+\tstrbuf_reset(&buf);\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\n+\tif ((res = stage(tmp_dir.buf, &buf, \"objects-local.txt\")))\n+\t\tgoto diagnose_cleanup;\n+\n \tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n \t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n \t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex b1745851e31..f2ec156d819 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -77,6 +77,8 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n \ttest_file_not_empty out\n '\n \n-- \ngitgitgadget\n\n"},{"id":"446925","messageId":"3a2cdce554a1755210592f3b9bb056d936ca98b6.1643186507.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH 5/5] scalar diagnose: show a spinner while staging content","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-01-26T08:41:47Z","receivedAt":"2022-01-26T08:42:04Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nIt can take a while to gather all the information that `scalar diagnose`\nwants to accumulate. Typically this happens when the user is in need of\nquick solutions and therefore their patience is tested already. By\nshowing a little spinner that spins around, we hope to help the user\nmuster just a tiny bit more patience until `scalar diagnose` is done.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c | 14 ++++++++++++++\n 1 file changed, 14 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex c0ad4948215..224329f38f5 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -259,12 +259,26 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static void spinner(void)\n+{\n+\tstatic const char whee[] = \"|\\010/\\010-\\010\\\\\\010\", *next = whee;\n+\n+\tif (!next)\n+\t\treturn;\n+\tif (write(2, next, 2) < 0)\n+\t\tnext = NULL;\n+\telse\n+\t\tnext = next[2] ? next + 2 : whee;\n+}\n+\n static int stage(const char *git_dir, struct strbuf *buf, const char *path)\n {\n \tstruct strbuf cacheinfo = STRBUF_INIT;\n \tstruct child_process cp = CHILD_PROCESS_INIT;\n \tint res;\n \n+\tspinner();\n+\n \tstrbuf_addstr(&cacheinfo, \"100644,\");\n \n \tcp.git_cmd = 1;\n-- \ngitgitgadget\n"},{"id":"446927","messageId":"f48cdebb-ff49-24b0-973f-b3e7954e11c8@web.de","threadId":"57313","inReplyTo":"ce85506e7a4313a4ae21ef712b84d8396ac45cdc.1643186507.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/5] Implement `scalar diagnose`","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-01-26T09:34:04Z","receivedAt":"2022-01-26T09:34:09Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 26.01.22 um 09:41 schrieb Johannes Schindelin via GitGitGadget:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> Over the course of Scalar's development, it became obvious that there is\n> a need for a command that can gather all kinds of useful information\n> that can help identify the most typical problems with large\n> worktrees/repositories.\n>\n> The `diagnose` command is the culmination of this hard-won knowledge: it\n> gathers the installed hooks, the config, a couple statistics describing\n> the data shape, among other pieces of information, and then wraps\n> everything up in a tidy, neat `.zip` archive.\n>\n> Note: originally, Scalar was implemented in C# using the .NET API, where\n> we had the luxury of a comprehensive standard library that includes\n> basic functionality such as writing a `.zip` file. In the C version, we\n> lack such a commodity. Rather than introducing a dependency on, say,\n> libzip, we slightly abuse Git's `archive` command: Instead of writing\n> the `.zip` file directly, we stage the file contents in a Git index of a\n> temporary, bare repository, only to let `git archive` have at it, and\n> finally removing the temporary repository.\n\ngit archive allows you to include untracked files in an archive with its\noption --add-file.  You can see an example in Git's Makefile; search for\nGIT_ARCHIVE_EXTRA_FILES.  It still requires a tree argument, but the\nempty tree object should suffice if you don't want to include any\ntracked files.  It doesn't currently support streaming, though, i.e.\nfiles are fully read into memory, so it's impractical for huge ones.\n\n> Also note: Due to the frequently-spawned `git hash-object` processes,\n> this command is quite a bit slow on Windows. Should it turn out to be a\n> big problem, the lack of a batch mode of the `hash-object` command could\n> potentially be worked around via using `git fast-import` with a crafted\n> `stdin`.\n\nOr we could add streaming support to git archive --add-file..\n\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  contrib/scalar/scalar.c          | 170 +++++++++++++++++++++++++++++++\n>  contrib/scalar/scalar.txt        |  12 +++\n>  contrib/scalar/t/t9099-scalar.sh |  13 +++\n>  3 files changed, 195 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 1ce9c2b00e8..13f2b0f4d5a 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -259,6 +259,108 @@ static int unregister_dir(void)\n>  \treturn res;\n>  }\n>\n> +static int stage(const char *git_dir, struct strbuf *buf, const char *path)\n> +{\n> +\tstruct strbuf cacheinfo = STRBUF_INIT;\n> +\tstruct child_process cp = CHILD_PROCESS_INIT;\n> +\tint res;\n> +\n> +\tstrbuf_addstr(&cacheinfo, \"100644,\");\n> +\n> +\tcp.git_cmd = 1;\n> +\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n> +\t\t     \"hash-object\", \"-w\", \"--stdin\", NULL);\n> +\tres = pipe_command(&cp, buf->buf, buf->len, &cacheinfo, 256, NULL, 0);\n> +\tif (!res) {\n> +\t\tstrbuf_rtrim(&cacheinfo);\n> +\t\tstrbuf_addch(&cacheinfo, ',');\n> +\t\t/* We cannot stage `.git`, use `_git` instead. */\n> +\t\tif (starts_with(path, \".git/\"))\n> +\t\t\tstrbuf_addf(&cacheinfo, \"_%s\", path + 1);\n> +\t\telse\n> +\t\t\tstrbuf_addstr(&cacheinfo, path);\n> +\n> +\t\tchild_process_init(&cp);\n> +\t\tcp.git_cmd = 1;\n> +\t\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n> +\t\t\t     \"update-index\", \"--add\", \"--cacheinfo\",\n> +\t\t\t     cacheinfo.buf, NULL);\n> +\t\tres = run_command(&cp);\n> +\t}\n> +\n> +\tstrbuf_release(&cacheinfo);\n> +\treturn res;\n> +}\n> +\n> +static int stage_file(const char *git_dir, const char *path)\n> +{\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tint res;\n> +\n> +\tif (strbuf_read_file(&buf, path, 0) < 0)\n> +\t\treturn error(_(\"could not read '%s'\"), path);\n> +\n> +\tres = stage(git_dir, &buf, path);\n> +\n> +\tstrbuf_release(&buf);\n> +\treturn res;\n> +}\n> +\n> +static int stage_directory(const char *git_dir, const char *path, int recurse)\n> +{\n> +\tint at_root = !*path;\n> +\tDIR *dir = opendir(at_root ? \".\" : path);\n> +\tstruct dirent *e;\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tsize_t len;\n> +\tint res = 0;\n> +\n> +\tif (!dir)\n> +\t\treturn error(_(\"could not open directory '%s'\"), path);\n> +\n> +\tif (!at_root)\n> +\t\tstrbuf_addf(&buf, \"%s/\", path);\n> +\tlen = buf.len;\n> +\n> +\twhile (!res && (e = readdir(dir))) {\n> +\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n> +\t\t\tcontinue;\n> +\n> +\t\tstrbuf_setlen(&buf, len);\n> +\t\tstrbuf_addstr(&buf, e->d_name);\n> +\n> +\t\tif ((e->d_type == DT_REG && stage_file(git_dir, buf.buf)) ||\n> +\t\t    (e->d_type == DT_DIR && recurse &&\n> +\t\t     stage_directory(git_dir, buf.buf, recurse)))\n> +\t\t\tres = -1;\n> +\t}\n> +\n> +\tclosedir(dir);\n> +\tstrbuf_release(&buf);\n> +\treturn res;\n> +}\n> +\n> +static int index_to_zip(const char *git_dir)\n> +{\n> +\tstruct child_process cp = CHILD_PROCESS_INIT;\n> +\tstruct strbuf oid = STRBUF_INIT;\n> +\n> +\tcp.git_cmd = 1;\n> +\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"write-tree\", NULL);\n> +\tif (pipe_command(&cp, NULL, 0, &oid, the_hash_algo->hexsz + 1,\n> +\t\t\t NULL, 0))\n> +\t\treturn error(_(\"could not write temporary tree object\"));\n> +\n> +\tstrbuf_rtrim(&oid);\n> +\tchild_process_init(&cp);\n> +\tcp.git_cmd = 1;\n> +\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"archive\", \"-o\", NULL);\n> +\tstrvec_pushf(&cp.args, \"%s.zip\", git_dir);\n> +\tstrvec_pushl(&cp.args, oid.buf, \"--\", NULL);\n> +\tstrbuf_release(&oid);\n> +\treturn run_command(&cp);\n> +}\n> +\n>  /* printf-style interface, expects `<key>=<value>` argument */\n>  static int set_config(const char *fmt, ...)\n>  {\n> @@ -499,6 +601,73 @@ cleanup:\n>  \treturn res;\n>  }\n>\n> +static int cmd_diagnose(int argc, const char **argv)\n> +{\n> +\tstruct option options[] = {\n> +\t\tOPT_END(),\n> +\t};\n> +\tconst char * const usage[] = {\n> +\t\tN_(\"scalar diagnose [<enlistment>]\"),\n> +\t\tNULL\n> +\t};\n> +\tstruct strbuf tmp_dir = STRBUF_INIT;\n> +\ttime_t now = time(NULL);\n> +\tstruct tm tm;\n> +\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n> +\tint res = 0;\n> +\n> +\targc = parse_options(argc, argv, NULL, options,\n> +\t\t\t     usage, 0);\n> +\n> +\tsetup_enlistment_directory(argc, argv, usage, options, &buf);\n> +\n> +\tstrbuf_addstr(&buf, \"/.scalarDiagnostics/scalar_\");\n> +\tstrbuf_addftime(&buf, \"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n> +\tif (run_git(\"init\", \"-q\", \"-b\", \"dummy\", \"--bare\", buf.buf, NULL)) {\n> +\t\tres = error(_(\"could not initialize temporary repository: %s\"),\n> +\t\t\t    buf.buf);\n> +\t\tgoto diagnose_cleanup;\n> +\t}\n> +\tstrbuf_realpath(&tmp_dir, buf.buf, 1);\n> +\n> +\tstrbuf_reset(&buf);\n> +\tstrbuf_addf(&buf, \"Collecting diagnostic info into temp folder %s\\n\\n\",\n> +\t\t    tmp_dir.buf);\n> +\n> +\tget_version_info(&buf, 1);\n> +\n> +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n> +\tfwrite(buf.buf, buf.len, 1, stdout);\n> +\n> +\tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n> +\t\tgoto diagnose_cleanup;\n> +\n> +\tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n> +\t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n> +\t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n> +\t    (res = stage_directory(tmp_dir.buf, \".git/logs\", 1)) ||\n> +\t    (res = stage_directory(tmp_dir.buf, \".git/objects/info\", 0)))\n> +\t\tgoto diagnose_cleanup;\n> +\n> +\tres = index_to_zip(tmp_dir.buf);\n> +\n> +\tif (!res)\n> +\t\tres = remove_dir_recursively(&tmp_dir, 0);\n> +\n> +\tif (!res)\n> +\t\tprintf(\"\\n\"\n> +\t\t       \"Diagnostics complete.\\n\"\n> +\t\t       \"All of the gathered info is captured in '%s.zip'\\n\",\n> +\t\t       tmp_dir.buf);\n> +\n> +diagnose_cleanup:\n> +\tstrbuf_release(&tmp_dir);\n> +\tstrbuf_release(&path);\n> +\tstrbuf_release(&buf);\n> +\n> +\treturn res;\n> +}\n> +\n>  static int cmd_list(int argc, const char **argv)\n>  {\n>  \tif (argc != 1)\n> @@ -800,6 +969,7 @@ static struct {\n>  \t{ \"reconfigure\", cmd_reconfigure },\n>  \t{ \"delete\", cmd_delete },\n>  \t{ \"version\", cmd_version },\n> +\t{ \"diagnose\", cmd_diagnose },\n>  \t{ NULL, NULL},\n>  };\n>\n> diff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\n> index f416d637289..22583fe046e 100644\n> --- a/contrib/scalar/scalar.txt\n> +++ b/contrib/scalar/scalar.txt\n> @@ -14,6 +14,7 @@ scalar register [<enlistment>]\n>  scalar unregister [<enlistment>]\n>  scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n>  scalar reconfigure [ --all | <enlistment> ]\n> +scalar diagnose [<enlistment>]\n>  scalar delete <enlistment>\n>\n>  DESCRIPTION\n> @@ -129,6 +130,17 @@ reconfigure the enlistment.\n>  With the `--all` option, all enlistments currently registered with Scalar\n>  will be reconfigured. Use this option after each Scalar upgrade.\n>\n> +Diagnose\n> +~~~~~~~~\n> +\n> +diagnose [<enlistment>]::\n> +    When reporting issues with Scalar, it is often helpful to provide the\n> +    information gathered by this command, including logs and certain\n> +    statistics describing the data shape of the current enlistment.\n> ++\n> +The output of this command is a `.zip` file that is written into\n> +a directory adjacent to the worktree in the `src` directory.\n> +\n>  Delete\n>  ~~~~~~\n>\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index 2e1502ad45e..ecd06e207c2 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -65,6 +65,19 @@ test_expect_success 'scalar clone' '\n>  \t)\n>  '\n>\n> +SQ=\"'\"\n> +test_expect_success UNZIP 'scalar diagnose' '\n> +\tscalar diagnose cloned >out &&\n> +\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n> +\tzip_path=$(cat zip_path) &&\n> +\ttest -n \"$zip_path\" &&\n> +\tunzip -v \"$zip_path\" &&\n> +\tfolder=${zip_path%.zip} &&\n> +\ttest_path_is_missing \"$folder\" &&\n> +\tunzip -p \"$zip_path\" diagnostics.log >out &&\n> +\ttest_file_not_empty out\n> +'\n> +\n>  test_expect_success 'scalar reconfigure' '\n>  \tgit init one/src &&\n>  \tscalar register one &&\n\n"},{"id":"447009","messageId":"YfHJHbMKA1u+A9LF@nand.local","threadId":"57313","inReplyTo":"f48cdebb-ff49-24b0-973f-b3e7954e11c8@web.de","subject":"Re: [PATCH 1/5] Implement `scalar diagnose`","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-01-26T22:20:13Z","receivedAt":"2022-01-26T22:20:17Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 26, 2022 at 10:34:04AM +0100, René Scharfe wrote:\n> Am 26.01.22 um 09:41 schrieb Johannes Schindelin via GitGitGadget:\n> > Note: originally, Scalar was implemented in C# using the .NET API, where\n> > we had the luxury of a comprehensive standard library that includes\n> > basic functionality such as writing a `.zip` file. In the C version, we\n> > lack such a commodity. Rather than introducing a dependency on, say,\n> > libzip, we slightly abuse Git's `archive` command: Instead of writing\n> > the `.zip` file directly, we stage the file contents in a Git index of a\n> > temporary, bare repository, only to let `git archive` have at it, and\n> > finally removing the temporary repository.\n>\n> git archive allows you to include untracked files in an archive with its\n> option --add-file.  You can see an example in Git's Makefile; search for\n> GIT_ARCHIVE_EXTRA_FILES.  It still requires a tree argument, but the\n> empty tree object should suffice if you don't want to include any\n> tracked files.  It doesn't currently support streaming, though, i.e.\n> files are fully read into memory, so it's impractical for huge ones.\n\nUsing `--add-file` would likely be preferable to setting up a temporary\nrepository just to invoke `git archive` in it. Johannes would be the\nexpert to ask whether or not big files are going to be a problem here\n(based on a cursory scan of the new functions in scalar.c, I don't\nexpect this to be the case).\n\nThe new stage_directory() function _could_ add `--add-file` arguments in\na loop around readdir(), but it might also be nice to add a new\n`--add-directory` function to `git archive` which would do the \"heavy\"\nlifting for us.\n\nThanks,\nTaylor\n"},{"id":"447011","messageId":"YfHOo8Mf3RP4j0Y6@nand.local","threadId":"57313","inReplyTo":"330b36de799f82425c22bec50e6e42f0e495cab8.1643186507.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 3/5] scalar: teach `diagnose` to gather packfile info","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-01-26T22:43:47Z","receivedAt":"2022-01-26T22:43:50Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 26, 2022 at 08:41:45AM +0000, Matthew John Cheetham via GitGitGadget wrote:\n> From: Matthew John Cheetham <mjcheetham@outlook.com>\n>\n> Teach the `scalar diagnose` command to gather file size information\n> about pack files.\n>\n> Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n> ---\n>  contrib/scalar/scalar.c          | 39 ++++++++++++++++++++++++++++++++\n>  contrib/scalar/t/t9099-scalar.sh |  2 ++\n>  2 files changed, 41 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index e26fb2fc018..690933ffdf3 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -653,6 +653,39 @@ cleanup:\n>  \treturn res;\n>  }\n>\n> +static void dir_file_stats(struct strbuf *buf, const char *path)\n> +{\n> +\tDIR *dir = opendir(path);\n> +\tstruct dirent *e;\n> +\tstruct stat e_stat;\n> +\tstruct strbuf file_path = STRBUF_INIT;\n> +\tsize_t base_path_len;\n> +\n> +\tif (!dir)\n> +\t\treturn;\n> +\n> +\tstrbuf_addstr(buf, \"Contents of \");\n> +\tstrbuf_add_absolute_path(buf, path);\n> +\tstrbuf_addstr(buf, \":\\n\");\n> +\n> +\tstrbuf_add_absolute_path(&file_path, path);\n> +\tstrbuf_addch(&file_path, '/');\n> +\tbase_path_len = file_path.len;\n> +\n> +\twhile ((e = readdir(dir)) != NULL)\n\nHmm. Is there a reason that this couldn't use\nfor_each_file_in_pack_dir() with a callback that just does the stat()\nand buffer manipulation?\n\nI don't think it's critical either way, but it would eliminate some of\nthe boilerplate that is shared between this implementation and the one\nthat already exists in for_each_file_in_pack_dir().\n\n> +\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG) {\n> +\t\t\tstrbuf_setlen(&file_path, base_path_len);\n> +\t\t\tstrbuf_addstr(&file_path, e->d_name);\n\nFor what it's worth, I think the callback would start here:\n\n> +\t\t\tif (!stat(file_path.buf, &e_stat))\n> +\t\t\t\tstrbuf_addf(buf, \"%-70s %16\"PRIuMAX\"\\n\",\n> +\t\t\t\t\t    e->d_name,\n> +\t\t\t\t\t    (uintmax_t)e_stat.st_size);\n\n...and end here.\n\nThanks,\nTaylor\n"},{"id":"447012","messageId":"YfHQSrdkieNuBEXT@nand.local","threadId":"57313","inReplyTo":"213f2c94b73f90fc758c2e3872804cf640cb2005.1643186507.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/5] scalar: teach `diagnose` to gather loose objects information","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2022-01-26T22:50:50Z","receivedAt":"2022-01-26T22:50:53Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Jan 26, 2022 at 08:41:46AM +0000, Matthew John Cheetham via GitGitGadget wrote:\n> +\twhile ((e = readdir(dir)) != NULL)\n> +\t\tif (!is_dot_or_dotdot(e->d_name) &&\n> +\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n> +\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n\nWhat is this call to hex_to_bytes() for? I assume it's checking to make\nsure the directory we're looking at is one of the shards of loose\nobjects.\n\nSimilar to my suggestion on the previous patch, I think that we could\nget rid of this function entirely and replace it with a call to\nfor_each_loose_file_in_objdir().\n\nWe'll pay a little bit of extra cost to parse out each loose object's\nOID, but it should be negligible since we're not actually opening up\neach object.\n\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index b1745851e31..f2ec156d819 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -77,6 +77,8 @@ test_expect_success UNZIP 'scalar diagnose' '\n>  \tunzip -p \"$zip_path\" diagnostics.log >out &&\n>  \ttest_file_not_empty out &&\n>  \tunzip -p \"$zip_path\" packs-local.txt >out &&\n> +\ttest_file_not_empty out &&\n\nA more comprehensive test (here, and in the earlier instances, too)\nmight be useful beyond just \"does this file exist in the archive\".\n\nConstructing an example repository where the number of loose objects is\nknown ahead of time, and then finding that number in the output of\nobjects-local.txt might be worthwhile to give us some extra confidence\nthat this is working as intended.\n\n> +\tunzip -p \"$zip_path\" objects-local.txt >out &&\n>  \ttest_file_not_empty out\n>  '\n\nThanks,\nTaylor\n"},{"id":"447085","messageId":"0a52155c-4605-d96f-965a-104a399ae86e@gmail.com","threadId":"57313","inReplyTo":"YfHOo8Mf3RP4j0Y6@nand.local","subject":"Re: [PATCH 3/5] scalar: teach `diagnose` to gather packfile info","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2022-01-27T15:14:28Z","receivedAt":"2022-01-27T15:14:33Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/26/2022 5:43 PM, Taylor Blau wrote:\n> On Wed, Jan 26, 2022 at 08:41:45AM +0000, Matthew John Cheetham via GitGitGadget wrote:\n>> From: Matthew John Cheetham <mjcheetham@outlook.com>\n>>\n>> Teach the `scalar diagnose` command to gather file size information\n>> about pack files.\n>>\n>> Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n>> ---\n>>  contrib/scalar/scalar.c          | 39 ++++++++++++++++++++++++++++++++\n>>  contrib/scalar/t/t9099-scalar.sh |  2 ++\n>>  2 files changed, 41 insertions(+)\n>>\n>> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n>> index e26fb2fc018..690933ffdf3 100644\n>> --- a/contrib/scalar/scalar.c\n>> +++ b/contrib/scalar/scalar.c\n>> @@ -653,6 +653,39 @@ cleanup:\n>>  \treturn res;\n>>  }\n>>\n>> +static void dir_file_stats(struct strbuf *buf, const char *path)\n>> +{\n>> +\tDIR *dir = opendir(path);\n>> +\tstruct dirent *e;\n>> +\tstruct stat e_stat;\n>> +\tstruct strbuf file_path = STRBUF_INIT;\n>> +\tsize_t base_path_len;\n>> +\n>> +\tif (!dir)\n>> +\t\treturn;\n>> +\n>> +\tstrbuf_addstr(buf, \"Contents of \");\n>> +\tstrbuf_add_absolute_path(buf, path);\n>> +\tstrbuf_addstr(buf, \":\\n\");\n>> +\n>> +\tstrbuf_add_absolute_path(&file_path, path);\n>> +\tstrbuf_addch(&file_path, '/');\n>> +\tbase_path_len = file_path.len;\n>> +\n>> +\twhile ((e = readdir(dir)) != NULL)\n> \n> Hmm. Is there a reason that this couldn't use\n> for_each_file_in_pack_dir() with a callback that just does the stat()\n> and buffer manipulation?\n> \n> I don't think it's critical either way, but it would eliminate some of\n> the boilerplate that is shared between this implementation and the one\n> that already exists in for_each_file_in_pack_dir().\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. This method is also extended in microsoft/git to\nscan the alternates directory (which we expect to exist as the\n\"shared objects cache).\n\nWe might want to modify the implementation in this series to\nrun dir_file_stats() on each odb in the_repository. This would\ngive us the data for the shared object cache for free while\nbeing more general to other Git repos. (It would require us to\ndo some reaction work in microsoft/git and be a change of\nbehavior, but we are the only ones who have looked at these\ndiagnose files before, so that change will be easy to manage.)\n\nThanks,\n-Stolee\n"},{"id":"447086","messageId":"027ebd30-77c7-1f09-9fe1-0523b9487319@gmail.com","threadId":"57313","inReplyTo":"YfHQSrdkieNuBEXT@nand.local","subject":"Re: [PATCH 4/5] scalar: teach `diagnose` to gather loose objects information","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2022-01-27T15:17:07Z","receivedAt":"2022-01-27T15:17:20Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/26/2022 5:50 PM, Taylor Blau wrote:\n> On Wed, Jan 26, 2022 at 08:41:46AM +0000, Matthew John Cheetham via GitGitGadget wrote:\n>> +\twhile ((e = readdir(dir)) != NULL)\n>> +\t\tif (!is_dot_or_dotdot(e->d_name) &&\n>> +\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n>> +\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n> \n> What is this call to hex_to_bytes() for? I assume it's checking to make\n> sure the directory we're looking at is one of the shards of loose\n> objects.\n> \n> Similar to my suggestion on the previous patch, I think that we could\n> get rid of this function entirely and replace it with a call to\n> for_each_loose_file_in_objdir().\n\nThere is a possibility that there are files other than loose objects\nin these directories, so summarizing those counts might be helpful\ninformation. For example: if somehow .git/objects/00/ was full of a\nbunch of non-objects, it would still slow down Git commands that ask\nfor a short-sha starting with \"00\".\n\nWhile this shouldn't be a normal case, the 'diagnose' command is\nbuilt to help us find these extremely odd scenarios because they\n_have_ happened before (typically because of a VFS for Git bug\ntaught us how to look for these situations).\n\nThanks,\n-Stolee\n"},{"id":"447087","messageId":"dda1b3c8-afc6-d9d5-1bfb-4a48ac87ca54@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/5] scalar: implement the subcommand \"diagnose\"","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2022-01-27T15:19:20Z","receivedAt":"2022-01-27T15:19:26Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 1/26/2022 3:41 AM, Johannes Schindelin via GitGitGadget wrote:\n> Over the course of the years, we developed a sub-command that gathers\n> diagnostic data into a .zip file that can then be attached to bug reports.\n> This sub-command turned out to be very useful in helping Scalar developers\n> identify and fix issues.\n\nFor historical context: The 'diagnose' command was implemented in VFS for\nGit and ported to the C# version of Scalar before 'git bugreport' existed,\nbut they serve very similar purposes.\n\nI wonder if 'scalar diagnose' could include some of the information\ncaptured by 'git bugreport' or whether this implementation of 'diagnose'\ncould help inform 'git bugreport' in any way.\n\nCC'ing Emily for thoughts.\n\nThanks,\n-Stolee\n"},{"id":"447111","messageId":"CABPp-BHeLzinXkX3WgqBNYntJwY_ZAm5D7VdOR7KQahvLOuV=w@mail.gmail.com","threadId":"57313","inReplyTo":"213f2c94b73f90fc758c2e3872804cf640cb2005.1643186507.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/5] scalar: teach `diagnose` to gather loose objects information","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-01-27T18:59:08Z","receivedAt":"2022-01-27T18:59:21Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Jan 26, 2022 at 3:37 PM Matthew John Cheetham via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Matthew John Cheetham <mjcheetham@outlook.com>\n>\n> When operating at the scale that Scalar wants to support, certain data\n> shapes are more likely to cause undesirable performance issues, such as\n> large numbers or large sizes of loose objects.\n\nMakes sense.\n\n> By including statistics about this, `scalar diagnose` now makes it\n> easier to identify such scenarios.\n>\n> Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n> ---\n>  contrib/scalar/scalar.c          | 60 ++++++++++++++++++++++++++++++++\n>  contrib/scalar/t/t9099-scalar.sh |  2 ++\n>  2 files changed, 62 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 690933ffdf3..c0ad4948215 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -686,6 +686,60 @@ static void dir_file_stats(struct strbuf *buf, const char *path)\n>         closedir(dir);\n>  }\n>\n> +static int count_files(char *path)\n> +{\n> +       DIR *dir = opendir(path);\n> +       struct dirent *e;\n> +       int count = 0;\n> +\n> +       if (!dir)\n> +               return 0;\n> +\n> +       while ((e = readdir(dir)) != NULL)\n> +               if (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n> +                       count++;\n> +\n> +       closedir(dir);\n> +       return count;\n> +}\n> +\n> +static void loose_objs_stats(struct strbuf *buf, const char *path)\n> +{\n> +       DIR *dir = opendir(path);\n> +       struct dirent *e;\n> +       int count;\n> +       int total = 0;\n> +       unsigned char c;\n> +       struct strbuf count_path = STRBUF_INIT;\n> +       size_t base_path_len;\n> +\n> +       if (!dir)\n> +               return;\n> +\n> +       strbuf_addstr(buf, \"Object directory stats for \");\n> +       strbuf_add_absolute_path(buf, path);\n> +       strbuf_addstr(buf, \":\\n\");\n> +\n> +       strbuf_add_absolute_path(&count_path, path);\n> +       strbuf_addch(&count_path, '/');\n> +       base_path_len = count_path.len;\n> +\n> +       while ((e = readdir(dir)) != NULL)\n> +               if (!is_dot_or_dotdot(e->d_name) &&\n> +                   e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n> +                   !hex_to_bytes(&c, e->d_name, 1)) {\n\nYou only recurse into directories, ignoring individual files.\n\n> +                       strbuf_setlen(&count_path, base_path_len);\n> +                       strbuf_addstr(&count_path, e->d_name);\n> +                       total += (count = count_files(count_path.buf));\n> +                       strbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n\nThis shows the number of files within a directory.\n\n> +               }\n> +\n> +       strbuf_addf(buf, \"Total: %d loose objects\", total);\n\nand this shows the total number of files across all the directories.\n\nBut the commit message suggested you also wanted to check for large\nsizes of loose objects.  Did that get ripped out at some point with\nthe commit message not being updated, or is it perhaps going to be\nincluded later?\n\n> +\n> +       strbuf_release(&count_path);\n> +       closedir(dir);\n> +}\n> +\n>  static int cmd_diagnose(int argc, const char **argv)\n>  {\n>         struct option options[] = {\n> @@ -734,6 +788,12 @@ static int cmd_diagnose(int argc, const char **argv)\n>         if ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n>                 goto diagnose_cleanup;\n>\n> +       strbuf_reset(&buf);\n> +       loose_objs_stats(&buf, \".git/objects\");\n> +\n> +       if ((res = stage(tmp_dir.buf, &buf, \"objects-local.txt\")))\n> +               goto diagnose_cleanup;\n> +\n>         if ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n>             (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n>             (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index b1745851e31..f2ec156d819 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -77,6 +77,8 @@ test_expect_success UNZIP 'scalar diagnose' '\n>         unzip -p \"$zip_path\" diagnostics.log >out &&\n>         test_file_not_empty out &&\n>         unzip -p \"$zip_path\" packs-local.txt >out &&\n> +       test_file_not_empty out &&\n> +       unzip -p \"$zip_path\" objects-local.txt >out &&\n>         test_file_not_empty out\n>  '\n>\n> --\n> gitgitgadget\n"},{"id":"447121","messageId":"CABPp-BHob22kAHRWBX-QLyQFKWn-682xQnx5oehqW=WYO4PBDQ@mail.gmail.com","threadId":"57313","inReplyTo":"ce85506e7a4313a4ae21ef712b84d8396ac45cdc.1643186507.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/5] Implement `scalar diagnose`","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-01-27T19:38:16Z","receivedAt":"2022-01-27T19:38:31Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Jan 26, 2022 at 3:37 PM Johannes Schindelin via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> Over the course of Scalar's development, it became obvious that there is\n> a need for a command that can gather all kinds of useful information\n> that can help identify the most typical problems with large\n> worktrees/repositories.\n>\n> The `diagnose` command is the culmination of this hard-won knowledge: it\n> gathers the installed hooks, the config, a couple statistics describing\n> the data shape, among other pieces of information, and then wraps\n> everything up in a tidy, neat `.zip` archive.\n>\n> Note: originally, Scalar was implemented in C# using the .NET API, where\n> we had the luxury of a comprehensive standard library that includes\n> basic functionality such as writing a `.zip` file. In the C version, we\n> lack such a commodity. Rather than introducing a dependency on, say,\n> libzip, we slightly abuse Git's `archive` command: Instead of writing\n> the `.zip` file directly, we stage the file contents in a Git index of a\n> temporary, bare repository, only to let `git archive` have at it, and\n> finally removing the temporary repository.\n>\n> Also note: Due to the frequently-spawned `git hash-object` processes,\n> this command is quite a bit slow on Windows. Should it turn out to be a\n> big problem, the lack of a batch mode of the `hash-object` command could\n> potentially be worked around via using `git fast-import` with a crafted\n> `stdin`.\n\nhash-object and update-index processes, right?  You spawn one of each\nfor each object.\n\nI was you investigate the fast-import idea because it gets rid of the\nN hash-object processes, the N update-index processes, and the\nwrite-tree process, instead giving you a single fast-import process as\na preliminary to calling out to git archive.  It'd also have the\nadvantage of providing just one pack instead of many loose objects.\n\nBut René's suggestion to use and extend archive's ability to handle\nuntracked files sounds like a better idea.\n\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  contrib/scalar/scalar.c          | 170 +++++++++++++++++++++++++++++++\n>  contrib/scalar/scalar.txt        |  12 +++\n>  contrib/scalar/t/t9099-scalar.sh |  13 +++\n>  3 files changed, 195 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 1ce9c2b00e8..13f2b0f4d5a 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -259,6 +259,108 @@ static int unregister_dir(void)\n>         return res;\n>  }\n>\n> +static int stage(const char *git_dir, struct strbuf *buf, const char *path)\n> +{\n> +       struct strbuf cacheinfo = STRBUF_INIT;\n> +       struct child_process cp = CHILD_PROCESS_INIT;\n> +       int res;\n> +\n> +       strbuf_addstr(&cacheinfo, \"100644,\");\n> +\n> +       cp.git_cmd = 1;\n> +       strvec_pushl(&cp.args, \"--git-dir\", git_dir,\n> +                    \"hash-object\", \"-w\", \"--stdin\", NULL);\n> +       res = pipe_command(&cp, buf->buf, buf->len, &cacheinfo, 256, NULL, 0);\n> +       if (!res) {\n> +               strbuf_rtrim(&cacheinfo);\n> +               strbuf_addch(&cacheinfo, ',');\n> +               /* We cannot stage `.git`, use `_git` instead. */\n> +               if (starts_with(path, \".git/\"))\n> +                       strbuf_addf(&cacheinfo, \"_%s\", path + 1);\n> +               else\n> +                       strbuf_addstr(&cacheinfo, path);\n> +\n> +               child_process_init(&cp);\n> +               cp.git_cmd = 1;\n> +               strvec_pushl(&cp.args, \"--git-dir\", git_dir,\n> +                            \"update-index\", \"--add\", \"--cacheinfo\",\n> +                            cacheinfo.buf, NULL);\n> +               res = run_command(&cp);\n> +       }\n> +\n> +       strbuf_release(&cacheinfo);\n> +       return res;\n> +}\n> +\n> +static int stage_file(const char *git_dir, const char *path)\n> +{\n> +       struct strbuf buf = STRBUF_INIT;\n> +       int res;\n> +\n> +       if (strbuf_read_file(&buf, path, 0) < 0)\n> +               return error(_(\"could not read '%s'\"), path);\n> +\n> +       res = stage(git_dir, &buf, path);\n> +\n> +       strbuf_release(&buf);\n> +       return res;\n> +}\n> +\n> +static int stage_directory(const char *git_dir, const char *path, int recurse)\n> +{\n> +       int at_root = !*path;\n> +       DIR *dir = opendir(at_root ? \".\" : path);\n> +       struct dirent *e;\n> +       struct strbuf buf = STRBUF_INIT;\n> +       size_t len;\n> +       int res = 0;\n> +\n> +       if (!dir)\n> +               return error(_(\"could not open directory '%s'\"), path);\n> +\n> +       if (!at_root)\n> +               strbuf_addf(&buf, \"%s/\", path);\n> +       len = buf.len;\n> +\n> +       while (!res && (e = readdir(dir))) {\n> +               if (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n> +                       continue;\n> +\n> +               strbuf_setlen(&buf, len);\n> +               strbuf_addstr(&buf, e->d_name);\n> +\n> +               if ((e->d_type == DT_REG && stage_file(git_dir, buf.buf)) ||\n> +                   (e->d_type == DT_DIR && recurse &&\n> +                    stage_directory(git_dir, buf.buf, recurse)))\n> +                       res = -1;\n> +       }\n> +\n> +       closedir(dir);\n> +       strbuf_release(&buf);\n> +       return res;\n> +}\n> +\n> +static int index_to_zip(const char *git_dir)\n> +{\n> +       struct child_process cp = CHILD_PROCESS_INIT;\n> +       struct strbuf oid = STRBUF_INIT;\n> +\n> +       cp.git_cmd = 1;\n> +       strvec_pushl(&cp.args, \"--git-dir\", git_dir, \"write-tree\", NULL);\n> +       if (pipe_command(&cp, NULL, 0, &oid, the_hash_algo->hexsz + 1,\n> +                        NULL, 0))\n> +               return error(_(\"could not write temporary tree object\"));\n> +\n> +       strbuf_rtrim(&oid);\n> +       child_process_init(&cp);\n> +       cp.git_cmd = 1;\n> +       strvec_pushl(&cp.args, \"--git-dir\", git_dir, \"archive\", \"-o\", NULL);\n> +       strvec_pushf(&cp.args, \"%s.zip\", git_dir);\n> +       strvec_pushl(&cp.args, oid.buf, \"--\", NULL);\n> +       strbuf_release(&oid);\n> +       return run_command(&cp);\n> +}\n> +\n>  /* printf-style interface, expects `<key>=<value>` argument */\n>  static int set_config(const char *fmt, ...)\n>  {\n> @@ -499,6 +601,73 @@ cleanup:\n>         return res;\n>  }\n>\n> +static int cmd_diagnose(int argc, const char **argv)\n> +{\n> +       struct option options[] = {\n> +               OPT_END(),\n> +       };\n> +       const char * const usage[] = {\n> +               N_(\"scalar diagnose [<enlistment>]\"),\n> +               NULL\n> +       };\n> +       struct strbuf tmp_dir = STRBUF_INIT;\n> +       time_t now = time(NULL);\n> +       struct tm tm;\n> +       struct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n> +       int res = 0;\n> +\n> +       argc = parse_options(argc, argv, NULL, options,\n> +                            usage, 0);\n> +\n> +       setup_enlistment_directory(argc, argv, usage, options, &buf);\n> +\n> +       strbuf_addstr(&buf, \"/.scalarDiagnostics/scalar_\");\n> +       strbuf_addftime(&buf, \"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n> +       if (run_git(\"init\", \"-q\", \"-b\", \"dummy\", \"--bare\", buf.buf, NULL)) {\n> +               res = error(_(\"could not initialize temporary repository: %s\"),\n> +                           buf.buf);\n> +               goto diagnose_cleanup;\n> +       }\n> +       strbuf_realpath(&tmp_dir, buf.buf, 1);\n> +\n> +       strbuf_reset(&buf);\n> +       strbuf_addf(&buf, \"Collecting diagnostic info into temp folder %s\\n\\n\",\n> +                   tmp_dir.buf);\n> +\n> +       get_version_info(&buf, 1);\n> +\n> +       strbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n> +       fwrite(buf.buf, buf.len, 1, stdout);\n> +\n> +       if ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n> +               goto diagnose_cleanup;\n> +\n> +       if ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n> +           (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n> +           (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n> +           (res = stage_directory(tmp_dir.buf, \".git/logs\", 1)) ||\n> +           (res = stage_directory(tmp_dir.buf, \".git/objects/info\", 0)))\n> +               goto diagnose_cleanup;\n> +\n> +       res = index_to_zip(tmp_dir.buf);\n> +\n> +       if (!res)\n> +               res = remove_dir_recursively(&tmp_dir, 0);\n> +\n> +       if (!res)\n> +               printf(\"\\n\"\n> +                      \"Diagnostics complete.\\n\"\n> +                      \"All of the gathered info is captured in '%s.zip'\\n\",\n> +                      tmp_dir.buf);\n> +\n> +diagnose_cleanup:\n> +       strbuf_release(&tmp_dir);\n> +       strbuf_release(&path);\n> +       strbuf_release(&buf);\n> +\n> +       return res;\n> +}\n> +\n>  static int cmd_list(int argc, const char **argv)\n>  {\n>         if (argc != 1)\n> @@ -800,6 +969,7 @@ static struct {\n>         { \"reconfigure\", cmd_reconfigure },\n>         { \"delete\", cmd_delete },\n>         { \"version\", cmd_version },\n> +       { \"diagnose\", cmd_diagnose },\n>         { NULL, NULL},\n>  };\n>\n> diff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\n> index f416d637289..22583fe046e 100644\n> --- a/contrib/scalar/scalar.txt\n> +++ b/contrib/scalar/scalar.txt\n> @@ -14,6 +14,7 @@ scalar register [<enlistment>]\n>  scalar unregister [<enlistment>]\n>  scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n>  scalar reconfigure [ --all | <enlistment> ]\n> +scalar diagnose [<enlistment>]\n>  scalar delete <enlistment>\n>\n>  DESCRIPTION\n> @@ -129,6 +130,17 @@ reconfigure the enlistment.\n>  With the `--all` option, all enlistments currently registered with Scalar\n>  will be reconfigured. Use this option after each Scalar upgrade.\n>\n> +Diagnose\n> +~~~~~~~~\n> +\n> +diagnose [<enlistment>]::\n> +    When reporting issues with Scalar, it is often helpful to provide the\n> +    information gathered by this command, including logs and certain\n> +    statistics describing the data shape of the current enlistment.\n> ++\n> +The output of this command is a `.zip` file that is written into\n> +a directory adjacent to the worktree in the `src` directory.\n> +\n>  Delete\n>  ~~~~~~\n>\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index 2e1502ad45e..ecd06e207c2 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -65,6 +65,19 @@ test_expect_success 'scalar clone' '\n>         )\n>  '\n>\n> +SQ=\"'\"\n> +test_expect_success UNZIP 'scalar diagnose' '\n> +       scalar diagnose cloned >out &&\n> +       sed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n> +       zip_path=$(cat zip_path) &&\n> +       test -n \"$zip_path\" &&\n> +       unzip -v \"$zip_path\" &&\n> +       folder=${zip_path%.zip} &&\n> +       test_path_is_missing \"$folder\" &&\n> +       unzip -p \"$zip_path\" diagnostics.log >out &&\n> +       test_file_not_empty out\n> +'\n> +\n>  test_expect_success 'scalar reconfigure' '\n>         git init one/src &&\n>         scalar register one &&\n> --\n> gitgitgadget\n>\n"},{"id":"447847","messageId":"nycvar.QRO.7.76.6.2202062213030.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"dda1b3c8-afc6-d9d5-1bfb-4a48ac87ca54@gmail.com","subject":"Re: [PATCH 0/5] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-06T21:13:58Z","receivedAt":"2022-02-06T21:19:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Stolee & Emily,\n\nOn Thu, 27 Jan 2022, Derrick Stolee wrote:\n\n> On 1/26/2022 3:41 AM, Johannes Schindelin via GitGitGadget wrote:\n> > Over the course of the years, we developed a sub-command that gathers\n> > diagnostic data into a .zip file that can then be attached to bug reports.\n> > This sub-command turned out to be very useful in helping Scalar developers\n> > identify and fix issues.\n>\n> For historical context: The 'diagnose' command was implemented in VFS for\n> Git and ported to the C# version of Scalar before 'git bugreport' existed,\n> but they serve very similar purposes.\n>\n> I wonder if 'scalar diagnose' could include some of the information\n> captured by 'git bugreport' or whether this implementation of 'diagnose'\n> could help inform 'git bugreport' in any way.\n\nIndeed, I think that the `bugreport` command could easily benefit from at\nleast the number of pack files and loose objects.\n\nCiao,\nDscho\n\n>\n> CC'ing Emily for thoughts.\n>\n> Thanks,\n> -Stolee\n>\n"},{"id":"447848","messageId":"nycvar.QRO.7.76.6.2202062214110.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"CABPp-BHeLzinXkX3WgqBNYntJwY_ZAm5D7VdOR7KQahvLOuV=w@mail.gmail.com","subject":"Re: [PATCH 4/5] scalar: teach `diagnose` to gather loose objects information","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-06T21:25:33Z","receivedAt":"2022-02-06T21:25:41Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Elijah,\n\nOn Thu, 27 Jan 2022, Elijah Newren wrote:\n\n> On Wed, Jan 26, 2022 at 3:37 PM Matthew John Cheetham via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: Matthew John Cheetham <mjcheetham@outlook.com>\n> >\n> > When operating at the scale that Scalar wants to support, certain data\n> > shapes are more likely to cause undesirable performance issues, such as\n> > large numbers or large sizes of loose objects.\n>\n> Makes sense.\n>\n> > By including statistics about this, `scalar diagnose` now makes it\n> > easier to identify such scenarios.\n> >\n> > Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n> > ---\n> >  contrib/scalar/scalar.c          | 60 ++++++++++++++++++++++++++++++++\n> >  contrib/scalar/t/t9099-scalar.sh |  2 ++\n> >  2 files changed, 62 insertions(+)\n> >\n> > diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> > index 690933ffdf3..c0ad4948215 100644\n> > --- a/contrib/scalar/scalar.c\n> > +++ b/contrib/scalar/scalar.c\n> > @@ -686,6 +686,60 @@ static void dir_file_stats(struct strbuf *buf, const char *path)\n> >         closedir(dir);\n> >  }\n> >\n> > +static int count_files(char *path)\n> > +{\n> > +       DIR *dir = opendir(path);\n> > +       struct dirent *e;\n> > +       int count = 0;\n> > +\n> > +       if (!dir)\n> > +               return 0;\n> > +\n> > +       while ((e = readdir(dir)) != NULL)\n> > +               if (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n> > +                       count++;\n> > +\n> > +       closedir(dir);\n> > +       return count;\n> > +}\n> > +\n> > +static void loose_objs_stats(struct strbuf *buf, const char *path)\n> > +{\n> > +       DIR *dir = opendir(path);\n> > +       struct dirent *e;\n> > +       int count;\n> > +       int total = 0;\n> > +       unsigned char c;\n> > +       struct strbuf count_path = STRBUF_INIT;\n> > +       size_t base_path_len;\n> > +\n> > +       if (!dir)\n> > +               return;\n> > +\n> > +       strbuf_addstr(buf, \"Object directory stats for \");\n> > +       strbuf_add_absolute_path(buf, path);\n> > +       strbuf_addstr(buf, \":\\n\");\n> > +\n> > +       strbuf_add_absolute_path(&count_path, path);\n> > +       strbuf_addch(&count_path, '/');\n> > +       base_path_len = count_path.len;\n> > +\n> > +       while ((e = readdir(dir)) != NULL)\n> > +               if (!is_dot_or_dotdot(e->d_name) &&\n> > +                   e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n> > +                   !hex_to_bytes(&c, e->d_name, 1)) {\n>\n> You only recurse into directories, ignoring individual files.\n>\n> > +                       strbuf_setlen(&count_path, base_path_len);\n> > +                       strbuf_addstr(&count_path, e->d_name);\n> > +                       total += (count = count_files(count_path.buf));\n> > +                       strbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n>\n> This shows the number of files within a directory.\n>\n> > +               }\n> > +\n> > +       strbuf_addf(buf, \"Total: %d loose objects\", total);\n>\n> and this shows the total number of files across all the directories.\n>\n> But the commit message suggested you also wanted to check for large\n> sizes of loose objects.  Did that get ripped out at some point with\n> the commit message not being updated, or is it perhaps going to be\n> included later?\n\nNo, there was no plan to include this information later, as the original\n.NET implementation of `scalar diagnose` did not provide that information,\neither (which I take as a strong sign that we never needed this type of\ninformation to help users, at least not up until this point).\n\nBesides, it would be kind of a difficult thing to say conclusively what\nmakes a loose file \"big\". Is it the zlib-compressed size on disk? Or the\nunpacked size? Should there be a configurable threshold to determine when\nan object is big? Should `core.bigFileThreshold` be co-opted for this?\n\nTogether with the fact that there was no need for this information in\npractice, it makes me doubt that we should add this type of information. I\nactually suspect that _iff_ information of that type would be helpful, a\nmore complete tool like git-sizer (https://github.com/github/git-sizer/)\nwould be needed, and I do not really want to subsume git-sizer's\nfunctionality in `scalar diagnose`.\n\nI rephrased the commit message.\n\nCiao,\nDscho\n\n>\n> > +\n> > +       strbuf_release(&count_path);\n> > +       closedir(dir);\n> > +}\n> > +\n> >  static int cmd_diagnose(int argc, const char **argv)\n> >  {\n> >         struct option options[] = {\n> > @@ -734,6 +788,12 @@ static int cmd_diagnose(int argc, const char **argv)\n> >         if ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n> >                 goto diagnose_cleanup;\n> >\n> > +       strbuf_reset(&buf);\n> > +       loose_objs_stats(&buf, \".git/objects\");\n> > +\n> > +       if ((res = stage(tmp_dir.buf, &buf, \"objects-local.txt\")))\n> > +               goto diagnose_cleanup;\n> > +\n> >         if ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n> >             (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n> >             (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n> > diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> > index b1745851e31..f2ec156d819 100755\n> > --- a/contrib/scalar/t/t9099-scalar.sh\n> > +++ b/contrib/scalar/t/t9099-scalar.sh\n> > @@ -77,6 +77,8 @@ test_expect_success UNZIP 'scalar diagnose' '\n> >         unzip -p \"$zip_path\" diagnostics.log >out &&\n> >         test_file_not_empty out &&\n> >         unzip -p \"$zip_path\" packs-local.txt >out &&\n> > +       test_file_not_empty out &&\n> > +       unzip -p \"$zip_path\" objects-local.txt >out &&\n> >         test_file_not_empty out\n> >  '\n> >\n> > --\n> > gitgitgadget\n>\n"},{"id":"447850","messageId":"nycvar.QRO.7.76.6.2202062225460.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"YfHJHbMKA1u+A9LF@nand.local","subject":"Re: [PATCH 1/5] Implement `scalar diagnose`","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-06T21:34:29Z","receivedAt":"2022-02-06T21:34:42Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi René & Taylor,\n\nOn Wed, 26 Jan 2022, Taylor Blau wrote:\n\n> On Wed, Jan 26, 2022 at 10:34:04AM +0100, René Scharfe wrote:\n> > Am 26.01.22 um 09:41 schrieb Johannes Schindelin via GitGitGadget:\n> > > Note: originally, Scalar was implemented in C# using the .NET API, where\n> > > we had the luxury of a comprehensive standard library that includes\n> > > basic functionality such as writing a `.zip` file. In the C version, we\n> > > lack such a commodity. Rather than introducing a dependency on, say,\n> > > libzip, we slightly abuse Git's `archive` command: Instead of writing\n> > > the `.zip` file directly, we stage the file contents in a Git index of a\n> > > temporary, bare repository, only to let `git archive` have at it, and\n> > > finally removing the temporary repository.\n> >\n> > git archive allows you to include untracked files in an archive with its\n> > option --add-file.  You can see an example in Git's Makefile; search for\n> > GIT_ARCHIVE_EXTRA_FILES.  It still requires a tree argument, but the\n> > empty tree object should suffice if you don't want to include any\n> > tracked files.  It doesn't currently support streaming, though, i.e.\n> > files are fully read into memory, so it's impractical for huge ones.\n\nThat's a good point.\n\nI did not want to invent any `fast-import`-like streaming protocol just\nfor the sake of supporting the \"funny\" use case of `scalar diagnose`, so I\ninvented a new option `--add-file-with-content=<path>:<content>` (with the\nobvious limitation that the `<path>` cannot contain any colon, if that is\ndesired, users will still need to write out untracked files).\n\n> Using `--add-file` would likely be preferable to setting up a temporary\n> repository just to invoke `git archive` in it. Johannes would be the\n> expert to ask whether or not big files are going to be a problem here\n> (based on a cursory scan of the new functions in scalar.c, I don't\n> expect this to be the case).\n\nIndeed, it is unlikely that any large files are included.\n\n> The new stage_directory() function _could_ add `--add-file` arguments in\n> a loop around readdir(), but it might also be nice to add a new\n> `--add-directory` function to `git archive` which would do the \"heavy\"\n> lifting for us.\n\nI went one step further and used `write_archive()` to do the\nheavy-lifting. That way, we truly avoid spawning any separate process let\nalone creating any throw-away repository.\n\nCiao,\nDscho\n"},{"id":"447851","messageId":"nycvar.QRO.7.76.6.2202062234520.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"0a52155c-4605-d96f-965a-104a399ae86e@gmail.com","subject":"Re: [PATCH 3/5] scalar: teach `diagnose` to gather packfile info","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-06T21:38:33Z","receivedAt":"2022-02-06T21:38:44Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Stolee & Taylor,\n\nOn Thu, 27 Jan 2022, Derrick Stolee wrote:\n\n> On 1/26/2022 5:43 PM, Taylor Blau wrote:\n> > On Wed, Jan 26, 2022 at 08:41:45AM +0000, Matthew John Cheetham via GitGitGadget wrote:\n> >> From: Matthew John Cheetham <mjcheetham@outlook.com>\n> >>\n> >> Teach the `scalar diagnose` command to gather file size information\n> >> about pack files.\n> >>\n> >> Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n> >> ---\n> >>  contrib/scalar/scalar.c          | 39 ++++++++++++++++++++++++++++++++\n> >>  contrib/scalar/t/t9099-scalar.sh |  2 ++\n> >>  2 files changed, 41 insertions(+)\n> >>\n> >> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> >> index e26fb2fc018..690933ffdf3 100644\n> >> --- a/contrib/scalar/scalar.c\n> >> +++ b/contrib/scalar/scalar.c\n> >> @@ -653,6 +653,39 @@ cleanup:\n> >>  \treturn res;\n> >>  }\n> >>\n> >> +static void dir_file_stats(struct strbuf *buf, const char *path)\n> >> +{\n> >> +\tDIR *dir = opendir(path);\n> >> +\tstruct dirent *e;\n> >> +\tstruct stat e_stat;\n> >> +\tstruct strbuf file_path = STRBUF_INIT;\n> >> +\tsize_t base_path_len;\n> >> +\n> >> +\tif (!dir)\n> >> +\t\treturn;\n> >> +\n> >> +\tstrbuf_addstr(buf, \"Contents of \");\n> >> +\tstrbuf_add_absolute_path(buf, path);\n> >> +\tstrbuf_addstr(buf, \":\\n\");\n> >> +\n> >> +\tstrbuf_add_absolute_path(&file_path, path);\n> >> +\tstrbuf_addch(&file_path, '/');\n> >> +\tbase_path_len = file_path.len;\n> >> +\n> >> +\twhile ((e = readdir(dir)) != NULL)\n> >\n> > Hmm. Is there a reason that this couldn't use\n> > for_each_file_in_pack_dir() with a callback that just does the stat()\n> > and buffer manipulation?\n> >\n> > I don't think it's critical either way, but it would eliminate some of\n> > the boilerplate that is shared between this implementation and the one\n> > that already exists in for_each_file_in_pack_dir().\n>\n> It's helpful to see if there are other crud files in the pack\n> directory. This method is also extended in microsoft/git to\n> scan the alternates directory (which we expect to exist as the\n> \"shared objects cache).\n>\n> We might want to modify the implementation in this series to\n> run dir_file_stats() on each odb in the_repository. This would\n> give us the data for the shared object cache for free while\n> being more general to other Git repos. (It would require us to\n> do some reaction work in microsoft/git and be a change of\n> behavior, but we are the only ones who have looked at these\n> diagnose files before, so that change will be easy to manage.)\n\nGood points all around. I went with the `for_each_file_in_pack_dir()`\napproach, and threw in the now very simple change to also enumerate the\nalternates, if there are any.\n\nAnd yes, that will require some reaction work in microsoft/git, but for an\nobvious improvement like this one, I don't grumble about the extra burden.\n\nCiao,\nDscho\n"},{"id":"447853","messageId":"600da8d465ef8e83c20cdf6fec85aff08eb65551.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 2/6] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:02Z","receivedAt":"2022-02-06T22:39:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..00dcd4b50ef 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..9d83fdf25e8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n \ttest_path_is_missing cloned\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"447854","messageId":"49ff3c1f2b32b16df2b4216aa016d715b6de46bc.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:01Z","receivedAt":"2022-02-06T22:39:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-file-with-content=<path>:<content>` option, `git\narchive` now supports use cases where relatively trivial files need to\nbe added that do not exist on disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 11 ++++++++\n archive.c                     | 51 +++++++++++++++++++++++++++++------\n t/t5003-archive-zip.sh        | 12 +++++++++\n 3 files changed, 66 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a7834..1b52a0a65a1 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -61,6 +61,17 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n \n+--add-file-with-content=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value for `--prefix` (if any) and the\n+\tbasename of <file>.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex a3bbb091256..172efd690c3 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n \t\tstrbuf_addstr(&path_in_archive, basename(path));\n \n \t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\tif (info->content)\n+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\telse if (strbuf_read_file(&content, path,\n+\t\t\t\t\t  info->stat.st_size) < 0)\n \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n@@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (strcmp(opt->long_name, \"add-file-with-content\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +586,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n+\t\t  N_(\"file\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 1e6d18b140e..8ff1257f1a0 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \ngitgitgadget\n\n"},{"id":"447855","messageId":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.git.1643186507.gitgitgadget@gmail.com","subject":"[PATCH v2 0/6] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:00Z","receivedAt":"2022-02-06T22:39:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nChanges since v1:\n\n * Instead of creating a throw-away repository, staging the contents of the\n   .zip file and then using git write-tree and git archive to write the .zip\n   file, the patch series now introduces a new option to git archive and\n   uses write_archive() directly (avoiding any separate process).\n * Since the command avoids separate processes, it is now blazing fast on\n   Windows, and I dropped the spinner() function because it's no longer\n   needed.\n * While reworking the test case, I noticed that scalar [...] <enlistment>\n   failed to verify that the specified directory exists, and would happily\n   \"traverse to its parent directory\" on its quest to find a Scalar\n   enlistment. That is of course incorrect, and has been fixed as a \"while\n   at it\" sort of preparatory commit.\n * I had forgotten to sign off on all the commits, which has been fixed.\n * Instead of some \"home-grown\" readdir()-based function, the code now uses\n   for_each_file_in_pack_dir() to look through the pack directories.\n * If any alternates are configured, their pack directories are now included\n   in the output.\n * The commit message that might be interpreted to promise information about\n   large loose files has been corrected to no longer promise that.\n * The test cases have been adjusted to test a little bit more (e.g.\n   verifying that specific paths are mentioned in the output, instead of\n   merely verifying that the output is non-empty).\n\nJohannes Schindelin (4):\n  archive: optionally add \"virtual\" files\n  scalar: validate the optional enlistment argument\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n Documentation/git-archive.txt    |  11 ++\n archive.c                        |  51 +++++-\n contrib/scalar/scalar.c          | 291 ++++++++++++++++++++++++++++++-\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  27 +++\n t/t5003-archive-zip.sh           |  12 ++\n 6 files changed, 394 insertions(+), 10 deletions(-)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n\nRange-diff vs v1:\n\n -:  ----------- > 1:  49ff3c1f2b3 archive: optionally add \"virtual\" files\n -:  ----------- > 2:  600da8d465e scalar: validate the optional enlistment argument\n 1:  ce85506e7a4 ! 3:  0d570137bb6 Implement `scalar diagnose`\n     @@ Commit message\n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## contrib/scalar/scalar.c ##\n     +@@\n     + #include \"dir.h\"\n     + #include \"packfile.h\"\n     + #include \"help.h\"\n     ++#include \"archive.h\"\n     + \n     + /*\n     +  * Remove the deepest subdirectory in the provided path string. Path must not\n      @@ contrib/scalar/scalar.c: static int unregister_dir(void)\n       \treturn res;\n       }\n       \n     -+static int stage(const char *git_dir, struct strbuf *buf, const char *path)\n     -+{\n     -+\tstruct strbuf cacheinfo = STRBUF_INIT;\n     -+\tstruct child_process cp = CHILD_PROCESS_INIT;\n     -+\tint res;\n     -+\n     -+\tstrbuf_addstr(&cacheinfo, \"100644,\");\n     -+\n     -+\tcp.git_cmd = 1;\n     -+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n     -+\t\t     \"hash-object\", \"-w\", \"--stdin\", NULL);\n     -+\tres = pipe_command(&cp, buf->buf, buf->len, &cacheinfo, 256, NULL, 0);\n     -+\tif (!res) {\n     -+\t\tstrbuf_rtrim(&cacheinfo);\n     -+\t\tstrbuf_addch(&cacheinfo, ',');\n     -+\t\t/* We cannot stage `.git`, use `_git` instead. */\n     -+\t\tif (starts_with(path, \".git/\"))\n     -+\t\t\tstrbuf_addf(&cacheinfo, \"_%s\", path + 1);\n     -+\t\telse\n     -+\t\t\tstrbuf_addstr(&cacheinfo, path);\n     -+\n     -+\t\tchild_process_init(&cp);\n     -+\t\tcp.git_cmd = 1;\n     -+\t\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir,\n     -+\t\t\t     \"update-index\", \"--add\", \"--cacheinfo\",\n     -+\t\t\t     cacheinfo.buf, NULL);\n     -+\t\tres = run_command(&cp);\n     -+\t}\n     -+\n     -+\tstrbuf_release(&cacheinfo);\n     -+\treturn res;\n     -+}\n     -+\n     -+static int stage_file(const char *git_dir, const char *path)\n     -+{\n     -+\tstruct strbuf buf = STRBUF_INIT;\n     -+\tint res;\n     -+\n     -+\tif (strbuf_read_file(&buf, path, 0) < 0)\n     -+\t\treturn error(_(\"could not read '%s'\"), path);\n     -+\n     -+\tres = stage(git_dir, &buf, path);\n     -+\n     -+\tstrbuf_release(&buf);\n     -+\treturn res;\n     -+}\n     -+\n     -+static int stage_directory(const char *git_dir, const char *path, int recurse)\n     ++static int add_directory_to_archiver(struct strvec *archiver_args,\n     ++\t\t\t\t\t  const char *path, int recurse)\n      +{\n      +\tint at_root = !*path;\n      +\tDIR *dir = opendir(at_root ? \".\" : path);\n     @@ contrib/scalar/scalar.c: static int unregister_dir(void)\n      +\tif (!at_root)\n      +\t\tstrbuf_addf(&buf, \"%s/\", path);\n      +\tlen = buf.len;\n     ++\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n      +\n      +\twhile (!res && (e = readdir(dir))) {\n      +\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n     @@ contrib/scalar/scalar.c: static int unregister_dir(void)\n      +\t\tstrbuf_setlen(&buf, len);\n      +\t\tstrbuf_addstr(&buf, e->d_name);\n      +\n     -+\t\tif ((e->d_type == DT_REG && stage_file(git_dir, buf.buf)) ||\n     -+\t\t    (e->d_type == DT_DIR && recurse &&\n     -+\t\t     stage_directory(git_dir, buf.buf, recurse)))\n     ++\t\tif (e->d_type == DT_REG)\n     ++\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n     ++\t\telse if (e->d_type != DT_DIR)\n      +\t\t\tres = -1;\n     ++\t\telse if (recurse)\n     ++\t\t     add_directory_to_archiver(archiver_args, buf.buf, recurse);\n      +\t}\n      +\n      +\tclosedir(dir);\n      +\tstrbuf_release(&buf);\n      +\treturn res;\n      +}\n     -+\n     -+static int index_to_zip(const char *git_dir)\n     -+{\n     -+\tstruct child_process cp = CHILD_PROCESS_INIT;\n     -+\tstruct strbuf oid = STRBUF_INIT;\n     -+\n     -+\tcp.git_cmd = 1;\n     -+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"write-tree\", NULL);\n     -+\tif (pipe_command(&cp, NULL, 0, &oid, the_hash_algo->hexsz + 1,\n     -+\t\t\t NULL, 0))\n     -+\t\treturn error(_(\"could not write temporary tree object\"));\n     -+\n     -+\tstrbuf_rtrim(&oid);\n     -+\tchild_process_init(&cp);\n     -+\tcp.git_cmd = 1;\n     -+\tstrvec_pushl(&cp.args, \"--git-dir\", git_dir, \"archive\", \"-o\", NULL);\n     -+\tstrvec_pushf(&cp.args, \"%s.zip\", git_dir);\n     -+\tstrvec_pushl(&cp.args, oid.buf, \"--\", NULL);\n     -+\tstrbuf_release(&oid);\n     -+\treturn run_command(&cp);\n     -+}\n      +\n       /* printf-style interface, expects `<key>=<value>` argument */\n       static int set_config(const char *fmt, ...)\n     @@ contrib/scalar/scalar.c: cleanup:\n      +\t\tN_(\"scalar diagnose [<enlistment>]\"),\n      +\t\tNULL\n      +\t};\n     -+\tstruct strbuf tmp_dir = STRBUF_INIT;\n     ++\tstruct strbuf zip_path = STRBUF_INIT;\n     ++\tstruct strvec archiver_args = STRVEC_INIT;\n     ++\tchar **argv_copy = NULL;\n     ++\tint stdout_fd = -1, archiver_fd = -1;\n      +\ttime_t now = time(NULL);\n      +\tstruct tm tm;\n      +\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n     ++\tsize_t off;\n      +\tint res = 0;\n      +\n      +\targc = parse_options(argc, argv, NULL, options,\n      +\t\t\t     usage, 0);\n      +\n     -+\tsetup_enlistment_directory(argc, argv, usage, options, &buf);\n     ++\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n     ++\n     ++\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n     ++\tstrbuf_addftime(&zip_path,\n     ++\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n     ++\tstrbuf_addstr(&zip_path, \".zip\");\n     ++\tswitch (safe_create_leading_directories(zip_path.buf)) {\n     ++\tcase SCLD_EXISTS:\n     ++\tcase SCLD_OK:\n     ++\t\tbreak;\n     ++\tdefault:\n     ++\t\terror_errno(_(\"could not create directory for '%s'\"),\n     ++\t\t\t    zip_path.buf);\n     ++\t\tgoto diagnose_cleanup;\n     ++\t}\n     ++\tstdout_fd = dup(1);\n     ++\tif (stdout_fd < 0) {\n     ++\t\tres = error_errno(_(\"could not duplicate stdout\"));\n     ++\t\tgoto diagnose_cleanup;\n     ++\t}\n      +\n     -+\tstrbuf_addstr(&buf, \"/.scalarDiagnostics/scalar_\");\n     -+\tstrbuf_addftime(&buf, \"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n     -+\tif (run_git(\"init\", \"-q\", \"-b\", \"dummy\", \"--bare\", buf.buf, NULL)) {\n     -+\t\tres = error(_(\"could not initialize temporary repository: %s\"),\n     -+\t\t\t    buf.buf);\n     ++\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n     ++\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n     ++\t\tres = error_errno(_(\"could not redirect output\"));\n      +\t\tgoto diagnose_cleanup;\n      +\t}\n     -+\tstrbuf_realpath(&tmp_dir, buf.buf, 1);\n      +\n     -+\tstrbuf_reset(&buf);\n     -+\tstrbuf_addf(&buf, \"Collecting diagnostic info into temp folder %s\\n\\n\",\n     -+\t\t    tmp_dir.buf);\n     ++\tinit_zip_archiver();\n     ++\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n      +\n     ++\tstrbuf_reset(&buf);\n     ++\tstrbuf_addstr(&buf,\n     ++\t\t      \"--add-file-with-content=diagnostics.log:\"\n     ++\t\t      \"Collecting diagnostic info\\n\\n\");\n      +\tget_version_info(&buf, 1);\n      +\n      +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n     -+\tfwrite(buf.buf, buf.len, 1, stdout);\n     -+\n     -+\tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n     -+\t\tgoto diagnose_cleanup;\n     -+\n     -+\tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n     -+\t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n     -+\t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n     -+\t    (res = stage_directory(tmp_dir.buf, \".git/logs\", 1)) ||\n     -+\t    (res = stage_directory(tmp_dir.buf, \".git/objects/info\", 0)))\n     ++\toff = strchr(buf.buf, ':') + 1 - buf.buf;\n     ++\twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     ++\tstrvec_push(&archiver_args, buf.buf);\n     ++\n     ++\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n     ++\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n     ++\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n     ++\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n     ++\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n      +\t\tgoto diagnose_cleanup;\n      +\n     -+\tres = index_to_zip(tmp_dir.buf);\n     ++\tstrvec_pushl(&archiver_args, \"--prefix=\",\n     ++\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n      +\n     -+\tif (!res)\n     -+\t\tres = remove_dir_recursively(&tmp_dir, 0);\n     ++\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n     ++\targv_copy = xmemdupz(archiver_args.v,\n     ++\t\t\t     sizeof(char *) * archiver_args.nr);\n     ++\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n     ++\t\t\t    the_repository, NULL, 0);\n     ++\tif (res) {\n     ++\t\terror(_(\"failed to write archive\"));\n     ++\t\tgoto diagnose_cleanup;\n     ++\t}\n      +\n      +\tif (!res)\n      +\t\tprintf(\"\\n\"\n      +\t\t       \"Diagnostics complete.\\n\"\n     -+\t\t       \"All of the gathered info is captured in '%s.zip'\\n\",\n     -+\t\t       tmp_dir.buf);\n     ++\t\t       \"All of the gathered info is captured in '%s'\\n\",\n     ++\t\t       zip_path.buf);\n      +\n      +diagnose_cleanup:\n     -+\tstrbuf_release(&tmp_dir);\n     ++\tif (archiver_fd >= 0) {\n     ++\t\tclose(1);\n     ++\t\tdup2(stdout_fd, 1);\n     ++\t}\n     ++\tfree(argv_copy);\n     ++\tstrvec_clear(&archiver_args);\n     ++\tstrbuf_release(&zip_path);\n      +\tstrbuf_release(&path);\n      +\tstrbuf_release(&buf);\n      +\n     @@ contrib/scalar/scalar.txt: reconfigure the enlistment.\n       \n      \n       ## contrib/scalar/t/t9099-scalar.sh ##\n     -@@ contrib/scalar/t/t9099-scalar.sh: test_expect_success 'scalar clone' '\n     - \t)\n     +@@ contrib/scalar/t/t9099-scalar.sh: test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n     + \tgrep \"cloned. does not exist\" err\n       '\n       \n      +SQ=\"'\"\n      +test_expect_success UNZIP 'scalar diagnose' '\n     ++\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n      +\tscalar diagnose cloned >out &&\n      +\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n      +\tzip_path=$(cat zip_path) &&\n     @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success 'scalar clone' '\n      +\ttest_file_not_empty out\n      +'\n      +\n     - test_expect_success 'scalar reconfigure' '\n     - \tgit init one/src &&\n     - \tscalar register one &&\n     + test_done\n 2:  f8885b27502 ! 4:  938e38b5a09 scalar diagnose: include disk space information\n     @@ Commit message\n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## contrib/scalar/scalar.c ##\n     -@@ contrib/scalar/scalar.c: static int index_to_zip(const char *git_dir)\n     - \treturn run_command(&cp);\n     +@@ contrib/scalar/scalar.c: static int add_directory_to_archiver(struct strvec *archiver_args,\n     + \treturn res;\n       }\n       \n      +#ifndef WIN32\n     @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n       \n       \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n      +\tget_disk_info(&buf);\n     - \tfwrite(buf.buf, buf.len, 1, stdout);\n     - \n     - \tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n     + \toff = strchr(buf.buf, ':') + 1 - buf.buf;\n     + \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     + \tstrvec_push(&archiver_args, buf.buf);\n     +\n     + ## contrib/scalar/t/t9099-scalar.sh ##\n     +@@ contrib/scalar/t/t9099-scalar.sh: SQ=\"'\"\n     + test_expect_success UNZIP 'scalar diagnose' '\n     + \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n     + \tscalar diagnose cloned >out &&\n     ++\tgrep \"Available space\" out &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n     + \tzip_path=$(cat zip_path) &&\n     + \ttest -n \"$zip_path\" &&\n 3:  330b36de799 ! 5:  bd9428919fa scalar: teach `diagnose` to gather packfile info\n     @@ Metadata\n       ## Commit message ##\n          scalar: teach `diagnose` to gather packfile info\n      \n     -    Teach the `scalar diagnose` command to gather file size information\n     -    about pack files.\n     +    It's helpful to see if there are other crud files in the pack\n     +    directory. Let's teach the `scalar diagnose` command to gather\n     +    file size information about pack files.\n     +\n     +    While at it, also enumerate the pack files in the alternate\n     +    object directories, if any are registered.\n      \n          Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n     +    Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## contrib/scalar/scalar.c ##\n     +@@\n     + #include \"packfile.h\"\n     + #include \"help.h\"\n     + #include \"archive.h\"\n     ++#include \"object-store.h\"\n     + \n     + /*\n     +  * Remove the deepest subdirectory in the provided path string. Path must not\n      @@ contrib/scalar/scalar.c: cleanup:\n       \treturn res;\n       }\n       \n     -+static void dir_file_stats(struct strbuf *buf, const char *path)\n     ++static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n     ++\t\t\t\t   const char *file_name, void *data)\n      +{\n     -+\tDIR *dir = opendir(path);\n     -+\tstruct dirent *e;\n     -+\tstruct stat e_stat;\n     -+\tstruct strbuf file_path = STRBUF_INIT;\n     -+\tsize_t base_path_len;\n     ++\tstruct strbuf *buf = data;\n     ++\tstruct stat st;\n      +\n     -+\tif (!dir)\n     -+\t\treturn;\n     ++\tif (!stat(full_path, &st))\n     ++\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n     ++\t\t\t    (uintmax_t)st.st_size);\n     ++}\n      +\n     -+\tstrbuf_addstr(buf, \"Contents of \");\n     -+\tstrbuf_add_absolute_path(buf, path);\n     -+\tstrbuf_addstr(buf, \":\\n\");\n     ++static int dir_file_stats(struct object_directory *object_dir, void *data)\n     ++{\n     ++\tstruct strbuf *buf = data;\n      +\n     -+\tstrbuf_add_absolute_path(&file_path, path);\n     -+\tstrbuf_addch(&file_path, '/');\n     -+\tbase_path_len = file_path.len;\n     ++\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n      +\n     -+\twhile ((e = readdir(dir)) != NULL)\n     -+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG) {\n     -+\t\t\tstrbuf_setlen(&file_path, base_path_len);\n     -+\t\t\tstrbuf_addstr(&file_path, e->d_name);\n     -+\t\t\tif (!stat(file_path.buf, &e_stat))\n     -+\t\t\t\tstrbuf_addf(buf, \"%-70s %16\"PRIuMAX\"\\n\",\n     -+\t\t\t\t\t    e->d_name,\n     -+\t\t\t\t\t    (uintmax_t)e_stat.st_size);\n     -+\t\t}\n     ++\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n     ++\t\t\t\t  data);\n      +\n     -+\tstrbuf_release(&file_path);\n     -+\tclosedir(dir);\n     ++\treturn 0;\n      +}\n      +\n       static int cmd_diagnose(int argc, const char **argv)\n       {\n       \tstruct option options[] = {\n      @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n     - \tif ((res = stage(tmp_dir.buf, &buf, \"diagnostics.log\")))\n     - \t\tgoto diagnose_cleanup;\n     + \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     + \tstrvec_push(&archiver_args, buf.buf);\n       \n      +\tstrbuf_reset(&buf);\n     -+\tdir_file_stats(&buf, \".git/objects/pack\");\n     -+\n     -+\tif ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n     -+\t\tgoto diagnose_cleanup;\n     ++\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n     ++\tdir_file_stats(the_repository->objects->odb, &buf);\n     ++\tforeach_alt_odb(dir_file_stats, &buf);\n     ++\tstrvec_push(&archiver_args, buf.buf);\n      +\n     - \tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n     - \t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n     - \t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n     + \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n     + \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n     + \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n      \n       ## contrib/scalar/t/t9099-scalar.sh ##\n     +@@ contrib/scalar/t/t9099-scalar.sh: test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n     + SQ=\"'\"\n     + test_expect_success UNZIP 'scalar diagnose' '\n     + \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n     ++\tgit repack &&\n     ++\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n     + \tscalar diagnose cloned >out &&\n     + \tgrep \"Available space\" out &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n      @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n       \tfolder=${zip_path%.zip} &&\n       \ttest_path_is_missing \"$folder\" &&\n       \tunzip -p \"$zip_path\" diagnostics.log >out &&\n     +-\ttest_file_not_empty out\n      +\ttest_file_not_empty out &&\n      +\tunzip -p \"$zip_path\" packs-local.txt >out &&\n     - \ttest_file_not_empty out\n     ++\tgrep \"$(pwd)/.git/objects\" out\n       '\n       \n     + test_done\n 4:  213f2c94b73 ! 6:  7a8875be425 scalar: teach `diagnose` to gather loose objects information\n     @@ Commit message\n      \n          When operating at the scale that Scalar wants to support, certain data\n          shapes are more likely to cause undesirable performance issues, such as\n     -    large numbers or large sizes of loose objects.\n     +    large numbers of loose objects.\n      \n          By including statistics about this, `scalar diagnose` now makes it\n          easier to identify such scenarios.\n      \n          Signed-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\n     +    Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## contrib/scalar/scalar.c ##\n     -@@ contrib/scalar/scalar.c: static void dir_file_stats(struct strbuf *buf, const char *path)\n     - \tclosedir(dir);\n     +@@ contrib/scalar/scalar.c: static int dir_file_stats(struct object_directory *object_dir, void *data)\n     + \treturn 0;\n       }\n       \n      +static int count_files(char *path)\n     @@ contrib/scalar/scalar.c: static void dir_file_stats(struct strbuf *buf, const ch\n       {\n       \tstruct option options[] = {\n      @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n     - \tif ((res = stage(tmp_dir.buf, &buf, \"packs-local.txt\")))\n     - \t\tgoto diagnose_cleanup;\n     + \tforeach_alt_odb(dir_file_stats, &buf);\n     + \tstrvec_push(&archiver_args, buf.buf);\n       \n      +\tstrbuf_reset(&buf);\n     ++\tstrbuf_addstr(&buf, \"--add-file-with-content=objects-local.txt:\");\n      +\tloose_objs_stats(&buf, \".git/objects\");\n     ++\tstrvec_push(&archiver_args, buf.buf);\n      +\n     -+\tif ((res = stage(tmp_dir.buf, &buf, \"objects-local.txt\")))\n     -+\t\tgoto diagnose_cleanup;\n     -+\n     - \tif ((res = stage_directory(tmp_dir.buf, \".git\", 0)) ||\n     - \t    (res = stage_directory(tmp_dir.buf, \".git/hooks\", 0)) ||\n     - \t    (res = stage_directory(tmp_dir.buf, \".git/info\", 0)) ||\n     + \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n     + \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n     + \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n      \n       ## contrib/scalar/t/t9099-scalar.sh ##\n     +@@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n     + \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n     + \tgit repack &&\n     + \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n     ++\ttest_commit -C cloned/src loose &&\n     + \tscalar diagnose cloned >out &&\n     + \tgrep \"Available space\" out &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n      @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n       \tunzip -p \"$zip_path\" diagnostics.log >out &&\n       \ttest_file_not_empty out &&\n       \tunzip -p \"$zip_path\" packs-local.txt >out &&\n     -+\ttest_file_not_empty out &&\n     +-\tgrep \"$(pwd)/.git/objects\" out\n     ++\tgrep \"$(pwd)/.git/objects\" out &&\n      +\tunzip -p \"$zip_path\" objects-local.txt >out &&\n     - \ttest_file_not_empty out\n     ++\tgrep \"^Total: [1-9]\" out\n       '\n       \n     + test_done\n 5:  3a2cdce554a < -:  ----------- scalar diagnose: show a spinner while staging content\n\n-- \ngitgitgadget\n"},{"id":"447856","messageId":"0d570137bb6aef675f4f5d74d140ace1dfba5eb7.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 3/6] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:03Z","receivedAt":"2022-02-06T22:39:18Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` command: Instead of writing\nthe `.zip` file directly, we stage the file contents in a Git index of a\ntemporary, bare repository, only to let `git archive` have at it, and\nfinally removing the temporary repository.\n\nAlso note: Due to the frequently-spawned `git hash-object` processes,\nthis command is quite a bit slow on Windows. Should it turn out to be a\nbig problem, the lack of a batch mode of the `hash-object` command could\npotentially be worked around via using `git fast-import` with a crafted\n`stdin`.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 143 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 169 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 00dcd4b50ef..30ce0799c7a 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -261,6 +262,44 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\tres = -1;\n+\t\telse if (recurse)\n+\t\t     add_directory_to_archiver(archiver_args, buf.buf, recurse);\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -501,6 +540,109 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tsize_t off;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf,\n+\t\t      \"--add-file-with-content=diagnostics.log:\"\n+\t\t      \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\toff = strchr(buf.buf, ':') + 1 - buf.buf;\n+\twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tprintf(\"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -802,6 +944,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 9d83fdf25e8..bbd07a44426 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"447857","messageId":"938e38b5a09c7cf1cef4faca577e969b455e1ea5.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 4/6] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:04Z","receivedAt":"2022-02-06T22:39:20Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 30ce0799c7a..fd666376109 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -300,6 +300,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -599,6 +651,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \toff = strchr(buf.buf, ':') + 1 - buf.buf;\n \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n \tstrvec_push(&archiver_args, buf.buf);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex bbd07a44426..f3d037823c8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -94,6 +94,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \ngitgitgadget\n\n"},{"id":"447858","messageId":"bd9428919fab6f4c600ed04ff0ebc5e0ef615ebc.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 5/6] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:05Z","receivedAt":"2022-02-06T22:39:30Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex fd666376109..331d48b2a80 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -592,6 +593,29 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -656,6 +680,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex f3d037823c8..e049221609d 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,6 +93,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n@@ -102,7 +104,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"447859","messageId":"7a8875be425b272becd6c08f4cd5b23c41304ae3.1644187146.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v2 6/6] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-02-06T22:39:06Z","receivedAt":"2022-02-06T22:39:32Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 331d48b2a80..537b97ae734 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -616,6 +616,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -686,6 +740,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex e049221609d..9b4eedbb0aa 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -95,6 +95,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n@@ -106,7 +107,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \ngitgitgadget\n"},{"id":"447903","messageId":"ef70b87e-989b-e99c-b4ca-2b91c05defcf@web.de","threadId":"57313","inReplyTo":"0d570137bb6aef675f4f5d74d140ace1dfba5eb7.1644187146.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 3/6] Implement `scalar diagnose`","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-07T19:55:12Z","receivedAt":"2022-02-07T20:02:24Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"\n\nAm 06.02.22 um 23:39 schrieb Johannes Schindelin via GitGitGadget:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> Over the course of Scalar's development, it became obvious that there is\n> a need for a command that can gather all kinds of useful information\n> that can help identify the most typical problems with large\n> worktrees/repositories.\n>\n> The `diagnose` command is the culmination of this hard-won knowledge: it\n> gathers the installed hooks, the config, a couple statistics describing\n> the data shape, among other pieces of information, and then wraps\n> everything up in a tidy, neat `.zip` archive.\n>\n> Note: originally, Scalar was implemented in C# using the .NET API, where\n> we had the luxury of a comprehensive standard library that includes\n> basic functionality such as writing a `.zip` file. In the C version, we\n> lack such a commodity. Rather than introducing a dependency on, say,\n> libzip, we slightly abuse Git's `archive` command: Instead of writing\n> the `.zip` file directly, we stage the file contents in a Git index of a\n> temporary, bare repository, only to let `git archive` have at it, and\n> finally removing the temporary repository.\n>\n> Also note: Due to the frequently-spawned `git hash-object` processes,\n> this command is quite a bit slow on Windows. Should it turn out to be a\n> big problem, the lack of a batch mode of the `hash-object` command could\n> potentially be worked around via using `git fast-import` with a crafted\n> `stdin`.\n\nThe two paragraphs above are not in sync with the patch.\n\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  contrib/scalar/scalar.c          | 143 +++++++++++++++++++++++++++++++\n>  contrib/scalar/scalar.txt        |  12 +++\n>  contrib/scalar/t/t9099-scalar.sh |  14 +++\n>  3 files changed, 169 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 00dcd4b50ef..30ce0799c7a 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -11,6 +11,7 @@\n>  #include \"dir.h\"\n>  #include \"packfile.h\"\n>  #include \"help.h\"\n> +#include \"archive.h\"\n>\n>  /*\n>   * Remove the deepest subdirectory in the provided path string. Path must not\n> @@ -261,6 +262,44 @@ static int unregister_dir(void)\n>  \treturn res;\n>  }\n>\n> +static int add_directory_to_archiver(struct strvec *archiver_args,\n> +\t\t\t\t\t  const char *path, int recurse)\n> +{\n> +\tint at_root = !*path;\n> +\tDIR *dir = opendir(at_root ? \".\" : path);\n> +\tstruct dirent *e;\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tsize_t len;\n> +\tint res = 0;\n> +\n> +\tif (!dir)\n> +\t\treturn error(_(\"could not open directory '%s'\"), path);\n> +\n> +\tif (!at_root)\n> +\t\tstrbuf_addf(&buf, \"%s/\", path);\n> +\tlen = buf.len;\n> +\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n> +\n> +\twhile (!res && (e = readdir(dir))) {\n> +\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n> +\t\t\tcontinue;\n> +\n> +\t\tstrbuf_setlen(&buf, len);\n> +\t\tstrbuf_addstr(&buf, e->d_name);\n> +\n> +\t\tif (e->d_type == DT_REG)\n> +\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n> +\t\telse if (e->d_type != DT_DIR)\n> +\t\t\tres = -1;\n> +\t\telse if (recurse)\n> +\t\t     add_directory_to_archiver(archiver_args, buf.buf, recurse);\n> +\t}\n> +\n> +\tclosedir(dir);\n> +\tstrbuf_release(&buf);\n> +\treturn res;\n> +}\n> +\n>  /* printf-style interface, expects `<key>=<value>` argument */\n>  static int set_config(const char *fmt, ...)\n>  {\n> @@ -501,6 +540,109 @@ cleanup:\n>  \treturn res;\n>  }\n>\n> +static int cmd_diagnose(int argc, const char **argv)\n> +{\n> +\tstruct option options[] = {\n> +\t\tOPT_END(),\n> +\t};\n> +\tconst char * const usage[] = {\n> +\t\tN_(\"scalar diagnose [<enlistment>]\"),\n> +\t\tNULL\n> +\t};\n> +\tstruct strbuf zip_path = STRBUF_INIT;\n> +\tstruct strvec archiver_args = STRVEC_INIT;\n> +\tchar **argv_copy = NULL;\n> +\tint stdout_fd = -1, archiver_fd = -1;\n> +\ttime_t now = time(NULL);\n> +\tstruct tm tm;\n> +\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n> +\tsize_t off;\n> +\tint res = 0;\n> +\n> +\targc = parse_options(argc, argv, NULL, options,\n> +\t\t\t     usage, 0);\n> +\n> +\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n> +\n> +\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n> +\tstrbuf_addftime(&zip_path,\n> +\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n> +\tstrbuf_addstr(&zip_path, \".zip\");\n> +\tswitch (safe_create_leading_directories(zip_path.buf)) {\n> +\tcase SCLD_EXISTS:\n> +\tcase SCLD_OK:\n> +\t\tbreak;\n> +\tdefault:\n> +\t\terror_errno(_(\"could not create directory for '%s'\"),\n> +\t\t\t    zip_path.buf);\n> +\t\tgoto diagnose_cleanup;\n> +\t}\n> +\tstdout_fd = dup(1);\n> +\tif (stdout_fd < 0) {\n> +\t\tres = error_errno(_(\"could not duplicate stdout\"));\n> +\t\tgoto diagnose_cleanup;\n> +\t}\n> +\n> +\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n> +\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n> +\t\tres = error_errno(_(\"could not redirect output\"));\n> +\t\tgoto diagnose_cleanup;\n> +\t}\n> +\n> +\tinit_zip_archiver();\n> +\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n> +\n> +\tstrbuf_reset(&buf);\n> +\tstrbuf_addstr(&buf,\n> +\t\t      \"--add-file-with-content=diagnostics.log:\"\n> +\t\t      \"Collecting diagnostic info\\n\\n\");\n> +\tget_version_info(&buf, 1);\n> +\n> +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n> +\toff = strchr(buf.buf, ':') + 1 - buf.buf;\n> +\twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n> +\tstrvec_push(&archiver_args, buf.buf);\n\nFun trick to reuse the buffer for both the ZIP entry and stdout. :)  I'd\nhave omitted the option from buf and added it like this, for simplicity:\n\n\tstrvec_pushf(&archiver_args,\n\t\t     \"--add-file-with-content=diagnostics.log:%s\", buf.buf);\n\nJust a thought.\n\n> +\n> +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n> +\t\tgoto diagnose_cleanup;\n> +\n> +\tstrvec_pushl(&archiver_args, \"--prefix=\",\n> +\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n> +\n> +\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n> +\targv_copy = xmemdupz(archiver_args.v,\n> +\t\t\t     sizeof(char *) * archiver_args.nr);\n\nLeaking the whole thing would be fine as well for this command, but\ncleaning up is tidier, of course.\n\n> +\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n> +\t\t\t    the_repository, NULL, 0);\n\nAh -- no shell means no command line length limits. :)\n\n> +\tif (res) {\n> +\t\terror(_(\"failed to write archive\"));\n> +\t\tgoto diagnose_cleanup;\n> +\t}\n> +\n> +\tif (!res)\n> +\t\tprintf(\"\\n\"\n> +\t\t       \"Diagnostics complete.\\n\"\n> +\t\t       \"All of the gathered info is captured in '%s'\\n\",\n> +\t\t       zip_path.buf);\n\nIs this message appended to the ZIP file or does it go to stdout?\n\nIn any case: mixing write(2) and stdio(3) is not a good idea.  Using\nfwrite(3) instead of write_or_die above and doing the stdout dup(2)\ndance only tightly around the write_archive call would help, I think.\n\n> +\n> +diagnose_cleanup:\n> +\tif (archiver_fd >= 0) {\n> +\t\tclose(1);\n> +\t\tdup2(stdout_fd, 1);\n> +\t}\n> +\tfree(argv_copy);\n> +\tstrvec_clear(&archiver_args);\n> +\tstrbuf_release(&zip_path);\n> +\tstrbuf_release(&path);\n> +\tstrbuf_release(&buf);\n> +\n> +\treturn res;\n> +}\n> +\n>  static int cmd_list(int argc, const char **argv)\n>  {\n>  \tif (argc != 1)\n> @@ -802,6 +944,7 @@ static struct {\n>  \t{ \"reconfigure\", cmd_reconfigure },\n>  \t{ \"delete\", cmd_delete },\n>  \t{ \"version\", cmd_version },\n> +\t{ \"diagnose\", cmd_diagnose },\n>  \t{ NULL, NULL},\n>  };\n>\n> diff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\n> index f416d637289..22583fe046e 100644\n> --- a/contrib/scalar/scalar.txt\n> +++ b/contrib/scalar/scalar.txt\n> @@ -14,6 +14,7 @@ scalar register [<enlistment>]\n>  scalar unregister [<enlistment>]\n>  scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n>  scalar reconfigure [ --all | <enlistment> ]\n> +scalar diagnose [<enlistment>]\n>  scalar delete <enlistment>\n>\n>  DESCRIPTION\n> @@ -129,6 +130,17 @@ reconfigure the enlistment.\n>  With the `--all` option, all enlistments currently registered with Scalar\n>  will be reconfigured. Use this option after each Scalar upgrade.\n>\n> +Diagnose\n> +~~~~~~~~\n> +\n> +diagnose [<enlistment>]::\n> +    When reporting issues with Scalar, it is often helpful to provide the\n> +    information gathered by this command, including logs and certain\n> +    statistics describing the data shape of the current enlistment.\n> ++\n> +The output of this command is a `.zip` file that is written into\n> +a directory adjacent to the worktree in the `src` directory.\n> +\n>  Delete\n>  ~~~~~~\n>\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index 9d83fdf25e8..bbd07a44426 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n>  \tgrep \"cloned. does not exist\" err\n>  '\n>\n> +SQ=\"'\"\n> +test_expect_success UNZIP 'scalar diagnose' '\n> +\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n> +\tscalar diagnose cloned >out &&\n> +\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n> +\tzip_path=$(cat zip_path) &&\n> +\ttest -n \"$zip_path\" &&\n> +\tunzip -v \"$zip_path\" &&\n> +\tfolder=${zip_path%.zip} &&\n> +\ttest_path_is_missing \"$folder\" &&\n> +\tunzip -p \"$zip_path\" diagnostics.log >out &&\n> +\ttest_file_not_empty out\n> +'\n> +\n>  test_done\n"},{"id":"447904","messageId":"d1e333b6-3ec1-8569-6ea9-4abd3dee1947@web.de","threadId":"57313","inReplyTo":"49ff3c1f2b32b16df2b4216aa016d715b6de46bc.1644187146.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-07T19:55:02Z","receivedAt":"2022-02-07T20:02:42Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 06.02.22 um 23:39 schrieb Johannes Schindelin via GitGitGadget:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> With the `--add-file-with-content=<path>:<content>` option, `git\n> archive` now supports use cases where relatively trivial files need to\n> be added that do not exist on disk.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  Documentation/git-archive.txt | 11 ++++++++\n>  archive.c                     | 51 +++++++++++++++++++++++++++++------\n>  t/t5003-archive-zip.sh        | 12 +++++++++\n>  3 files changed, 66 insertions(+), 8 deletions(-)\n>\n> diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> index bc4e76a7834..1b52a0a65a1 100644\n> --- a/Documentation/git-archive.txt\n> +++ b/Documentation/git-archive.txt\n> @@ -61,6 +61,17 @@ OPTIONS\n>  \tby concatenating the value for `--prefix` (if any) and the\n>  \tbasename of <file>.\n>\n> +--add-file-with-content=<path>:<content>::\n> +\tAdd the specified contents to the archive.  Can be repeated to add\n> +\tmultiple files.  The path of the file in the archive is built\n> +\tby concatenating the value for `--prefix` (if any) and the\n> +\tbasename of <file>.\n> ++\n> +The `<path>` cannot contain any colon, the file mode is limited to\n> +a regular file, and the option may be subject platform-dependent\n\ns/subject/& to/\n\n> +command-line limits. For non-trivial cases, write an untracked file\n> +and use `--add-file` instead.\n> +\n\nWe could use that option in Git's own Makefile to add the file named\n\"version\", which contains $GIT_VERSION.  Hmm, but it also contains a\nterminating newline, which would be a bit tricky (but not impossible) to\nadd.  Would it make sense to add one automatically if it's missing (e.g.\nwith strbuf_complete_line)?  Not sure.\n\n>  --worktree-attributes::\n>  \tLook for attributes in .gitattributes files in the working tree\n>  \tas well (see <<ATTRIBUTES>>).\n> diff --git a/archive.c b/archive.c\n> index a3bbb091256..172efd690c3 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n>  struct extra_file_info {\n>  \tchar *base;\n>  \tstruct stat stat;\n> +\tvoid *content;\n>  };\n>\n>  int write_archive_entries(struct archiver_args *args,\n> @@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n>  \t\tstrbuf_addstr(&path_in_archive, basename(path));\n>\n>  \t\tstrbuf_reset(&content);\n> -\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n> +\t\tif (info->content)\n> +\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> +\t\t\t\t\t  path_in_archive.len,\n> +\t\t\t\t\t  info->stat.st_mode,\n> +\t\t\t\t\t  info->content, info->stat.st_size);\n> +\t\telse if (strbuf_read_file(&content, path,\n> +\t\t\t\t\t  info->stat.st_size) < 0)\n>  \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n>  \t\telse\n>  \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> @@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n>  {\n>  \tstruct extra_file_info *info = util;\n>  \tfree(info->base);\n> +\tfree(info->content);\n>  \tfree(info);\n>  }\n>\n> @@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n>  \tif (!arg)\n>  \t\treturn -1;\n>\n> -\tpath = prefix_filename(args->prefix, arg);\n> -\titem = string_list_append_nodup(&args->extra_files, path);\n> -\titem->util = info = xmalloc(sizeof(*info));\n> +\tinfo = xmalloc(sizeof(*info));\n>  \tinfo->base = xstrdup_or_null(base);\n> -\tif (stat(path, &info->stat))\n> -\t\tdie(_(\"File not found: %s\"), path);\n> -\tif (!S_ISREG(info->stat.st_mode))\n> -\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\n> +\tif (strcmp(opt->long_name, \"add-file-with-content\")) {\n\nEquivalent to:\n\n\tif (!strcmp(opt->long_name, \"add-file\")) {\n\nI mention that because the inequality check confused me a bit at first.\n\n> +\t\tpath = prefix_filename(args->prefix, arg);\n> +\t\tif (stat(path, &info->stat))\n> +\t\t\tdie(_(\"File not found: %s\"), path);\n> +\t\tif (!S_ISREG(info->stat.st_mode))\n> +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\t\tinfo->content = NULL; /* read the file later */\n> +\t} else {\n> +\t\tconst char *colon = strchr(arg, ':');\n> +\t\tchar *p;\n> +\n> +\t\tif (!colon)\n> +\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n> +\n> +\t\tp = xstrndup(arg, colon - arg);\n> +\t\tif (!args->prefix)\n> +\t\t\tpath = p;\n> +\t\telse {\n> +\t\t\tpath = prefix_filename(args->prefix, p);\n> +\t\t\tfree(p);\n> +\t\t}\n> +\t\tmemset(&info->stat, 0, sizeof(info->stat));\n> +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n> +\t\tinfo->content = xstrdup(colon + 1);\n> +\t\tinfo->stat.st_size = strlen(info->content);\n> +\t}\n> +\titem = string_list_append_nodup(&args->extra_files, path);\n> +\titem->util = info;\n> +\n>  \treturn 0;\n>  }\n>\n> @@ -554,6 +586,9 @@ static int parse_archive_args(int argc, const char **argv,\n>  \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n>  \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n>  \t\t  (intptr_t)&base },\n> +\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n> +\t\t  N_(\"file\"), N_(\"add untracked file to archive\"), 0,\n                      ^^^^\n\"<file>\" seems wrong, because there is no actual file.  It should rather\nbe \"<name>:<content>\" for the virtual one, right?\n\n> +\t\t  add_file_cb, (intptr_t)&base },\n>  \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n>  \t\t\tN_(\"write the archive to this file\")),\n>  \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\n> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> index 1e6d18b140e..8ff1257f1a0 100755\n> --- a/t/t5003-archive-zip.sh\n> +++ b/t/t5003-archive-zip.sh\n> @@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n>  check_zip with_untracked\n>  check_added with_untracked untracked untracked\n>\n> +test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n> +\tgit archive --format=zip >with_file_with_content.zip \\\n> +\t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n> +\ttest_when_finished \"rm -rf tmp-unpack\" &&\n> +\tmkdir tmp-unpack && (\n> +\t\tcd tmp-unpack &&\n> +\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> +\t\ttest_path_is_file hello &&\n> +\t\ttest world = $(cat hello)\n> +\t)\n> +'\n> +\n>  test_expect_success 'git archive --format=zip --add-file twice' '\n>  \techo untracked >untracked &&\n>  \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n"},{"id":"447929","messageId":"xmqqbkzigspr.fsf@gitster.g","threadId":"57313","inReplyTo":"d1e333b6-3ec1-8569-6ea9-4abd3dee1947@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-07T23:30:40Z","receivedAt":"2022-02-08T01:06:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> We could use that option in Git's own Makefile to add the file named\n> \"version\", which contains $GIT_VERSION.  Hmm, but it also contains a\n> terminating newline, which would be a bit tricky (but not impossible) to\n> add.  Would it make sense to add one automatically if it's missing (e.g.\n> with strbuf_complete_line)?  Not sure.\n\nI do not think it is a good UI to give raw file content from the\ncommand line, which will be usable only for trivial, even single\nliner files, and forces people to learn two parallel option, one\nfor trivial ones and the other for contents with meaningful size.\n\n\"--add-blob=<path>:<blob-object-name>\" may be another option, useful\nwhen you have done \"hash-object -w\" already, and can be used to add\nsingle-liner, or an entire novel.\n\nIn any case, \"--add-file=<file>\", which we already have, would be\nmore appropriate feature to use to record our \"version\" file, so\nthere is no need to change our Makefile for it.\n\n"},{"id":"447973","messageId":"nycvar.QRO.7.76.6.2202081303230.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"ef70b87e-989b-e99c-b4ca-2b91c05defcf@web.de","subject":"Re: [PATCH v2 3/6] Implement `scalar diagnose`","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-08T12:08:19Z","receivedAt":"2022-02-08T13:15:15Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi René,\n\nOn Mon, 7 Feb 2022, René Scharfe wrote:\n\n> > Note: originally, Scalar was implemented in C# using the .NET API, where\n> > we had the luxury of a comprehensive standard library that includes\n> > basic functionality such as writing a `.zip` file. In the C version, we\n> > lack such a commodity. Rather than introducing a dependency on, say,\n> > libzip, we slightly abuse Git's `archive` command: Instead of writing\n> > the `.zip` file directly, we stage the file contents in a Git index of a\n> > temporary, bare repository, only to let `git archive` have at it, and\n> > finally removing the temporary repository.\n> >\n> > Also note: Due to the frequently-spawned `git hash-object` processes,\n> > this command is quite a bit slow on Windows. Should it turn out to be a\n> > big problem, the lack of a batch mode of the `hash-object` command could\n> > potentially be worked around via using `git fast-import` with a crafted\n> > `stdin`.\n>\n> The two paragraphs above are not in sync with the patch.\n\nWhoopsie!\n\n> > +\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n> > +\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n> > +\t\tres = error_errno(_(\"could not redirect output\"));\n> > +\t\tgoto diagnose_cleanup;\n> > +\t}\n> > +\n> > +\tinit_zip_archiver();\n> > +\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n> > +\n> > +\tstrbuf_reset(&buf);\n> > +\tstrbuf_addstr(&buf,\n> > +\t\t      \"--add-file-with-content=diagnostics.log:\"\n> > +\t\t      \"Collecting diagnostic info\\n\\n\");\n> > +\tget_version_info(&buf, 1);\n> > +\n> > +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n> > +\toff = strchr(buf.buf, ':') + 1 - buf.buf;\n> > +\twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n> > +\tstrvec_push(&archiver_args, buf.buf);\n>\n> Fun trick to reuse the buffer for both the ZIP entry and stdout. :)  I'd\n> have omitted the option from buf and added it like this, for simplicity:\n>\n> \tstrvec_pushf(&archiver_args,\n> \t\t     \"--add-file-with-content=diagnostics.log:%s\", buf.buf);\n>\n> Just a thought.\n\nOh, that's even better. I did not like that `off` pattern at all but\nforgot to think of `pushf()`. Thanks!\n\n> > +\n> > +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n> > +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n> > +\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n> > +\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n> > +\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n> > +\t\tgoto diagnose_cleanup;\n> > +\n> > +\tstrvec_pushl(&archiver_args, \"--prefix=\",\n> > +\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n> > +\n> > +\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n> > +\targv_copy = xmemdupz(archiver_args.v,\n> > +\t\t\t     sizeof(char *) * archiver_args.nr);\n>\n> Leaking the whole thing would be fine as well for this command, but\n> cleaning up is tidier, of course.\n>\n> > +\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n> > +\t\t\t    the_repository, NULL, 0);\n>\n> Ah -- no shell means no command line length limits. :)\n\nYes!!!\n\nIt also makes the command a ridiculous amount faster on Windows.\n\n> > +\tif (res) {\n> > +\t\terror(_(\"failed to write archive\"));\n> > +\t\tgoto diagnose_cleanup;\n> > +\t}\n> > +\n> > +\tif (!res)\n> > +\t\tprintf(\"\\n\"\n> > +\t\t       \"Diagnostics complete.\\n\"\n> > +\t\t       \"All of the gathered info is captured in '%s'\\n\",\n> > +\t\t       zip_path.buf);\n>\n> Is this message appended to the ZIP file or does it go to stdout?\n\nIt goes to `stdout`, this is for the user who runs `scalar diagnose`.\n\nHmm.\n\nNow that you pointed it out, I think I want it to go to `stderr` instead.\n\n> In any case: mixing write(2) and stdio(3) is not a good idea.  Using\n> fwrite(3) instead of write_or_die above and doing the stdout dup(2)\n> dance only tightly around the write_archive call would help, I think.\n\nSure, but let's print this message to `stderr` instead, that'll be much\ncleaner, right?\n\nAlternatively, I think I'd rather move the `printf()` below...\n\n>\n> > +\n> > +diagnose_cleanup:\n> > +\tif (archiver_fd >= 0) {\n> > +\t\tclose(1);\n> > +\t\tdup2(stdout_fd, 1);\n> > +\t}\n\n... this re-redirection.\n\nWhat do you think? `stdout` or `stderr`?\n\nThank you for your review!\nDscho\n"},{"id":"447987","messageId":"nycvar.QRO.7.76.6.2202081310480.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"d1e333b6-3ec1-8569-6ea9-4abd3dee1947@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-08T12:54:53Z","receivedAt":"2022-02-08T13:15:25Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi René,\n\nOn Mon, 7 Feb 2022, René Scharfe wrote:\n\n> Am 06.02.22 um 23:39 schrieb Johannes Schindelin via GitGitGadget:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >\n> > With the `--add-file-with-content=<path>:<content>` option, `git\n> > archive` now supports use cases where relatively trivial files need to\n> > be added that do not exist on disk.\n> >\n> > Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > ---\n> >  Documentation/git-archive.txt | 11 ++++++++\n> >  archive.c                     | 51 +++++++++++++++++++++++++++++------\n> >  t/t5003-archive-zip.sh        | 12 +++++++++\n> >  3 files changed, 66 insertions(+), 8 deletions(-)\n> >\n> > diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> > index bc4e76a7834..1b52a0a65a1 100644\n> > --- a/Documentation/git-archive.txt\n> > +++ b/Documentation/git-archive.txt\n> > @@ -61,6 +61,17 @@ OPTIONS\n> >  \tby concatenating the value for `--prefix` (if any) and the\n> >  \tbasename of <file>.\n> >\n> > +--add-file-with-content=<path>:<content>::\n> > +\tAdd the specified contents to the archive.  Can be repeated to add\n> > +\tmultiple files.  The path of the file in the archive is built\n> > +\tby concatenating the value for `--prefix` (if any) and the\n> > +\tbasename of <file>.\n> > ++\n> > +The `<path>` cannot contain any colon, the file mode is limited to\n> > +a regular file, and the option may be subject platform-dependent\n>\n> s/subject/& to/\n\nThanks.\n\n> > +command-line limits. For non-trivial cases, write an untracked file\n> > +and use `--add-file` instead.\n> > +\n>\n> We could use that option in Git's own Makefile to add the file named\n> \"version\", which contains $GIT_VERSION.\n\nWe could do that, that opportunity is a side effect of this patch series.\n\n> Hmm, but it also contains a terminating newline, which would be a bit\n> tricky (but not impossible) to add.  Would it make sense to add one\n> automatically if it's missing (e.g. with strbuf_complete_line)?  Not\n> sure.\n\nIt is really easy:\n\n\tLF='\n\t'\n\n\tgit archive --add-file-with-content=version:\"$GIT_VERSION$LF\" ...\n\n(That's shell script, in the Makefile it would need those `\\`\ncontinuations.)\n\n> >  --worktree-attributes::\n> >  \tLook for attributes in .gitattributes files in the working tree\n> >  \tas well (see <<ATTRIBUTES>>).\n> > diff --git a/archive.c b/archive.c\n> > index a3bbb091256..172efd690c3 100644\n> > --- a/archive.c\n> > +++ b/archive.c\n> > @@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n> >  struct extra_file_info {\n> >  \tchar *base;\n> >  \tstruct stat stat;\n> > +\tvoid *content;\n> >  };\n> >\n> >  int write_archive_entries(struct archiver_args *args,\n> > @@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n> >  \t\tstrbuf_addstr(&path_in_archive, basename(path));\n> >\n> >  \t\tstrbuf_reset(&content);\n> > -\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n> > +\t\tif (info->content)\n> > +\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> > +\t\t\t\t\t  path_in_archive.len,\n> > +\t\t\t\t\t  info->stat.st_mode,\n> > +\t\t\t\t\t  info->content, info->stat.st_size);\n> > +\t\telse if (strbuf_read_file(&content, path,\n> > +\t\t\t\t\t  info->stat.st_size) < 0)\n> >  \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n> >  \t\telse\n> >  \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> > @@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n> >  {\n> >  \tstruct extra_file_info *info = util;\n> >  \tfree(info->base);\n> > +\tfree(info->content);\n> >  \tfree(info);\n> >  }\n> >\n> > @@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n> >  \tif (!arg)\n> >  \t\treturn -1;\n> >\n> > -\tpath = prefix_filename(args->prefix, arg);\n> > -\titem = string_list_append_nodup(&args->extra_files, path);\n> > -\titem->util = info = xmalloc(sizeof(*info));\n> > +\tinfo = xmalloc(sizeof(*info));\n> >  \tinfo->base = xstrdup_or_null(base);\n> > -\tif (stat(path, &info->stat))\n> > -\t\tdie(_(\"File not found: %s\"), path);\n> > -\tif (!S_ISREG(info->stat.st_mode))\n> > -\t\tdie(_(\"Not a regular file: %s\"), path);\n> > +\n> > +\tif (strcmp(opt->long_name, \"add-file-with-content\")) {\n>\n> Equivalent to:\n>\n> \tif (!strcmp(opt->long_name, \"add-file\")) {\n>\n> I mention that because the inequality check confused me a bit at first.\n\nGood point. For some reason I thought it would be clearer to handle\neverything but `--add-file-with-content` here, but that \"everything but\"\nis only `--add-file`, so I sowed more confusion. Sorry about that.\n\n>\n> > +\t\tpath = prefix_filename(args->prefix, arg);\n> > +\t\tif (stat(path, &info->stat))\n> > +\t\t\tdie(_(\"File not found: %s\"), path);\n> > +\t\tif (!S_ISREG(info->stat.st_mode))\n> > +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n> > +\t\tinfo->content = NULL; /* read the file later */\n> > +\t} else {\n> > +\t\tconst char *colon = strchr(arg, ':');\n> > +\t\tchar *p;\n> > +\n> > +\t\tif (!colon)\n> > +\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n> > +\n> > +\t\tp = xstrndup(arg, colon - arg);\n> > +\t\tif (!args->prefix)\n> > +\t\t\tpath = p;\n> > +\t\telse {\n> > +\t\t\tpath = prefix_filename(args->prefix, p);\n> > +\t\t\tfree(p);\n> > +\t\t}\n> > +\t\tmemset(&info->stat, 0, sizeof(info->stat));\n> > +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n> > +\t\tinfo->content = xstrdup(colon + 1);\n> > +\t\tinfo->stat.st_size = strlen(info->content);\n> > +\t}\n> > +\titem = string_list_append_nodup(&args->extra_files, path);\n> > +\titem->util = info;\n> > +\n> >  \treturn 0;\n> >  }\n> >\n> > @@ -554,6 +586,9 @@ static int parse_archive_args(int argc, const char **argv,\n> >  \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n> >  \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n> >  \t\t  (intptr_t)&base },\n> > +\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n> > +\t\t  N_(\"file\"), N_(\"add untracked file to archive\"), 0,\n>                       ^^^^\n> \"<file>\" seems wrong, because there is no actual file.  It should rather\n> be \"<name>:<content>\" for the virtual one, right?\n\nOr `<path>:<content>`. Yes.\n\nAgain, thank you for your clear and helpful review,\nDscho\n"},{"id":"447988","messageId":"nycvar.QRO.7.76.6.2202081406520.347@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"xmqqbkzigspr.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-02-08T13:12:21Z","receivedAt":"2022-02-08T13:15:25Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Mon, 7 Feb 2022, Junio C Hamano wrote:\n\n> René Scharfe <l.s.r@web.de> writes:\n>\n> > We could use that option in Git's own Makefile to add the file named\n> > \"version\", which contains $GIT_VERSION.  Hmm, but it also contains a\n> > terminating newline, which would be a bit tricky (but not impossible) to\n> > add.  Would it make sense to add one automatically if it's missing (e.g.\n> > with strbuf_complete_line)?  Not sure.\n>\n> I do not think it is a good UI to give raw file content from the\n> command line, which will be usable only for trivial, even single\n> liner files, and forces people to learn two parallel option, one\n> for trivial ones and the other for contents with meaningful size.\n\nNevertheless, it is still the most elegant way that I can think of to\ngenerate a diagnostic `.zip` file without messing up the very things that\nare to be diagnosed: the repository and the worktree.\n\n> \"--add-blob=<path>:<blob-object-name>\" may be another option, useful\n> when you have done \"hash-object -w\" already, and can be used to add\n> single-liner, or an entire novel.\n\nThis would mess with the repository. Granted, it is unlikely that adding a\ntiny blob will all of a sudden work around a bug that the user wanted to\nreport, but less big mutations have been known to subtly change a bug's\nmanifested symptoms.\n\nSo I really do not want to do that, not in `scalar diagnose.\n\n> In any case, \"--add-file=<file>\", which we already have, would be\n> more appropriate feature to use to record our \"version\" file, so\n> there is no need to change our Makefile for it.\n\nSame here. It is bad enough that `scalar diagnose` has to create a\ndirectory in the current enlistment. Let's not make the situation even\nworse.\n\nThe most elegant solution would have been that streaming `--add-file` mode\nsuggested by René, I think, but that's too involved to implement just to\nbenefit `scalar diagnose`. It's not like we can simply stream the contents\nvia `stdin`, as there are more than one \"virtual\" file we need to add to\nthat `.zip` file.\n\nCiao,\nDscho\n"},{"id":"447992","messageId":"xmqqbkzhdzib.fsf@gitster.g","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2202081406520.347@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-08T17:44:28Z","receivedAt":"2022-02-08T17:44:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> > We could use that option in Git's own Makefile to add the file named\n>> > \"version\", which contains $GIT_VERSION.  Hmm, but it also contains a\n>> > terminating newline, which would be a bit tricky (but not impossible) to\n>> > add.  Would it make sense to add one automatically if it's missing (e.g.\n>> > with strbuf_complete_line)?  Not sure.\n>>\n>> I do not think it is a good UI to give raw file content from the\n>> command line, which will be usable only for trivial, even single\n>> liner files, and forces people to learn two parallel option, one\n>> for trivial ones and the other for contents with meaningful size.\n>\n> Nevertheless, it is still the most elegant way that I can think of to\n> generate a diagnostic `.zip` file without messing up the very things that\n> are to be diagnosed: the repository and the worktree.\n\nPuzzled.  Are you feeding contents of a .zip file from the command\nline?\n\nI was mostly worried about busting command line argument limit by\ntrying to feed too many bytes, as the ceiling is fairly low on some\nplatforms.  Another worry was that when <contents> can have\narbitrary bytes, with --opt=<path>:<contents> syntax, the input\nbecomes ambiguous (i.e. \"which colon is the <path> separator?\"),\nwithout some way to escape a colon in the payload.\n\nFor a single-liner, --add-file-with-contents=<path>:<contents> would\nbe an OK way, and my comment was not a strong objection against this\nnew option existing.  It was primarily an objection against changing\nthe way to add the 'version' file in our \"make dist\" procedure to\nuse it anyway.\n\nBut now I think about it more, I am becoming less happy about it\nexisting in the first place.\n\nThis will throw another monkey wrench to Konstantin's plan [*] to\nmake \"git archive\" output verifiable with the signature on original\nGit objects, but it is not a new problem ;-)\n\n\n[Reference]\n\n* https://lore.kernel.org/git/20220207213449.ljqjhdx4f45a3lx5@meerkat.local/\n"},{"id":"448003","messageId":"b49d396d-a433-51a4-2d19-55e175af571a@web.de","threadId":"57313","inReplyTo":"xmqqbkzhdzib.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-08T20:58:16Z","receivedAt":"2022-02-08T22:25:17Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 08.02.22 um 18:44 schrieb Junio C Hamano:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n>\n>>>> We could use that option in Git's own Makefile to add the file named\n>>>> \"version\", which contains $GIT_VERSION.  Hmm, but it also contains a\n>>>> terminating newline, which would be a bit tricky (but not impossible) to\n>>>> add.  Would it make sense to add one automatically if it's missing (e.g.\n>>>> with strbuf_complete_line)?  Not sure.\n>>>\n>>> I do not think it is a good UI to give raw file content from the\n>>> command line, which will be usable only for trivial, even single\n>>> liner files, and forces people to learn two parallel option, one\n>>> for trivial ones and the other for contents with meaningful size.\n>>\n>> Nevertheless, it is still the most elegant way that I can think of to\n>> generate a diagnostic `.zip` file without messing up the very things that\n>> are to be diagnosed: the repository and the worktree.\n>\n> Puzzled.  Are you feeding contents of a .zip file from the command\n> line?\n\nKind of.  Command line arguments are built and handed to write_archive()\nin-process.  It's done by patch 3 and extended by 5 and 6.\n\nThe number of files is relatively low and they aren't huge, right?\nStaging their content in the object database would be messy, but $TMPDIR\nmight be able to take them with a low impact.  Unless the problem to\ndiagnose is that this directory is full -- but you don't need a fancy\nreport for that. :)\n\nCurrently there is no easy way to write a temporary file with a chosen\nname.  diff.c would benefit from such a thing when running an external\ndiff program; currently it adds a random prefix.  git archive --add-file\nalso uses the filename (and discards the directory part).  The patch\nbelow adds a function to create temporary files with a chosen name.\nPerhaps it would be useful here as well, instead of the new option?\n\n> I was mostly worried about busting command line argument limit by\n> trying to feed too many bytes, as the ceiling is fairly low on some\n> platforms.\n\nCommand line length limits don't apply to the way scalar uses the new\noption.\n\n> Another worry was that when <contents> can have\n> arbitrary bytes, with --opt=<path>:<contents> syntax, the input\n> becomes ambiguous (i.e. \"which colon is the <path> separator?\"),\n> without some way to escape a colon in the payload.\n\nThe first colon is the separator here.\n\n> For a single-liner, --add-file-with-contents=<path>:<contents> would\n> be an OK way, and my comment was not a strong objection against this\n> new option existing.  It was primarily an objection against changing\n> the way to add the 'version' file in our \"make dist\" procedure to\n> use it anyway.\n>\n> But now I think about it more, I am becoming less happy about it\n> existing in the first place.\n>\n> This will throw another monkey wrench to Konstantin's plan [*] to\n> make \"git archive\" output verifiable with the signature on original\n> Git objects, but it is not a new problem ;-)\n>\n>\n> [Reference]\n>\n> * https://lore.kernel.org/git/20220207213449.ljqjhdx4f45a3lx5@meerkat.local/\n\nI don't see the conflict: If an untracked file is added to an archive\nusing --add-file, --add-file-with-content, or ZIP or tar then we'd\n*want* the verification against a signed commit or tag to fail, no?  A\ndifferent signature would be required for the non-tracked parts.\n\nRené\n\n\n--- >8 ---\nSubject: [PATCH] tempfile: add mks_tempfile_dt()\n\nAdd a function to create a temporary file with a certain name in a\ntemporary directory created using mkdtemp(3).  Its result is more\nsightly than the paths created by mks_tempfile_ts(), which include\na random prefix.  That's useful for files passed to a program that\ndisplays their name, e.g. an external diff tool.\n\nSigned-off-by: René Scharfe <l.s.r@web.de>\n---\n tempfile.c | 63 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n tempfile.h | 13 +++++++++++\n 2 files changed, 76 insertions(+)\n\ndiff --git a/tempfile.c b/tempfile.c\nindex 94aa18f3f7..2024c82691 100644\n--- a/tempfile.c\n+++ b/tempfile.c\n@@ -56,6 +56,20 @@\n\n static VOLATILE_LIST_HEAD(tempfile_list);\n\n+static void remove_template_directory(struct tempfile *tempfile,\n+\t\t\t\t      int in_signal_handler)\n+{\n+\tif (tempfile->directorylen > 0 &&\n+\t    tempfile->directorylen < tempfile->filename.len &&\n+\t    tempfile->filename.buf[tempfile->directorylen] == '/') {\n+\t\tstrbuf_setlen(&tempfile->filename, tempfile->directorylen);\n+\t\tif (in_signal_handler)\n+\t\t\trmdir(tempfile->filename.buf);\n+\t\telse\n+\t\t\trmdir_or_warn(tempfile->filename.buf);\n+\t}\n+}\n+\n static void remove_tempfiles(int in_signal_handler)\n {\n \tpid_t me = getpid();\n@@ -74,6 +88,7 @@ static void remove_tempfiles(int in_signal_handler)\n \t\t\tunlink(p->filename.buf);\n \t\telse\n \t\t\tunlink_or_warn(p->filename.buf);\n+\t\tremove_template_directory(p, in_signal_handler);\n\n \t\tp->active = 0;\n \t}\n@@ -100,6 +115,7 @@ static struct tempfile *new_tempfile(void)\n \ttempfile->owner = 0;\n \tINIT_LIST_HEAD(&tempfile->list);\n \tstrbuf_init(&tempfile->filename, 0);\n+\ttempfile->directorylen = 0;\n \treturn tempfile;\n }\n\n@@ -198,6 +214,52 @@ struct tempfile *mks_tempfile_tsm(const char *filename_template, int suffixlen,\n \treturn tempfile;\n }\n\n+struct tempfile *mks_tempfile_dt(const char *directory_template,\n+\t\t\t\t const char *filename)\n+{\n+\tstruct tempfile *tempfile;\n+\tconst char *tmpdir;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tint fd;\n+\tsize_t directorylen;\n+\n+\tif (!ends_with(directory_template, \"XXXXXX\")) {\n+\t\terrno = EINVAL;\n+\t\treturn NULL;\n+\t}\n+\n+\ttmpdir = getenv(\"TMPDIR\");\n+\tif (!tmpdir)\n+\t\ttmpdir = \"/tmp\";\n+\n+\tstrbuf_addf(&sb, \"%s/%s\", tmpdir, directory_template);\n+\tdirectorylen = sb.len;\n+\tif (!mkdtemp(sb.buf)) {\n+\t\tint orig_errno = errno;\n+\t\tstrbuf_release(&sb);\n+\t\terrno = orig_errno;\n+\t\treturn NULL;\n+\t}\n+\n+\tstrbuf_addf(&sb, \"/%s\", filename);\n+\tfd = open(sb.buf, O_CREAT | O_EXCL | O_RDWR, 0600);\n+\tif (fd < 0) {\n+\t\tint orig_errno = errno;\n+\t\tstrbuf_setlen(&sb, directorylen);\n+\t\trmdir(sb.buf);\n+\t\tstrbuf_release(&sb);\n+\t\terrno = orig_errno;\n+\t\treturn NULL;\n+\t}\n+\n+\ttempfile = new_tempfile();\n+\tstrbuf_swap(&tempfile->filename, &sb);\n+\ttempfile->directorylen = directorylen;\n+\ttempfile->fd = fd;\n+\tactivate_tempfile(tempfile);\n+\treturn tempfile;\n+}\n+\n struct tempfile *xmks_tempfile_m(const char *filename_template, int mode)\n {\n \tstruct tempfile *tempfile;\n@@ -316,6 +378,7 @@ void delete_tempfile(struct tempfile **tempfile_p)\n\n \tclose_tempfile_gently(tempfile);\n \tunlink_or_warn(tempfile->filename.buf);\n+\tremove_template_directory(tempfile, 0);\n \tdeactivate_tempfile(tempfile);\n \t*tempfile_p = NULL;\n }\ndiff --git a/tempfile.h b/tempfile.h\nindex 4de3bc77d2..d7804a214a 100644\n--- a/tempfile.h\n+++ b/tempfile.h\n@@ -82,6 +82,7 @@ struct tempfile {\n \tFILE *volatile fp;\n \tvolatile pid_t owner;\n \tstruct strbuf filename;\n+\tsize_t directorylen;\n };\n\n /*\n@@ -198,6 +199,18 @@ static inline struct tempfile *xmks_tempfile(const char *filename_template)\n \treturn xmks_tempfile_m(filename_template, 0600);\n }\n\n+/*\n+ * Attempt to create a temporary directory in $TMPDIR and to create and\n+ * open a file in that new directory. Derive the directory name from the\n+ * template in the manner of mkdtemp(). Arrange for directory and file\n+ * to be deleted if the program exits before they are deleted\n+ * explicitly. On success return a tempfile whose \"filename\" member\n+ * contains the full path of the file and its \"fd\" member is open for\n+ * writing the file. On error return NULL and set errno appropriately.\n+ */\n+struct tempfile *mks_tempfile_dt(const char *directory_template,\n+\t\t\t\t const char *filename);\n+\n /*\n  * Associate a stdio stream with the temporary file (which must still\n  * be open). Return `NULL` (*without* deleting the file) on error. The\n--\n2.35.1\n"},{"id":"448076","messageId":"xmqqk0e364h7.fsf@gitster.g","threadId":"57313","inReplyTo":"b49d396d-a433-51a4-2d19-55e175af571a@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-09T22:48:52Z","receivedAt":"2022-02-09T22:48:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n>>> Nevertheless, it is still the most elegant way that I can think of to\n>>> generate a diagnostic `.zip` file without messing up the very things that\n>>> are to be diagnosed: the repository and the worktree.\n>>\n>> Puzzled.  Are you feeding contents of a .zip file from the command\n>> line?\n>\n> Kind of.  Command line arguments are built and handed to write_archive()\n> in-process.  It's done by patch 3 and extended by 5 and 6.\n\nI meant to ask if this is doing\n\n    git archive --store-contents-at-path=\"report.zip:$(cat diag.zip)\"\n\nas I misunderstood what 'the diagnostic .zip file' referred to.\nThat was a reference to the output of the \"git archive\" command.\n\n> The number of files is relatively low and they aren't huge, right?\n\nAs long as it is expected to fit on the command line, that's fine.\nBut if the question is \"it is OK to add a new option with known\nlimitation\", then it should be stated a bit differently.\n\n\"We add this option for use cases where we handle only small number\nof one-liner files\", and it is OK.  We may however want to do\nsomething imilar to what we do to the \"-m '<message>'\" option used\nby \"git commit\" and \"git merge\", i.e. add the final LF when it is\nmissing to make it a complete line, to hint the fact that this is\nmeant to add a small number of single liner files.\n\n>> Another worry was that when <contents> can have\n>> arbitrary bytes, with --opt=<path>:<contents> syntax, the input\n>> becomes ambiguous (i.e. \"which colon is the <path> separator?\"),\n>> without some way to escape a colon in the payload.\n>\n> The first colon is the separator here.\n\nMeaning you cannot have a colon in the path, which is not exactly\npleasing limitation.  I know you may not be able to do so on Windows\nor CIFS mounted on non-Windows, but we do not limit ourselves to\nportable filename character set (POSIX.1 3.282), either.\n\n>> This will throw another monkey wrench to Konstantin's plan [*] to\n>> make \"git archive\" output verifiable with the signature on original\n>> Git objects, but it is not a new problem ;-)\n>>\n>>\n>> [Reference]\n>>\n>> * https://lore.kernel.org/git/20220207213449.ljqjhdx4f45a3lx5@meerkat.local/\n>\n> I don't see the conflict: If an untracked file is added to an archive\n> using --add-file, --add-file-with-content, or ZIP or tar then we'd\n> *want* the verification against a signed commit or tag to fail, no?  A\n> different signature would be required for the non-tracked parts.\n\nYes, which is exactly how this (and existing --add-file) makes\nKonstantin's plan much less useful.\n\nThanks.\n"},{"id":"448182","messageId":"6f3d288a-8c2f-0d63-ea17-f6c038a9fa3e@web.de","threadId":"57313","inReplyTo":"xmqqk0e364h7.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-10T19:10:35Z","receivedAt":"2022-02-10T19:10:46Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 09.02.22 um 23:48 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>> The number of files is relatively low and they aren't huge, right?\n>\n> As long as it is expected to fit on the command line, that's fine.\n> But if the question is \"it is OK to add a new option with known\n> limitation\", then it should be stated a bit differently.\n\nI asked this question to find out if writing the files to $TMPDIR and\nadding them with --add-file instead of with --add-file-with-content\nwould be feasible in patches 3 to 6.  git archive would not have to be\nchanged in that case.\n\n>>> This will throw another monkey wrench to Konstantin's plan [*] to\n>>> make \"git archive\" output verifiable with the signature on original\n>>> Git objects, but it is not a new problem ;-)\n>>>\n>>>\n>>> [Reference]\n>>>\n>>> * https://lore.kernel.org/git/20220207213449.ljqjhdx4f45a3lx5@meerkat.local/\n>>\n>> I don't see the conflict: If an untracked file is added to an archive\n>> using --add-file, --add-file-with-content, or ZIP or tar then we'd\n>> *want* the verification against a signed commit or tag to fail, no?  A\n>> different signature would be required for the non-tracked parts.\n>\n> Yes, which is exactly how this (and existing --add-file) makes\n> Konstantin's plan much less useful.\nPeople added untracked files to archives before --add-file existed.\n\n--add-file-with-content could be used to add the .GIT_ARCHIVE_SIG file.\n\nAdditional untracked files would need a manifest to specify which files\nare (not) covered by the signed commit/tag.  Or the .GIT_ARCHIVE_SIG\nfiles could be added just after the signed files as a rule, before any\nother untracked files, as some kind of a separator.\n\nJust listing untracked files and verifying the others might still be\nuseful.  Warning about untracked files shadowing tracked ones would be\nvery useful.\n\nSome equivalent to the .GIT_ARCHIVE_SIG file containing a signature of\nthe untracked files could optionally be added at the end to allow full\nverification -- but would require signing at archive creation time.\n\nRené\n"},{"id":"448187","messageId":"xmqqk0e2frux.fsf@gitster.g","threadId":"57313","inReplyTo":"6f3d288a-8c2f-0d63-ea17-f6c038a9fa3e@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-10T19:23:34Z","receivedAt":"2022-02-10T19:23:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n>> Yes, which is exactly how this (and existing --add-file) makes\n>> Konstantin's plan much less useful.\n> People added untracked files to archives before --add-file existed.\n>\n> --add-file-with-content could be used to add the .GIT_ARCHIVE_SIG file.\n>\n> Additional untracked files would need a manifest to specify which files\n> are (not) covered by the signed commit/tag.  Or the .GIT_ARCHIVE_SIG\n> files could be added just after the signed files as a rule, before any\n> other untracked files, as some kind of a separator.\n\nOr if people do not _exclude_ tracked files from the archive, then\nthe verifier who has a tarball and a Git tree object can consult the\ntree object to see which ones are added untracked cruft.\n\n> Just listing untracked files and verifying the others might still be\n> useful.  Warning about untracked files shadowing tracked ones would be\n> very useful.\n\nYup.\n\n> Some equivalent to the .GIT_ARCHIVE_SIG file containing a signature of\n> the untracked files could optionally be added at the end to allow full\n> verification -- but would require signing at archive creation time.\n\nYeah, and at that point, it is not much more convenient than just\nsigning the whole archive (sans the SIG part, obviously), which is\nwhat people have always done ;-)\n"},{"id":"448240","messageId":"f83ed995-6dff-bc41-8782-48ac9f1a2651@web.de","threadId":"57313","inReplyTo":"xmqqk0e2frux.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-11T19:16:43Z","receivedAt":"2022-02-11T19:22:00Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 10.02.22 um 20:23 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>>> Yes, which is exactly how this (and existing --add-file) makes\n>>> Konstantin's plan much less useful.\n\nA harder obstacle to verification would be end-of-line conversion.\nRetrying a failed signature check after applying convert_to_git() might\nwork, but not for files that have mixed line endings in the repository\nand end up being homogenized during checkout (and thus archiving).\n\n>> People added untracked files to archives before --add-file existed.\n>>\n>> --add-file-with-content could be used to add the .GIT_ARCHIVE_SIG file.\n>>\n>> Additional untracked files would need a manifest to specify which files\n>> are (not) covered by the signed commit/tag.  Or the .GIT_ARCHIVE_SIG\n>> files could be added just after the signed files as a rule, before any\n>> other untracked files, as some kind of a separator.\n>\n> Or if people do not _exclude_ tracked files from the archive, then\n> the verifier who has a tarball and a Git tree object can consult the\n> tree object to see which ones are added untracked cruft.\n\nTrue, but if you have the tree objects then you probably also have the\nblobs and don't need the archive?  Or is this some kind of sparse\ncheckout scenario?\n\n>> Some equivalent to the .GIT_ARCHIVE_SIG file containing a signature of\n>> the untracked files could optionally be added at the end to allow full\n>> verification -- but would require signing at archive creation time.\n>\n> Yeah, and at that point, it is not much more convenient than just\n> signing the whole archive (sans the SIG part, obviously), which is\n> what people have always done ;-)\n\nIndeed.\n\nRené\n"},{"id":"448285","messageId":"xmqqk0e19jrp.fsf@gitster.g","threadId":"57313","inReplyTo":"f83ed995-6dff-bc41-8782-48ac9f1a2651@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-11T21:27:06Z","receivedAt":"2022-02-11T21:27:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n>> Or if people do not _exclude_ tracked files from the archive, then\n>> the verifier who has a tarball and a Git tree object can consult the\n>> tree object to see which ones are added untracked cruft.\n>\n> True, but if you have the tree objects then you probably also have the\n> blobs and don't need the archive?  Or is this some kind of sparse\n> checkout scenario?\n\nMy phrasing was too loose.  This is a \"how to verify a distro\ntarball\" (without having a copy of the project repository, but with\nsome common tools like \"git\") scenario.\n\nThe verifier has a tarball.  In addition, the verifier knows the\nobject name of the Git tree object the tarball was taken from, and\nsomehow trusts that the object name is genuine.  We can do either\n\"untar + git-add . && git write-tree\" or its equivalent to see how\nthe contents hashes to the expected tree (or not).\n\nHow the verifier trusts the object name is out of scope (it may come\nfrom a copy of a signed tag object and a copy of the commit object\nthat the tag points at and the contents of signed tag object, with\nits known format, would allow you to write a stand alone tool to\nverify the PGP signature).\n\nLine-end normalization and smudge filter rules may get in the way,\nif we truly did \"untar\" to the filesystem, but I thought \"git\narchive\" didn't do smudge conversion and core.crlf handling when\ncreating the archive?\n\n\n"},{"id":"448294","messageId":"b05f916c-4b04-4db6-d203-10be0a8eb615@web.de","threadId":"57313","inReplyTo":"xmqqk0e19jrp.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-12T09:12:30Z","receivedAt":"2022-02-12T09:12:41Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 11.02.22 um 22:27 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>>> Or if people do not _exclude_ tracked files from the archive, then\n>>> the verifier who has a tarball and a Git tree object can consult the\n>>> tree object to see which ones are added untracked cruft.\n>>\n>> True, but if you have the tree objects then you probably also have the\n>> blobs and don't need the archive?  Or is this some kind of sparse\n>> checkout scenario?\n>\n> My phrasing was too loose.  This is a \"how to verify a distro\n> tarball\" (without having a copy of the project repository, but with\n> some common tools like \"git\") scenario.\n>\n> The verifier has a tarball.  In addition, the verifier knows the\n> object name of the Git tree object the tarball was taken from, and\n> somehow trusts that the object name is genuine.  We can do either\n> \"untar + git-add . && git write-tree\" or its equivalent to see how\n> the contents hashes to the expected tree (or not).\n>\n> How the verifier trusts the object name is out of scope (it may come\n> from a copy of a signed tag object and a copy of the commit object\n> that the tag points at and the contents of signed tag object, with\n> its known format, would allow you to write a stand alone tool to\n> verify the PGP signature).\n\nRight, but the tree hash does not directly allow to see which objects\nare tracked or not.  This information is necessary to reconstruct the\nsigned tree.  (Having tracked files first, then the signature file and\nthen untracked files in the archive would be an easy way to transmit\nit.)\n\n> Line-end normalization and smudge filter rules may get in the way,\n> if we truly did \"untar\" to the filesystem, but I thought \"git\n> archive\" didn't do smudge conversion and core.crlf handling when\n> creating the archive?\n\ngit archive uses convert_to_working_tree() to archive the same file\ncontents as tar or zip would.\n\nRené\n"},{"id":"448336","messageId":"xmqqfson706p.fsf@gitster.g","threadId":"57313","inReplyTo":"b05f916c-4b04-4db6-d203-10be0a8eb615@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-13T06:25:18Z","receivedAt":"2022-02-13T06:25:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n>> The verifier has a tarball.  In addition, the verifier knows the\n>> object name of the Git tree object the tarball was taken from, and\n>> somehow trusts that the object name is genuine.  We can do either\n>> \"untar + git-add . && git write-tree\" or its equivalent to see how\n>> the contents hashes to the expected tree (or not).\n> ...\n> Right, but the tree hash does not directly allow to see which objects\n> are tracked or not.\n\nAh, of course---it was silly of me to overlook this obvious fact X-<.\nSo we do need some extra \"manifest\" to declare what's untracked etc.,\nif we allow --add-file etc. to munge the tree when creating a tarball\nout of it.\n\n"},{"id":"448338","messageId":"2d4358db-fc6e-dc89-e647-b1b810817873@web.de","threadId":"57313","inReplyTo":"xmqqfson706p.fsf@gitster.g","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-02-13T09:02:07Z","receivedAt":"2022-02-13T09:02:18Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 13.02.22 um 07:25 schrieb Junio C Hamano:\n>\n> So we do need some extra \"manifest\" to declare what's untracked etc.,\n> if we allow --add-file etc. to munge the tree when creating a tarball\n> out of it.\n\nRight, or get that information from the order of files in the archive,\nby having tracked files come first, then the signature file with a\ncertain name and then untracked files.\n\nRené\n"},{"id":"448368","messageId":"xmqqczjp4b2x.fsf@gitster.g","threadId":"57313","inReplyTo":"2d4358db-fc6e-dc89-e647-b1b810817873@web.de","subject":"Re: [PATCH v2 1/6] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-02-14T17:22:46Z","receivedAt":"2022-02-14T17:22:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 13.02.22 um 07:25 schrieb Junio C Hamano:\n>>\n>> So we do need some extra \"manifest\" to declare what's untracked etc.,\n>> if we allow --add-file etc. to munge the tree when creating a tarball\n>> out of it.\n>\n> Right, or get that information from the order of files in the archive,\n> by having tracked files come first, then the signature file with a\n> certain name and then untracked files.\n\nThat sounds like a workable approach, modulo that the details of the\n\"signature file with a certain name\" part needs to be worked out.\n\nWe should make sure that we clearly document that \"--add-file=\" and\nfriends add their material after the contents that come from the\ntree-ish, and make sure that the program does so and will stay doing\nso.  Otherwise users cannot easily create an archive that follows\nthe above rule.\n\nThanks.\n"},{"id":"454817","messageId":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v2.git.1644187146.gitgitgadget@gmail.com","subject":"[PATCH v3 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:12Z","receivedAt":"2022-05-04T15:26:04Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nChanges since v2:\n\n * Clarified in the commit message what the biggest benefit of\n   --add-file-with-content is.\n * The <path> part of the -add-file-with-content argument can now contain\n   colons. To do this, the path needs to start and end in double-quote\n   characters (which are stripped), and the backslash serves as escape\n   character in that case (to allow the path to contain both colons and\n   double-quotes).\n * Fixed incorrect grammar.\n * Instead of strcmp(<what-we-don't-want>), we now say\n   !strcmp(<what-we-want>).\n * The help text for --add-file-with-content was improved a tiny bit.\n * Adjusted the commit message that still talked about spawning plenty of\n   processes and about a throw-away repository for the sake of generating a\n   .zip file.\n * Simplified the code that shows the diagnostics and adds them to the .zip\n   file.\n * The final message that reports that the archive is complete is now\n   printed to stderr instead of stdout.\n\nChanges since v1:\n\n * Instead of creating a throw-away repository, staging the contents of the\n   .zip file and then using git write-tree and git archive to write the .zip\n   file, the patch series now introduces a new option to git archive and\n   uses write_archive() directly (avoiding any separate process).\n * Since the command avoids separate processes, it is now blazing fast on\n   Windows, and I dropped the spinner() function because it's no longer\n   needed.\n * While reworking the test case, I noticed that scalar [...] <enlistment>\n   failed to verify that the specified directory exists, and would happily\n   \"traverse to its parent directory\" on its quest to find a Scalar\n   enlistment. That is of course incorrect, and has been fixed as a \"while\n   at it\" sort of preparatory commit.\n * I had forgotten to sign off on all the commits, which has been fixed.\n * Instead of some \"home-grown\" readdir()-based function, the code now uses\n   for_each_file_in_pack_dir() to look through the pack directories.\n * If any alternates are configured, their pack directories are now included\n   in the output.\n * The commit message that might be interpreted to promise information about\n   large loose files has been corrected to no longer promise that.\n * The test cases have been adjusted to test a little bit more (e.g.\n   verifying that specific paths are mentioned in the output, instead of\n   merely verifying that the output is non-empty).\n\nJohannes Schindelin (5):\n  archive: optionally add \"virtual\" files\n  archive --add-file-with-contents: allow paths containing colons\n  scalar: validate the optional enlistment argument\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n Documentation/git-archive.txt    |  16 ++\n archive.c                        |  75 +++++++-\n contrib/scalar/scalar.c          | 289 ++++++++++++++++++++++++++++++-\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  27 +++\n t/t5003-archive-zip.sh           |  20 +++\n 6 files changed, 429 insertions(+), 10 deletions(-)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n\nRange-diff vs v2:\n\n 1:  49ff3c1f2b3 ! 1:  45662cf582a archive: optionally add \"virtual\" files\n     @@ Commit message\n          archive` now supports use cases where relatively trivial files need to\n          be added that do not exist on disk.\n      \n     +    This will allow us to generate `.zip` files with generated content,\n     +    without having to add said content to the object database and without\n     +    having to write it out to disk.\n     +\n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## Documentation/git-archive.txt ##\n     @@ Documentation/git-archive.txt: OPTIONS\n      +\tbasename of <file>.\n      ++\n      +The `<path>` cannot contain any colon, the file mode is limited to\n     -+a regular file, and the option may be subject platform-dependent\n     ++a regular file, and the option may be subject to platform-dependent\n      +command-line limits. For non-trivial cases, write an untracked file\n      +and use `--add-file` instead.\n      +\n     @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int\n      -\tif (!S_ISREG(info->stat.st_mode))\n      -\t\tdie(_(\"Not a regular file: %s\"), path);\n      +\n     -+\tif (strcmp(opt->long_name, \"add-file-with-content\")) {\n     ++\tif (!strcmp(opt->long_name, \"add-file\")) {\n      +\t\tpath = prefix_filename(args->prefix, arg);\n      +\t\tif (stat(path, &info->stat))\n      +\t\t\tdie(_(\"File not found: %s\"), path);\n     @@ archive.c: static int parse_archive_args(int argc, const char **argv,\n       \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n       \t\t  (intptr_t)&base },\n      +\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n     -+\t\t  N_(\"file\"), N_(\"add untracked file to archive\"), 0,\n     ++\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n      +\t\t  add_file_cb, (intptr_t)&base },\n       \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n       \t\t\tN_(\"write the archive to this file\")),\n -:  ----------- > 2:  ce4b1b680c9 archive --add-file-with-contents: allow paths containing colons\n 2:  600da8d465e = 3:  5a3eeb55409 scalar: validate the optional enlistment argument\n 3:  0d570137bb6 ! 4:  dfe821d10fe Implement `scalar diagnose`\n     @@ Commit message\n          we had the luxury of a comprehensive standard library that includes\n          basic functionality such as writing a `.zip` file. In the C version, we\n          lack such a commodity. Rather than introducing a dependency on, say,\n     -    libzip, we slightly abuse Git's `archive` command: Instead of writing\n     -    the `.zip` file directly, we stage the file contents in a Git index of a\n     -    temporary, bare repository, only to let `git archive` have at it, and\n     -    finally removing the temporary repository.\n     -\n     -    Also note: Due to the frequently-spawned `git hash-object` processes,\n     -    this command is quite a bit slow on Windows. Should it turn out to be a\n     -    big problem, the lack of a batch mode of the `hash-object` command could\n     -    potentially be worked around via using `git fast-import` with a crafted\n     -    `stdin`.\n     +    libzip, we slightly abuse Git's `archive` machinery: we write out a\n     +    `.zip` of the empty try, augmented by a couple files that are added via\n     +    the `--add-file*` options. We are careful trying not to modify the\n     +    current repository in any way lest the very circumstances that required\n     +    `scalar diagnose` to be run are changed by the `diagnose` run itself.\n      \n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n     @@ contrib/scalar/scalar.c: cleanup:\n      +\ttime_t now = time(NULL);\n      +\tstruct tm tm;\n      +\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n     -+\tsize_t off;\n      +\tint res = 0;\n      +\n      +\targc = parse_options(argc, argv, NULL, options,\n     @@ contrib/scalar/scalar.c: cleanup:\n      +\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n      +\n      +\tstrbuf_reset(&buf);\n     -+\tstrbuf_addstr(&buf,\n     -+\t\t      \"--add-file-with-content=diagnostics.log:\"\n     -+\t\t      \"Collecting diagnostic info\\n\\n\");\n     ++\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n      +\tget_version_info(&buf, 1);\n      +\n      +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n     -+\toff = strchr(buf.buf, ':') + 1 - buf.buf;\n     -+\twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     -+\tstrvec_push(&archiver_args, buf.buf);\n     ++\twrite_or_die(stdout_fd, buf.buf, buf.len);\n     ++\tstrvec_pushf(&archiver_args,\n     ++\t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n     ++\t\t     (int)buf.len, buf.buf);\n      +\n      +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n      +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n     @@ contrib/scalar/scalar.c: cleanup:\n      +\t}\n      +\n      +\tif (!res)\n     -+\t\tprintf(\"\\n\"\n     ++\t\tfprintf(stderr, \"\\n\"\n      +\t\t       \"Diagnostics complete.\\n\"\n      +\t\t       \"All of the gathered info is captured in '%s'\\n\",\n      +\t\t       zip_path.buf);\n 4:  938e38b5a09 ! 5:  bb162abd383 scalar diagnose: include disk space information\n     @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n       \n       \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n      +\tget_disk_info(&buf);\n     - \toff = strchr(buf.buf, ':') + 1 - buf.buf;\n     - \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     - \tstrvec_push(&archiver_args, buf.buf);\n     + \twrite_or_die(stdout_fd, buf.buf, buf.len);\n     + \tstrvec_pushf(&archiver_args,\n     + \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n      \n       ## contrib/scalar/t/t9099-scalar.sh ##\n      @@ contrib/scalar/t/t9099-scalar.sh: SQ=\"'\"\n 5:  bd9428919fa ! 6:  32aaad7cce1 scalar: teach `diagnose` to gather packfile info\n     @@ contrib/scalar/scalar.c: cleanup:\n       {\n       \tstruct option options[] = {\n      @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n     - \twrite_or_die(stdout_fd, buf.buf + off, buf.len - off);\n     - \tstrvec_push(&archiver_args, buf.buf);\n     + \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n     + \t\t     (int)buf.len, buf.buf);\n       \n      +\tstrbuf_reset(&buf);\n      +\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n 6:  7a8875be425 = 7:  322932f0bb8 scalar: teach `diagnose` to gather loose objects information\n\n-- \ngitgitgadget\n"},{"id":"454818","messageId":"ce4b1b680c98d0f55d4d307b8c746a81d90ffa06.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:14Z","receivedAt":"2022-05-04T15:26:05Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nBy allowing the path to be enclosed in double-quotes, we can avoid\nthe limitation that paths cannot contain colons.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 13 +++++++++----\n archive.c                     | 34 +++++++++++++++++++++++++++++-----\n t/t5003-archive-zip.sh        |  8 ++++++++\n 3 files changed, 46 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex a0edc9167b2..1789ce4c232 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -67,10 +67,15 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n +\n-The `<path>` cannot contain any colon, the file mode is limited to\n-a regular file, and the option may be subject to platform-dependent\n-command-line limits. For non-trivial cases, write an untracked file\n-and use `--add-file` instead.\n+The `<path>` argument can start and end with a literal double-quote\n+character. In this case, the backslash is interpreted as escape\n+character. The path must be quoted if it contains a colon, to avoid\n+the colon from being misinterpreted as the separator between the\n+path and the contents.\n++\n+The file mode is limited to a regular file, and the option may be\n+subject to platform-dependent command-line limits. For non-trivial\n+cases, write an untracked file and use `--add-file` instead.\n \n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\ndiff --git a/archive.c b/archive.c\nindex d798624cd5f..3b751027143 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -533,13 +533,37 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n \t} else {\n-\t\tconst char *colon = strchr(arg, ':');\n \t\tchar *p;\n \n-\t\tif (!colon)\n-\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\t\tif (*arg != '\"') {\n+\t\t\tconst char *colon = strchr(arg, ':');\n+\n+\t\t\tif (!colon)\n+\t\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\t\t\tp = xstrndup(arg, colon - arg);\n+\t\t\targ = colon + 1;\n+\t\t} else {\n+\t\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\t\tconst char *orig = arg;\n+\n+\t\t\tfor (;;) {\n+\t\t\t\tif (!*(++arg))\n+\t\t\t\t\tdie(_(\"unclosed quote: '%s'\"), orig);\n+\t\t\t\tif (*arg == '\"')\n+\t\t\t\t\tbreak;\n+\t\t\t\tif (*arg == '\\\\' && *(++arg) == '\\0')\n+\t\t\t\t\tdie(_(\"trailing backslash: '%s\"), orig);\n+\t\t\t\telse\n+\t\t\t\t\tstrbuf_addch(&buf, *arg);\n+\t\t\t}\n+\n+\t\t\tif (*(++arg) != ':')\n+\t\t\t\tdie(_(\"missing colon: '%s'\"), orig);\n+\n+\t\t\tp = strbuf_detach(&buf, NULL);\n+\t\t\targ++;\n+\t\t}\n \n-\t\tp = xstrndup(arg, colon - arg);\n \t\tif (!args->prefix)\n \t\t\tpath = p;\n \t\telse {\n@@ -548,7 +572,7 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t}\n \t\tmemset(&info->stat, 0, sizeof(info->stat));\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n-\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->content = xstrdup(arg);\n \t\tinfo->stat.st_size = strlen(info->content);\n \t}\n \titem = string_list_append_nodup(&args->extra_files, path);\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 8ff1257f1a0..5b8bbfc2692 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -207,13 +207,21 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n+\tif test_have_prereq FUNNYNAMES\n+\tthen\n+\t\tQUOTED=quoted:colon\n+\telse\n+\t\tQUOTED=quoted\n+\tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n+\t\ttest_path_is_file $QUOTED &&\n \t\ttest world = $(cat hello)\n \t)\n '\n-- \ngitgitgadget\n\n"},{"id":"454819","messageId":"45662cf582ab7c8b1c32f55c9a34f4d73a28b71d.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 1/7] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:13Z","receivedAt":"2022-05-04T15:26:07Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-file-with-content=<path>:<content>` option, `git\narchive` now supports use cases where relatively trivial files need to\nbe added that do not exist on disk.\n\nThis will allow us to generate `.zip` files with generated content,\nwithout having to add said content to the object database and without\nhaving to write it out to disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 11 ++++++++\n archive.c                     | 51 +++++++++++++++++++++++++++++------\n t/t5003-archive-zip.sh        | 12 +++++++++\n 3 files changed, 66 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a7834..a0edc9167b2 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -61,6 +61,17 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n \n+--add-file-with-content=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value for `--prefix` (if any) and the\n+\tbasename of <file>.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject to platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex a3bbb091256..d798624cd5f 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n \t\tstrbuf_addstr(&path_in_archive, basename(path));\n \n \t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\tif (info->content)\n+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\telse if (strbuf_read_file(&content, path,\n+\t\t\t\t\t  info->stat.st_size) < 0)\n \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n@@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (!strcmp(opt->long_name, \"add-file\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +586,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n+\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 1e6d18b140e..8ff1257f1a0 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \ngitgitgadget\n\n"},{"id":"454820","messageId":"5a3eeb5540943279d1677c1338df86b58239abc1.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:15Z","receivedAt":"2022-05-04T15:26:12Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..00dcd4b50ef 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..9d83fdf25e8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n \ttest_path_is_missing cloned\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"454821","messageId":"bb162abd383efd57c2953812fe44ac9fad838523.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 5/7] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:17Z","receivedAt":"2022-05-04T15:26:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex a290e52e1d2..df44902c909 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -300,6 +300,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -596,6 +648,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \twrite_or_die(stdout_fd, buf.buf, buf.len);\n \tstrvec_pushf(&archiver_args,\n \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex bbd07a44426..f3d037823c8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -94,6 +94,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \ngitgitgadget\n\n"},{"id":"454822","messageId":"322932f0bb87db2e11fc7b8447c3fc0fe134ae9a.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 7/7] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:19Z","receivedAt":"2022-05-04T15:26:15Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 9adde8cf4b9..f2fe3858eca 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -616,6 +616,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -684,6 +738,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex e049221609d..9b4eedbb0aa 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -95,6 +95,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n@@ -106,7 +107,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \ngitgitgadget\n"},{"id":"454823","messageId":"32aaad7cce1436ee1c6a7607cc8656b1275aac2f.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 6/7] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:18Z","receivedAt":"2022-05-04T15:26:18Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex df44902c909..9adde8cf4b9 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -592,6 +593,29 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -654,6 +678,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n \t\t     (int)buf.len, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex f3d037823c8..e049221609d 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,6 +93,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n@@ -102,7 +104,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"454824","messageId":"dfe821d10fe8b02ec0a2d210531e56fa16aefcd1.1651677919.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v3 4/7] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-04T15:25:16Z","receivedAt":"2022-05-04T15:26:19Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` machinery: we write out a\n`.zip` of the empty try, augmented by a couple files that are added via\nthe `--add-file*` options. We are careful trying not to modify the\ncurrent repository in any way lest the very circumstances that required\n`scalar diagnose` to be run are changed by the `diagnose` run itself.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 141 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 167 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 00dcd4b50ef..a290e52e1d2 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -261,6 +262,44 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\tres = -1;\n+\t\telse if (recurse)\n+\t\t     add_directory_to_archiver(archiver_args, buf.buf, recurse);\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -501,6 +540,107 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\twrite_or_die(stdout_fd, buf.buf, buf.len);\n+\tstrvec_pushf(&archiver_args,\n+\t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n+\t\t     (int)buf.len, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tfprintf(stderr, \"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -802,6 +942,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 9d83fdf25e8..bbd07a44426 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"454959","messageId":"CABPp-BEpA52Zq=43j3D=io7h4ooSVGd-644iJkGU5+MmsJjDZw@mail.gmail.com","threadId":"57313","inReplyTo":"ce4b1b680c98d0f55d4d307b8c746a81d90ffa06.1651677919.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-05-07T02:06:50Z","receivedAt":"2022-05-07T02:07:09Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, May 4, 2022 at 8:25 AM Johannes Schindelin via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> By allowing the path to be enclosed in double-quotes, we can avoid\n> the limitation that paths cannot contain colons.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  Documentation/git-archive.txt | 13 +++++++++----\n>  archive.c                     | 34 +++++++++++++++++++++++++++++-----\n>  t/t5003-archive-zip.sh        |  8 ++++++++\n>  3 files changed, 46 insertions(+), 9 deletions(-)\n>\n> diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> index a0edc9167b2..1789ce4c232 100644\n> --- a/Documentation/git-archive.txt\n> +++ b/Documentation/git-archive.txt\n> @@ -67,10 +67,15 @@ OPTIONS\n>         by concatenating the value for `--prefix` (if any) and the\n>         basename of <file>.\n>  +\n> -The `<path>` cannot contain any colon, the file mode is limited to\n> -a regular file, and the option may be subject to platform-dependent\n> -command-line limits. For non-trivial cases, write an untracked file\n> -and use `--add-file` instead.\n> +The `<path>` argument can start and end with a literal double-quote\n> +character. In this case, the backslash is interpreted as escape\n> +character. The path must be quoted if it contains a colon, to avoid\n> +the colon from being misinterpreted as the separator between the\n> +path and the contents.\n\nThe path must also be quoted if it begins or ends with a double-quote, right?\n\nAlso, would people want to be able to pass a pathname from the output\nof e.g. `git ls-files -o`, which may quote additional characters?\n\n> ++\n> +The file mode is limited to a regular file, and the option may be\n> +subject to platform-dependent command-line limits. For non-trivial\n> +cases, write an untracked file and use `--add-file` instead.\n>\n>  --worktree-attributes::\n>         Look for attributes in .gitattributes files in the working tree\n> diff --git a/archive.c b/archive.c\n> index d798624cd5f..3b751027143 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -533,13 +533,37 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n>                         die(_(\"Not a regular file: %s\"), path);\n>                 info->content = NULL; /* read the file later */\n>         } else {\n> -               const char *colon = strchr(arg, ':');\n>                 char *p;\n>\n> -               if (!colon)\n> -                       die(_(\"missing colon: '%s'\"), arg);\n> +               if (*arg != '\"') {\n> +                       const char *colon = strchr(arg, ':');\n> +\n> +                       if (!colon)\n> +                               die(_(\"missing colon: '%s'\"), arg);\n> +                       p = xstrndup(arg, colon - arg);\n> +                       arg = colon + 1;\n> +               } else {\n> +                       struct strbuf buf = STRBUF_INIT;\n> +                       const char *orig = arg;\n> +\n> +                       for (;;) {\n> +                               if (!*(++arg))\n> +                                       die(_(\"unclosed quote: '%s'\"), orig);\n> +                               if (*arg == '\"')\n> +                                       break;\n> +                               if (*arg == '\\\\' && *(++arg) == '\\0')\n> +                                       die(_(\"trailing backslash: '%s\"), orig);\n> +                               else\n> +                                       strbuf_addch(&buf, *arg);\n> +                       }\n> +\n> +                       if (*(++arg) != ':')\n> +                               die(_(\"missing colon: '%s'\"), orig);\n> +\n> +                       p = strbuf_detach(&buf, NULL);\n> +                       arg++;\n> +               }\n\nShould we use unquote_c_style() here instead of rolling another parser\nto do unquoting?  That would have the added benefit of allowing people\nto use filenames from the output of various git commands that do\nspecial quoting -- such as octal sequences for non-ascii characters.\n\n>\n> -               p = xstrndup(arg, colon - arg);\n>                 if (!args->prefix)\n>                         path = p;\n>                 else {\n> @@ -548,7 +572,7 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n>                 }\n>                 memset(&info->stat, 0, sizeof(info->stat));\n>                 info->stat.st_mode = S_IFREG | 0644;\n> -               info->content = xstrdup(colon + 1);\n> +               info->content = xstrdup(arg);\n>                 info->stat.st_size = strlen(info->content);\n>         }\n>         item = string_list_append_nodup(&args->extra_files, path);\n> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> index 8ff1257f1a0..5b8bbfc2692 100755\n> --- a/t/t5003-archive-zip.sh\n> +++ b/t/t5003-archive-zip.sh\n> @@ -207,13 +207,21 @@ check_zip with_untracked\n>  check_added with_untracked untracked untracked\n>\n>  test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n> +       if test_have_prereq FUNNYNAMES\n> +       then\n> +               QUOTED=quoted:colon\n> +       else\n> +               QUOTED=quoted\n> +       fi &&\n>         git archive --format=zip >with_file_with_content.zip \\\n> +               --add-file-with-content=\\\"$QUOTED\\\": \\\n>                 --add-file-with-content=hello:world $EMPTY_TREE &&\n>         test_when_finished \"rm -rf tmp-unpack\" &&\n>         mkdir tmp-unpack && (\n>                 cd tmp-unpack &&\n>                 \"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>                 test_path_is_file hello &&\n> +               test_path_is_file $QUOTED &&\n>                 test world = $(cat hello)\n>         )\n>  '\n> --\n> gitgitgadget\n"},{"id":"454961","messageId":"CABPp-BHbVFkqbw=zoBmJPNkUv3e_+i72A=9sCEU7gWHGfDsEsg@mail.gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-05-07T02:23:22Z","receivedAt":"2022-05-07T02:23:39Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, May 4, 2022 at 8:25 AM Johannes Schindelin via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Over the course of the years, we developed a sub-command that gathers\n> diagnostic data into a .zip file that can then be attached to bug reports.\n> This sub-command turned out to be very useful in helping Scalar developers\n> identify and fix issues.\n>\n> Changes since v2:\n>\n>  * Clarified in the commit message what the biggest benefit of\n>    --add-file-with-content is.\n>  * The <path> part of the -add-file-with-content argument can now contain\n>    colons. To do this, the path needs to start and end in double-quote\n>    characters (which are stripped), and the backslash serves as escape\n>    character in that case (to allow the path to contain both colons and\n>    double-quotes).\n\nYou addressed all my previous feedback from an earlier round.  The\nonly thing I noticed in this round is I wonder if we should use\nunquote_c_style() for this, as commented on the patch in question.\n\n>  * Fixed incorrect grammar.\n>  * Instead of strcmp(<what-we-don't-want>), we now say\n>    !strcmp(<what-we-want>).\n>  * The help text for --add-file-with-content was improved a tiny bit.\n>  * Adjusted the commit message that still talked about spawning plenty of\n>    processes and about a throw-away repository for the sake of generating a\n>    .zip file.\n>  * Simplified the code that shows the diagnostics and adds them to the .zip\n>    file.\n>  * The final message that reports that the archive is complete is now\n>    printed to stderr instead of stdout.\n>\n> Changes since v1:\n>\n>  * Instead of creating a throw-away repository, staging the contents of the\n>    .zip file and then using git write-tree and git archive to write the .zip\n>    file, the patch series now introduces a new option to git archive and\n>    uses write_archive() directly (avoiding any separate process).\n>  * Since the command avoids separate processes, it is now blazing fast on\n>    Windows, and I dropped the spinner() function because it's no longer\n>    needed.\n>  * While reworking the test case, I noticed that scalar [...] <enlistment>\n>    failed to verify that the specified directory exists, and would happily\n>    \"traverse to its parent directory\" on its quest to find a Scalar\n>    enlistment. That is of course incorrect, and has been fixed as a \"while\n>    at it\" sort of preparatory commit.\n>  * I had forgotten to sign off on all the commits, which has been fixed.\n>  * Instead of some \"home-grown\" readdir()-based function, the code now uses\n>    for_each_file_in_pack_dir() to look through the pack directories.\n>  * If any alternates are configured, their pack directories are now included\n>    in the output.\n>  * The commit message that might be interpreted to promise information about\n>    large loose files has been corrected to no longer promise that.\n>  * The test cases have been adjusted to test a little bit more (e.g.\n>    verifying that specific paths are mentioned in the output, instead of\n>    merely verifying that the output is non-empty).\n>\n> Johannes Schindelin (5):\n>   archive: optionally add \"virtual\" files\n>   archive --add-file-with-contents: allow paths containing colons\n>   scalar: validate the optional enlistment argument\n>   Implement `scalar diagnose`\n>   scalar diagnose: include disk space information\n>\n> Matthew John Cheetham (2):\n>   scalar: teach `diagnose` to gather packfile info\n>   scalar: teach `diagnose` to gather loose objects information\n>\n>  Documentation/git-archive.txt    |  16 ++\n>  archive.c                        |  75 +++++++-\n>  contrib/scalar/scalar.c          | 289 ++++++++++++++++++++++++++++++-\n>  contrib/scalar/scalar.txt        |  12 ++\n>  contrib/scalar/t/t9099-scalar.sh |  27 +++\n>  t/t5003-archive-zip.sh           |  20 +++\n>  6 files changed, 429 insertions(+), 10 deletions(-)\n>\n>\n> base-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v3\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v3\n> Pull-Request: https://github.com/gitgitgadget/git/pull/1128\n>\n> Range-diff vs v2:\n>\n>  1:  49ff3c1f2b3 ! 1:  45662cf582a archive: optionally add \"virtual\" files\n>      @@ Commit message\n>           archive` now supports use cases where relatively trivial files need to\n>           be added that do not exist on disk.\n>\n>      +    This will allow us to generate `.zip` files with generated content,\n>      +    without having to add said content to the object database and without\n>      +    having to write it out to disk.\n>      +\n>           Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n>        ## Documentation/git-archive.txt ##\n>      @@ Documentation/git-archive.txt: OPTIONS\n>       + basename of <file>.\n>       ++\n>       +The `<path>` cannot contain any colon, the file mode is limited to\n>      -+a regular file, and the option may be subject platform-dependent\n>      ++a regular file, and the option may be subject to platform-dependent\n>       +command-line limits. For non-trivial cases, write an untracked file\n>       +and use `--add-file` instead.\n>       +\n>      @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int\n>       - if (!S_ISREG(info->stat.st_mode))\n>       -         die(_(\"Not a regular file: %s\"), path);\n>       +\n>      -+ if (strcmp(opt->long_name, \"add-file-with-content\")) {\n>      ++ if (!strcmp(opt->long_name, \"add-file\")) {\n>       +         path = prefix_filename(args->prefix, arg);\n>       +         if (stat(path, &info->stat))\n>       +                 die(_(\"File not found: %s\"), path);\n>      @@ archive.c: static int parse_archive_args(int argc, const char **argv,\n>                   N_(\"add untracked file to archive\"), 0, add_file_cb,\n>                   (intptr_t)&base },\n>       +         { OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n>      -+           N_(\"file\"), N_(\"add untracked file to archive\"), 0,\n>      ++           N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n>       +           add_file_cb, (intptr_t)&base },\n>                 OPT_STRING('o', \"output\", &output, N_(\"file\"),\n>                         N_(\"write the archive to this file\")),\n>  -:  ----------- > 2:  ce4b1b680c9 archive --add-file-with-contents: allow paths containing colons\n>  2:  600da8d465e = 3:  5a3eeb55409 scalar: validate the optional enlistment argument\n>  3:  0d570137bb6 ! 4:  dfe821d10fe Implement `scalar diagnose`\n>      @@ Commit message\n>           we had the luxury of a comprehensive standard library that includes\n>           basic functionality such as writing a `.zip` file. In the C version, we\n>           lack such a commodity. Rather than introducing a dependency on, say,\n>      -    libzip, we slightly abuse Git's `archive` command: Instead of writing\n>      -    the `.zip` file directly, we stage the file contents in a Git index of a\n>      -    temporary, bare repository, only to let `git archive` have at it, and\n>      -    finally removing the temporary repository.\n>      -\n>      -    Also note: Due to the frequently-spawned `git hash-object` processes,\n>      -    this command is quite a bit slow on Windows. Should it turn out to be a\n>      -    big problem, the lack of a batch mode of the `hash-object` command could\n>      -    potentially be worked around via using `git fast-import` with a crafted\n>      -    `stdin`.\n>      +    libzip, we slightly abuse Git's `archive` machinery: we write out a\n>      +    `.zip` of the empty try, augmented by a couple files that are added via\n>      +    the `--add-file*` options. We are careful trying not to modify the\n>      +    current repository in any way lest the very circumstances that required\n>      +    `scalar diagnose` to be run are changed by the `diagnose` run itself.\n>\n>           Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n>      @@ contrib/scalar/scalar.c: cleanup:\n>       + time_t now = time(NULL);\n>       + struct tm tm;\n>       + struct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n>      -+ size_t off;\n>       + int res = 0;\n>       +\n>       + argc = parse_options(argc, argv, NULL, options,\n>      @@ contrib/scalar/scalar.c: cleanup:\n>       + strvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n>       +\n>       + strbuf_reset(&buf);\n>      -+ strbuf_addstr(&buf,\n>      -+               \"--add-file-with-content=diagnostics.log:\"\n>      -+               \"Collecting diagnostic info\\n\\n\");\n>      ++ strbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n>       + get_version_info(&buf, 1);\n>       +\n>       + strbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n>      -+ off = strchr(buf.buf, ':') + 1 - buf.buf;\n>      -+ write_or_die(stdout_fd, buf.buf + off, buf.len - off);\n>      -+ strvec_push(&archiver_args, buf.buf);\n>      ++ write_or_die(stdout_fd, buf.buf, buf.len);\n>      ++ strvec_pushf(&archiver_args,\n>      ++              \"--add-file-with-content=diagnostics.log:%.*s\",\n>      ++              (int)buf.len, buf.buf);\n>       +\n>       + if ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n>       +     (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n>      @@ contrib/scalar/scalar.c: cleanup:\n>       + }\n>       +\n>       + if (!res)\n>      -+         printf(\"\\n\"\n>      ++         fprintf(stderr, \"\\n\"\n>       +                \"Diagnostics complete.\\n\"\n>       +                \"All of the gathered info is captured in '%s'\\n\",\n>       +                zip_path.buf);\n>  4:  938e38b5a09 ! 5:  bb162abd383 scalar diagnose: include disk space information\n>      @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n>\n>         strbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n>       + get_disk_info(&buf);\n>      -  off = strchr(buf.buf, ':') + 1 - buf.buf;\n>      -  write_or_die(stdout_fd, buf.buf + off, buf.len - off);\n>      -  strvec_push(&archiver_args, buf.buf);\n>      +  write_or_die(stdout_fd, buf.buf, buf.len);\n>      +  strvec_pushf(&archiver_args,\n>      +               \"--add-file-with-content=diagnostics.log:%.*s\",\n>\n>        ## contrib/scalar/t/t9099-scalar.sh ##\n>       @@ contrib/scalar/t/t9099-scalar.sh: SQ=\"'\"\n>  5:  bd9428919fa ! 6:  32aaad7cce1 scalar: teach `diagnose` to gather packfile info\n>      @@ contrib/scalar/scalar.c: cleanup:\n>        {\n>         struct option options[] = {\n>       @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n>      -  write_or_die(stdout_fd, buf.buf + off, buf.len - off);\n>      -  strvec_push(&archiver_args, buf.buf);\n>      +               \"--add-file-with-content=diagnostics.log:%.*s\",\n>      +               (int)buf.len, buf.buf);\n>\n>       + strbuf_reset(&buf);\n>       + strbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n>  6:  7a8875be425 = 7:  322932f0bb8 scalar: teach `diagnose` to gather loose objects information\n>\n> --\n> gitgitgadget\n"},{"id":"455054","messageId":"nycvar.QRO.7.76.6.2205092302550.346@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"CABPp-BEpA52Zq=43j3D=io7h4ooSVGd-644iJkGU5+MmsJjDZw@mail.gmail.com","subject":"Re: [PATCH v3 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-09T21:04:45Z","receivedAt":"2022-05-09T21:04:56Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Elijah,\n\nOn Fri, 6 May 2022, Elijah Newren wrote:\n\n> On Wed, May 4, 2022 at 8:25 AM Johannes Schindelin via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> >\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >\n> > By allowing the path to be enclosed in double-quotes, we can avoid\n> > the limitation that paths cannot contain colons.\n> >\n> > Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > ---\n> >  Documentation/git-archive.txt | 13 +++++++++----\n> >  archive.c                     | 34 +++++++++++++++++++++++++++++-----\n> >  t/t5003-archive-zip.sh        |  8 ++++++++\n> >  3 files changed, 46 insertions(+), 9 deletions(-)\n> >\n> > diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> > index a0edc9167b2..1789ce4c232 100644\n> > --- a/Documentation/git-archive.txt\n> > +++ b/Documentation/git-archive.txt\n> > @@ -67,10 +67,15 @@ OPTIONS\n> >         by concatenating the value for `--prefix` (if any) and the\n> >         basename of <file>.\n> >  +\n> > -The `<path>` cannot contain any colon, the file mode is limited to\n> > -a regular file, and the option may be subject to platform-dependent\n> > -command-line limits. For non-trivial cases, write an untracked file\n> > -and use `--add-file` instead.\n> > +The `<path>` argument can start and end with a literal double-quote\n> > +character. In this case, the backslash is interpreted as escape\n> > +character. The path must be quoted if it contains a colon, to avoid\n> > +the colon from being misinterpreted as the separator between the\n> > +path and the contents.\n>\n> The path must also be quoted if it begins or ends with a double-quote, right?\n\nTrue.\n\n> Also, would people want to be able to pass a pathname from the output\n> of e.g. `git ls-files -o`, which may quote additional characters?\n\nAlso true.\n\n> > ++\n> > +The file mode is limited to a regular file, and the option may be\n> > +subject to platform-dependent command-line limits. For non-trivial\n> > +cases, write an untracked file and use `--add-file` instead.\n> >\n> >  --worktree-attributes::\n> >         Look for attributes in .gitattributes files in the working tree\n> > diff --git a/archive.c b/archive.c\n> > index d798624cd5f..3b751027143 100644\n> > --- a/archive.c\n> > +++ b/archive.c\n> > @@ -533,13 +533,37 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n> >                         die(_(\"Not a regular file: %s\"), path);\n> >                 info->content = NULL; /* read the file later */\n> >         } else {\n> > -               const char *colon = strchr(arg, ':');\n> >                 char *p;\n> >\n> > -               if (!colon)\n> > -                       die(_(\"missing colon: '%s'\"), arg);\n> > +               if (*arg != '\"') {\n> > +                       const char *colon = strchr(arg, ':');\n> > +\n> > +                       if (!colon)\n> > +                               die(_(\"missing colon: '%s'\"), arg);\n> > +                       p = xstrndup(arg, colon - arg);\n> > +                       arg = colon + 1;\n> > +               } else {\n> > +                       struct strbuf buf = STRBUF_INIT;\n> > +                       const char *orig = arg;\n> > +\n> > +                       for (;;) {\n> > +                               if (!*(++arg))\n> > +                                       die(_(\"unclosed quote: '%s'\"), orig);\n> > +                               if (*arg == '\"')\n> > +                                       break;\n> > +                               if (*arg == '\\\\' && *(++arg) == '\\0')\n> > +                                       die(_(\"trailing backslash: '%s\"), orig);\n> > +                               else\n> > +                                       strbuf_addch(&buf, *arg);\n> > +                       }\n> > +\n> > +                       if (*(++arg) != ':')\n> > +                               die(_(\"missing colon: '%s'\"), orig);\n> > +\n> > +                       p = strbuf_detach(&buf, NULL);\n> > +                       arg++;\n> > +               }\n>\n> Should we use unquote_c_style() here instead of rolling another parser\n> to do unquoting?  That would have the added benefit of allowing people\n> to use filenames from the output of various git commands that do\n> special quoting -- such as octal sequences for non-ascii characters.\n\nYep, let's do that. I somehow missed that function while glimpsing at\n`quote.h`.\n\nThank you for your review!\nDscho\n\n> >\n> > -               p = xstrndup(arg, colon - arg);\n> >                 if (!args->prefix)\n> >                         path = p;\n> >                 else {\n> > @@ -548,7 +572,7 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n> >                 }\n> >                 memset(&info->stat, 0, sizeof(info->stat));\n> >                 info->stat.st_mode = S_IFREG | 0644;\n> > -               info->content = xstrdup(colon + 1);\n> > +               info->content = xstrdup(arg);\n> >                 info->stat.st_size = strlen(info->content);\n> >         }\n> >         item = string_list_append_nodup(&args->extra_files, path);\n> > diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> > index 8ff1257f1a0..5b8bbfc2692 100755\n> > --- a/t/t5003-archive-zip.sh\n> > +++ b/t/t5003-archive-zip.sh\n> > @@ -207,13 +207,21 @@ check_zip with_untracked\n> >  check_added with_untracked untracked untracked\n> >\n> >  test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n> > +       if test_have_prereq FUNNYNAMES\n> > +       then\n> > +               QUOTED=quoted:colon\n> > +       else\n> > +               QUOTED=quoted\n> > +       fi &&\n> >         git archive --format=zip >with_file_with_content.zip \\\n> > +               --add-file-with-content=\\\"$QUOTED\\\": \\\n> >                 --add-file-with-content=hello:world $EMPTY_TREE &&\n> >         test_when_finished \"rm -rf tmp-unpack\" &&\n> >         mkdir tmp-unpack && (\n> >                 cd tmp-unpack &&\n> >                 \"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> >                 test_path_is_file hello &&\n> > +               test_path_is_file $QUOTED &&\n> >                 test world = $(cat hello)\n> >         )\n> >  '\n> > --\n> > gitgitgadget\n>\n>\n"},{"id":"455104","messageId":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v3.git.1651677919.gitgitgadget@gmail.com","subject":"[PATCH v4 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:26:57Z","receivedAt":"2022-05-10T19:27:20Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nChanges since v3:\n\n * We're now using unquote_c_style() instead of rolling our own unquoter.\n * Fixed the added regression test.\n * As pointed out by Scalar's Functional Tests, the\n   add_directory_to_archiver() function should not fail when scalar diagnose\n   encounters FSMonitor's Unix socket, but only warn instead.\n * Related: add_directory_to_archiver() needs to propagate errors from\n   processing subdirectories so that the top-level call returns an error,\n   too.\n\nChanges since v2:\n\n * Clarified in the commit message what the biggest benefit of\n   --add-file-with-content is.\n * The <path> part of the -add-file-with-content argument can now contain\n   colons. To do this, the path needs to start and end in double-quote\n   characters (which are stripped), and the backslash serves as escape\n   character in that case (to allow the path to contain both colons and\n   double-quotes).\n * Fixed incorrect grammar.\n * Instead of strcmp(<what-we-don't-want>), we now say\n   !strcmp(<what-we-want>).\n * The help text for --add-file-with-content was improved a tiny bit.\n * Adjusted the commit message that still talked about spawning plenty of\n   processes and about a throw-away repository for the sake of generating a\n   .zip file.\n * Simplified the code that shows the diagnostics and adds them to the .zip\n   file.\n * The final message that reports that the archive is complete is now\n   printed to stderr instead of stdout.\n\nChanges since v1:\n\n * Instead of creating a throw-away repository, staging the contents of the\n   .zip file and then using git write-tree and git archive to write the .zip\n   file, the patch series now introduces a new option to git archive and\n   uses write_archive() directly (avoiding any separate process).\n * Since the command avoids separate processes, it is now blazing fast on\n   Windows, and I dropped the spinner() function because it's no longer\n   needed.\n * While reworking the test case, I noticed that scalar [...] <enlistment>\n   failed to verify that the specified directory exists, and would happily\n   \"traverse to its parent directory\" on its quest to find a Scalar\n   enlistment. That is of course incorrect, and has been fixed as a \"while\n   at it\" sort of preparatory commit.\n * I had forgotten to sign off on all the commits, which has been fixed.\n * Instead of some \"home-grown\" readdir()-based function, the code now uses\n   for_each_file_in_pack_dir() to look through the pack directories.\n * If any alternates are configured, their pack directories are now included\n   in the output.\n * The commit message that might be interpreted to promise information about\n   large loose files has been corrected to no longer promise that.\n * The test cases have been adjusted to test a little bit more (e.g.\n   verifying that specific paths are mentioned in the output, instead of\n   merely verifying that the output is non-empty).\n\nJohannes Schindelin (5):\n  archive: optionally add \"virtual\" files\n  archive --add-file-with-contents: allow paths containing colons\n  scalar: validate the optional enlistment argument\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n Documentation/git-archive.txt    |  17 ++\n archive.c                        |  61 ++++++-\n contrib/scalar/scalar.c          | 292 ++++++++++++++++++++++++++++++-\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  27 +++\n t/t5003-archive-zip.sh           |  20 +++\n 6 files changed, 419 insertions(+), 10 deletions(-)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n\nRange-diff vs v3:\n\n 1:  45662cf582a = 1:  45662cf582a archive: optionally add \"virtual\" files\n 2:  ce4b1b680c9 ! 2:  fdba4ed6f4d archive --add-file-with-contents: allow paths containing colons\n     @@ Documentation/git-archive.txt: OPTIONS\n      -command-line limits. For non-trivial cases, write an untracked file\n      -and use `--add-file` instead.\n      +The `<path>` argument can start and end with a literal double-quote\n     -+character. In this case, the backslash is interpreted as escape\n     -+character. The path must be quoted if it contains a colon, to avoid\n     -+the colon from being misinterpreted as the separator between the\n     -+path and the contents.\n     ++character; The contained file name is interpreted as a C-style string,\n     ++i.e. the backslash is interpreted as escape character. The path must\n     ++be quoted if it contains a colon, to avoid the colon from being\n     ++misinterpreted as the separator between the path and the contents, or\n     ++if the path begins or ends with a double-quote character.\n      ++\n      +The file mode is limited to a regular file, and the option may be\n      +subject to platform-dependent command-line limits. For non-trivial\n     @@ Documentation/git-archive.txt: OPTIONS\n       \tLook for attributes in .gitattributes files in the working tree\n      \n       ## archive.c ##\n     +@@\n     + #include \"parse-options.h\"\n     + #include \"unpack-trees.h\"\n     + #include \"dir.h\"\n     ++#include \"quote.h\"\n     + \n     + static char const * const archive_usage[] = {\n     + \tN_(\"git archive [<options>] <tree-ish> [<path>...]\"),\n      @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int unset)\n       \t\t\tdie(_(\"Not a regular file: %s\"), path);\n       \t\tinfo->content = NULL; /* read the file later */\n       \t} else {\n      -\t\tconst char *colon = strchr(arg, ':');\n     - \t\tchar *p;\n     +-\t\tchar *p;\n     ++\t\tstruct strbuf buf = STRBUF_INIT;\n     ++\t\tconst char *p = arg;\n     ++\n     ++\t\tif (*p != '\"')\n     ++\t\t\tp = strchr(p, ':');\n     ++\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n     ++\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n       \n      -\t\tif (!colon)\n     --\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n     -+\t\tif (*arg != '\"') {\n     -+\t\t\tconst char *colon = strchr(arg, ':');\n     -+\n     -+\t\t\tif (!colon)\n     -+\t\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n     -+\t\t\tp = xstrndup(arg, colon - arg);\n     -+\t\t\targ = colon + 1;\n     -+\t\t} else {\n     -+\t\t\tstruct strbuf buf = STRBUF_INIT;\n     -+\t\t\tconst char *orig = arg;\n     -+\n     -+\t\t\tfor (;;) {\n     -+\t\t\t\tif (!*(++arg))\n     -+\t\t\t\t\tdie(_(\"unclosed quote: '%s'\"), orig);\n     -+\t\t\t\tif (*arg == '\"')\n     -+\t\t\t\t\tbreak;\n     -+\t\t\t\tif (*arg == '\\\\' && *(++arg) == '\\0')\n     -+\t\t\t\t\tdie(_(\"trailing backslash: '%s\"), orig);\n     -+\t\t\t\telse\n     -+\t\t\t\t\tstrbuf_addch(&buf, *arg);\n     -+\t\t\t}\n     -+\n     -+\t\t\tif (*(++arg) != ':')\n     -+\t\t\t\tdie(_(\"missing colon: '%s'\"), orig);\n     -+\n     -+\t\t\tp = strbuf_detach(&buf, NULL);\n     -+\t\t\targ++;\n     -+\t\t}\n     ++\t\tif (!p || *p != ':')\n     + \t\t\tdie(_(\"missing colon: '%s'\"), arg);\n       \n      -\t\tp = xstrndup(arg, colon - arg);\n     - \t\tif (!args->prefix)\n     - \t\t\tpath = p;\n     - \t\telse {\n     -@@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int unset)\n     +-\t\tif (!args->prefix)\n     +-\t\t\tpath = p;\n     +-\t\telse {\n     +-\t\t\tpath = prefix_filename(args->prefix, p);\n     +-\t\t\tfree(p);\n     ++\t\tif (p == arg)\n     ++\t\t\tdie(_(\"empty file name: '%s'\"), arg);\n     ++\n     ++\t\tpath = buf.len ?\n     ++\t\t\tstrbuf_detach(&buf, NULL) : xstrndup(arg, p - arg);\n     ++\n     ++\t\tif (args->prefix) {\n     ++\t\t\tchar *save = path;\n     ++\t\t\tpath = prefix_filename(args->prefix, path);\n     ++\t\t\tfree(save);\n       \t\t}\n       \t\tmemset(&info->stat, 0, sizeof(info->stat));\n       \t\tinfo->stat.st_mode = S_IFREG | 0644;\n      -\t\tinfo->content = xstrdup(colon + 1);\n     -+\t\tinfo->content = xstrdup(arg);\n     ++\t\tinfo->content = xstrdup(p + 1);\n       \t\tinfo->stat.st_size = strlen(info->content);\n       \t}\n       \titem = string_list_append_nodup(&args->extra_files, path);\n 3:  5a3eeb55409 = 3:  da9f52a8240 scalar: validate the optional enlistment argument\n 4:  dfe821d10fe ! 4:  87bdc22322b Implement `scalar diagnose`\n     @@ contrib/scalar/scalar.c: static int unregister_dir(void)\n      +\t\tif (e->d_type == DT_REG)\n      +\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n      +\t\telse if (e->d_type != DT_DIR)\n     ++\t\t\twarning(_(\"skipping '%s', which is neither file nor \"\n     ++\t\t\t\t  \"directory\"), buf.buf);\n     ++\t\telse if (recurse &&\n     ++\t\t\t add_directory_to_archiver(archiver_args,\n     ++\t\t\t\t\t\t   buf.buf, recurse) < 0)\n      +\t\t\tres = -1;\n     -+\t\telse if (recurse)\n     -+\t\t     add_directory_to_archiver(archiver_args, buf.buf, recurse);\n      +\t}\n      +\n      +\tclosedir(dir);\n     @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success '`scalar [...] <dir>` erro\n      +SQ=\"'\"\n      +test_expect_success UNZIP 'scalar diagnose' '\n      +\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n     -+\tscalar diagnose cloned >out &&\n     -+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n     ++\tscalar diagnose cloned >out 2>err &&\n     ++\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n      +\tzip_path=$(cat zip_path) &&\n      +\ttest -n \"$zip_path\" &&\n      +\tunzip -v \"$zip_path\" &&\n 5:  bb162abd383 ! 5:  3f63b197d42 scalar diagnose: include disk space information\n     @@ contrib/scalar/t/t9099-scalar.sh\n      @@ contrib/scalar/t/t9099-scalar.sh: SQ=\"'\"\n       test_expect_success UNZIP 'scalar diagnose' '\n       \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n     - \tscalar diagnose cloned >out &&\n     + \tscalar diagnose cloned >out 2>err &&\n      +\tgrep \"Available space\" out &&\n     - \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n       \tzip_path=$(cat zip_path) &&\n       \ttest -n \"$zip_path\" &&\n 6:  32aaad7cce1 ! 6:  fc1319338fc scalar: teach `diagnose` to gather packfile info\n     @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success '`scalar [...] <dir>` erro\n       \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n      +\tgit repack &&\n      +\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n     - \tscalar diagnose cloned >out &&\n     + \tscalar diagnose cloned >out 2>err &&\n       \tgrep \"Available space\" out &&\n     - \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n      @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n       \tfolder=${zip_path%.zip} &&\n       \ttest_path_is_missing \"$folder\" &&\n 7:  322932f0bb8 ! 7:  e8f5b42f7b7 scalar: teach `diagnose` to gather loose objects information\n     @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n       \tgit repack &&\n       \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n      +\ttest_commit -C cloned/src loose &&\n     - \tscalar diagnose cloned >out &&\n     + \tscalar diagnose cloned >out 2>err &&\n       \tgrep \"Available space\" out &&\n     - \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <out >zip_path &&\n     + \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n      @@ contrib/scalar/t/t9099-scalar.sh: test_expect_success UNZIP 'scalar diagnose' '\n       \tunzip -p \"$zip_path\" diagnostics.log >out &&\n       \ttest_file_not_empty out &&\n\n-- \ngitgitgadget\n"},{"id":"455105","messageId":"fdba4ed6f4d5ed4f78404e0a0c5b338c22678533.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:26:59Z","receivedAt":"2022-05-10T19:27:21Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nBy allowing the path to be enclosed in double-quotes, we can avoid\nthe limitation that paths cannot contain colons.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 14 ++++++++++----\n archive.c                     | 30 ++++++++++++++++++++----------\n t/t5003-archive-zip.sh        |  8 ++++++++\n 3 files changed, 38 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex a0edc9167b2..21eab5690ad 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -67,10 +67,16 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n +\n-The `<path>` cannot contain any colon, the file mode is limited to\n-a regular file, and the option may be subject to platform-dependent\n-command-line limits. For non-trivial cases, write an untracked file\n-and use `--add-file` instead.\n+The `<path>` argument can start and end with a literal double-quote\n+character; The contained file name is interpreted as a C-style string,\n+i.e. the backslash is interpreted as escape character. The path must\n+be quoted if it contains a colon, to avoid the colon from being\n+misinterpreted as the separator between the path and the contents, or\n+if the path begins or ends with a double-quote character.\n++\n+The file mode is limited to a regular file, and the option may be\n+subject to platform-dependent command-line limits. For non-trivial\n+cases, write an untracked file and use `--add-file` instead.\n \n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\ndiff --git a/archive.c b/archive.c\nindex d798624cd5f..477eba60ac3 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -9,6 +9,7 @@\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n #include \"dir.h\"\n+#include \"quote.h\"\n \n static char const * const archive_usage[] = {\n \tN_(\"git archive [<options>] <tree-ish> [<path>...]\"),\n@@ -533,22 +534,31 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n \t} else {\n-\t\tconst char *colon = strchr(arg, ':');\n-\t\tchar *p;\n+\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\tconst char *p = arg;\n+\n+\t\tif (*p != '\"')\n+\t\t\tp = strchr(p, ':');\n+\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n+\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n \n-\t\tif (!colon)\n+\t\tif (!p || *p != ':')\n \t\t\tdie(_(\"missing colon: '%s'\"), arg);\n \n-\t\tp = xstrndup(arg, colon - arg);\n-\t\tif (!args->prefix)\n-\t\t\tpath = p;\n-\t\telse {\n-\t\t\tpath = prefix_filename(args->prefix, p);\n-\t\t\tfree(p);\n+\t\tif (p == arg)\n+\t\t\tdie(_(\"empty file name: '%s'\"), arg);\n+\n+\t\tpath = buf.len ?\n+\t\t\tstrbuf_detach(&buf, NULL) : xstrndup(arg, p - arg);\n+\n+\t\tif (args->prefix) {\n+\t\t\tchar *save = path;\n+\t\t\tpath = prefix_filename(args->prefix, path);\n+\t\t\tfree(save);\n \t\t}\n \t\tmemset(&info->stat, 0, sizeof(info->stat));\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n-\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->content = xstrdup(p + 1);\n \t\tinfo->stat.st_size = strlen(info->content);\n \t}\n \titem = string_list_append_nodup(&args->extra_files, path);\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 8ff1257f1a0..5b8bbfc2692 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -207,13 +207,21 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n+\tif test_have_prereq FUNNYNAMES\n+\tthen\n+\t\tQUOTED=quoted:colon\n+\telse\n+\t\tQUOTED=quoted\n+\tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n+\t\ttest_path_is_file $QUOTED &&\n \t\ttest world = $(cat hello)\n \t)\n '\n-- \ngitgitgadget\n\n"},{"id":"455106","messageId":"da9f52a82406ffc909e9c5f2b6b5e77818d972c0.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:27:00Z","receivedAt":"2022-05-10T19:27:23Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..00dcd4b50ef 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..9d83fdf25e8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n \ttest_path_is_missing cloned\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455107","messageId":"87bdc22322b0f58bf153b963207cffe4f41c9ae9.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 4/7] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:27:01Z","receivedAt":"2022-05-10T19:27:27Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` machinery: we write out a\n`.zip` of the empty try, augmented by a couple files that are added via\nthe `--add-file*` options. We are careful trying not to modify the\ncurrent repository in any way lest the very circumstances that required\n`scalar diagnose` to be run are changed by the `diagnose` run itself.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 144 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 170 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 00dcd4b50ef..367a2c50e25 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -261,6 +262,47 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\twarning(_(\"skipping '%s', which is neither file nor \"\n+\t\t\t\t  \"directory\"), buf.buf);\n+\t\telse if (recurse &&\n+\t\t\t add_directory_to_archiver(archiver_args,\n+\t\t\t\t\t\t   buf.buf, recurse) < 0)\n+\t\t\tres = -1;\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -501,6 +543,107 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\twrite_or_die(stdout_fd, buf.buf, buf.len);\n+\tstrvec_pushf(&archiver_args,\n+\t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n+\t\t     (int)buf.len, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tfprintf(stderr, \"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -802,6 +945,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 9d83fdf25e8..6802d317258 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out 2>err &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455108","messageId":"fc1319338fc3e44584e92cb92bff2f6c99c59859.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 6/7] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:27:03Z","receivedAt":"2022-05-10T19:27:30Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 34cbec59b45..e8e0a5ec473 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -595,6 +596,29 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -657,6 +681,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n \t\t     (int)buf.len, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 934b2485d91..3dd5650cceb 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,6 +93,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -102,7 +104,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"455109","messageId":"3f63b197d420c35f7606d823a981b067964876a6.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 5/7] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:27:02Z","receivedAt":"2022-05-10T19:27:32Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 367a2c50e25..34cbec59b45 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -303,6 +303,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -599,6 +651,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \twrite_or_die(stdout_fd, buf.buf, buf.len);\n \tstrvec_pushf(&archiver_args,\n \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 6802d317258..934b2485d91 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -94,6 +94,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out 2>err &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \ngitgitgadget\n\n"},{"id":"455110","messageId":"45662cf582ab7c8b1c32f55c9a34f4d73a28b71d.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:26:58Z","receivedAt":"2022-05-10T19:27:34Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-file-with-content=<path>:<content>` option, `git\narchive` now supports use cases where relatively trivial files need to\nbe added that do not exist on disk.\n\nThis will allow us to generate `.zip` files with generated content,\nwithout having to add said content to the object database and without\nhaving to write it out to disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 11 ++++++++\n archive.c                     | 51 +++++++++++++++++++++++++++++------\n t/t5003-archive-zip.sh        | 12 +++++++++\n 3 files changed, 66 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a7834..a0edc9167b2 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -61,6 +61,17 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n \n+--add-file-with-content=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value for `--prefix` (if any) and the\n+\tbasename of <file>.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject to platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex a3bbb091256..d798624cd5f 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n \t\tstrbuf_addstr(&path_in_archive, basename(path));\n \n \t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\tif (info->content)\n+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\telse if (strbuf_read_file(&content, path,\n+\t\t\t\t\t  info->stat.st_size) < 0)\n \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n@@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (!strcmp(opt->long_name, \"add-file\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +586,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n+\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 1e6d18b140e..8ff1257f1a0 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \ngitgitgadget\n\n"},{"id":"455111","messageId":"e8f5b42f7b7a89ee5dabdd58da7449f31ebc2a4c.1652210824.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v4 7/7] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-10T19:27:04Z","receivedAt":"2022-05-10T19:27:42Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex e8e0a5ec473..03da7452d83 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -619,6 +619,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -687,6 +741,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-file-with-content=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 3dd5650cceb..72023a1ca1d 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -95,6 +95,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -106,7 +107,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \ngitgitgadget\n"},{"id":"455120","messageId":"xmqqtu9x6ovh.fsf@gitster.g","threadId":"57313","inReplyTo":"45662cf582ab7c8b1c32f55c9a34f4d73a28b71d.1652210824.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-10T21:48:02Z","receivedAt":"2022-05-10T21:48:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n> @@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n>  \tif (!arg)\n>  \t\treturn -1;\n>  \n> -\tpath = prefix_filename(args->prefix, arg);\n> -\titem = string_list_append_nodup(&args->extra_files, path);\n> -\titem->util = info = xmalloc(sizeof(*info));\n> +\tinfo = xmalloc(sizeof(*info));\n>  \tinfo->base = xstrdup_or_null(base);\n> -\tif (stat(path, &info->stat))\n> -\t\tdie(_(\"File not found: %s\"), path);\n> -\tif (!S_ISREG(info->stat.st_mode))\n> -\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\n> +\tif (!strcmp(opt->long_name, \"add-file\")) {\n> +\t\tpath = prefix_filename(args->prefix, arg);\n> +\t\tif (stat(path, &info->stat))\n> +\t\t\tdie(_(\"File not found: %s\"), path);\n> +\t\tif (!S_ISREG(info->stat.st_mode))\n> +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\t\tinfo->content = NULL; /* read the file later */\n> +\t} else {\n\nThis pretends that this new one will stay to be the only other\noption that uses the same callback in the future.  To be more\ndefensive, it should do\n\n\t} else if (!strcmp(opt->long_name, \"...\")) {\n\nand end the if/else if/else cascade with\n\n\t} else {\n\t\tBUG(\"add_file_cb called for unknown option\");\n\t}\n\n> +\t\tconst char *colon = strchr(arg, ':');\n> +\t\tchar *p;\n> +\n> +\t\tif (!colon)\n> +\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n> +\n> +\t\tp = xstrndup(arg, colon - arg);\n> +\t\tif (!args->prefix)\n> +\t\t\tpath = p;\n> +\t\telse {\n> +\t\t\tpath = prefix_filename(args->prefix, p);\n> +\t\t\tfree(p);\n> +\t\t}\n> +\t\tmemset(&info->stat, 0, sizeof(info->stat));\n> +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n\nI can sympathize with the desire to omit the mode bits because it\nmay not be useful for the immediate purpose of \"scalar diagnose\"\nwhere the extracting end won't care what the file's permission bits\nare, but by letting this \"mode is hardcoded\" thing squat here would\nlater make it more work when other people want to add an option that\ntruely lets the caller add a \"vitual\" file, in response to end-user\ncomplaints that they cannot use the existing one to add an\nexectuable file, for example.  I do not care too much about the\npathname limitation that does not allow a colon in it, simply\nbecause it is unusual enough, but I am not sure about hardcoded\npermission bits.\n\nIf we did \"--add-virtual-file=<path>:0644:<contents>\" instead from\nday one, it certainly adds a few more lines of logic to this patch,\nand the calling \"scalar diagnose\" may have to pass a few more bytes,\nbut I suspect that such a change would help the project in the\nlonger run.\n\nThanks.\n"},{"id":"455121","messageId":"xmqqmtfp6ohc.fsf@gitster.g","threadId":"57313","inReplyTo":"fdba4ed6f4d5ed4f78404e0a0c5b338c22678533.1652210824.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-10T21:56:31Z","receivedAt":"2022-05-10T21:56:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> By allowing the path to be enclosed in double-quotes, we can avoid\n> the limitation that paths cannot contain colons.\n> ...\n> +\t\tstruct strbuf buf = STRBUF_INIT;\n> +\t\tconst char *p = arg;\n> +\n> +\t\tif (*p != '\"')\n> +\t\t\tp = strchr(p, ':');\n> +\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n> +\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n\nEven though I do not think people necessarily would want to use\ncolons in their pathnames (it has problems interoperating with other\nsystems), lifting the limitation is a good thing to do.  I totally\nforgot that we designed unquote_c_style() to self terminate and\nreturn the end pointer to the caller so the caller does not have to\nworry, which is very nice. \n\nEven if this step weren't here in the series, I would have thought\nthe mode bits issue was more serious than \"no colons in path\"\nlimitation, but given that we address this unusual corner case\nlimitation, I would think we should address the hardcoded mode bits\nat the same time.\n\n> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> index 8ff1257f1a0..5b8bbfc2692 100755\n> --- a/t/t5003-archive-zip.sh\n> +++ b/t/t5003-archive-zip.sh\n> @@ -207,13 +207,21 @@ check_zip with_untracked\n>  check_added with_untracked untracked untracked\n>  \n>  test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n> +\tif test_have_prereq FUNNYNAMES\n> +\tthen\n> +\t\tQUOTED=quoted:colon\n> +\telse\n> +\t\tQUOTED=quoted\n> +\tfi &&\n\n;-)\n\n>  \tgit archive --format=zip >with_file_with_content.zip \\\n> +\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n>  \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n>  \tmkdir tmp-unpack && (\n>  \t\tcd tmp-unpack &&\n>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>  \t\ttest_path_is_file hello &&\n> +\t\ttest_path_is_file $QUOTED &&\n\nLooks OK, even though it probably is a good idea to have dq around\n$QUOTED, so that future developers can easily insert SP into its\nvalue to use a bit more common but still a bit more problematic\npathnames in the test.\n\nThanks.\n"},{"id":"455123","messageId":"03d701d864ba$46d15c10$d4741430$@nexbridge.com","threadId":"57313","inReplyTo":"xmqqtu9x6ovh.fsf@gitster.g","subject":"RE: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-10T22:06:56Z","receivedAt":"2022-05-10T22:07:15Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 10, 2022 5:48 PM, Junio C Hamano wrote:\n>\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n>writes:\n>\n>> @@ -514,14 +522,38 @@ static int add_file_cb(const struct option *opt, const\n>char *arg, int unset)\n>>  \tif (!arg)\n>>  \t\treturn -1;\n>>\n>> -\tpath = prefix_filename(args->prefix, arg);\n>> -\titem = string_list_append_nodup(&args->extra_files, path);\n>> -\titem->util = info = xmalloc(sizeof(*info));\n>> +\tinfo = xmalloc(sizeof(*info));\n>>  \tinfo->base = xstrdup_or_null(base);\n>> -\tif (stat(path, &info->stat))\n>> -\t\tdie(_(\"File not found: %s\"), path);\n>> -\tif (!S_ISREG(info->stat.st_mode))\n>> -\t\tdie(_(\"Not a regular file: %s\"), path);\n>> +\n>> +\tif (!strcmp(opt->long_name, \"add-file\")) {\n>> +\t\tpath = prefix_filename(args->prefix, arg);\n>> +\t\tif (stat(path, &info->stat))\n>> +\t\t\tdie(_(\"File not found: %s\"), path);\n>> +\t\tif (!S_ISREG(info->stat.st_mode))\n>> +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n>> +\t\tinfo->content = NULL; /* read the file later */\n>> +\t} else {\n>\n>This pretends that this new one will stay to be the only other option that uses the\n>same callback in the future.  To be more defensive, it should do\n>\n>\t} else if (!strcmp(opt->long_name, \"...\")) {\n>\n>and end the if/else if/else cascade with\n>\n>\t} else {\n>\t\tBUG(\"add_file_cb called for unknown option\");\n>\t}\n>\n>> +\t\tconst char *colon = strchr(arg, ':');\n>> +\t\tchar *p;\n>> +\n>> +\t\tif (!colon)\n>> +\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n>> +\n>> +\t\tp = xstrndup(arg, colon - arg);\n>> +\t\tif (!args->prefix)\n>> +\t\t\tpath = p;\n>> +\t\telse {\n>> +\t\t\tpath = prefix_filename(args->prefix, p);\n>> +\t\t\tfree(p);\n>> +\t\t}\n>> +\t\tmemset(&info->stat, 0, sizeof(info->stat));\n>> +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n>\n>I can sympathize with the desire to omit the mode bits because it may not be\n>useful for the immediate purpose of \"scalar diagnose\"\n>where the extracting end won't care what the file's permission bits are, but by\n>letting this \"mode is hardcoded\" thing squat here would later make it more work\n>when other people want to add an option that truely lets the caller add a \"vitual\"\n>file, in response to end-user complaints that they cannot use the existing one to\n>add an exectuable file, for example.  I do not care too much about the pathname\n>limitation that does not allow a colon in it, simply because it is unusual enough, but\n>I am not sure about hardcoded permission bits.\n>\n>If we did \"--add-virtual-file=<path>:0644:<contents>\" instead from day one, it\n>certainly adds a few more lines of logic to this patch, and the calling \"scalar\n>diagnose\" may have to pass a few more bytes, but I suspect that such a change\n>would help the project in the longer run.\n\nWould not core.filemode=false somewhat simulate this? The consumer-client would not care/do anything with the mode anyway. Or am I missing something?\n--Randall\n\n"},{"id":"455125","messageId":"03d801d864bc$85fe62a0$91fb27e0$@nexbridge.com","threadId":"57313","inReplyTo":"xmqqmtfp6ohc.fsf@gitster.g","subject":"RE: [PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-10T22:23:01Z","receivedAt":"2022-05-10T22:23:18Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 10, 2022 5:57 PM, Junio C Hamano wrote:\n>\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n>writes:\n>\n>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>>\n>> By allowing the path to be enclosed in double-quotes, we can avoid the\n>> limitation that paths cannot contain colons.\n>> ...\n>> +\t\tstruct strbuf buf = STRBUF_INIT;\n>> +\t\tconst char *p = arg;\n>> +\n>> +\t\tif (*p != '\"')\n>> +\t\t\tp = strchr(p, ':');\n>> +\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n>> +\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n>\n>Even though I do not think people necessarily would want to use colons in their\n>pathnames (it has problems interoperating with other systems), lifting the\n>limitation is a good thing to do.  I totally forgot that we designed\n>unquote_c_style() to self terminate and return the end pointer to the caller so the\n>caller does not have to worry, which is very nice.\n>\n>Even if this step weren't here in the series, I would have thought the mode bits\n>issue was more serious than \"no colons in path\"\n>limitation, but given that we address this unusual corner case limitation, I would\n>think we should address the hardcoded mode bits at the same time.\n>\n>> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh index\n>> 8ff1257f1a0..5b8bbfc2692 100755\n>> --- a/t/t5003-archive-zip.sh\n>> +++ b/t/t5003-archive-zip.sh\n>> @@ -207,13 +207,21 @@ check_zip with_untracked  check_added\n>> with_untracked untracked untracked\n>>\n>>  test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n>> +\tif test_have_prereq FUNNYNAMES\n>> +\tthen\n>> +\t\tQUOTED=quoted:colon\n>> +\telse\n>> +\t\tQUOTED=quoted\n>> +\tfi &&\n>\n>;-)\n>\n>>  \tgit archive --format=zip >with_file_with_content.zip \\\n>> +\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n>>  \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n>>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n>>  \tmkdir tmp-unpack && (\n>>  \t\tcd tmp-unpack &&\n>>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>>  \t\ttest_path_is_file hello &&\n>> +\t\ttest_path_is_file $QUOTED &&\n>\n>Looks OK, even though it probably is a good idea to have dq around $QUOTED, so\n>that future developers can easily insert SP into its value to use a bit more common\n>but still a bit more problematic pathnames in the test.\n\nA test case for .gitignore in this would be good too. People on our exotic platform do this stuff as a matter of course. As an example, a name of $Z3P4:12399334 being used as a named pipe (associated with the unique name of a process) actually has been seen in the wild recently. My solution was to wild card this and/or contain it in an ignored directory.\nRegards,\nRandall\n\n"},{"id":"455130","messageId":"xmqq8rr955zf.fsf@gitster.g","threadId":"57313","inReplyTo":"03d701d864ba$46d15c10$d4741430$@nexbridge.com","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-10T23:21:24Z","receivedAt":"2022-05-10T23:21:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"<rsbecker@nexbridge.com> writes:\n\n>>If we did \"--add-virtual-file=<path>:0644:<contents>\" instead from day one, it\n>>certainly adds a few more lines of logic to this patch, and the calling \"scalar\n>>diagnose\" may have to pass a few more bytes, but I suspect that such a change\n>>would help the project in the longer run.\n>\n> Would not core.filemode=false somewhat simulate this? The\n> consumer-client would not care/do anything with the mode\n> anyway. Or am I missing something?\n\nOr I must be missing something.  This is part of \"git archive\" where\nits output is a tarball (or a zipfile) in which each entry knows its\npermission bits (or at least, if it is executable).  Running \"tar xf\" \nor \"unzip\" on the receiving end of the output of this command should\nset the executable bit (and other permission bits) correctly I would\ncertainly hope, so it does matter, no?\n\nI did say \"scalar diagnose\" may not care.  But a patch to \"git\narchive\" will affect other people, and among them there would be\npeople who say \"gee, now I can add a handful of files from the\ncommand line with their contents, without actually having them in\nthrow-away untracked files, when running 'git archive'.  That's\nhandy!\", try it out and get disappointed by their inability to\ncreate executable files that way.  And obviously I care more about\n\"git archive\" than \"scalar diagnose\".  I very welcome to enhance the\nformer to support the need for the latter.  I do not see a good\nreason to stop at a half-feature added to the former, even that\nadded half is enough to satisfy the latter, when the other half is\nnot all that hard to add, and it is reasonably expected that users\nother than \"scalar diagnose\" would naturally want the other half,\ntoo.\n\n\n"},{"id":"455165","messageId":"3cf6e4f8-9151-6d68-21ca-b94d6a7557e6@web.de","threadId":"57313","inReplyTo":"xmqq8rr955zf.fsf@gitster.g","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-11T16:14:12Z","receivedAt":"2022-05-11T16:14:42Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 11.05.22 um 01:21 schrieb Junio C Hamano:\n> <rsbecker@nexbridge.com> writes:\n>\n>>> If we did \"--add-virtual-file=<path>:0644:<contents>\" instead from day one, it\n>>> certainly adds a few more lines of logic to this patch, and the calling \"scalar\n>>> diagnose\" may have to pass a few more bytes, but I suspect that such a change\n>>> would help the project in the longer run.\n\n> I did say \"scalar diagnose\" may not care.  But a patch to \"git\n> archive\" will affect other people, and among them there would be\n> people who say \"gee, now I can add a handful of files from the\n> command line with their contents, without actually having them in\n> throw-away untracked files, when running 'git archive'.  That's\n> handy!\", try it out and get disappointed by their inability to\n> create executable files that way.\n\nWhich might motivate them to contribute a patch to add that feature.\nGive them a chance! :)\n\n> And obviously I care more about\n> \"git archive\" than \"scalar diagnose\".  I very welcome to enhance the\n> former to support the need for the latter.  I do not see a good\n> reason to stop at a half-feature added to the former, even that\n> added half is enough to satisfy the latter, when the other half is\n> not all that hard to add, and it is reasonably expected that users\n> other than \"scalar diagnose\" would naturally want the other half,\n> too.\n\nFWIW, I'd already be satisfied by a convincing outline of a way towards\na complete solution to accept the partial feature, just to be sure we\ndon't paint ourselves into a corner.  But I'm bad at both strategy and\nsaying no, so that's that.\n\nRegarding file modes: We only effectively support the executable bit,\nso an additional option --add-virtual-executable-file=<path>:<contents>\nwould suffice.  It would also prevent the false impression that\narbitrary file modes can be used (\"I said 0123 and got 0644, bug!\").\nAnd it would not even be the longest Git option..\n\nRené\n"},{"id":"455173","messageId":"xmqqzgjnkgy0.fsf@gitster.g","threadId":"57313","inReplyTo":"3cf6e4f8-9151-6d68-21ca-b94d6a7557e6@web.de","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-11T19:27:51Z","receivedAt":"2022-05-11T19:28:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 11.05.22 um 01:21 schrieb Junio C Hamano:\n>> <rsbecker@nexbridge.com> writes:\n>>\n>>>> If we did \"--add-virtual-file=<path>:0644:<contents>\" instead from day one, it\n>>>> certainly adds a few more lines of logic to this patch, and the calling \"scalar\n>>>> diagnose\" may have to pass a few more bytes, but I suspect that such a change\n>>>> would help the project in the longer run.\n>\n>> I did say \"scalar diagnose\" may not care.  But a patch to \"git\n>> archive\" will affect other people, and among them there would be\n>> people who say \"gee, now I can add a handful of files from the\n>> command line with their contents, without actually having them in\n>> throw-away untracked files, when running 'git archive'.  That's\n>> handy!\", try it out and get disappointed by their inability to\n>> create executable files that way.\n>\n> Which might motivate them to contribute a patch to add that feature.\n> Give them a chance! :)\n\nYes, but there is no way to reuse the same option in a backward\ncompatible way to later add the mode information, and that is why we\nwant to be careful before a half-feature squats on an option.\n\n> FWIW, I'd already be satisfied by a convincing outline of a way towards\n> a complete solution to accept the partial feature, just to be sure we\n> don't paint ourselves into a corner.\n\nExactly.  As you say, an extra and separate option can be used.  I\ndo not know if that is a workaround because we didn't design the\nfirst option to take an additional option, or a welcome feature.\n\n> Regarding file modes: We only effectively support the executable bit,\n> so an additional option --add-virtual-executable-file=<path>:<contents>\n> would suffice.\n\nWhile I do not think we want to support more than one \"is it\nexecutable or not?\" bit, I am not so sure about what the current\ncode does, though, for these \"not from a tree, but added as extra\nfiles\" entries.\n\nIf you add an extra file from an on-disk untracked file, the\nadd_file_cb() callback picks up the full st.st_mode for the file,\nand write_archive_entries() in its loop over args->extra_files pass\nthe full info->stat.st_mode down to write_entry(), which is used by\narchive-tar.c::write_tar_entry() to obtain mode bits pretty much\nas-is.  For tracked paths, we probably are normalizing the blobs\nbetween 0644 and 0755 way before the values are passed as \"mode\"\nparameter to the write_entry() functions, but for these extra files,\nthere is no such massaging.\n\nSo, I am OK with --add-virtual-executable=<path>:<contents> (but the\npoint still stands that the way the code in the patch squats in the\ncodepath makes it necessary to first refator it before it can\nhappen) as a separate option.  We may want to massage the mode bit\nwe grab from these extra files, if we were to go that route, though.\n\nThanks.\n\n"},{"id":"455210","messageId":"47ed5a2f-f4aa-1ec1-27c9-9b0b70eb8bca@web.de","threadId":"57313","inReplyTo":"xmqqzgjnkgy0.fsf@gitster.g","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-12T16:16:50Z","receivedAt":"2022-05-12T16:17:20Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 11.05.22 um 21:27 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>> Regarding file modes: We only effectively support the executable bit,\n>> so an additional option --add-virtual-executable-file=<path>:<contents>\n>> would suffice.\n>\n> While I do not think we want to support more than one \"is it\n> executable or not?\" bit, I am not so sure about what the current\n> code does, though, for these \"not from a tree, but added as extra\n> files\" entries.\n>\n> If you add an extra file from an on-disk untracked file, the\n> add_file_cb() callback picks up the full st.st_mode for the file,\n> and write_archive_entries() in its loop over args->extra_files pass\n> the full info->stat.st_mode down to write_entry(), which is used by\n> archive-tar.c::write_tar_entry() to obtain mode bits pretty much\n> as-is.\n\nGood point.  write_tar_entry() actually normalizes the permission bits\nand applies tar.umask (0002 by default):\n\n\tif (S_ISDIR(mode) || S_ISGITLINK(mode)) {\n\t\t*header.typeflag = TYPEFLAG_DIR;\n\t\tmode = (mode | 0777) & ~tar_umask;\n\t} else if (S_ISLNK(mode)) {\n\t\t*header.typeflag = TYPEFLAG_LNK;\n\t\tmode |= 0777;\n\t} else if (S_ISREG(mode)) {\n\t\t*header.typeflag = TYPEFLAG_REG;\n\t\tmode = (mode | ((mode & 0100) ? 0777 : 0666)) & ~tar_umask;\n\nBut write_zip_entry() only normalizes (drops) the permission bits of\nnon-executable files:\n\n                attr2 = S_ISLNK(mode) ? ((mode | 0777) << 16) :\n                        (mode & 0111) ? ((mode) << 16) : 0;\n                if (S_ISLNK(mode) || (mode & 0111))\n                        creator_version = 0x0317;\n\nattr2 corresponds to the field \"external file attributes\" mentioned in\nthe ZIP format specification, APPNOTE.TXT.  It's interpreted based on\nthe \"version made by\" (creator_version here); that 0x03 part above\nmeans \"UNIX\".  The default is MS-DOS (FAT filesystem), with effectivly\nno support for file permissions.\n\nSo we currently leak permission bits of executable files into ZIP\narchives, but not tar files. :-|  Normalizing those to 0755 would be\nmore consistent.\n\n> For tracked paths, we probably are normalizing the blobs\n> between 0644 and 0755 way before the values are passed as \"mode\"\n> parameter to the write_entry() functions, but for these extra files,\n> there is no such massaging.\n\nRight, mode values from read_tree() pass through canon_mode(), so only\nuntracked files (those appended with --add-file) are affected by the\nleakage mentioned above.\n\nRené\n"},{"id":"455213","messageId":"xmqqfslefwie.fsf@gitster.g","threadId":"57313","inReplyTo":"47ed5a2f-f4aa-1ec1-27c9-9b0b70eb8bca@web.de","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-12T18:15:05Z","receivedAt":"2022-05-12T18:15:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Good point.  write_tar_entry() actually normalizes the permission bits\n> and applies tar.umask (0002 by default):\n>\n> \tif (S_ISDIR(mode) || S_ISGITLINK(mode)) {\n> \t\t*header.typeflag = TYPEFLAG_DIR;\n> \t\tmode = (mode | 0777) & ~tar_umask;\n> \t} else if (S_ISLNK(mode)) {\n> \t\t*header.typeflag = TYPEFLAG_LNK;\n> \t\tmode |= 0777;\n> \t} else if (S_ISREG(mode)) {\n> \t\t*header.typeflag = TYPEFLAG_REG;\n> \t\tmode = (mode | ((mode & 0100) ? 0777 : 0666)) & ~tar_umask;\n\nYeah, this side seems to care only about u+x bit, so\n\"add-executable\" as a separate option would fly we..\n\n> But write_zip_entry() only normalizes (drops) the permission bits of\n> non-executable files:\n>\n>                 attr2 = S_ISLNK(mode) ? ((mode | 0777) << 16) :\n>                         (mode & 0111) ? ((mode) << 16) : 0;\n>                 if (S_ISLNK(mode) || (mode & 0111))\n>                         creator_version = 0x0317;\n>\n> attr2 corresponds to the field \"external file attributes\" mentioned in\n> the ZIP format specification, APPNOTE.TXT.  It's interpreted based on\n> the \"version made by\" (creator_version here); that 0x03 part above\n> means \"UNIX\".  The default is MS-DOS (FAT filesystem), with effectivly\n> no support for file permissions.\n>\n> So we currently leak permission bits of executable files into ZIP\n> archives, but not tar files. :-|  Normalizing those to 0755 would be\n> more consistent.\n\nYup.\n\n>> For tracked paths, we probably are normalizing the blobs\n>> between 0644 and 0755 way before the values are passed as \"mode\"\n>> parameter to the write_entry() functions, but for these extra files,\n>> there is no such massaging.\n>\n> Right, mode values from read_tree() pass through canon_mode(), so only\n> untracked files (those appended with --add-file) are affected by the\n> leakage mentioned above.\n\nThanks for sanity-checking.\n\n"},{"id":"455219","messageId":"xmqqmtfme8v6.fsf@gitster.g","threadId":"57313","inReplyTo":"xmqqfslefwie.fsf@gitster.g","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-12T21:31:09Z","receivedAt":"2022-05-12T21:31:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n>> So we currently leak permission bits of executable files into ZIP\n>> archives, but not tar files. :-|  Normalizing those to 0755 would be\n>> more consistent.\n\nToday, I was scanning the \"What's cooking\" draft and saw too many\ntopics that are marked with \"Expecting a reroll\".  It turns out that\nthis \"mode bits\" thing will not be a blocker to make us wait for a\nreroll of the topic, so let's handle it separately, before we\nforget, as an independent fix outside the series under discussion.\n\nThanks.\n\n--- >8 ---\nSubject: [PATCH] archive: do not let on-disk mode leak to zip archives\n\nWhen the \"--add-file\" option is used to add the contents from an\nuntracked file to the archive, the permission mode bits for these\nfiles are sent to the archive-backend specific \"write_entry()\"\nmethod as-is.  We normalize the mode bits for tracked files way\nbefore we pass them to the write_entry() method; we should do the\nsame here.\n\nThis is not strictly needed for \"tar\" archive-backend, as it has its\nown code to further clean them up, but \"zip\" archive-backend is not\nso well prepared.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n archive.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/archive.c b/archive.c\nindex e29d0e00f6..12a08af531 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -342,7 +342,7 @@ int write_archive_entries(struct archiver_args *args,\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n \t\t\t\t\t  path_in_archive.len,\n-\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  canon_mode(info->stat.st_mode),\n \t\t\t\t\t  content.buf, content.len);\n \t\tif (err)\n \t\t\tbreak;\n-- \n2.36.1-338-g1c7f76a54c\n\n"},{"id":"455222","messageId":"xmqqee0ye62w.fsf_-_@gitster.g","threadId":"57313","inReplyTo":"xmqqtu9x6ovh.fsf@gitster.g","subject":"[PATCH] fixup! archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-12T22:31:19Z","receivedAt":"2022-05-12T22:31:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Do not let add_file_cb() assume that two existing callers are the\nonly ones, and checking that the caller is not one of them is\nsufficient to determine it is the other one.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * To be squashed to the commit with the title in the series.\n\n   The \"What's cooking\" report is getting crowded with too many\n   topics marked as \"Expecting a reroll\", and I'm trying to do\n   easier ones myself to see how much reduction we can make.\n\n archive.c | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/archive.c b/archive.c\nindex 477eba60ac..98c7449ea1 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -533,7 +533,7 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\tif (!S_ISREG(info->stat.st_mode))\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n-\t} else {\n+\t} else if (!strcmp(opt->long_name, \"add-file-with-content\")) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tconst char *p = arg;\n \n@@ -560,6 +560,8 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n \t\tinfo->content = xstrdup(p + 1);\n \t\tinfo->stat.st_size = strlen(info->content);\n+\t} else {\n+\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n \t}\n \titem = string_list_append_nodup(&args->extra_files, path);\n \titem->util = info;\n-- \n2.36.1-338-g1c7f76a54c\n\n"},{"id":"455284","messageId":"20179ba6-33d5-e95e-cf90-10aabc81734b@web.de","threadId":"57313","inReplyTo":"xmqqmtfme8v6.fsf@gitster.g","subject":"Re: [PATCH v4 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-14T07:06:07Z","receivedAt":"2022-05-14T07:06:39Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 12.05.22 um 23:31 schrieb Junio C Hamano:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>>> So we currently leak permission bits of executable files into ZIP\n>>> archives, but not tar files. :-|  Normalizing those to 0755 would be\n>>> more consistent.\n>\n> Today, I was scanning the \"What's cooking\" draft and saw too many\n> topics that are marked with \"Expecting a reroll\".  It turns out that\n> this \"mode bits\" thing will not be a blocker to make us wait for a\n> reroll of the topic, so let's handle it separately, before we\n> forget, as an independent fix outside the series under discussion.\n>\n> Thanks.\n>\n> --- >8 ---\n> Subject: [PATCH] archive: do not let on-disk mode leak to zip archives\n>\n> When the \"--add-file\" option is used to add the contents from an\n> untracked file to the archive, the permission mode bits for these\n> files are sent to the archive-backend specific \"write_entry()\"\n> method as-is.  We normalize the mode bits for tracked files way\n> before we pass them to the write_entry() method; we should do the\n> same here.\n>\n> This is not strictly needed for \"tar\" archive-backend, as it has its\n> own code to further clean them up, but \"zip\" archive-backend is not\n> so well prepared.\n>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  archive.c | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git a/archive.c b/archive.c\n> index e29d0e00f6..12a08af531 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -342,7 +342,7 @@ int write_archive_entries(struct archiver_args *args,\n>  \t\telse\n>  \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n>  \t\t\t\t\t  path_in_archive.len,\n> -\t\t\t\t\t  info->stat.st_mode,\n> +\t\t\t\t\t  canon_mode(info->stat.st_mode),\n>  \t\t\t\t\t  content.buf, content.len);\n>  \t\tif (err)\n>  \t\t\tbreak;\n\nLooks good to me, thank you!\n\nRené\n"},{"id":"455385","messageId":"220517.867d6k6wjr.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"da9f52a82406ffc909e9c5f2b6b5e77818d972c0.1652210824.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-17T14:51:59Z","receivedAt":"2022-05-17T14:53:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 10 2022, Johannes Schindelin via GitGitGadget wrote:\n\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> The `scalar` command needs a Scalar enlistment for many subcommands, and\n> looks in the current directory for such an enlistment (traversing the\n> parent directories until it finds one).\n>\n> These is subcommands can also be called with an optional argument\n> specifying the enlistment. Here, too, we traverse parent directories as\n> needed, until we find an enlistment.\n>\n> However, if the specified directory does not even exist, or is not a\n> directory, we should stop right there, with an error message.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  contrib/scalar/scalar.c          | 6 ++++--\n>  contrib/scalar/t/t9099-scalar.sh | 5 +++++\n>  2 files changed, 9 insertions(+), 2 deletions(-)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 1ce9c2b00e8..00dcd4b50ef 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n>  \t\tusage_with_options(usagestr, options);\n>  \n>  \t/* find the worktree, determine its corresponding root */\n> -\tif (argc == 1)\n> +\tif (argc == 1) {\n>  \t\tstrbuf_add_absolute_path(&path, argv[0]);\n> -\telse if (strbuf_getcwd(&path) < 0)\n> +\t\tif (!is_directory(path.buf))\n> +\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n> +\t} else if (strbuf_getcwd(&path) < 0)\n>  \t\tdie(_(\"need a working directory\"));\n>  \n>  \tstrbuf_trim_trailing_dir_sep(&path);\n> diff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\n> index 2e1502ad45e..9d83fdf25e8 100755\n> --- a/contrib/scalar/t/t9099-scalar.sh\n> +++ b/contrib/scalar/t/t9099-scalar.sh\n> @@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n>  \ttest_path_is_missing cloned\n>  '\n>  \n> +test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n> +\t! scalar run config cloned 2>err &&\n\nNeeds to use test_must_fail, not !\n"},{"id":"455386","messageId":"220517.8635h86w7h.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"87bdc22322b0f58bf153b963207cffe4f41c9ae9.1652210824.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 4/7] Implement `scalar diagnose`","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-17T14:53:17Z","receivedAt":"2022-05-17T15:00:12Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 10 2022, Johannes Schindelin via GitGitGadget wrote:\n\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> Over the course of Scalar's development, it became obvious that there is\n> a need for a command that can gather all kinds of useful information\n> that can help identify the most typical problems with large\n> worktrees/repositories.\n>\n> The `diagnose` command is the culmination of this hard-won knowledge: it\n> gathers the installed hooks, the config, a couple statistics describing\n> the data shape, among other pieces of information, and then wraps\n> everything up in a tidy, neat `.zip` archive.\n>\n> Note: originally, Scalar was implemented in C# using the .NET API, where\n> we had the luxury of a comprehensive standard library that includes\n> basic functionality such as writing a `.zip` file. In the C version, we\n> lack such a commodity. Rather than introducing a dependency on, say,\n> libzip, we slightly abuse Git's `archive` machinery: we write out a\n> `.zip` of the empty try, augmented by a couple files that are added via\n> the `--add-file*` options. We are careful trying not to modify the\n> current repository in any way lest the very circumstances that required\n> `scalar diagnose` to be run are changed by the `diagnose` run itself.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  contrib/scalar/scalar.c          | 144 +++++++++++++++++++++++++++++++\n>  contrib/scalar/scalar.txt        |  12 +++\n>  contrib/scalar/t/t9099-scalar.sh |  14 +++\n>  3 files changed, 170 insertions(+)\n>\n> diff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\n> index 00dcd4b50ef..367a2c50e25 100644\n> --- a/contrib/scalar/scalar.c\n> +++ b/contrib/scalar/scalar.c\n> @@ -11,6 +11,7 @@\n>  #include \"dir.h\"\n>  #include \"packfile.h\"\n>  #include \"help.h\"\n> +#include \"archive.h\"\n>  \n>  /*\n>   * Remove the deepest subdirectory in the provided path string. Path must not\n> @@ -261,6 +262,47 @@ static int unregister_dir(void)\n>  \treturn res;\n>  }\n>  \n> +static int add_directory_to_archiver(struct strvec *archiver_args,\n> +\t\t\t\t\t  const char *path, int recurse)\n> +{\n> +\tint at_root = !*path;\n> +\tDIR *dir = opendir(at_root ? \".\" : path);\n> +\tstruct dirent *e;\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tsize_t len;\n> +\tint res = 0;\n> +\n> +\tif (!dir)\n> +\t\treturn error(_(\"could not open directory '%s'\"), path);\n\n\ns/error/error_errno/, surely?\n\n> +\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n> +\tstrbuf_addftime(&zip_path,\n> +\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n\nWould we be worse off if we stole this timestamp from some known file\n(or HEAD), and thus made a second run of this reproducable?\n"},{"id":"455387","messageId":"220517.86y1z05gja.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-17T15:03:27Z","receivedAt":"2022-05-17T15:24:03Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 10 2022, Johannes Schindelin via GitGitGadget wrote:\n\n> Over the course of the years, we developed a sub-command that gathers\n> diagnostic data into a .zip file that can then be attached to bug reports.\n> This sub-command turned out to be very useful in helping Scalar developers\n> identify and fix issues.\n\nI don't mind this as some intermediate step, but re the context of the\nplan for scalar \"eventually going away\" (discussed in previous threads)\nI wonder why (especially re the earlier thread upthread at [1]) this\nisn't being added to \"git bugreport\".\n\nIs the plan to integrate this into \"git bugreport\" eventually?\n\n1. https://lore.kernel.org/git/nycvar.QRO.7.76.6.2202062213030.347@tvgsbejvaqbjf.bet/\n"},{"id":"455388","messageId":"00f201d86a02$c6399150$52acb3f0$@nexbridge.com","threadId":"57313","inReplyTo":"220517.86y1z05gja.gmgdl@evledraar.gmail.com","subject":"RE: [PATCH v4 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-17T15:28:29Z","receivedAt":"2022-05-17T15:28:45Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 17, 2022 11:03 AM, Ævar Arnfjörð Bjarmason wrote:\n>On Tue, May 10 2022, Johannes Schindelin via GitGitGadget wrote:\n>\n>> Over the course of the years, we developed a sub-command that gathers\n>> diagnostic data into a .zip file that can then be attached to bug reports.\n>> This sub-command turned out to be very useful in helping Scalar\n>> developers identify and fix issues.\n>\n>I don't mind this as some intermediate step, but re the context of the plan for\n>scalar \"eventually going away\" (discussed in previous threads) I wonder why\n>(especially re the earlier thread upthread at [1]) this isn't being added to \"git\n>bugreport\".\n>\n>Is the plan to integrate this into \"git bugreport\" eventually?\n>\n>1.\n>https://lore.kernel.org/git/nycvar.QRO.7.76.6.2202062213030.347@tvgsbejvaqbjf.\n>bet/\n\nCould this also not be useful in fsck, as --diagnose? That's the go-to command when there are issues for many users.\n--Randall\n\n"},{"id":"455406","messageId":"xmqqbkvuwxps.fsf@gitster.g","threadId":"57313","inReplyTo":"220517.867d6k6wjr.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-18T17:35:11Z","receivedAt":"2022-05-18T17:35:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> +test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n>> +\t! scalar run config cloned 2>err &&\n>\n> Needs to use test_must_fail, not !\n\nGood eyes and careful reading are very much appreciated, but in this\ncase, doesn't such an improvement depend on an update to teach\ntest_must_fail_acceptable about scalar being whitelisted?\n\n\n"},{"id":"455513","messageId":"nycvar.QRO.7.76.6.2205192004490.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"xmqqmtfp6ohc.fsf@gitster.g","subject":"Re: [PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-19T18:09:40Z","receivedAt":"2022-05-19T18:10:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Tue, 10 May 2022, Junio C Hamano wrote:\n\n> \"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n> writes:\n>\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >\n> > By allowing the path to be enclosed in double-quotes, we can avoid\n> > the limitation that paths cannot contain colons.\n> > ...\n> > +\t\tstruct strbuf buf = STRBUF_INIT;\n> > +\t\tconst char *p = arg;\n> > +\n> > +\t\tif (*p != '\"')\n> > +\t\t\tp = strchr(p, ':');\n> > +\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n> > +\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n>\n> Even though I do not think people necessarily would want to use\n> colons in their pathnames (it has problems interoperating with other\n> systems), lifting the limitation is a good thing to do.  I totally\n> forgot that we designed unquote_c_style() to self terminate and\n> return the end pointer to the caller so the caller does not have to\n> worry, which is very nice.\n>\n> Even if this step weren't here in the series, I would have thought\n> the mode bits issue was more serious than \"no colons in path\"\n> limitation, but given that we address this unusual corner case\n> limitation, I would think we should address the hardcoded mode bits\n> at the same time.\n>\n> > diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> > index 8ff1257f1a0..5b8bbfc2692 100755\n> > --- a/t/t5003-archive-zip.sh\n> > +++ b/t/t5003-archive-zip.sh\n> > @@ -207,13 +207,21 @@ check_zip with_untracked\n> >  check_added with_untracked untracked untracked\n> >\n> >  test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n> > +\tif test_have_prereq FUNNYNAMES\n> > +\tthen\n> > +\t\tQUOTED=quoted:colon\n> > +\telse\n> > +\t\tQUOTED=quoted\n> > +\tfi &&\n>\n> ;-)\n>\n> >  \tgit archive --format=zip >with_file_with_content.zip \\\n> > +\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n> >  \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n> >  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n> >  \tmkdir tmp-unpack && (\n> >  \t\tcd tmp-unpack &&\n> >  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> >  \t\ttest_path_is_file hello &&\n> > +\t\ttest_path_is_file $QUOTED &&\n>\n> Looks OK, even though it probably is a good idea to have dq around\n> $QUOTED, so that future developers can easily insert SP into its\n> value to use a bit more common but still a bit more problematic\n> pathnames in the test.\n\nI actually decided against this because reading\n\n\t\"$QUOTED\"\n\nwould mislead future me to think that the double quotes that enclose\n$QUOTED are the quotes that the variable's name talks about. But the\nquotes are actually the escaped ones that are passed to `git archive`\nabove.\n\nSo, to help future Dscho should they read this code six months from now or\neven later, I wanted to specifically only add quotes to the `git archive`\ncall to make the intention abundantly clear.\n\nCiao,\nDscho\n"},{"id":"455514","messageId":"nycvar.QRO.7.76.6.2205192010210.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"03d801d864bc$85fe62a0$91fb27e0$@nexbridge.com","subject":"RE: [PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-19T18:12:15Z","receivedAt":"2022-05-19T18:12:41Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Randall,\n\nOn Tue, 10 May 2022, rsbecker@nexbridge.com wrote:\n\n> On May 10, 2022 5:57 PM, Junio C Hamano wrote:\n> >\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n> >writes:\n> >\n> >>  \tgit archive --format=zip >with_file_with_content.zip \\\n> >> +\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n> >>  \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n> >>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n> >>  \tmkdir tmp-unpack && (\n> >>  \t\tcd tmp-unpack &&\n> >>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> >>  \t\ttest_path_is_file hello &&\n> >> +\t\ttest_path_is_file $QUOTED &&\n> >\n> >Looks OK, even though it probably is a good idea to have dq around $QUOTED, so\n> >that future developers can easily insert SP into its value to use a bit more common\n> >but still a bit more problematic pathnames in the test.\n>\n> A test case for .gitignore in this would be good too. People on our\n> exotic platform do this stuff as a matter of course. As an example, a\n> name of $Z3P4:12399334 being used as a named pipe (associated with the\n> unique name of a process) actually has been seen in the wild recently.\n> My solution was to wild card this and/or contain it in an ignored\n> directory.\n\nThe `--add-file-with-content` option, which this test case is all about,\nspecifically does not heed `.gitignore`. Is this what you want to test? If\nso, I don't think that's necessary. Unless you expect some future version\nto introduce a patch by mistake that makes `--add-file-with-content`\nsubject to the `.gitignore` rules.\n\nCiao,\nDscho\n"},{"id":"455515","messageId":"nycvar.QRO.7.76.6.2205192013140.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"00f201d86a02$c6399150$52acb3f0$@nexbridge.com","subject":"RE: [PATCH v4 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-19T18:17:05Z","receivedAt":"2022-05-19T18:17:23Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Randall and Ævar,\n\nOn Tue, 17 May 2022, rsbecker@nexbridge.com wrote:\n\n> On May 17, 2022 11:03 AM, Ævar Arnfjörð Bjarmason wrote:\n> >On Tue, May 10 2022, Johannes Schindelin via GitGitGadget wrote:\n> >\n> >> Over the course of the years, we developed a sub-command that gathers\n> >> diagnostic data into a .zip file that can then be attached to bug\n> >> reports. This sub-command turned out to be very useful in helping\n> >> Scalar developers identify and fix issues.\n> >\n> >I don't mind this as some intermediate step, but re the context of the\n> >plan for scalar \"eventually going away\" (discussed in previous threads)\n> >I wonder why (especially re the earlier thread upthread at [1]) this\n> >isn't being added to \"git bugreport\".\n> >\n> >Is the plan to integrate this into \"git bugreport\" eventually?\n\nPotentially a variation of the `scalar diagnose` code could be useful in\n`git bugreport`, opt-in via a new option.\n\nBut that's not the purpose of this patch series.\n\n> Could this also not be useful in fsck, as --diagnose? That's the go-to\n> command when there are issues for many users.\n\nI can see where you're coming from, but `fsck`'s mission is to verify the\nintegrity of the local Git database. That is very different from the\nmission of `scalar diagnose`, which is to help diagnose issues (whether\nthey are truly bugs or usage patterns causing unfortunate performance).\n\nCiao,\nDscho\n"},{"id":"455516","messageId":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v4.git.1652210824.gitgitgadget@gmail.com","subject":"[PATCH v5 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:17:56Z","receivedAt":"2022-05-19T18:18:14Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nChanges since v4:\n\n * Squashed in Junio's suggested fixups\n * Renamed the option from --add-file-with-content=<name>:<content> to\n   --add-virtual-file=<name>:<content>\n * Fixed one instance where I had used error() instead of error_errno().\n\nChanges since v3:\n\n * We're now using unquote_c_style() instead of rolling our own unquoter.\n * Fixed the added regression test.\n * As pointed out by Scalar's Functional Tests, the\n   add_directory_to_archiver() function should not fail when scalar diagnose\n   encounters FSMonitor's Unix socket, but only warn instead.\n * Related: add_directory_to_archiver() needs to propagate errors from\n   processing subdirectories so that the top-level call returns an error,\n   too.\n\nChanges since v2:\n\n * Clarified in the commit message what the biggest benefit of\n   --add-file-with-content is.\n * The <path> part of the -add-file-with-content argument can now contain\n   colons. To do this, the path needs to start and end in double-quote\n   characters (which are stripped), and the backslash serves as escape\n   character in that case (to allow the path to contain both colons and\n   double-quotes).\n * Fixed incorrect grammar.\n * Instead of strcmp(<what-we-don't-want>), we now say\n   !strcmp(<what-we-want>).\n * The help text for --add-file-with-content was improved a tiny bit.\n * Adjusted the commit message that still talked about spawning plenty of\n   processes and about a throw-away repository for the sake of generating a\n   .zip file.\n * Simplified the code that shows the diagnostics and adds them to the .zip\n   file.\n * The final message that reports that the archive is complete is now\n   printed to stderr instead of stdout.\n\nChanges since v1:\n\n * Instead of creating a throw-away repository, staging the contents of the\n   .zip file and then using git write-tree and git archive to write the .zip\n   file, the patch series now introduces a new option to git archive and\n   uses write_archive() directly (avoiding any separate process).\n * Since the command avoids separate processes, it is now blazing fast on\n   Windows, and I dropped the spinner() function because it's no longer\n   needed.\n * While reworking the test case, I noticed that scalar [...] <enlistment>\n   failed to verify that the specified directory exists, and would happily\n   \"traverse to its parent directory\" on its quest to find a Scalar\n   enlistment. That is of course incorrect, and has been fixed as a \"while\n   at it\" sort of preparatory commit.\n * I had forgotten to sign off on all the commits, which has been fixed.\n * Instead of some \"home-grown\" readdir()-based function, the code now uses\n   for_each_file_in_pack_dir() to look through the pack directories.\n * If any alternates are configured, their pack directories are now included\n   in the output.\n * The commit message that might be interpreted to promise information about\n   large loose files has been corrected to no longer promise that.\n * The test cases have been adjusted to test a little bit more (e.g.\n   verifying that specific paths are mentioned in the output, instead of\n   merely verifying that the output is non-empty).\n\nJohannes Schindelin (5):\n  archive: optionally add \"virtual\" files\n  archive --add-file-with-contents: allow paths containing colons\n  scalar: validate the optional enlistment argument\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n Documentation/git-archive.txt    |  17 ++\n archive.c                        |  63 ++++++-\n contrib/scalar/scalar.c          | 292 ++++++++++++++++++++++++++++++-\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  27 +++\n t/t5003-archive-zip.sh           |  20 +++\n 6 files changed, 421 insertions(+), 10 deletions(-)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n\nRange-diff vs v4:\n\n 1:  45662cf582a ! 1:  42e73fb0aac archive: optionally add \"virtual\" files\n     @@ Documentation/git-archive.txt: OPTIONS\n       \tby concatenating the value for `--prefix` (if any) and the\n       \tbasename of <file>.\n       \n     -+--add-file-with-content=<path>:<content>::\n     ++--add-virtual-file=<path>:<content>::\n      +\tAdd the specified contents to the archive.  Can be repeated to add\n      +\tmultiple files.  The path of the file in the archive is built\n      +\tby concatenating the value for `--prefix` (if any) and the\n     @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int\n      +\t\tif (!S_ISREG(info->stat.st_mode))\n      +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n      +\t\tinfo->content = NULL; /* read the file later */\n     -+\t} else {\n     ++\t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n      +\t\tconst char *colon = strchr(arg, ':');\n      +\t\tchar *p;\n      +\n     @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int\n      +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n      +\t\tinfo->content = xstrdup(colon + 1);\n      +\t\tinfo->stat.st_size = strlen(info->content);\n     ++\t} else {\n     ++\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n      +\t}\n      +\titem = string_list_append_nodup(&args->extra_files, path);\n      +\titem->util = info;\n     @@ archive.c: static int parse_archive_args(int argc, const char **argv,\n       \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n       \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n       \t\t  (intptr_t)&base },\n     -+\t\t{ OPTION_CALLBACK, 0, \"add-file-with-content\", args,\n     ++\t\t{ OPTION_CALLBACK, 0, \"add-virtual-file\", args,\n      +\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n      +\t\t  add_file_cb, (intptr_t)&base },\n       \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n     @@ t/t5003-archive-zip.sh: test_expect_success 'git archive --format=zip --add-file\n       check_zip with_untracked\n       check_added with_untracked untracked untracked\n       \n     -+test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n     ++test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n      +\tgit archive --format=zip >with_file_with_content.zip \\\n     -+\t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n     ++\t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n      +\ttest_when_finished \"rm -rf tmp-unpack\" &&\n      +\tmkdir tmp-unpack && (\n      +\t\tcd tmp-unpack &&\n 2:  fdba4ed6f4d ! 2:  b5ebd61066a archive --add-file-with-contents: allow paths containing colons\n     @@ archive.c\n      @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int unset)\n       \t\t\tdie(_(\"Not a regular file: %s\"), path);\n       \t\tinfo->content = NULL; /* read the file later */\n     - \t} else {\n     + \t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n      -\t\tconst char *colon = strchr(arg, ':');\n      -\t\tchar *p;\n      +\t\tstruct strbuf buf = STRBUF_INIT;\n     @@ archive.c: static int add_file_cb(const struct option *opt, const char *arg, int\n      -\t\tinfo->content = xstrdup(colon + 1);\n      +\t\tinfo->content = xstrdup(p + 1);\n       \t\tinfo->stat.st_size = strlen(info->content);\n     - \t}\n     - \titem = string_list_append_nodup(&args->extra_files, path);\n     + \t} else {\n     + \t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n      \n       ## t/t5003-archive-zip.sh ##\n      @@ t/t5003-archive-zip.sh: check_zip with_untracked\n       check_added with_untracked untracked untracked\n       \n     - test_expect_success UNZIP 'git archive --format=zip --add-file-with-content' '\n     + test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n      +\tif test_have_prereq FUNNYNAMES\n      +\tthen\n      +\t\tQUOTED=quoted:colon\n     @@ t/t5003-archive-zip.sh: check_zip with_untracked\n      +\t\tQUOTED=quoted\n      +\tfi &&\n       \tgit archive --format=zip >with_file_with_content.zip \\\n     -+\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n     - \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n     ++\t\t--add-virtual-file=\\\"$QUOTED\\\": \\\n     + \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n       \ttest_when_finished \"rm -rf tmp-unpack\" &&\n       \tmkdir tmp-unpack && (\n       \t\tcd tmp-unpack &&\n 3:  da9f52a8240 = 3:  f1ba69c02d7 scalar: validate the optional enlistment argument\n 4:  87bdc22322b ! 4:  3fb90194744 Implement `scalar diagnose`\n     @@ contrib/scalar/scalar.c: static int unregister_dir(void)\n      +\tint res = 0;\n      +\n      +\tif (!dir)\n     -+\t\treturn error(_(\"could not open directory '%s'\"), path);\n     ++\t\treturn error_errno(_(\"could not open directory '%s'\"), path);\n      +\n      +\tif (!at_root)\n      +\t\tstrbuf_addf(&buf, \"%s/\", path);\n     @@ contrib/scalar/scalar.c: cleanup:\n      +\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n      +\twrite_or_die(stdout_fd, buf.buf, buf.len);\n      +\tstrvec_pushf(&archiver_args,\n     -+\t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n     ++\t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n      +\t\t     (int)buf.len, buf.buf);\n      +\n      +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n 5:  3f63b197d42 ! 5:  2e645b08a9e scalar diagnose: include disk space information\n     @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n      +\tget_disk_info(&buf);\n       \twrite_or_die(stdout_fd, buf.buf, buf.len);\n       \tstrvec_pushf(&archiver_args,\n     - \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n     + \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n      \n       ## contrib/scalar/t/t9099-scalar.sh ##\n      @@ contrib/scalar/t/t9099-scalar.sh: SQ=\"'\"\n 6:  fc1319338fc ! 6:  0fa20d73750 scalar: teach `diagnose` to gather packfile info\n     @@ contrib/scalar/scalar.c: cleanup:\n       {\n       \tstruct option options[] = {\n      @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n     - \t\t     \"--add-file-with-content=diagnostics.log:%.*s\",\n     + \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n       \t\t     (int)buf.len, buf.buf);\n       \n      +\tstrbuf_reset(&buf);\n     -+\tstrbuf_addstr(&buf, \"--add-file-with-content=packs-local.txt:\");\n     ++\tstrbuf_addstr(&buf, \"--add-virtual-file=packs-local.txt:\");\n      +\tdir_file_stats(the_repository->objects->odb, &buf);\n      +\tforeach_alt_odb(dir_file_stats, &buf);\n      +\tstrvec_push(&archiver_args, buf.buf);\n 7:  e8f5b42f7b7 ! 7:  62e173b47cf scalar: teach `diagnose` to gather loose objects information\n     @@ contrib/scalar/scalar.c: static int cmd_diagnose(int argc, const char **argv)\n       \tstrvec_push(&archiver_args, buf.buf);\n       \n      +\tstrbuf_reset(&buf);\n     -+\tstrbuf_addstr(&buf, \"--add-file-with-content=objects-local.txt:\");\n     ++\tstrbuf_addstr(&buf, \"--add-virtual-file=objects-local.txt:\");\n      +\tloose_objs_stats(&buf, \".git/objects\");\n      +\tstrvec_push(&archiver_args, buf.buf);\n      +\n\n-- \ngitgitgadget\n"},{"id":"455517","messageId":"42e73fb0aaca1f2498ed817c517859103d72d32b.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 1/7] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:17:57Z","receivedAt":"2022-05-19T18:18:21Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-file-with-content=<path>:<content>` option, `git\narchive` now supports use cases where relatively trivial files need to\nbe added that do not exist on disk.\n\nThis will allow us to generate `.zip` files with generated content,\nwithout having to add said content to the object database and without\nhaving to write it out to disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 11 ++++++++\n archive.c                     | 53 +++++++++++++++++++++++++++++------\n t/t5003-archive-zip.sh        | 12 ++++++++\n 3 files changed, 68 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a7834..893cb1075bf 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -61,6 +61,17 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n \n+--add-virtual-file=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value for `--prefix` (if any) and the\n+\tbasename of <file>.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject to platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex a3bbb091256..d20e16fa819 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n \t\tstrbuf_addstr(&path_in_archive, basename(path));\n \n \t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\tif (info->content)\n+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\telse if (strbuf_read_file(&content, path,\n+\t\t\t\t\t  info->stat.st_size) < 0)\n \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n@@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +522,40 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (!strcmp(opt->long_name, \"add-file\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t} else {\n+\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +588,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-virtual-file\", args,\n+\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 1e6d18b140e..ebc26e89a9b 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \ngitgitgadget\n\n"},{"id":"455518","messageId":"b5ebd61066ab4e1ef978605c376c1925a270e091.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:17:58Z","receivedAt":"2022-05-19T18:18:25Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nBy allowing the path to be enclosed in double-quotes, we can avoid\nthe limitation that paths cannot contain colons.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 14 ++++++++++----\n archive.c                     | 30 ++++++++++++++++++++----------\n t/t5003-archive-zip.sh        |  8 ++++++++\n 3 files changed, 38 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 893cb1075bf..54de945a84e 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -67,10 +67,16 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n +\n-The `<path>` cannot contain any colon, the file mode is limited to\n-a regular file, and the option may be subject to platform-dependent\n-command-line limits. For non-trivial cases, write an untracked file\n-and use `--add-file` instead.\n+The `<path>` argument can start and end with a literal double-quote\n+character; The contained file name is interpreted as a C-style string,\n+i.e. the backslash is interpreted as escape character. The path must\n+be quoted if it contains a colon, to avoid the colon from being\n+misinterpreted as the separator between the path and the contents, or\n+if the path begins or ends with a double-quote character.\n++\n+The file mode is limited to a regular file, and the option may be\n+subject to platform-dependent command-line limits. For non-trivial\n+cases, write an untracked file and use `--add-file` instead.\n \n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\ndiff --git a/archive.c b/archive.c\nindex d20e16fa819..b7756b91200 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -9,6 +9,7 @@\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n #include \"dir.h\"\n+#include \"quote.h\"\n \n static char const * const archive_usage[] = {\n \tN_(\"git archive [<options>] <tree-ish> [<path>...]\"),\n@@ -533,22 +534,31 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n \t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n-\t\tconst char *colon = strchr(arg, ':');\n-\t\tchar *p;\n+\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\tconst char *p = arg;\n+\n+\t\tif (*p != '\"')\n+\t\t\tp = strchr(p, ':');\n+\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n+\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n \n-\t\tif (!colon)\n+\t\tif (!p || *p != ':')\n \t\t\tdie(_(\"missing colon: '%s'\"), arg);\n \n-\t\tp = xstrndup(arg, colon - arg);\n-\t\tif (!args->prefix)\n-\t\t\tpath = p;\n-\t\telse {\n-\t\t\tpath = prefix_filename(args->prefix, p);\n-\t\t\tfree(p);\n+\t\tif (p == arg)\n+\t\t\tdie(_(\"empty file name: '%s'\"), arg);\n+\n+\t\tpath = buf.len ?\n+\t\t\tstrbuf_detach(&buf, NULL) : xstrndup(arg, p - arg);\n+\n+\t\tif (args->prefix) {\n+\t\t\tchar *save = path;\n+\t\t\tpath = prefix_filename(args->prefix, path);\n+\t\t\tfree(save);\n \t\t}\n \t\tmemset(&info->stat, 0, sizeof(info->stat));\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n-\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->content = xstrdup(p + 1);\n \t\tinfo->stat.st_size = strlen(info->content);\n \t} else {\n \t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex ebc26e89a9b..50932a866c9 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -207,13 +207,21 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tif test_have_prereq FUNNYNAMES\n+\tthen\n+\t\tQUOTED=quoted:colon\n+\telse\n+\t\tQUOTED=quoted\n+\tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=\\\"$QUOTED\\\": \\\n \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n+\t\ttest_path_is_file $QUOTED &&\n \t\ttest world = $(cat hello)\n \t)\n '\n-- \ngitgitgadget\n\n"},{"id":"455519","messageId":"2e645b08a9e3c8505d4c525dadddd1f3eac77201.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 5/7] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:18:01Z","receivedAt":"2022-05-19T18:18:30Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 53213f9a3b9..0a9e25a57f8 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -303,6 +303,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -599,6 +651,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \twrite_or_die(stdout_fd, buf.buf, buf.len);\n \tstrvec_pushf(&archiver_args,\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 6802d317258..934b2485d91 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -94,6 +94,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out 2>err &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \ngitgitgadget\n\n"},{"id":"455521","messageId":"3fb90194744998e5dd8f2e5b46610a7cc773a09b.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 4/7] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:18:00Z","receivedAt":"2022-05-19T18:18:35Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` machinery: we write out a\n`.zip` of the empty try, augmented by a couple files that are added via\nthe `--add-file*` options. We are careful trying not to modify the\ncurrent repository in any way lest the very circumstances that required\n`scalar diagnose` to be run are changed by the `diagnose` run itself.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 144 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 170 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 00dcd4b50ef..53213f9a3b9 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -261,6 +262,47 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error_errno(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\twarning(_(\"skipping '%s', which is neither file nor \"\n+\t\t\t\t  \"directory\"), buf.buf);\n+\t\telse if (recurse &&\n+\t\t\t add_directory_to_archiver(archiver_args,\n+\t\t\t\t\t\t   buf.buf, recurse) < 0)\n+\t\t\tres = -1;\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -501,6 +543,107 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\twrite_or_die(stdout_fd, buf.buf, buf.len);\n+\tstrvec_pushf(&archiver_args,\n+\t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n+\t\t     (int)buf.len, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tfprintf(stderr, \"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -802,6 +945,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 9d83fdf25e8..6802d317258 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out 2>err &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455520","messageId":"f1ba69c02d705ffd1d1ccbc96e2801adf470c6f1.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:17:59Z","receivedAt":"2022-05-19T18:18:36Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..00dcd4b50ef 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..9d83fdf25e8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n \ttest_path_is_missing cloned\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455522","messageId":"0fa20d7375009999d978da682c3f98a604cda9e4.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 6/7] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:18:02Z","receivedAt":"2022-05-19T18:18:39Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 0a9e25a57f8..d302c27e114 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -595,6 +596,29 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -657,6 +681,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n \t\t     (int)buf.len, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 934b2485d91..3dd5650cceb 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,6 +93,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -102,7 +104,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"455523","messageId":"62e173b47cf2550a8e458c4d2cf667402b0c8ee3.1652984283.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v5 7/7] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-19T18:18:03Z","receivedAt":"2022-05-19T18:18:40Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex d302c27e114..0c278681758 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -619,6 +619,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -687,6 +741,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 3dd5650cceb..72023a1ca1d 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -95,6 +95,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -106,7 +107,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \ngitgitgadget\n"},{"id":"455525","messageId":"xmqqh75lnz07.fsf@gitster.g","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2205192004490.352@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v4 2/7] archive --add-file-with-contents: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T18:44:24Z","receivedAt":"2022-05-19T18:44:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> >  \tgit archive --format=zip >with_file_with_content.zip \\\n>> > +\t\t--add-file-with-content=\\\"$QUOTED\\\": \\\n>> >  \t\t--add-file-with-content=hello:world $EMPTY_TREE &&\n>> >  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n>> >  \tmkdir tmp-unpack && (\n>> >  \t\tcd tmp-unpack &&\n>> >  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>> >  \t\ttest_path_is_file hello &&\n>> > +\t\ttest_path_is_file $QUOTED &&\n>>\n>> Looks OK, even though it probably is a good idea to have dq around\n>> $QUOTED, so that future developers can easily insert SP into its\n>> value to use a bit more common but still a bit more problematic\n>> pathnames in the test.\n>\n> I actually decided against this because reading\n>\n> \t\"$QUOTED\"\n>\n> would mislead future me to think that the double quotes that enclose\n> $QUOTED are the quotes that the variable's name talks about. But the\n> quotes are actually the escaped ones that are passed to `git archive`\n> above.\n\n>\n> So, to help future Dscho should they read this code six months from now or\n> even later, I wanted to specifically only add quotes to the `git archive`\n> call to make the intention abundantly clear.\n\nIf you find \"$QUOTED\" misleads any reader to think QUOTED may have\nsome quote characters in there, you could rename it, of course, to\nsignal what the value is (e.g. $PATHNAME) better.\n\nBut I think you misunderstood my comment completely.\n\nWhat I meant was to write these lines like:\n\n\t--add-file-with-content=\\\"\"$QUOTED\"\\\":\n\ttest_path_is_file \"$QUOTED\"\n\nBecause the value in QUOTED can have $IFS whitespaces in it (after\nall, allowing random letters like colon, quotes and whitespaces is\nwhy we are adding this unquote_c_style() call), and without the\nextra double quotes to protect the parameter expansion of $QUOTED,\nthe command line is broken.\n\nSo, don't decide against it; the reasoning behind that decision is\nsimply wrong.\n\nThanks.\n\n"},{"id":"455528","messageId":"xmqq7d6hnx6g.fsf@gitster.g","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-19T19:23:51Z","receivedAt":"2022-05-19T19:24:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n> Changes since v4:\n>\n>  * Squashed in Junio's suggested fixups\n>  * Renamed the option from --add-file-with-content=<name>:<content> to\n>    --add-virtual-file=<name>:<content>\n\n;-)  5 letters shorter and is a good name.\n\n>  * Fixed one instance where I had used error() instead of error_errno().\n\nLooks good.\n\nThanks.  Will replace and queue.\n"},{"id":"455553","messageId":"220520.86fsl43bkf.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"xmqqbkvuwxps.fsf@gitster.g","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-20T07:30:22Z","receivedAt":"2022-05-20T07:31:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 18 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>>> +test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n>>> +\t! scalar run config cloned 2>err &&\n>>\n>> Needs to use test_must_fail, not !\n>\n> Good eyes and careful reading are very much appreciated, but in this\n> case, doesn't such an improvement depend on an update to teach\n> test_must_fail_acceptable about scalar being whitelisted?\n\nYes, I think so (but haven't tested it just now), but it's a relatively\nsmall change to t/test-lib-functions.sh.\n\nI was just noting the potential hidden segfault etc., the issue remains\nin v5.\n"},{"id":"455569","messageId":"99ffe394-42bb-6c04-b8ab-2d67c38d7756@web.de","threadId":"57313","inReplyTo":"42e73fb0aaca1f2498ed817c517859103d72d32b.1652984283.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-20T14:41:07Z","receivedAt":"2022-05-20T14:41:43Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 19.05.22 um 20:17 schrieb Johannes Schindelin via GitGitGadget:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> With the `--add-file-with-content=<path>:<content>` option, `git\n            ^^^^^^^^^^^^^^^^^^^^^^^\nThat's still the old option name.  Same in the subject of patch 2.\n\n> archive` now supports use cases where relatively trivial files need to\n> be added that do not exist on disk.\n>\n> This will allow us to generate `.zip` files with generated content,\n> without having to add said content to the object database and without\n> having to write it out to disk.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  Documentation/git-archive.txt | 11 ++++++++\n>  archive.c                     | 53 +++++++++++++++++++++++++++++------\n>  t/t5003-archive-zip.sh        | 12 ++++++++\n>  3 files changed, 68 insertions(+), 8 deletions(-)\n>\n> diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> index bc4e76a7834..893cb1075bf 100644\n> --- a/Documentation/git-archive.txt\n> +++ b/Documentation/git-archive.txt\n> @@ -61,6 +61,17 @@ OPTIONS\n>  \tby concatenating the value for `--prefix` (if any) and the\n>  \tbasename of <file>.\n>\n> +--add-virtual-file=<path>:<content>::\n> +\tAdd the specified contents to the archive.  Can be repeated to add\n> +\tmultiple files.  The path of the file in the archive is built\n> +\tby concatenating the value for `--prefix` (if any) and the\n> +\tbasename of <file>.\n> ++\n> +The `<path>` cannot contain any colon, the file mode is limited to\n> +a regular file, and the option may be subject to platform-dependent\n> +command-line limits. For non-trivial cases, write an untracked file\n> +and use `--add-file` instead.\n> +\n>  --worktree-attributes::\n>  \tLook for attributes in .gitattributes files in the working tree\n>  \tas well (see <<ATTRIBUTES>>).\n> diff --git a/archive.c b/archive.c\n> index a3bbb091256..d20e16fa819 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n>  struct extra_file_info {\n>  \tchar *base;\n>  \tstruct stat stat;\n> +\tvoid *content;\n>  };\n>\n>  int write_archive_entries(struct archiver_args *args,\n> @@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n>  \t\tstrbuf_addstr(&path_in_archive, basename(path));\n>\n>  \t\tstrbuf_reset(&content);\n> -\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n> +\t\tif (info->content)\n> +\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> +\t\t\t\t\t  path_in_archive.len,\n> +\t\t\t\t\t  info->stat.st_mode,\n> +\t\t\t\t\t  info->content, info->stat.st_size);\n> +\t\telse if (strbuf_read_file(&content, path,\n> +\t\t\t\t\t  info->stat.st_size) < 0)\n>  \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n>  \t\telse\n>  \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n> @@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n>  {\n>  \tstruct extra_file_info *info = util;\n>  \tfree(info->base);\n> +\tfree(info->content);\n>  \tfree(info);\n>  }\n>\n> @@ -514,14 +522,40 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n>  \tif (!arg)\n>  \t\treturn -1;\n>\n> -\tpath = prefix_filename(args->prefix, arg);\n> -\titem = string_list_append_nodup(&args->extra_files, path);\n> -\titem->util = info = xmalloc(sizeof(*info));\n> +\tinfo = xmalloc(sizeof(*info));\n>  \tinfo->base = xstrdup_or_null(base);\n> -\tif (stat(path, &info->stat))\n> -\t\tdie(_(\"File not found: %s\"), path);\n> -\tif (!S_ISREG(info->stat.st_mode))\n> -\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\n> +\tif (!strcmp(opt->long_name, \"add-file\")) {\n> +\t\tpath = prefix_filename(args->prefix, arg);\n> +\t\tif (stat(path, &info->stat))\n> +\t\t\tdie(_(\"File not found: %s\"), path);\n> +\t\tif (!S_ISREG(info->stat.st_mode))\n> +\t\t\tdie(_(\"Not a regular file: %s\"), path);\n> +\t\tinfo->content = NULL; /* read the file later */\n> +\t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n> +\t\tconst char *colon = strchr(arg, ':');\n> +\t\tchar *p;\n> +\n> +\t\tif (!colon)\n> +\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n> +\n> +\t\tp = xstrndup(arg, colon - arg);\n> +\t\tif (!args->prefix)\n> +\t\t\tpath = p;\n> +\t\telse {\n> +\t\t\tpath = prefix_filename(args->prefix, p);\n> +\t\t\tfree(p);\n> +\t\t}\n> +\t\tmemset(&info->stat, 0, sizeof(info->stat));\n> +\t\tinfo->stat.st_mode = S_IFREG | 0644;\n> +\t\tinfo->content = xstrdup(colon + 1);\n> +\t\tinfo->stat.st_size = strlen(info->content);\n> +\t} else {\n> +\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n> +\t}\n> +\titem = string_list_append_nodup(&args->extra_files, path);\n> +\titem->util = info;\n> +\n>  \treturn 0;\n>  }\n>\n> @@ -554,6 +588,9 @@ static int parse_archive_args(int argc, const char **argv,\n>  \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n>  \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n>  \t\t  (intptr_t)&base },\n> +\t\t{ OPTION_CALLBACK, 0, \"add-virtual-file\", args,\n> +\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n> +\t\t  add_file_cb, (intptr_t)&base },\n>  \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n>  \t\t\tN_(\"write the archive to this file\")),\n>  \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\n> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> index 1e6d18b140e..ebc26e89a9b 100755\n> --- a/t/t5003-archive-zip.sh\n> +++ b/t/t5003-archive-zip.sh\n> @@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n>  check_zip with_untracked\n>  check_added with_untracked untracked untracked\n>\n> +test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> +\tgit archive --format=zip >with_file_with_content.zip \\\n> +\t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n> +\ttest_when_finished \"rm -rf tmp-unpack\" &&\n> +\tmkdir tmp-unpack && (\n> +\t\tcd tmp-unpack &&\n> +\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> +\t\ttest_path_is_file hello &&\n> +\t\ttest world = $(cat hello)\n> +\t)\n> +'\n> +\n>  test_expect_success 'git archive --format=zip --add-file twice' '\n>  \techo untracked >untracked &&\n>  \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n"},{"id":"455571","messageId":"nycvar.QRO.7.76.6.2205201753300.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"220520.86fsl43bkf.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-20T15:55:04Z","receivedAt":"2022-05-20T15:55:24Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Ævar,\n\nOn Fri, 20 May 2022, Ævar Arnfjörð Bjarmason wrote:\n\n>\n> On Wed, May 18 2022, Junio C Hamano wrote:\n>\n> > Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> >\n> >>> +test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n> >>> +\t! scalar run config cloned 2>err &&\n> >>\n> >> Needs to use test_must_fail, not !\n> >\n> > Good eyes and careful reading are very much appreciated, but in this\n> > case, doesn't such an improvement depend on an update to teach\n> > test_must_fail_acceptable about scalar being whitelisted?\n>\n> Yes, I think so (but haven't tested it just now), but it's a relatively\n> small change to t/test-lib-functions.sh.\n\nLet it be noted that I fully agree with Junio that good eyes and careful\nreading are very much appreciated. And that in this case, that would have\nimplied noticing that `test_must_fail` is reserved for Git commands.\n\nScalar is not (yet?) a Git command.\n\nCiao,\nJohannes\n"},{"id":"455573","messageId":"xmqqmtfcjhso.fsf@gitster.g","threadId":"57313","inReplyTo":"99ffe394-42bb-6c04-b8ab-2d67c38d7756@web.de","subject":"Re: [PATCH v5 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-20T16:21:59Z","receivedAt":"2022-05-20T16:22:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 19.05.22 um 20:17 schrieb Johannes Schindelin via GitGitGadget:\n>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>>\n>> With the `--add-file-with-content=<path>:<content>` option, `git\n>             ^^^^^^^^^^^^^^^^^^^^^^^\n> That's still the old option name.  Same in the subject of patch 2.\n\nGood eyes, and thanks for catching what I missed---the risk of\nrelying too much on the range-diff X-<.\n"},{"id":"455667","messageId":"220521.86leuv199g.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2205201753300.352@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-21T09:54:42Z","receivedAt":"2022-05-21T10:16:11Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, May 20 2022, Johannes Schindelin wrote:\n\n> Hi Ævar,\n>\n> On Fri, 20 May 2022, Ævar Arnfjörð Bjarmason wrote:\n>\n>>\n>> On Wed, May 18 2022, Junio C Hamano wrote:\n>>\n>> > Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>> >\n>> >>> +test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n>> >>> +\t! scalar run config cloned 2>err &&\n>> >>\n>> >> Needs to use test_must_fail, not !\n>> >\n>> > Good eyes and careful reading are very much appreciated, but in this\n>> > case, doesn't such an improvement depend on an update to teach\n>> > test_must_fail_acceptable about scalar being whitelisted?\n>>\n>> Yes, I think so (but haven't tested it just now), but it's a relatively\n>> small change to t/test-lib-functions.sh.\n>\n> Let it be noted that I fully agree with Junio that good eyes and careful\n> reading are very much appreciated. And that in this case, that would have\n> implied noticing that `test_must_fail` is reserved for Git commands.\n>\n> Scalar is not (yet?) a Git command.\n\n\"test-tool\" isn't \"git\" either, so I think this argument is a\nnon-starter.\n\nAs the documentation for \"test_must_fail\" notes the distinction is\nwhether something is \"system-supplied\". I.e. we're not going to test\nwhether \"grep\" segfaults, but we should test our own code to see if it\nsegfaults.\n\nThe scalar code is code we ship and test, so we should use the helper\nthat doesn't hide a segfault.\n\nI don't understand why you wouldn't think that's the obvious fix here,\nadding \"scalar\" to that whitelist is a one-line fix, and clearly yields\na more useful end result than a test silently hiding segfaults.\n"},{"id":"455690","messageId":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v5.git.1652984283.gitgitgadget@gmail.com","subject":"[PATCH v6 0/7] scalar: implement the subcommand \"diagnose\"","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:09Z","receivedAt":"2022-05-21T15:08:24Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Over the course of the years, we developed a sub-command that gathers\ndiagnostic data into a .zip file that can then be attached to bug reports.\nThis sub-command turned out to be very useful in helping Scalar developers\nidentify and fix issues.\n\nChanges since v5:\n\n * Reworded the missed mentions of the old name of the --add-virtual-file\n   option (thanks René!).\n * Renamed misleading variable name from $QUOTED to $PATHNAME (thanks\n   Junio!).\n\nChanges since v4:\n\n * Squashed in Junio's suggested fixups\n * Renamed the option from --add-file-with-content=<name>:<content> to\n   --add-virtual-file=<name>:<content>\n * Fixed one instance where I had used error() instead of error_errno().\n\nChanges since v3:\n\n * We're now using unquote_c_style() instead of rolling our own unquoter.\n * Fixed the added regression test.\n * As pointed out by Scalar's Functional Tests, the\n   add_directory_to_archiver() function should not fail when scalar diagnose\n   encounters FSMonitor's Unix socket, but only warn instead.\n * Related: add_directory_to_archiver() needs to propagate errors from\n   processing subdirectories so that the top-level call returns an error,\n   too.\n\nChanges since v2:\n\n * Clarified in the commit message what the biggest benefit of\n   --add-file-with-content is.\n * The <path> part of the -add-file-with-content argument can now contain\n   colons. To do this, the path needs to start and end in double-quote\n   characters (which are stripped), and the backslash serves as escape\n   character in that case (to allow the path to contain both colons and\n   double-quotes).\n * Fixed incorrect grammar.\n * Instead of strcmp(<what-we-don't-want>), we now say\n   !strcmp(<what-we-want>).\n * The help text for --add-file-with-content was improved a tiny bit.\n * Adjusted the commit message that still talked about spawning plenty of\n   processes and about a throw-away repository for the sake of generating a\n   .zip file.\n * Simplified the code that shows the diagnostics and adds them to the .zip\n   file.\n * The final message that reports that the archive is complete is now\n   printed to stderr instead of stdout.\n\nChanges since v1:\n\n * Instead of creating a throw-away repository, staging the contents of the\n   .zip file and then using git write-tree and git archive to write the .zip\n   file, the patch series now introduces a new option to git archive and\n   uses write_archive() directly (avoiding any separate process).\n * Since the command avoids separate processes, it is now blazing fast on\n   Windows, and I dropped the spinner() function because it's no longer\n   needed.\n * While reworking the test case, I noticed that scalar [...] <enlistment>\n   failed to verify that the specified directory exists, and would happily\n   \"traverse to its parent directory\" on its quest to find a Scalar\n   enlistment. That is of course incorrect, and has been fixed as a \"while\n   at it\" sort of preparatory commit.\n * I had forgotten to sign off on all the commits, which has been fixed.\n * Instead of some \"home-grown\" readdir()-based function, the code now uses\n   for_each_file_in_pack_dir() to look through the pack directories.\n * If any alternates are configured, their pack directories are now included\n   in the output.\n * The commit message that might be interpreted to promise information about\n   large loose files has been corrected to no longer promise that.\n * The test cases have been adjusted to test a little bit more (e.g.\n   verifying that specific paths are mentioned in the output, instead of\n   merely verifying that the output is non-empty).\n\nJohannes Schindelin (5):\n  archive: optionally add \"virtual\" files\n  archive --add-virtual-file: allow paths containing colons\n  scalar: validate the optional enlistment argument\n  Implement `scalar diagnose`\n  scalar diagnose: include disk space information\n\nMatthew John Cheetham (2):\n  scalar: teach `diagnose` to gather packfile info\n  scalar: teach `diagnose` to gather loose objects information\n\n Documentation/git-archive.txt    |  17 ++\n archive.c                        |  63 ++++++-\n contrib/scalar/scalar.c          | 292 ++++++++++++++++++++++++++++++-\n contrib/scalar/scalar.txt        |  12 ++\n contrib/scalar/t/t9099-scalar.sh |  27 +++\n t/t5003-archive-zip.sh           |  20 +++\n 6 files changed, 421 insertions(+), 10 deletions(-)\n\n\nbase-commit: ddc35d833dd6f9e8946b09cecd3311b8aa18d295\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-1128%2Fdscho%2Fscalar-diagnose-v6\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-1128/dscho/scalar-diagnose-v6\nPull-Request: https://github.com/gitgitgadget/git/pull/1128\n\nRange-diff vs v5:\n\n 1:  42e73fb0aac ! 1:  0005cfae31d archive: optionally add \"virtual\" files\n     @@ Metadata\n       ## Commit message ##\n          archive: optionally add \"virtual\" files\n      \n     -    With the `--add-file-with-content=<path>:<content>` option, `git\n     -    archive` now supports use cases where relatively trivial files need to\n     -    be added that do not exist on disk.\n     +    With the `--add-virtual-file=<path>:<content>` option, `git archive` now\n     +    supports use cases where relatively trivial files need to be added that\n     +    do not exist on disk.\n      \n          This will allow us to generate `.zip` files with generated content,\n          without having to add said content to the object database and without\n 2:  b5ebd61066a ! 2:  7eebcf27b45 archive --add-file-with-contents: allow paths containing colons\n     @@ Metadata\n      Author: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n      \n       ## Commit message ##\n     -    archive --add-file-with-contents: allow paths containing colons\n     +    archive --add-virtual-file: allow paths containing colons\n      \n          By allowing the path to be enclosed in double-quotes, we can avoid\n          the limitation that paths cannot contain colons.\n     @@ t/t5003-archive-zip.sh: check_zip with_untracked\n       test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n      +\tif test_have_prereq FUNNYNAMES\n      +\tthen\n     -+\t\tQUOTED=quoted:colon\n     ++\t\tPATHNAME=quoted:colon\n      +\telse\n     -+\t\tQUOTED=quoted\n     ++\t\tPATHNAME=quoted\n      +\tfi &&\n       \tgit archive --format=zip >with_file_with_content.zip \\\n     -+\t\t--add-virtual-file=\\\"$QUOTED\\\": \\\n     ++\t\t--add-virtual-file=\\\"$PATHNAME\\\": \\\n       \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n       \ttest_when_finished \"rm -rf tmp-unpack\" &&\n       \tmkdir tmp-unpack && (\n       \t\tcd tmp-unpack &&\n       \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n       \t\ttest_path_is_file hello &&\n     -+\t\ttest_path_is_file $QUOTED &&\n     ++\t\ttest_path_is_file $PATHNAME &&\n       \t\ttest world = $(cat hello)\n       \t)\n       '\n 3:  f1ba69c02d7 = 3:  ca83ddd5eed scalar: validate the optional enlistment argument\n 4:  3fb90194744 = 4:  89c13a45e00 Implement `scalar diagnose`\n 5:  2e645b08a9e = 5:  8ffbaad3086 scalar diagnose: include disk space information\n 6:  0fa20d73750 = 6:  15cd7f17896 scalar: teach `diagnose` to gather packfile info\n 7:  62e173b47cf = 7:  a4a74d5ef58 scalar: teach `diagnose` to gather loose objects information\n\n-- \ngitgitgadget\n"},{"id":"455691","messageId":"0005cfae31d52a157d4df5ba3db9f9f5b2167ddc.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:10Z","receivedAt":"2022-05-21T15:08:26Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-virtual-file=<path>:<content>` option, `git archive` now\nsupports use cases where relatively trivial files need to be added that\ndo not exist on disk.\n\nThis will allow us to generate `.zip` files with generated content,\nwithout having to add said content to the object database and without\nhaving to write it out to disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 11 ++++++++\n archive.c                     | 53 +++++++++++++++++++++++++++++------\n t/t5003-archive-zip.sh        | 12 ++++++++\n 3 files changed, 68 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a7834..893cb1075bf 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -61,6 +61,17 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n \n+--add-virtual-file=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value for `--prefix` (if any) and the\n+\tbasename of <file>.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject to platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex a3bbb091256..d20e16fa819 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n \t\tstrbuf_addstr(&path_in_archive, basename(path));\n \n \t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\tif (info->content)\n+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t  info->stat.st_mode,\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\telse if (strbuf_read_file(&content, path,\n+\t\t\t\t\t  info->stat.st_size) < 0)\n \t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n \t\telse\n \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n@@ -493,6 +500,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +522,40 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (!strcmp(opt->long_name, \"add-file\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t} else {\n+\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +588,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-virtual-file\", args,\n+\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex 1e6d18b140e..ebc26e89a9b 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \ngitgitgadget\n\n"},{"id":"455692","messageId":"ca83ddd5eed8ec9946d36ec0289ad41c2837bdc8.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:12Z","receivedAt":"2022-05-21T15:08:30Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 1ce9c2b00e8..00dcd4b50ef 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2e1502ad45e..9d83fdf25e8 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -85,4 +85,9 @@ test_expect_success 'scalar delete with enlistment' '\n \ttest_path_is_missing cloned\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455693","messageId":"7eebcf27b45eb13541d4abae70a374a0e35ab6b8.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:11Z","receivedAt":"2022-05-21T15:08:34Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nBy allowing the path to be enclosed in double-quotes, we can avoid\nthe limitation that paths cannot contain colons.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/git-archive.txt | 14 ++++++++++----\n archive.c                     | 30 ++++++++++++++++++++----------\n t/t5003-archive-zip.sh        |  8 ++++++++\n 3 files changed, 38 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 893cb1075bf..54de945a84e 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -67,10 +67,16 @@ OPTIONS\n \tby concatenating the value for `--prefix` (if any) and the\n \tbasename of <file>.\n +\n-The `<path>` cannot contain any colon, the file mode is limited to\n-a regular file, and the option may be subject to platform-dependent\n-command-line limits. For non-trivial cases, write an untracked file\n-and use `--add-file` instead.\n+The `<path>` argument can start and end with a literal double-quote\n+character; The contained file name is interpreted as a C-style string,\n+i.e. the backslash is interpreted as escape character. The path must\n+be quoted if it contains a colon, to avoid the colon from being\n+misinterpreted as the separator between the path and the contents, or\n+if the path begins or ends with a double-quote character.\n++\n+The file mode is limited to a regular file, and the option may be\n+subject to platform-dependent command-line limits. For non-trivial\n+cases, write an untracked file and use `--add-file` instead.\n \n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\ndiff --git a/archive.c b/archive.c\nindex d20e16fa819..b7756b91200 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -9,6 +9,7 @@\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n #include \"dir.h\"\n+#include \"quote.h\"\n \n static char const * const archive_usage[] = {\n \tN_(\"git archive [<options>] <tree-ish> [<path>...]\"),\n@@ -533,22 +534,31 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n \t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n-\t\tconst char *colon = strchr(arg, ':');\n-\t\tchar *p;\n+\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\tconst char *p = arg;\n+\n+\t\tif (*p != '\"')\n+\t\t\tp = strchr(p, ':');\n+\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n+\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n \n-\t\tif (!colon)\n+\t\tif (!p || *p != ':')\n \t\t\tdie(_(\"missing colon: '%s'\"), arg);\n \n-\t\tp = xstrndup(arg, colon - arg);\n-\t\tif (!args->prefix)\n-\t\t\tpath = p;\n-\t\telse {\n-\t\t\tpath = prefix_filename(args->prefix, p);\n-\t\t\tfree(p);\n+\t\tif (p == arg)\n+\t\t\tdie(_(\"empty file name: '%s'\"), arg);\n+\n+\t\tpath = buf.len ?\n+\t\t\tstrbuf_detach(&buf, NULL) : xstrndup(arg, p - arg);\n+\n+\t\tif (args->prefix) {\n+\t\t\tchar *save = path;\n+\t\t\tpath = prefix_filename(args->prefix, path);\n+\t\t\tfree(save);\n \t\t}\n \t\tmemset(&info->stat, 0, sizeof(info->stat));\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n-\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->content = xstrdup(p + 1);\n \t\tinfo->stat.st_size = strlen(info->content);\n \t} else {\n \t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex ebc26e89a9b..3a5a052e8ce 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -207,13 +207,21 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tif test_have_prereq FUNNYNAMES\n+\tthen\n+\t\tPATHNAME=quoted:colon\n+\telse\n+\t\tPATHNAME=quoted\n+\tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=\\\"$PATHNAME\\\": \\\n \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n+\t\ttest_path_is_file $PATHNAME &&\n \t\ttest world = $(cat hello)\n \t)\n '\n-- \ngitgitgadget\n\n"},{"id":"455694","messageId":"15cd7f1789601662c0db979e4c12f13c632efb3c.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 6/7] scalar: teach `diagnose` to gather packfile info","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:15Z","receivedAt":"2022-05-21T15:08:42Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 0a9e25a57f8..d302c27e114 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -595,6 +596,29 @@ cleanup:\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -657,6 +681,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n \t\t     (int)buf.len, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 934b2485d91..3dd5650cceb 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,6 +93,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -102,7 +104,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \ngitgitgadget\n\n"},{"id":"455696","messageId":"89c13a45e00deb1df45113a9249d3a125c7b4fee.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 4/7] Implement `scalar diagnose`","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:13Z","receivedAt":"2022-05-21T15:08:46Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` machinery: we write out a\n`.zip` of the empty try, augmented by a couple files that are added via\nthe `--add-file*` options. We are careful trying not to modify the\ncurrent repository in any way lest the very circumstances that required\n`scalar diagnose` to be run are changed by the `diagnose` run itself.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 144 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 170 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 00dcd4b50ef..53213f9a3b9 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -261,6 +262,47 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error_errno(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\twarning(_(\"skipping '%s', which is neither file nor \"\n+\t\t\t\t  \"directory\"), buf.buf);\n+\t\telse if (recurse &&\n+\t\t\t add_directory_to_archiver(archiver_args,\n+\t\t\t\t\t\t   buf.buf, recurse) < 0)\n+\t\t\tres = -1;\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -501,6 +543,107 @@ cleanup:\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\twrite_or_die(stdout_fd, buf.buf, buf.len);\n+\tstrvec_pushf(&archiver_args,\n+\t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n+\t\t     (int)buf.len, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tfprintf(stderr, \"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -802,6 +945,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex f416d637289..22583fe046e 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -129,6 +130,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 9d83fdf25e8..6802d317258 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -90,4 +90,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out 2>err &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"455695","messageId":"8ffbaad30869ae03e8ba0b1eae4c23aa7d83759e.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 5/7] scalar diagnose: include disk space information","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:14Z","receivedAt":"2022-05-21T15:08:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 53213f9a3b9..0a9e25a57f8 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -303,6 +303,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -599,6 +651,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \twrite_or_die(stdout_fd, buf.buf, buf.len);\n \tstrvec_pushf(&archiver_args,\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 6802d317258..934b2485d91 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -94,6 +94,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out 2>err &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \ngitgitgadget\n\n"},{"id":"455697","messageId":"a4a74d5ef58f9c80fd8f53003f34958e9f7df2a4.1653145696.git.gitgitgadget@gmail.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6 7/7] scalar: teach `diagnose` to gather loose objects information","fromName":"Matthew John Cheetham via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2022-05-21T15:08:16Z","receivedAt":"2022-05-21T15:08:52Z","isPatch":true,"sender":{"key":"mjcheetham@outlook.com","avatar":"https://avatars.githubusercontent.com/u/5658207?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex d302c27e114..0c278681758 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -619,6 +619,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -687,6 +741,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 3dd5650cceb..72023a1ca1d 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -95,6 +95,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -106,7 +107,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \ngitgitgadget\n"},{"id":"455729","messageId":"xmqqleuuazew.fsf@gitster.g","threadId":"57313","inReplyTo":"220521.86leuv199g.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-22T05:50:47Z","receivedAt":"2022-05-22T05:51:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> Scalar is not (yet?) a Git command.\n>\n> \"test-tool\" isn't \"git\" either, so I think this argument is a\n> non-starter.\n>\n> As the documentation for \"test_must_fail\" notes the distinction is\n> whether something is \"system-supplied\". I.e. we're not going to test\n> whether \"grep\" segfaults, but we should test our own code to see if it\n> segfaults.\n>\n> The scalar code is code we ship and test, so we should use the helper\n> that doesn't hide a segfault.\n>\n> I don't understand why you wouldn't think that's the obvious fix here,\n> adding \"scalar\" to that whitelist is a one-line fix, and clearly yields\n> a more useful end result than a test silently hiding segfaults.\n\nFWIW, I don't, either.\n\n"},{"id":"455893","messageId":"nycvar.QRO.7.76.6.2205241423260.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"xmqqleuuazew.fsf@gitster.g","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-24T12:25:12Z","receivedAt":"2022-05-24T12:25:32Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Sat, 21 May 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n> >> Scalar is not (yet?) a Git command.\n> >\n> > \"test-tool\" isn't \"git\" either, so I think this argument is a\n> > non-starter.\n> >\n> > As the documentation for \"test_must_fail\" notes the distinction is\n> > whether something is \"system-supplied\". I.e. we're not going to test\n> > whether \"grep\" segfaults, but we should test our own code to see if it\n> > segfaults.\n> >\n> > The scalar code is code we ship and test, so we should use the helper\n> > that doesn't hide a segfault.\n> >\n> > I don't understand why you wouldn't think that's the obvious fix here,\n> > adding \"scalar\" to that whitelist is a one-line fix, and clearly yields\n> > a more useful end result than a test silently hiding segfaults.\n>\n> FWIW, I don't, either.\n\nBecause we are still talking about code that lives as much encapsulated\ninside `contrib/scalar/` as possible.\n\nThe `! scalar` call is in `contrib/scalar/t/t9099-scalar.sh`.\n\nTo make it work with Git's test suite, you would have to bleed an\nimplementation detail of something inside `contrib/` into\n`t/test-lib-functions.sh`.\n\nNot what we want, at this stage.\n\nCiao,\nDscho\n"},{"id":"455921","messageId":"220524.86h75ex01s.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2205241423260.352@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-24T18:11:57Z","receivedAt":"2022-05-24T18:23:01Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, May 24 2022, Johannes Schindelin wrote:\n\n> Hi Junio,\n>\n> On Sat, 21 May 2022, Junio C Hamano wrote:\n>\n>> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>>\n>> >> Scalar is not (yet?) a Git command.\n>> >\n>> > \"test-tool\" isn't \"git\" either, so I think this argument is a\n>> > non-starter.\n>> >\n>> > As the documentation for \"test_must_fail\" notes the distinction is\n>> > whether something is \"system-supplied\". I.e. we're not going to test\n>> > whether \"grep\" segfaults, but we should test our own code to see if it\n>> > segfaults.\n>> >\n>> > The scalar code is code we ship and test, so we should use the helper\n>> > that doesn't hide a segfault.\n>> >\n>> > I don't understand why you wouldn't think that's the obvious fix here,\n>> > adding \"scalar\" to that whitelist is a one-line fix, and clearly yields\n>> > a more useful end result than a test silently hiding segfaults.\n>>\n>> FWIW, I don't, either.\n>\n> Because we are still talking about code that lives as much encapsulated\n> inside `contrib/scalar/` as possible.\n>\n> The `! scalar` call is in `contrib/scalar/t/t9099-scalar.sh`.\n>\n> To make it work with Git's test suite, you would have to bleed an\n> implementation detail of something inside `contrib/` into\n> `t/test-lib-functions.sh`.\n\nThe \"scalar\" command is already built by the top-level Makefile, so I\ndon't think the distinction you're trying to maintain here even exists\nin practice.\n\nI.e. if we ran with this strict reasoning then surely \"scalar\" belongs\non there just as much as \"test-tool\" does.\n\nBoth are built by our main build process, and thus should have\ncorresponding adjustments in our main test code, just as is already the\ncase for both \"git\" and \"test-tool\".\n\nBut even if that wasn't the case I'd still be of the view that we should\nadd \"scalar\" to that list.\n\nIt's just a matter of potential time sinks in the future. If we\nintroduce a hidden segfault in the scalar code and don't notice for some\ntime because we're using that test pattern that's going to suck, and\nlikely to waste a lot of time. We might even ship a broken command to\nusers.\n\nWhereas having \"scalar\" on that list is going to be a relatively easy\nmatter of grepping and doing some boilerplate changes if and when we\never \"git rm\" it entirely, or \"promote it\" from contrib or whatever.\n\nI also think that just getting rid of that whitelist entirely is an\nacceptable solution. Perhaps it's just being overzealous in forbidding\neverything except \"git\", we should still not use it for the likes of\n\"grep\", but we could just leave that to the documentation.\n\nBut I suspect Junio would disagree with that, so in lieu of that ...\n"},{"id":"455936","messageId":"xmqqsfoyhgqe.fsf@gitster.g","threadId":"57313","inReplyTo":"220524.86h75ex01s.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-24T19:29:13Z","receivedAt":"2022-05-24T19:29:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> Both are built by our main build process, and thus should have\n> corresponding adjustments in our main test code, just as is already the\n> case for both \"git\" and \"test-tool\".\n>\n> But even if that wasn't the case I'd still be of the view that we should\n> add \"scalar\" to that list.\n>\n> It's just a matter of potential time sinks in the future. If we\n> introduce a hidden segfault in the scalar code and don't notice for some\n> time because we're using that test pattern that's going to suck, and\n> likely to waste a lot of time. We might even ship a broken command to\n> users.\n>\n> Whereas having \"scalar\" on that list is going to be a relatively easy\n> matter of grepping and doing some boilerplate changes if and when we\n> ever \"git rm\" it entirely, or \"promote it\" from contrib or whatever.\n\nIn addition, it already is an actual time sink that causes us send a\nlot more bytes back and forth than the number of bytes necessary to\nsend a reroll that adds one liner to the same step.\n\n> I also think that just getting rid of that whitelist entirely is an\n> acceptable solution. Perhaps it's just being overzealous in forbidding\n> everything except \"git\", we should still not use it for the likes of\n> \"grep\", but we could just leave that to the documentation.\n\nIt indeed is tempting entry into a slippery slope, and I'd see it as\na change bigger than we could comfortably make as a \"while at it\"\nchange.\n\nWe can stop arguing and instead send in a reroll that squashes in\nsomething like this, which shouldn't be controversial, I would say.\n\n t/test-lib-functions.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git i/t/test-lib-functions.sh w/t/test-lib-functions.sh\nindex 93c03380d4..8899eaabed 100644\n--- i/t/test-lib-functions.sh\n+++ w/t/test-lib-functions.sh\n@@ -1106,7 +1106,7 @@ test_must_fail_acceptable () {\n \tfi\n \n \tcase \"$1\" in\n-\tgit|__git*|test-tool|test_terminal)\n+\tgit|__git*|scalar|test-tool|test_terminal)\n \t\treturn 0\n \t\t;;\n \t*)\n\n\n\n"},{"id":"456042","messageId":"nycvar.QRO.7.76.6.2205251224560.352@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"xmqqsfoyhgqe.fsf@gitster.g","subject":"Re: [PATCH v4 3/7] scalar: validate the optional enlistment argument","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-25T10:31:14Z","receivedAt":"2022-05-25T10:31:39Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Tue, 24 May 2022, Junio C Hamano wrote:\n\n> We can stop arguing and instead send in a reroll that squashes in\n> something like this, which shouldn't be controversial, I would say.\n>\n>  t/test-lib-functions.sh | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git i/t/test-lib-functions.sh w/t/test-lib-functions.sh\n> index 93c03380d4..8899eaabed 100644\n> --- i/t/test-lib-functions.sh\n> +++ w/t/test-lib-functions.sh\n> @@ -1106,7 +1106,7 @@ test_must_fail_acceptable () {\n>  \tfi\n>\n>  \tcase \"$1\" in\n> -\tgit|__git*|test-tool|test_terminal)\n> +\tgit|__git*|scalar|test-tool|test_terminal)\n>  \t\treturn 0\n>  \t\t;;\n>  \t*)\n>\n>\n>\n>\n\nIt is still wrong to adjust Git's test suite for a user that is not part\nof Git proper. But if your pragmatism says that this is the only way we\ncan venture on to more productive venues, I won't argue against that :-)\n\nCiao,\nDscho\n"},{"id":"456114","messageId":"xmqq5ylt7473.fsf@gitster.g","threadId":"57313","inReplyTo":"7eebcf27b45eb13541d4abae70a374a0e35ab6b8.1653145696.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-25T20:22:24Z","receivedAt":"2022-05-25T20:22:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> +\tif test_have_prereq FUNNYNAMES\n> +\tthen\n> +\t\tPATHNAME=quoted:colon\n> +\telse\n> +\t\tPATHNAME=quoted\n> +\tfi &&\n>  \tgit archive --format=zip >with_file_with_content.zip \\\n> +\t\t--add-virtual-file=\\\"$PATHNAME\\\": \\\n\nThe name is better, but this still limits what can be in PATHNAME.\n\nWrite either one of these:\n\n\t\t--add-virtual-file=\"\\\"$PATHNAME\\\":\" \\\n\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n\nto signal the intention better to future readers.  We are showing an\nexplicit dq-pair we want to pass to the c-unquote machinery, and we\nare showing that we are not being unnecessarily loose by protecting\nthe string from getting word split.\n\nEither is fine, but leaving it unquoted is not.\n\n> +\t\ttest_path_is_file $PATHNAME &&\n\nDitto.  There is no reason to forbid future developers from futzing\nthe test to include space in the PATHNAME variable.  \n\nIOW, I want us to be better than saying\n\n    I know there is no $IFS whitespace now because I just wrote it.\n    Because I do not think there is any need to test with a string\n    with whitespace in it, I will leave the variable unquoted.\n    Anybody who changes the variable and breaks this assumption have\n    only themselves to blame for breaking the tests.  It is not my\n    fault and it is not my problem.\n\nwhich is the signal our readers would get from this patch (I would,\nif I were reading this commit as a third-party), especially once\nthey become aware of the fact that this exact issue was already\npointed out during the review discussion.\n\nUsing double-quote appropriately sends a strong signal to reviewers\nand future developers that we care about details.\n\nA valid alternative is to write the assumption out where we\ncurrently assign to PATHNAME.\n\n\t# The PATHNAME variable is used without quote in the code\n\t# below for such and such reasons, so you cannot use a $IFS\n\t# whitespace in it.\n\tif test_have_prereq FUNNYNAMES\n\tthen\n\t\t...\n\nIf the \"defensive\" measure that is necessary to avoid a limitation\nis too onerous, such an approach may be very much more preferrable\nthan preparing for future changes.  \"for such and such reasons\" is\na good place to justify why we avoid unnecessarily complex defensive\nmeasure and restrict future changes in the documented way.\n\nBut in _this_ particular case, the \"defensive\" measure necessary is\nmerely just to quote the shell variables properly, which nobody\nsensible would say too onerous.  I couldn't come up with anything\nremotely plausible to fill \"for such and such reasons\" myself when I\ntried to justify leaving the variables unquoted.\n\nRegardless of the quoting issue, we probably want to comment on what\nvalue exactly is in PATHNAME before the assignment, by the way.\n\nE.g.\n\n\t# The PATHNAME variable holds a filename encoded like a\n\t# string constant in C language (e.g. \"\\060\" is digit \"0\")\n\tif test_have_prereq FUNNYNAMES\n\tthen\n\t\tPATHNAME=quoted:colon:\\\\060zero\n\telse\n\t\tPATHNAME=quoted\\\\060zero\n\tfi\n\nThat would not just protect only one aspect (i.e. we can pass a\ncolon into the resulting filename) this change but the path goes\nthrough the c-unquoting rules.\n\nThanks.\n"},{"id":"456119","messageId":"xmqqfskx5ndd.fsf@gitster.g","threadId":"57313","inReplyTo":"0005cfae31d52a157d4df5ba3db9f9f5b2167ddc.1653145696.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-25T21:11:10Z","receivedAt":"2022-05-25T21:11:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n> @@ -61,6 +61,17 @@ OPTIONS\n>  \tby concatenating the value for `--prefix` (if any) and the\n>  \tbasename of <file>.\n>  \n> +--add-virtual-file=<path>:<content>::\n> +\tAdd the specified contents to the archive.  Can be repeated to add\n> +\tmultiple files.  The path of the file in the archive is built\n> +\tby concatenating the value for `--prefix` (if any) and the\n> +\tbasename of <file>.\n\nThis sentence was copy-pasted from --add-file without adjusting.\nThere is no <file>; this new feature gives <path>.\n\nAlso, I suspect that the feature is losing end-user supplied\ninformation without a good reason.  --add-file=<file> may have\nprepared an input in a randomly named temporary directory and it\nwould make quite a lot of sense to strip the leading directory\ncomponents from <file> and use only the basename part.  But the\n<path> given to \"--add-virtual-file\" does not refer to anything on\nthe filesystem.  Its ONLY use is to be used as the path in the\narchive to store the content.  There is no justification why we\nwould discard the leading path components from it.  I am not\ndecided, but I am inclined to say that we should not honor\n\"--prefix\".\n\n   $ git archive --prefix=2.36.0 v2.36.0\n\nwould be a way to create a single directory and put everything in\nthe tree-ish in there, but there probably are cases where the user\nof an \"extra file\" feature wants to add untracked cruft _in_ that\ndirectory, and there are other cases where an extra file wants to go\nto the top-level next to the 2.36.0 directory.  A user can use the\nsame string as --prefix=<base> in front of <path> if the extra file\nshould go next to the top-level of the tree-ish, or without such\nprefixing to place the extra file at the top-level.\n\nHence\n\n\tAdd the specified contents to the archive.  Can be repeated\n\tto add multiple files.  `<path>` is used as the path of the\n\tfile in the archive.\n\t\nwould be what I would expect in a version of this feature that is\nreasonably designed.\n\n> ++\n> +The `<path>` cannot contain any colon, the file mode is limited to\n> +a regular file, and the option may be subject to platform-dependent\n> +command-line limits. For non-trivial cases, write an untracked file\n> +and use `--add-file` instead.\n\nOK.\n\n> diff --git a/archive.c b/archive.c\n> index a3bbb091256..d20e16fa819 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n>  struct extra_file_info {\n>  \tchar *base;\n>  \tstruct stat stat;\n> +\tvoid *content;\n>  };\n>  \n>  int write_archive_entries(struct archiver_args *args,\n> @@ -337,7 +338,13 @@ int write_archive_entries(struct archiver_args *args,\n>  \t\tstrbuf_addstr(&path_in_archive, basename(path));\n>  \n>  \t\tstrbuf_reset(&content);\n> -\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n> +\t\tif (info->content)\n\nWe ended up with the problematic \"leading <path> components are\ndiscarded\" design only because the implementation reuses the logic\npath_in_archive computation (the last line is seen in precontext),\nwhich is a bit unfortunate.  I think we could rewrite the inside of\nthat \"for each extra file\" loop like so, instead:\n\n\tfor (i = 0; i < args->extra_files.nr; i++) {\n\t\tstruct string_list_item *item = args->extra_files.items + i;\n\t\tchar *path = item->string;\n\t\tstruct extra_file_info *info = item->util;\n\n\t\tput_be64(fake_oid.hash, i + 1);\n\n\t\tif (!info->content) {\n\t\t\tstrbuf_reset(&path_in_archive);\n\t\t\tif (info->base)\n\t\t\t\tstrbuf_addstr(&path_in_archive, info->base);\n\t\t\tstrbuf_addstr(&path_in_archive, basename(path));\n\n\t\t\tstrbuf_reset(&content);\n\t\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n\t\t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n\t\t\telse\n\t\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n\t\t\t\t\t\t  path_in_archive.len,\n\t\t\t\t\t\t  info->stat.st_mode,\n\t\t\t\t\t\t  content.buf, content.len);\n\t\t} else {\n\t\t\terr = write_entry(args, &fake_oid,\n\t\t\t\t\t  path, strlen(path),\n\t\t\t\t\t  info->stat.st_mode,\n\t\t\t\t\t  info->content, info->stat.st_size);\n\t\t}\n\n\t\tif (err)\n\t\t\tbreak;\n\t}\n\nThe first half is the original code for \"--add-file\", which clears\ninfo->content to NULL.  We mangle the filename to come up with the\nname in the archive (i.e. take basename and prefix with info->base).\n\nThe \"else\" side is the new code.  \"--add-virtual-file\" has the\n\"<path>\" thing in item->string, and info has the contents, so we\njust write it out.\n"},{"id":"456124","messageId":"xmqq4k1d5lwj.fsf@gitster.g","threadId":"57313","inReplyTo":"xmqq5ylt7473.fsf@gitster.g","subject":"Re: [PATCH v6 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-25T21:42:52Z","receivedAt":"2022-05-25T21:42:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> But in _this_ particular case, the \"defensive\" measure necessary is\n> merely just to quote the shell variables properly, which nobody\n> sensible would say too onerous.  I couldn't come up with anything\n> remotely plausible to fill \"for such and such reasons\" myself when I\n> tried to justify leaving the variables unquoted.\n>\n> Regardless of the quoting issue, we probably want to comment on what\n> value exactly is in PATHNAME before the assignment, by the way.\n>\n> E.g.\n>\n> \t# The PATHNAME variable holds a filename encoded like a\n> \t# string constant in C language (e.g. \"\\060\" is digit \"0\")\n> \tif test_have_prereq FUNNYNAMES\n> \tthen\n> \t\tPATHNAME=quoted:colon:\\\\060zero\n> \telse\n> \t\tPATHNAME=quoted\\\\060zero\n> \tfi\n>\n> That would not just protect only one aspect (i.e. we can pass a\n> colon into the resulting filename) this change but the path goes\n> through the c-unquoting rules.\n\nActually, I _think_ that pushes us beyond the \"reasonably defensive\nfor the current need\".  We'd need to prepare how the pathname is\nexpected to be unquoted for the later test\n\n\ttest_path_is_file \"$PATHNAME\"\n\nto work.  So here is what I queued as a fixup for this step on top\nof the series.\n\n t/t5003-archive-zip.sh | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git c/t/t5003-archive-zip.sh w/t/t5003-archive-zip.sh\nindex 3a5a052e8c..6addb6c684 100755\n--- c/t/t5003-archive-zip.sh\n+++ w/t/t5003-archive-zip.sh\n@@ -209,19 +209,19 @@ check_added with_untracked untracked untracked\n test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n \tif test_have_prereq FUNNYNAMES\n \tthen\n-\t\tPATHNAME=quoted:colon\n+\t\tPATHNAME=\"pathname with : colon\"\n \telse\n-\t\tPATHNAME=quoted\n+\t\tPATHNAME=\"pathname without colon\"\n \tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n-\t\t--add-virtual-file=\\\"$PATHNAME\\\": \\\n+\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n-\t\ttest_path_is_file $PATHNAME &&\n+\t\ttest_path_is_file \"$PATHNAME\" &&\n \t\ttest world = $(cat hello)\n \t)\n '\n"},{"id":"456128","messageId":"xmqqpmk144yq.fsf@gitster.g","threadId":"57313","inReplyTo":"xmqq4k1d5lwj.fsf@gitster.g","subject":"Re: [PATCH v6 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-25T22:34:05Z","receivedAt":"2022-05-25T22:34:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n>> \t# The PATHNAME variable holds a filename encoded like a\n>> \t# string constant in C language (e.g. \"\\060\" is digit \"0\")\n>> \tif test_have_prereq FUNNYNAMES\n>> \tthen\n>> \t\tPATHNAME=quoted:colon:\\\\060zero\n>> \t...\n> Actually, I _think_ that pushes us beyond the \"reasonably defensive\n> for the current need\".  We'd need to prepare how the pathname is\n> expected to be unquoted for the later test\n>\n> \ttest_path_is_file \"$PATHNAME\"\n>\n> to work.\n\nIOW, I would need to add a new test-tool (attached) and then start\nthis test like so:\n\n\tif ...\n\tthen\n\t\tPATHNAME=quoted:colon:\\\\060zero\n\telse\n\t\tPATHNAME=quoted\\\\060zero\n\tfi\n\tUQPATHNAME=$(test-tool unquote-c-style \\\"\"$PATHNAME\"\\\")\n\nand change the last test to\n\n\ttest_path_is_file \"$UQPATHNAME\"\n\nif we really wanted to test that the the PATHNAME is treated as a\nc-style quoted string.\n\nI am on the fence.  We do not have an immediate need, in the sense\nthat nobody needs to encode \"0\" as \"\\060\" and trigger the unquote\ncodepath in real life.  But it does feel prudent to make sure we can\ngrok C-quoted pathname as we claim in the documentation.\n\nAnd the resulting change to the test does not look _too_ bad (and\nthe new test-tool certainly does not hurt, either).\n\nSo...\n\n\n Makefile               |  1 +\n t/helper/test-quoted.c | 34 ++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c   |  2 ++\n t/helper/test-tool.h   |  2 ++\n 4 files changed, 39 insertions(+)\n\ndiff --git c/Makefile w/Makefile\nindex 298becd5a5..1d544ad46a 100644\n--- c/Makefile\n+++ w/Makefile\n@@ -749,6 +749,7 @@ TEST_BUILTINS_OBJS += test-pkt-line.o\n TEST_BUILTINS_OBJS += test-prio-queue.o\n TEST_BUILTINS_OBJS += test-proc-receive.o\n TEST_BUILTINS_OBJS += test-progress.o\n+TEST_BUILTINS_OBJS += test-quoted.o\n TEST_BUILTINS_OBJS += test-reach.o\n TEST_BUILTINS_OBJS += test-read-cache.o\n TEST_BUILTINS_OBJS += test-read-graph.o\ndiff --git c/t/helper/test-quoted.c w/t/helper/test-quoted.c\nnew file mode 100644\nindex 0000000000..15baa55e43\n--- /dev/null\n+++ w/t/helper/test-quoted.c\n@@ -0,0 +1,34 @@\n+#include \"test-tool.h\"\n+#include \"cache.h\"\n+#include \"quote.h\"\n+\n+int cmd__unquote_c_style(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\twhile (*++argv) {\n+\t\tconst char *p = *argv;\n+\n+\t\tif (unquote_c_style(&buf, p, &p) < 0)\n+\t\t\terror(\"cannot unquote '%s'\", *argv);\n+\t\telse\n+\t\t\tprintf(\"%s\\n\", buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\treturn 0;\n+}\n+\n+int cmd__quote_c_style(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\twhile (*++argv) {\n+\t\tconst char *p = *argv;\n+\n+\t\tquote_c_style(p, &buf, NULL, 0);\n+\t\tprintf(\"%s\\n\", buf.buf);\n+\t\tstrbuf_reset(&buf);\n+\t}\n+\treturn 0;\n+}\n+\ndiff --git c/t/helper/test-tool.c w/t/helper/test-tool.c\nindex d2eacd302d..5633c98569 100644\n--- c/t/helper/test-tool.c\n+++ w/t/helper/test-tool.c\n@@ -58,6 +58,7 @@ static struct test_cmd cmds[] = {\n \t{ \"prio-queue\", cmd__prio_queue },\n \t{ \"proc-receive\", cmd__proc_receive },\n \t{ \"progress\", cmd__progress },\n+\t{ \"quote-c-style\", cmd__quote_c_style },\n \t{ \"reach\", cmd__reach },\n \t{ \"read-cache\", cmd__read_cache },\n \t{ \"read-graph\", cmd__read_graph },\n@@ -81,6 +82,7 @@ static struct test_cmd cmds[] = {\n \t{ \"submodule-nested-repo-config\", cmd__submodule_nested_repo_config },\n \t{ \"subprocess\", cmd__subprocess },\n \t{ \"trace2\", cmd__trace2 },\n+\t{ \"unquote-c-style\", cmd__unquote_c_style },\n \t{ \"userdiff\", cmd__userdiff },\n \t{ \"urlmatch-normalization\", cmd__urlmatch_normalization },\n \t{ \"xml-encode\", cmd__xml_encode },\ndiff --git c/t/helper/test-tool.h w/t/helper/test-tool.h\nindex 960cc27ef7..f5e8929009 100644\n--- c/t/helper/test-tool.h\n+++ w/t/helper/test-tool.h\n@@ -48,6 +48,7 @@ int cmd__pkt_line(int argc, const char **argv);\n int cmd__prio_queue(int argc, const char **argv);\n int cmd__proc_receive(int argc, const char **argv);\n int cmd__progress(int argc, const char **argv);\n+int cmd__quote_c_style(int argc, const char **argv);\n int cmd__reach(int argc, const char **argv);\n int cmd__read_cache(int argc, const char **argv);\n int cmd__read_graph(int argc, const char **argv);\n@@ -71,6 +72,7 @@ int cmd__submodule_config(int argc, const char **argv);\n int cmd__submodule_nested_repo_config(int argc, const char **argv);\n int cmd__subprocess(int argc, const char **argv);\n int cmd__trace2(int argc, const char **argv);\n+int cmd__unquote_c_style(int argc, const char **argv);\n int cmd__userdiff(int argc, const char **argv);\n int cmd__urlmatch_normalization(int argc, const char **argv);\n int cmd__xml_encode(int argc, const char **argv);\n"},{"id":"456155","messageId":"7815a07a-da2f-d348-4179-6dc5b1d5fee6@web.de","threadId":"57313","inReplyTo":"xmqqfskx5ndd.fsf@gitster.g","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-26T09:09:15Z","receivedAt":"2022-05-26T09:09:49Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 25.05.22 um 23:11 schrieb Junio C Hamano:\n> \"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n> writes:\n>\n>> @@ -61,6 +61,17 @@ OPTIONS\n>>  \tby concatenating the value for `--prefix` (if any) and the\n>>  \tbasename of <file>.\n>>\n>> +--add-virtual-file=<path>:<content>::\n>> +\tAdd the specified contents to the archive.  Can be repeated to add\n>> +\tmultiple files.  The path of the file in the archive is built\n>> +\tby concatenating the value for `--prefix` (if any) and the\n>> +\tbasename of <file>.\n>\n> This sentence was copy-pasted from --add-file without adjusting.\n> There is no <file>; this new feature gives <path>.\n>\n> Also, I suspect that the feature is losing end-user supplied\n> information without a good reason.  --add-file=<file> may have\n> prepared an input in a randomly named temporary directory and it\n> would make quite a lot of sense to strip the leading directory\n> components from <file> and use only the basename part.  But the\n> <path> given to \"--add-virtual-file\" does not refer to anything on\n> the filesystem.  Its ONLY use is to be used as the path in the\n> archive to store the content.  There is no justification why we\n> would discard the leading path components from it.\n\nGood point.\n\n>  I am not\n> decided, but I am inclined to say that we should not honor\n> \"--prefix\".\n>\n>    $ git archive --prefix=2.36.0 v2.36.0\n>\n> would be a way to create a single directory and put everything in\n> the tree-ish in there, but there probably are cases where the user\n> of an \"extra file\" feature wants to add untracked cruft _in_ that\n> directory, and there are other cases where an extra file wants to go\n> to the top-level next to the 2.36.0 directory.  A user can use the\n> same string as --prefix=<base> in front of <path> if the extra file\n> should go next to the top-level of the tree-ish, or without such\n> prefixing to place the extra file at the top-level.\n\nIf the prefix is applied then a prefix-less extra file can by had by\nusing --prefix= or --no-prefix for it and --prefix=... for the tree,\ne.g.:\n\n   $ git archive --add-file=extra --prefix=dir/ v2.36.0\n\nputs \"extra\" at the root and the rest under \"dir\".  The order of\narguments matters here, and the default prefix is the empty string.\n\nSo extra files can be put anywhere even if --prefix is honored.\n\nKeeping the whole path from --add-virtual-file makes sense to me; I\nslightly prefer applying --prefix on top of that for consistency.\n\nRené\n"},{"id":"456192","messageId":"xmqqee0g1aoz.fsf@gitster.g","threadId":"57313","inReplyTo":"7815a07a-da2f-d348-4179-6dc5b1d5fee6@web.de","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-26T17:10:52Z","receivedAt":"2022-05-26T17:11:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> If the prefix is applied then a prefix-less extra file can by had by\n> using --prefix= or --no-prefix for it and --prefix=... for the tree,\n> e.g.:\n>\n>    $ git archive --add-file=extra --prefix=dir/ v2.36.0\n>\n> puts \"extra\" at the root and the rest under \"dir\".  The order of\n> arguments matters here, and the default prefix is the empty string.\n\nThis was the part of the design for the original \"--add-file\" that I\nwas moderately unhappy with.  If \"--add-file\" were the only feature\nthat used \"--prefix\", I wouldn't have been unhappy, but this rule:\n\n        The value of \"--prefix\" most recently seen at the point of\n        \"--add-file\" is prepended.  (By the way, it is not clearly\n        documented what happens when you give multiple prefix and\n        when you give prefix before or after add-file)\n\nmakes the original use of \"--prefix\":\n\n\tThe value given to \"--prefix\" is prepended to each filename\n\tin the archive.  (IOW \"git archive --prefix=git-2.36.0/\n\tv2.36.0\" is a way to prefix each and every path in the\n\ttree-ish with the given prefix)\n\nconfusing.  Does\n\n\tgit archive --prefix=bonus-files/ --add-file=extra v2.36.0\n\nplace the main part of the archive also in bonus-files/ or at the\ntop level?  One reasonable interpretation is \"yes\", if we imagine\nthat each invocation of --add-file will consume and reset the prefix.\nAnother reasonable interpretation is \"no\", if we imagine that the\nprefix last specified will stay around and equally affect both extra\nones and main part of the archive.\n\nUnfortunately what the implmentation does is the latter, and those\nwho want to put the main part of the archive at the top-level must\nadd \"--prefix=''\" at the end (before the tree-ish).\n\nBecause of this potential for confusion ...\n\n> So extra files can be put anywhere even if --prefix is honored.\n>\n> Keeping the whole path from --add-virtual-file makes sense to me; I\n> slightly prefer applying --prefix on top of that for consistency.\n\n... I was hoping that we can releave users from having to worry\nabout the interaction between \"prefix\" and contents coming from\noutside the tree-ish by ignoring the \"prefix\".\n\nBut either is fine by me.\n\nThanks.\n\n\n"},{"id":"456215","messageId":"ed95b26a-2fa3-d1f7-3142-05719a44a8f7@web.de","threadId":"57313","inReplyTo":"xmqqee0g1aoz.fsf@gitster.g","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-26T18:57:49Z","receivedAt":"2022-05-26T18:58:21Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 26.05.22 um 19:10 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>> If the prefix is applied then a prefix-less extra file can by had by\n>> using --prefix= or --no-prefix for it and --prefix=... for the tree,\n>> e.g.:\n>>\n>>    $ git archive --add-file=extra --prefix=dir/ v2.36.0\n>>\n>> puts \"extra\" at the root and the rest under \"dir\".  The order of\n>> arguments matters here, and the default prefix is the empty string.\n>\n> This was the part of the design for the original \"--add-file\" that I\n> was moderately unhappy with.  If \"--add-file\" were the only feature\n> that used \"--prefix\", I wouldn't have been unhappy, but this rule:\n>\n>         The value of \"--prefix\" most recently seen at the point of\n>         \"--add-file\" is prepended.  (By the way, it is not clearly\n>         documented what happens when you give multiple prefix and\n>         when you give prefix before or after add-file)\n\nRegarding documentation: I wonder what's missing; a guess is below.\n\n>\n> makes the original use of \"--prefix\":\n>\n> \tThe value given to \"--prefix\" is prepended to each filename\n> \tin the archive.  (IOW \"git archive --prefix=git-2.36.0/\n> \tv2.36.0\" is a way to prefix each and every path in the\n> \ttree-ish with the given prefix)\n>\n> confusing.  Does\n>\n> \tgit archive --prefix=bonus-files/ --add-file=extra v2.36.0\n>\n> place the main part of the archive also in bonus-files/ or at the\n> top level?  One reasonable interpretation is \"yes\", if we imagine\n> that each invocation of --add-file will consume and reset the prefix.\n> Another reasonable interpretation is \"no\", if we imagine that the\n> prefix last specified will stay around and equally affect both extra\n> ones and main part of the archive.\n>\n> Unfortunately what the implmentation does is the latter, and those\n> who want to put the main part of the archive at the top-level must\n> add \"--prefix=''\" at the end (before the tree-ish).\n\nA one-shot --prefix would be surprising -- usually options keep their\nvalue until they are specified again with a different value or negated\n(--no-...).  That surprise could be documented away by using a\ndifferent name like --next-prefix or --single-use-prefix.  But a\nsub-option to a single option like that would probably be better baked\ninto that option, e.g. allow --add-file=<path_in_archive>:<path_in_fs>.\n\n>\n> Because of this potential for confusion ...\n>\n>> So extra files can be put anywhere even if --prefix is honored.\n>>\n>> Keeping the whole path from --add-virtual-file makes sense to me; I\n>> slightly prefer applying --prefix on top of that for consistency.\n>\n> ... I was hoping that we can releave users from having to worry\n> about the interaction between \"prefix\" and contents coming from\n> outside the tree-ish by ignoring the \"prefix\".\n>\n> But either is fine by me.\n\nThe unusual thing about the current --prefix implementation is that its\ncurrent value is captured along the way instead of just using its\nright-most value.  Not sure ignoring it for one of the three archive\ncontent sources helps.  (Really, it's hard for me to put me in the shoes\nof someone who doesn't know how these options are supposed to be used.)\n\n\n--- >8 ---\nSubject: [PATCH] archive: improve documentation of --prefix\n\nDocument the interaction between --add-file and --prefix by giving an\nexample.\n\nSigned-off-by: René Scharfe <l.s.r@web.de>\n---\n Documentation/git-archive.txt | 14 +++++++++++---\n 1 file changed, 11 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a783..10a48ab5f8 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -49,7 +49,9 @@ OPTIONS\n \tReport progress to stderr.\n\n --prefix=<prefix>/::\n-\tPrepend <prefix>/ to each filename in the archive.\n+\tPrepend <prefix>/ to each filename in the archive.  Can be\n+\tspecified multiple times; the last one seen when reading from\n+\tleft to right is applied.\n\n -o <file>::\n --output=<file>::\n@@ -58,8 +60,8 @@ OPTIONS\n --add-file=<file>::\n \tAdd a non-tracked file to the archive.  Can be repeated to add\n \tmultiple files.  The path of the file in the archive is built\n-\tby concatenating the value for `--prefix` (if any) and the\n-\tbasename of <file>.\n+\tby concatenating the current value for `--prefix` (if any) and\n+\tthe basename of <file>.\n\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n@@ -194,6 +196,12 @@ EXAMPLES\n \tcommit on the current branch. Note that the output format is\n \tinferred by the extension of the output file.\n\n+`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n+\n+\tCreates a tar archive that contains the contents of the latest\n+\tcommit on the current branch with no prefix and the untracked\n+\tfile 'configure' with the prefix 'build/'.\n+\n `git config tar.tar.xz.command \"xz -c\"`::\n\n \tConfigure a \"tar.xz\" format for making LZMA-compressed tarfiles.\n--\n2.35.3\n"},{"id":"456227","messageId":"xmqqfskwxd6j.fsf@gitster.g","threadId":"57313","inReplyTo":"ed95b26a-2fa3-d1f7-3142-05719a44a8f7@web.de","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-26T20:16:04Z","receivedAt":"2022-05-26T20:16:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n> index bc4e76a783..10a48ab5f8 100644\n> --- a/Documentation/git-archive.txt\n> +++ b/Documentation/git-archive.txt\n> @@ -49,7 +49,9 @@ OPTIONS\n>  \tReport progress to stderr.\n>\n>  --prefix=<prefix>/::\n> -\tPrepend <prefix>/ to each filename in the archive.\n> +\tPrepend <prefix>/ to each filename in the archive.  Can be\n> +\tspecified multiple times; the last one seen when reading from\n> +\tleft to right is applied.\n\nThat can be read to mean that we will use C consistently,\n\n$ cmd --prefix=A other-args --prefix=B other-args --prefix=C other-args\n\nwhich was what I am worried to be a source of confusion.\n\n>  -o <file>::\n>  --output=<file>::\n> @@ -58,8 +60,8 @@ OPTIONS\n>  --add-file=<file>::\n>  \tAdd a non-tracked file to the archive.  Can be repeated to add\n>  \tmultiple files.  The path of the file in the archive is built\n> -\tby concatenating the value for `--prefix` (if any) and the\n> -\tbasename of <file>.\n> +\tby concatenating the current value for `--prefix` (if any) and\n> +\tthe basename of <file>.\n\n\"the current value for `--prefix` (if any)\" would work well once we\nsomehow make the reader form a mental model that there is \"the\ncurrent\" for the \"prefix\", which starts with an empty string, and\ngets updated every time the \"--prefix=<prefix>/\" option is given.\n\nSo, perhaps with\n\n\t--prefix=<prefix>/::\n\t\tThe paths of the files in the tree being archived,\n\t\tand untracked contents added via the `--add-file`\n\t\tand `--add-virtual-file` options, can be modified by\n\t\tprepending the \"prefix\" value that is in effect when\n\t\tthese options or the tree object is seen on the\n\t\tcommand line.  The \"prefix\" value initially starts\n\t\tas an empty string, and it gets updated every time\n\t\tthis option is given on the command line.\n\nor something like that, with something like\n\n> +\tby concatenating the current value for \"prefix\" (see `--prefix`\n> +\tabove) and the basename of <file>.\n\nhere, it might make it less misunderstanding-prone, hopefully?\n\n> +`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n> +\n> +\tCreates a tar archive that contains the contents of the latest\n> +\tcommit on the current branch with no prefix and the untracked\n> +\tfile 'configure' with the prefix 'build/'.\n\nGreat to have this example.\n\nThanks.\n"},{"id":"456303","messageId":"71ae5983-6ef8-fe28-46ab-1675e819ce8b@web.de","threadId":"57313","inReplyTo":"xmqqfskwxd6j.fsf@gitster.g","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-27T17:02:09Z","receivedAt":"2022-05-27T17:02:45Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 26.05.22 um 22:16 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>> diff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\n>> index bc4e76a783..10a48ab5f8 100644\n>> --- a/Documentation/git-archive.txt\n>> +++ b/Documentation/git-archive.txt\n>> @@ -49,7 +49,9 @@ OPTIONS\n>>  \tReport progress to stderr.\n>>\n>>  --prefix=<prefix>/::\n>> -\tPrepend <prefix>/ to each filename in the archive.\n>> +\tPrepend <prefix>/ to each filename in the archive.  Can be\n>> +\tspecified multiple times; the last one seen when reading from\n>> +\tleft to right is applied.\n>\n> That can be read to mean that we will use C consistently,\n>\n> $ cmd --prefix=A other-args --prefix=B other-args --prefix=C other-args\n>\n> which was what I am worried to be a source of confusion.\n>\n>>  -o <file>::\n>>  --output=<file>::\n>> @@ -58,8 +60,8 @@ OPTIONS\n>>  --add-file=<file>::\n>>  \tAdd a non-tracked file to the archive.  Can be repeated to add\n>>  \tmultiple files.  The path of the file in the archive is built\n>> -\tby concatenating the value for `--prefix` (if any) and the\n>> -\tbasename of <file>.\n>> +\tby concatenating the current value for `--prefix` (if any) and\n>> +\tthe basename of <file>.\n>\n> \"the current value for `--prefix` (if any)\" would work well once we\n> somehow make the reader form a mental model that there is \"the\n> current\" for the \"prefix\", which starts with an empty string, and\n> gets updated every time the \"--prefix=<prefix>/\" option is given.\n\nRight, \"current\" has a well-known meaning, but its not enough to convey\nthat the non-standard concept of capturing option values in the middle of\nthe argument list is used here.\n\n>\n> So, perhaps with\n>\n> \t--prefix=<prefix>/::\n> \t\tThe paths of the files in the tree being archived,\n> \t\tand untracked contents added via the `--add-file`\n> \t\tand `--add-virtual-file` options, can be modified by\n> \t\tprepending the \"prefix\" value that is in effect when\n> \t\tthese options or the tree object is seen on the\n> \t\tcommand line.  The \"prefix\" value initially starts\n> \t\tas an empty string, and it gets updated every time\n> \t\tthis option is given on the command line.\n>\n> or something like that, with something like\n>\n>> +\tby concatenating the current value for \"prefix\" (see `--prefix`\n>> +\tabove) and the basename of <file>.\n>\n> here, it might make it less misunderstanding-prone, hopefully?\n\nSo how about this, which avoids mentioning the idea of a \"current\"\noption, or of updating its value (which implies an order that might not\nbe obvious)?\n\n--- >8 ---\nSubject: [PATCH v2] archive: improve documentation of --prefix\n\nDocument the interaction between --add-file and --prefix by giving an\nexample.\n\nSigned-off-by: René Scharfe <l.s.r@web.de>\n---\n Documentation/git-archive.txt | 15 ++++++++++++---\n 1 file changed, 12 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a783..9c0e306c03 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -49,7 +49,9 @@ OPTIONS\n \tReport progress to stderr.\n\n --prefix=<prefix>/::\n-\tPrepend <prefix>/ to each filename in the archive.\n+\tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n+\tleftmost value is used for all tracked files.  See below which\n+\tvalue gets used by `--add-file`.\n\n -o <file>::\n --output=<file>::\n@@ -58,8 +60,9 @@ OPTIONS\n --add-file=<file>::\n \tAdd a non-tracked file to the archive.  Can be repeated to add\n \tmultiple files.  The path of the file in the archive is built\n-\tby concatenating the value for `--prefix` (if any) and the\n-\tbasename of <file>.\n+\tby concatenating the value of the leftmost `--prefix` option to\n+\tthe right of this `--add-file` (if any) and the basename of\n+\t<file>.\n\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n@@ -194,6 +197,12 @@ EXAMPLES\n \tcommit on the current branch. Note that the output format is\n \tinferred by the extension of the output file.\n\n+`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n+\n+\tCreates a tar archive that contains the contents of the latest\n+\tcommit on the current branch with no prefix and the untracked\n+\tfile 'configure' with the prefix 'build/'.\n+\n `git config tar.tar.xz.command \"xz -c\"`::\n\n \tConfigure a \"tar.xz\" format for making LZMA-compressed tarfiles.\n--\n2.35.3\n\n"},{"id":"456311","messageId":"xmqqh75a3imc.fsf@gitster.g","threadId":"57313","inReplyTo":"71ae5983-6ef8-fe28-46ab-1675e819ce8b@web.de","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-27T19:01:15Z","receivedAt":"2022-05-27T19:01:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n>  --prefix=<prefix>/::\n> -\tPrepend <prefix>/ to each filename in the archive.\n> +\tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n> +\tleftmost value is used for all tracked files.  See below which\n> +\tvalue gets used by `--add-file`.\n\nDoesn't \"the last one wins\" take the rightmost one?\n\n> @@ -58,8 +60,9 @@ OPTIONS\n>  --add-file=<file>::\n>  \tAdd a non-tracked file to the archive.  Can be repeated to add\n>  \tmultiple files.  The path of the file in the archive is built\n> -\tby concatenating the value for `--prefix` (if any) and the\n> -\tbasename of <file>.\n> +\tby concatenating the value of the leftmost `--prefix` option to\n> +\tthe right of this `--add-file` (if any) and the basename of\n> +\t<file>.\n\nIt is not what archive.c::add_file_cb() seems to be doing, though\n\nIt is passed the pointer to \"base\" that is on-stack of\nparse_archive_args(), which is the same variable that is used to\nremember the latest value that was given to \"--prefix\".  Then it\nconcatenates the argument it received after that base value, so\n\n    by concatenating the value of the last \"--prefix\" seen on the\n    command line (if any) before this `--add-file` and the basename\n    of <file>.\n\nprobably.  I always get my left and right mixed up X-<.\n\n> @@ -194,6 +197,12 @@ EXAMPLES\n>  \tcommit on the current branch. Note that the output format is\n>  \tinferred by the extension of the output file.\n>\n> +`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n> +\n> +\tCreates a tar archive that contains the contents of the latest\n> +\tcommit on the current branch with no prefix and the untracked\n> +\tfile 'configure' with the prefix 'build/'.\n> +\n>  `git config tar.tar.xz.command \"xz -c\"`::\n>\n>  \tConfigure a \"tar.xz\" format for making LZMA-compressed tarfiles.\n\nThanks.\n\nThis patch probably needs to come before the \"scalar diagnose\"\nseries, which we haven't heard much about recently (no, I am not\ncomplaining---we all heard that Dscho is busy).\n\n\n"},{"id":"456343","messageId":"6ef7f836-45f6-8386-03c0-dc18b125ec67@web.de","threadId":"57313","inReplyTo":"xmqqh75a3imc.fsf@gitster.g","subject":"Re: [PATCH v6 1/7] archive: optionally add \"virtual\" files","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-05-28T06:57:46Z","receivedAt":"2022-05-28T06:58:45Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 27.05.22 um 21:01 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n>\n>>  --prefix=<prefix>/::\n>> -\tPrepend <prefix>/ to each filename in the archive.\n>> +\tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n>> +\tleftmost value is used for all tracked files.  See below which\n>> +\tvalue gets used by `--add-file`.\n>\n> Doesn't \"the last one wins\" take the rightmost one?\n\nHa ha!  Classic mistake, I do that all the time, especially when in a\nhurry. >_<\n\n>\n>> @@ -58,8 +60,9 @@ OPTIONS\n>>  --add-file=<file>::\n>>  \tAdd a non-tracked file to the archive.  Can be repeated to add\n>>  \tmultiple files.  The path of the file in the archive is built\n>> -\tby concatenating the value for `--prefix` (if any) and the\n>> -\tbasename of <file>.\n>> +\tby concatenating the value of the leftmost `--prefix` option to\n>> +\tthe right of this `--add-file` (if any) and the basename of\n>> +\t<file>.\n>\n> It is not what archive.c::add_file_cb() seems to be doing, though\n>\n> It is passed the pointer to \"base\" that is on-stack of\n> parse_archive_args(), which is the same variable that is used to\n> remember the latest value that was given to \"--prefix\".  Then it\n> concatenates the argument it received after that base value, so\n>\n>     by concatenating the value of the last \"--prefix\" seen on the\n>     command line (if any) before this `--add-file` and the basename\n>     of <file>.\n>\n> probably.  I always get my left and right mixed up X-<.\n\nYou too?  So yeah, avoiding the terms is appealing.\n\n>\n>> @@ -194,6 +197,12 @@ EXAMPLES\n>>  \tcommit on the current branch. Note that the output format is\n>>  \tinferred by the extension of the output file.\n>>\n>> +`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n>> +\n>> +\tCreates a tar archive that contains the contents of the latest\n>> +\tcommit on the current branch with no prefix and the untracked\n>> +\tfile 'configure' with the prefix 'build/'.\n>> +\n>>  `git config tar.tar.xz.command \"xz -c\"`::\n>>\n>>  \tConfigure a \"tar.xz\" format for making LZMA-compressed tarfiles.\n>\n> Thanks.\n>\n> This patch probably needs to come before the \"scalar diagnose\"\n> series, which we haven't heard much about recently (no, I am not\n> complaining---we all heard that Dscho is busy).\n>\n>\n\n--- >8 ---\nSubject: [PATCH v3] archive: improve documentation of --prefix\n\nDocument the interaction between --add-file and --prefix by giving an\nexample.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: René Scharfe <l.s.r@web.de>\n---\n Documentation/git-archive.txt | 16 ++++++++++++----\n 1 file changed, 12 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex bc4e76a783..94519aae23 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -49,7 +49,9 @@ OPTIONS\n \tReport progress to stderr.\n\n --prefix=<prefix>/::\n-\tPrepend <prefix>/ to each filename in the archive.\n+\tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n+\trightmost value is used for all tracked files.  See below which\n+\tvalue gets used by `--add-file`.\n\n -o <file>::\n --output=<file>::\n@@ -57,9 +59,9 @@ OPTIONS\n\n --add-file=<file>::\n \tAdd a non-tracked file to the archive.  Can be repeated to add\n-\tmultiple files.  The path of the file in the archive is built\n-\tby concatenating the value for `--prefix` (if any) and the\n-\tbasename of <file>.\n+\tmultiple files.  The path of the file in the archive is built by\n+\tconcatenating the value of the last `--prefix` option (if any)\n+\tbefore this `--add-file` and the basename of <file>.\n\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n@@ -194,6 +196,12 @@ EXAMPLES\n \tcommit on the current branch. Note that the output format is\n \tinferred by the extension of the output file.\n\n+`git archive -o latest.tar --prefix=build/ --add-file=configure --prefix= HEAD`::\n+\n+\tCreates a tar archive that contains the contents of the latest\n+\tcommit on the current branch with no prefix and the untracked\n+\tfile 'configure' with the prefix 'build/'.\n+\n `git config tar.tar.xz.command \"xz -c\"`::\n\n \tConfigure a \"tar.xz\" format for making LZMA-compressed tarfiles.\n--\n2.35.3\n"},{"id":"456347","messageId":"20220528231118.3504387-1-gitster@pobox.com","threadId":"57313","inReplyTo":"pull.1128.v6.git.1653145696.gitgitgadget@gmail.com","subject":"[PATCH v6+ 0/7] js/scalar-diagnose rebased","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:11Z","receivedAt":"2022-05-28T23:11:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Recent document clarification on the \"--prefix\" option of the \"git\narchive\" command from René serves as a good basis for the\ndocumentation of the \"--add-virtual-file\" option added by this\nseries, so here is my attempt to rebase js/scalar-diagnose topic\non it to hopefully help reduce Dscho's workload ;-)\n\nAside from obvious adjustments needed while rebasing onto the\nupdated documentation, there are only a couple of changes:\n\n - The way the <path> in --add-virtual-file=<path>:<contents> is\n   used has been corrected.  Earlier, leading directory components\n   of the <path> were all discarded and used nowhere, which made no\n   sense.  The <path> is used as a whole, but for consistency with\n   --add-file=<path>, <prefix> is still applied.\n\n - Overly loose quoting of variables in test scripts has been\n   corrected.\n\nBoth changes have been in 'seen' from before the rebase.\n\n1:  510f6b226b ! 1:  61522a0866 archive: optionally add \"virtual\" files\n    @@ Commit message\n     \n      ## Documentation/git-archive.txt ##\n     @@ Documentation/git-archive.txt: OPTIONS\n    - \tby concatenating the value for `--prefix` (if any) and the\n    - \tbasename of <file>.\n    + --prefix=<prefix>/::\n    + \tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n    + \trightmost value is used for all tracked files.  See below which\n    +-\tvalue gets used by `--add-file`.\n    ++\tvalue gets used by `--add-file` and `--add-virtual-file`.\n    + \n    + -o <file>::\n    + --output=<file>::\n    +@@ Documentation/git-archive.txt: OPTIONS\n    + \tconcatenating the value of the last `--prefix` option (if any)\n    + \tbefore this `--add-file` and the basename of <file>.\n      \n     +--add-virtual-file=<path>:<content>::\n     +\tAdd the specified contents to the archive.  Can be repeated to add\n     +\tmultiple files.  The path of the file in the archive is built\n    -+\tby concatenating the value for `--prefix` (if any) and the\n    -+\tbasename of <file>.\n    ++\tby concatenating the value of the last `--prefix` option (if any)\n    ++\tbefore this `--add-virtual-file` and `<path>`.\n     ++\n     +The `<path>` cannot contain any colon, the file mode is limited to\n     +a regular file, and the option may be subject to platform-dependent\n    @@ archive.c: static int queue_or_write_archive_entry(const struct object_id *oid,\n      \n      int write_archive_entries(struct archiver_args *args,\n     @@ archive.c: int write_archive_entries(struct archiver_args *args,\n    - \t\tstrbuf_addstr(&path_in_archive, basename(path));\n      \n    - \t\tstrbuf_reset(&content);\n    + \t\tput_be64(fake_oid.hash, i + 1);\n    + \n    +-\t\tstrbuf_reset(&path_in_archive);\n    +-\t\tif (info->base)\n    +-\t\t\tstrbuf_addstr(&path_in_archive, info->base);\n    +-\t\tstrbuf_addstr(&path_in_archive, basename(path));\n    +-\n    +-\t\tstrbuf_reset(&content);\n     -\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n    -+\t\tif (info->content)\n    -+\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n    -+\t\t\t\t\t  path_in_archive.len,\n    -+\t\t\t\t\t  canon_mode(info->stat.st_mode),\n    +-\t\t\terr = error_errno(_(\"cannot read '%s'\"), path);\n    +-\t\telse\n    +-\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n    +-\t\t\t\t\t  path_in_archive.len,\n    ++\t\tif (!info->content) {\n    ++\t\t\tstrbuf_reset(&path_in_archive);\n    ++\t\t\tif (info->base)\n    ++\t\t\t\tstrbuf_addstr(&path_in_archive, info->base);\n    ++\t\t\tstrbuf_addstr(&path_in_archive, basename(path));\n    ++\n    ++\t\t\tstrbuf_reset(&content);\n    ++\t\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n    ++\t\t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n    ++\t\t\telse\n    ++\t\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n    ++\t\t\t\t\t\t  path_in_archive.len,\n    ++\t\t\t\t\t\t  canon_mode(info->stat.st_mode),\n    ++\t\t\t\t\t\t  content.buf, content.len);\n    ++\t\t} else {\n    ++\t\t\terr = write_entry(args, &fake_oid,\n    ++\t\t\t\t\t  path, strlen(path),\n    + \t\t\t\t\t  canon_mode(info->stat.st_mode),\n    +-\t\t\t\t\t  content.buf, content.len);\n     +\t\t\t\t\t  info->content, info->stat.st_size);\n    -+\t\telse if (strbuf_read_file(&content, path,\n    -+\t\t\t\t\t  info->stat.st_size) < 0)\n    - \t\t\terr = error_errno(_(\"cannot read '%s'\"), path);\n    - \t\telse\n    - \t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n    ++\t\t}\n    ++\n    + \t\tif (err)\n    + \t\t\tbreak;\n    + \t}\n     @@ archive.c: static void extra_file_info_clear(void *util, const char *str)\n      {\n      \tstruct extra_file_info *info = util;\n2:  208f4aad5f ! 2:  5e9d19a70f archive --add-virtual-file: allow paths containing colons\n    @@ Commit message\n     \n      ## Documentation/git-archive.txt ##\n     @@ Documentation/git-archive.txt: OPTIONS\n    - \tby concatenating the value for `--prefix` (if any) and the\n    - \tbasename of <file>.\n    + \tby concatenating the value of the last `--prefix` option (if any)\n    + \tbefore this `--add-virtual-file` and `<path>`.\n      +\n     -The `<path>` cannot contain any colon, the file mode is limited to\n     -a regular file, and the option may be subject to platform-dependent\n     -command-line limits. For non-trivial cases, write an untracked file\n     -and use `--add-file` instead.\n     +The `<path>` argument can start and end with a literal double-quote\n    -+character; The contained file name is interpreted as a C-style string,\n    ++character; the contained file name is interpreted as a C-style string,\n     +i.e. the backslash is interpreted as escape character. The path must\n     +be quoted if it contains a colon, to avoid the colon from being\n     +misinterpreted as the separator between the path and the contents, or\n    @@ t/t5003-archive-zip.sh: check_zip with_untracked\n      test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n     +\tif test_have_prereq FUNNYNAMES\n     +\tthen\n    -+\t\tPATHNAME=quoted:colon\n    ++\t\tPATHNAME=\"pathname with : colon\"\n     +\telse\n    -+\t\tPATHNAME=quoted\n    ++\t\tPATHNAME=\"pathname without colon\"\n     +\tfi &&\n      \tgit archive --format=zip >with_file_with_content.zip \\\n    -+\t\t--add-virtual-file=\\\"$PATHNAME\\\": \\\n    ++\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n      \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n      \ttest_when_finished \"rm -rf tmp-unpack\" &&\n      \tmkdir tmp-unpack && (\n      \t\tcd tmp-unpack &&\n      \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n      \t\ttest_path_is_file hello &&\n    -+\t\ttest_path_is_file $PATHNAME &&\n    ++\t\ttest_path_is_file \"$PATHNAME\" &&\n      \t\ttest world = $(cat hello)\n      \t)\n      '\n3:  bc1164404f = 3:  4f5b3aa775 scalar: validate the optional enlistment argument\n4:  69daeb7d9d ! 4:  f4f070df8e Implement `scalar diagnose`\n    @@ Metadata\n     Author: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n     \n      ## Commit message ##\n    -    Implement `scalar diagnose`\n    +    scalar: implement `scalar diagnose`\n     \n         Over the course of Scalar's development, it became obvious that there is\n         a need for a command that can gather all kinds of useful information\n5:  5c1ef19524 = 5:  0417d8abe4 scalar diagnose: include disk space information\n6:  0325b9c3ab = 6:  5531b65ddb scalar: teach `diagnose` to gather packfile info\n7:  8fee365b07 = 7:  ce9eba5e32 scalar: teach `diagnose` to gather loose objects information\n\n\n"},{"id":"456348","messageId":"20220528231118.3504387-2-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 1/7] archive: optionally add \"virtual\" files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:12Z","receivedAt":"2022-05-28T23:11:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWith the `--add-virtual-file=<path>:<content>` option, `git archive` now\nsupports use cases where relatively trivial files need to be added that\ndo not exist on disk.\n\nThis will allow us to generate `.zip` files with generated content,\nwithout having to add said content to the object database and without\nhaving to write it out to disk.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n[jc: tweaked <path> handling]\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * The changes to the way how leading components of the <path> are\n   not discarded and used made the \"extra entries\" handling into two\n   separate code to independently come up with the path stored in\n   the archive, as well as the contents stored in the archive.\n\n   The explanation of how --prefix and --add-file interacts also\n   applies to the new option.\n\n Documentation/git-archive.txt | 13 +++++-\n archive.c                     | 77 ++++++++++++++++++++++++++---------\n t/t5003-archive-zip.sh        | 12 ++++++\n 3 files changed, 82 insertions(+), 20 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex 94519aae23..b41cc5bc2e 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -51,7 +51,7 @@ OPTIONS\n --prefix=<prefix>/::\n \tPrepend <prefix>/ to paths in the archive.  Can be repeated; its\n \trightmost value is used for all tracked files.  See below which\n-\tvalue gets used by `--add-file`.\n+\tvalue gets used by `--add-file` and `--add-virtual-file`.\n \n -o <file>::\n --output=<file>::\n@@ -63,6 +63,17 @@ OPTIONS\n \tconcatenating the value of the last `--prefix` option (if any)\n \tbefore this `--add-file` and the basename of <file>.\n \n+--add-virtual-file=<path>:<content>::\n+\tAdd the specified contents to the archive.  Can be repeated to add\n+\tmultiple files.  The path of the file in the archive is built\n+\tby concatenating the value of the last `--prefix` option (if any)\n+\tbefore this `--add-virtual-file` and `<path>`.\n++\n+The `<path>` cannot contain any colon, the file mode is limited to\n+a regular file, and the option may be subject to platform-dependent\n+command-line limits. For non-trivial cases, write an untracked file\n+and use `--add-file` instead.\n+\n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\n \tas well (see <<ATTRIBUTES>>).\ndiff --git a/archive.c b/archive.c\nindex e2121ebefb..d26f4ef945 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -263,6 +263,7 @@ static int queue_or_write_archive_entry(const struct object_id *oid,\n struct extra_file_info {\n \tchar *base;\n \tstruct stat stat;\n+\tvoid *content;\n };\n \n int write_archive_entries(struct archiver_args *args,\n@@ -331,19 +332,27 @@ int write_archive_entries(struct archiver_args *args,\n \n \t\tput_be64(fake_oid.hash, i + 1);\n \n-\t\tstrbuf_reset(&path_in_archive);\n-\t\tif (info->base)\n-\t\t\tstrbuf_addstr(&path_in_archive, info->base);\n-\t\tstrbuf_addstr(&path_in_archive, basename(path));\n-\n-\t\tstrbuf_reset(&content);\n-\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n-\t\t\terr = error_errno(_(\"cannot read '%s'\"), path);\n-\t\telse\n-\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n-\t\t\t\t\t  path_in_archive.len,\n+\t\tif (!info->content) {\n+\t\t\tstrbuf_reset(&path_in_archive);\n+\t\t\tif (info->base)\n+\t\t\t\tstrbuf_addstr(&path_in_archive, info->base);\n+\t\t\tstrbuf_addstr(&path_in_archive, basename(path));\n+\n+\t\t\tstrbuf_reset(&content);\n+\t\t\tif (strbuf_read_file(&content, path, info->stat.st_size) < 0)\n+\t\t\t\terr = error_errno(_(\"could not read '%s'\"), path);\n+\t\t\telse\n+\t\t\t\terr = write_entry(args, &fake_oid, path_in_archive.buf,\n+\t\t\t\t\t\t  path_in_archive.len,\n+\t\t\t\t\t\t  canon_mode(info->stat.st_mode),\n+\t\t\t\t\t\t  content.buf, content.len);\n+\t\t} else {\n+\t\t\terr = write_entry(args, &fake_oid,\n+\t\t\t\t\t  path, strlen(path),\n \t\t\t\t\t  canon_mode(info->stat.st_mode),\n-\t\t\t\t\t  content.buf, content.len);\n+\t\t\t\t\t  info->content, info->stat.st_size);\n+\t\t}\n+\n \t\tif (err)\n \t\t\tbreak;\n \t}\n@@ -493,6 +502,7 @@ static void extra_file_info_clear(void *util, const char *str)\n {\n \tstruct extra_file_info *info = util;\n \tfree(info->base);\n+\tfree(info->content);\n \tfree(info);\n }\n \n@@ -514,14 +524,40 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \tif (!arg)\n \t\treturn -1;\n \n-\tpath = prefix_filename(args->prefix, arg);\n-\titem = string_list_append_nodup(&args->extra_files, path);\n-\titem->util = info = xmalloc(sizeof(*info));\n+\tinfo = xmalloc(sizeof(*info));\n \tinfo->base = xstrdup_or_null(base);\n-\tif (stat(path, &info->stat))\n-\t\tdie(_(\"File not found: %s\"), path);\n-\tif (!S_ISREG(info->stat.st_mode))\n-\t\tdie(_(\"Not a regular file: %s\"), path);\n+\n+\tif (!strcmp(opt->long_name, \"add-file\")) {\n+\t\tpath = prefix_filename(args->prefix, arg);\n+\t\tif (stat(path, &info->stat))\n+\t\t\tdie(_(\"File not found: %s\"), path);\n+\t\tif (!S_ISREG(info->stat.st_mode))\n+\t\t\tdie(_(\"Not a regular file: %s\"), path);\n+\t\tinfo->content = NULL; /* read the file later */\n+\t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n+\t\tconst char *colon = strchr(arg, ':');\n+\t\tchar *p;\n+\n+\t\tif (!colon)\n+\t\t\tdie(_(\"missing colon: '%s'\"), arg);\n+\n+\t\tp = xstrndup(arg, colon - arg);\n+\t\tif (!args->prefix)\n+\t\t\tpath = p;\n+\t\telse {\n+\t\t\tpath = prefix_filename(args->prefix, p);\n+\t\t\tfree(p);\n+\t\t}\n+\t\tmemset(&info->stat, 0, sizeof(info->stat));\n+\t\tinfo->stat.st_mode = S_IFREG | 0644;\n+\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->stat.st_size = strlen(info->content);\n+\t} else {\n+\t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\n+\t}\n+\titem = string_list_append_nodup(&args->extra_files, path);\n+\titem->util = info;\n+\n \treturn 0;\n }\n \n@@ -554,6 +590,9 @@ static int parse_archive_args(int argc, const char **argv,\n \t\t{ OPTION_CALLBACK, 0, \"add-file\", args, N_(\"file\"),\n \t\t  N_(\"add untracked file to archive\"), 0, add_file_cb,\n \t\t  (intptr_t)&base },\n+\t\t{ OPTION_CALLBACK, 0, \"add-virtual-file\", args,\n+\t\t  N_(\"path:content\"), N_(\"add untracked file to archive\"), 0,\n+\t\t  add_file_cb, (intptr_t)&base },\n \t\tOPT_STRING('o', \"output\", &output, N_(\"file\"),\n \t\t\tN_(\"write the archive to this file\")),\n \t\tOPT_BOOL(0, \"worktree-attributes\", &worktree_attributes,\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex d726964307..d6027189e2 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -206,6 +206,18 @@ test_expect_success 'git archive --format=zip --add-file' '\n check_zip with_untracked\n check_added with_untracked untracked untracked\n \n+test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n+\ttest_when_finished \"rm -rf tmp-unpack\" &&\n+\tmkdir tmp-unpack && (\n+\t\tcd tmp-unpack &&\n+\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n+\t\ttest_path_is_file hello &&\n+\t\ttest world = $(cat hello)\n+\t)\n+'\n+\n test_expect_success 'git archive --format=zip --add-file twice' '\n \techo untracked >untracked &&\n \tgit archive --format=zip --prefix=one/ --add-file=untracked \\\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456349","messageId":"20220528231118.3504387-4-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 3/7] scalar: validate the optional enlistment argument","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:14Z","receivedAt":"2022-05-28T23:11:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe `scalar` command needs a Scalar enlistment for many subcommands, and\nlooks in the current directory for such an enlistment (traversing the\nparent directories until it finds one).\n\nThese is subcommands can also be called with an optional argument\nspecifying the enlistment. Here, too, we traverse parent directories as\nneeded, until we find an enlistment.\n\nHowever, if the specified directory does not even exist, or is not a\ndirectory, we should stop right there, with an error message.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n contrib/scalar/scalar.c          | 6 ++++--\n contrib/scalar/t/t9099-scalar.sh | 5 +++++\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 58ca0e56f1..6d58c7a698 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -43,9 +43,11 @@ static void setup_enlistment_directory(int argc, const char **argv,\n \t\tusage_with_options(usagestr, options);\n \n \t/* find the worktree, determine its corresponding root */\n-\tif (argc == 1)\n+\tif (argc == 1) {\n \t\tstrbuf_add_absolute_path(&path, argv[0]);\n-\telse if (strbuf_getcwd(&path) < 0)\n+\t\tif (!is_directory(path.buf))\n+\t\t\tdie(_(\"'%s' does not exist\"), path.buf);\n+\t} else if (strbuf_getcwd(&path) < 0)\n \t\tdie(_(\"need a working directory\"));\n \n \tstrbuf_trim_trailing_dir_sep(&path);\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 89781568f4..bb42354a8b 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -93,4 +93,9 @@ test_expect_success 'scalar supports -c/-C' '\n \ttest true = \"$(git -C sub config core.preloadIndex)\"\n '\n \n+test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n+\t! scalar run config cloned 2>err &&\n+\tgrep \"cloned. does not exist\" err\n+'\n+\n test_done\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456350","messageId":"20220528231118.3504387-3-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:13Z","receivedAt":"2022-05-28T23:11:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nBy allowing the path to be enclosed in double-quotes, we can avoid\nthe limitation that paths cannot contain colons.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n * Tightened shell variable quoting\n\n Documentation/git-archive.txt | 14 ++++++++++----\n archive.c                     | 30 ++++++++++++++++++++----------\n t/t5003-archive-zip.sh        |  8 ++++++++\n 3 files changed, 38 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/git-archive.txt b/Documentation/git-archive.txt\nindex b41cc5bc2e..56989a2f34 100644\n--- a/Documentation/git-archive.txt\n+++ b/Documentation/git-archive.txt\n@@ -69,10 +69,16 @@ OPTIONS\n \tby concatenating the value of the last `--prefix` option (if any)\n \tbefore this `--add-virtual-file` and `<path>`.\n +\n-The `<path>` cannot contain any colon, the file mode is limited to\n-a regular file, and the option may be subject to platform-dependent\n-command-line limits. For non-trivial cases, write an untracked file\n-and use `--add-file` instead.\n+The `<path>` argument can start and end with a literal double-quote\n+character; the contained file name is interpreted as a C-style string,\n+i.e. the backslash is interpreted as escape character. The path must\n+be quoted if it contains a colon, to avoid the colon from being\n+misinterpreted as the separator between the path and the contents, or\n+if the path begins or ends with a double-quote character.\n++\n+The file mode is limited to a regular file, and the option may be\n+subject to platform-dependent command-line limits. For non-trivial\n+cases, write an untracked file and use `--add-file` instead.\n \n --worktree-attributes::\n \tLook for attributes in .gitattributes files in the working tree\ndiff --git a/archive.c b/archive.c\nindex d26f4ef945..48aba4ac46 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -9,6 +9,7 @@\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n #include \"dir.h\"\n+#include \"quote.h\"\n \n static char const * const archive_usage[] = {\n \tN_(\"git archive [<options>] <tree-ish> [<path>...]\"),\n@@ -535,22 +536,31 @@ static int add_file_cb(const struct option *opt, const char *arg, int unset)\n \t\t\tdie(_(\"Not a regular file: %s\"), path);\n \t\tinfo->content = NULL; /* read the file later */\n \t} else if (!strcmp(opt->long_name, \"add-virtual-file\")) {\n-\t\tconst char *colon = strchr(arg, ':');\n-\t\tchar *p;\n+\t\tstruct strbuf buf = STRBUF_INIT;\n+\t\tconst char *p = arg;\n+\n+\t\tif (*p != '\"')\n+\t\t\tp = strchr(p, ':');\n+\t\telse if (unquote_c_style(&buf, p, &p) < 0)\n+\t\t\tdie(_(\"unclosed quote: '%s'\"), arg);\n \n-\t\tif (!colon)\n+\t\tif (!p || *p != ':')\n \t\t\tdie(_(\"missing colon: '%s'\"), arg);\n \n-\t\tp = xstrndup(arg, colon - arg);\n-\t\tif (!args->prefix)\n-\t\t\tpath = p;\n-\t\telse {\n-\t\t\tpath = prefix_filename(args->prefix, p);\n-\t\t\tfree(p);\n+\t\tif (p == arg)\n+\t\t\tdie(_(\"empty file name: '%s'\"), arg);\n+\n+\t\tpath = buf.len ?\n+\t\t\tstrbuf_detach(&buf, NULL) : xstrndup(arg, p - arg);\n+\n+\t\tif (args->prefix) {\n+\t\t\tchar *save = path;\n+\t\t\tpath = prefix_filename(args->prefix, path);\n+\t\t\tfree(save);\n \t\t}\n \t\tmemset(&info->stat, 0, sizeof(info->stat));\n \t\tinfo->stat.st_mode = S_IFREG | 0644;\n-\t\tinfo->content = xstrdup(colon + 1);\n+\t\tinfo->content = xstrdup(p + 1);\n \t\tinfo->stat.st_size = strlen(info->content);\n \t} else {\n \t\tBUG(\"add_file_cb() called for %s\", opt->long_name);\ndiff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\nindex d6027189e2..3992d08158 100755\n--- a/t/t5003-archive-zip.sh\n+++ b/t/t5003-archive-zip.sh\n@@ -207,13 +207,21 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n+\tif test_have_prereq FUNNYNAMES\n+\tthen\n+\t\tPATHNAME=\"pathname with : colon\"\n+\telse\n+\t\tPATHNAME=\"pathname without colon\"\n+\tfi &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n+\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n \ttest_when_finished \"rm -rf tmp-unpack\" &&\n \tmkdir tmp-unpack && (\n \t\tcd tmp-unpack &&\n \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n \t\ttest_path_is_file hello &&\n+\t\ttest_path_is_file \"$PATHNAME\" &&\n \t\ttest world = $(cat hello)\n \t)\n '\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456351","messageId":"20220528231118.3504387-5-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 4/7] scalar: implement `scalar diagnose`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:15Z","receivedAt":"2022-05-28T23:11:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOver the course of Scalar's development, it became obvious that there is\na need for a command that can gather all kinds of useful information\nthat can help identify the most typical problems with large\nworktrees/repositories.\n\nThe `diagnose` command is the culmination of this hard-won knowledge: it\ngathers the installed hooks, the config, a couple statistics describing\nthe data shape, among other pieces of information, and then wraps\neverything up in a tidy, neat `.zip` archive.\n\nNote: originally, Scalar was implemented in C# using the .NET API, where\nwe had the luxury of a comprehensive standard library that includes\nbasic functionality such as writing a `.zip` file. In the C version, we\nlack such a commodity. Rather than introducing a dependency on, say,\nlibzip, we slightly abuse Git's `archive` machinery: we write out a\n`.zip` of the empty try, augmented by a couple files that are added via\nthe `--add-file*` options. We are careful trying not to modify the\ncurrent repository in any way lest the very circumstances that required\n`scalar diagnose` to be run are changed by the `diagnose` run itself.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n contrib/scalar/scalar.c          | 144 +++++++++++++++++++++++++++++++\n contrib/scalar/scalar.txt        |  12 +++\n contrib/scalar/t/t9099-scalar.sh |  14 +++\n 3 files changed, 170 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex 6d58c7a698..a1e05a2146 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"packfile.h\"\n #include \"help.h\"\n+#include \"archive.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -260,6 +261,47 @@ static int unregister_dir(void)\n \treturn res;\n }\n \n+static int add_directory_to_archiver(struct strvec *archiver_args,\n+\t\t\t\t\t  const char *path, int recurse)\n+{\n+\tint at_root = !*path;\n+\tDIR *dir = opendir(at_root ? \".\" : path);\n+\tstruct dirent *e;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tsize_t len;\n+\tint res = 0;\n+\n+\tif (!dir)\n+\t\treturn error_errno(_(\"could not open directory '%s'\"), path);\n+\n+\tif (!at_root)\n+\t\tstrbuf_addf(&buf, \"%s/\", path);\n+\tlen = buf.len;\n+\tstrvec_pushf(archiver_args, \"--prefix=%s\", buf.buf);\n+\n+\twhile (!res && (e = readdir(dir))) {\n+\t\tif (!strcmp(\".\", e->d_name) || !strcmp(\"..\", e->d_name))\n+\t\t\tcontinue;\n+\n+\t\tstrbuf_setlen(&buf, len);\n+\t\tstrbuf_addstr(&buf, e->d_name);\n+\n+\t\tif (e->d_type == DT_REG)\n+\t\t\tstrvec_pushf(archiver_args, \"--add-file=%s\", buf.buf);\n+\t\telse if (e->d_type != DT_DIR)\n+\t\t\twarning(_(\"skipping '%s', which is neither file nor \"\n+\t\t\t\t  \"directory\"), buf.buf);\n+\t\telse if (recurse &&\n+\t\t\t add_directory_to_archiver(archiver_args,\n+\t\t\t\t\t\t   buf.buf, recurse) < 0)\n+\t\t\tres = -1;\n+\t}\n+\n+\tclosedir(dir);\n+\tstrbuf_release(&buf);\n+\treturn res;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -500,6 +542,107 @@ static int cmd_clone(int argc, const char **argv)\n \treturn res;\n }\n \n+static int cmd_diagnose(int argc, const char **argv)\n+{\n+\tstruct option options[] = {\n+\t\tOPT_END(),\n+\t};\n+\tconst char * const usage[] = {\n+\t\tN_(\"scalar diagnose [<enlistment>]\"),\n+\t\tNULL\n+\t};\n+\tstruct strbuf zip_path = STRBUF_INIT;\n+\tstruct strvec archiver_args = STRVEC_INIT;\n+\tchar **argv_copy = NULL;\n+\tint stdout_fd = -1, archiver_fd = -1;\n+\ttime_t now = time(NULL);\n+\tstruct tm tm;\n+\tstruct strbuf path = STRBUF_INIT, buf = STRBUF_INIT;\n+\tint res = 0;\n+\n+\targc = parse_options(argc, argv, NULL, options,\n+\t\t\t     usage, 0);\n+\n+\tsetup_enlistment_directory(argc, argv, usage, options, &zip_path);\n+\n+\tstrbuf_addstr(&zip_path, \"/.scalarDiagnostics/scalar_\");\n+\tstrbuf_addftime(&zip_path,\n+\t\t\t\"%Y%m%d_%H%M%S\", localtime_r(&now, &tm), 0, 0);\n+\tstrbuf_addstr(&zip_path, \".zip\");\n+\tswitch (safe_create_leading_directories(zip_path.buf)) {\n+\tcase SCLD_EXISTS:\n+\tcase SCLD_OK:\n+\t\tbreak;\n+\tdefault:\n+\t\terror_errno(_(\"could not create directory for '%s'\"),\n+\t\t\t    zip_path.buf);\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\tstdout_fd = dup(1);\n+\tif (stdout_fd < 0) {\n+\t\tres = error_errno(_(\"could not duplicate stdout\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tarchiver_fd = xopen(zip_path.buf, O_CREAT | O_WRONLY | O_TRUNC, 0666);\n+\tif (archiver_fd < 0 || dup2(archiver_fd, 1) < 0) {\n+\t\tres = error_errno(_(\"could not redirect output\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tinit_zip_archiver();\n+\tstrvec_pushl(&archiver_args, \"scalar-diagnose\", \"--format=zip\", NULL);\n+\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"Collecting diagnostic info\\n\\n\");\n+\tget_version_info(&buf, 1);\n+\n+\tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\twrite_or_die(stdout_fd, buf.buf, buf.len);\n+\tstrvec_pushf(&archiver_args,\n+\t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n+\t\t     (int)buf.len, buf.buf);\n+\n+\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n+\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n+\t\tgoto diagnose_cleanup;\n+\n+\tstrvec_pushl(&archiver_args, \"--prefix=\",\n+\t\t     oid_to_hex(the_hash_algo->empty_tree), \"--\", NULL);\n+\n+\t/* `write_archive()` modifies the `argv` passed to it. Let it. */\n+\targv_copy = xmemdupz(archiver_args.v,\n+\t\t\t     sizeof(char *) * archiver_args.nr);\n+\tres = write_archive(archiver_args.nr, (const char **)argv_copy, NULL,\n+\t\t\t    the_repository, NULL, 0);\n+\tif (res) {\n+\t\terror(_(\"failed to write archive\"));\n+\t\tgoto diagnose_cleanup;\n+\t}\n+\n+\tif (!res)\n+\t\tfprintf(stderr, \"\\n\"\n+\t\t       \"Diagnostics complete.\\n\"\n+\t\t       \"All of the gathered info is captured in '%s'\\n\",\n+\t\t       zip_path.buf);\n+\n+diagnose_cleanup:\n+\tif (archiver_fd >= 0) {\n+\t\tclose(1);\n+\t\tdup2(stdout_fd, 1);\n+\t}\n+\tfree(argv_copy);\n+\tstrvec_clear(&archiver_args);\n+\tstrbuf_release(&zip_path);\n+\tstrbuf_release(&path);\n+\tstrbuf_release(&buf);\n+\n+\treturn res;\n+}\n+\n static int cmd_list(int argc, const char **argv)\n {\n \tif (argc != 1)\n@@ -801,6 +944,7 @@ static struct {\n \t{ \"reconfigure\", cmd_reconfigure },\n \t{ \"delete\", cmd_delete },\n \t{ \"version\", cmd_version },\n+\t{ \"diagnose\", cmd_diagnose },\n \t{ NULL, NULL},\n };\n \ndiff --git a/contrib/scalar/scalar.txt b/contrib/scalar/scalar.txt\nindex cf4e5b889c..c0425e0653 100644\n--- a/contrib/scalar/scalar.txt\n+++ b/contrib/scalar/scalar.txt\n@@ -14,6 +14,7 @@ scalar register [<enlistment>]\n scalar unregister [<enlistment>]\n scalar run ( all | config | commit-graph | fetch | loose-objects | pack-files ) [<enlistment>]\n scalar reconfigure [ --all | <enlistment> ]\n+scalar diagnose [<enlistment>]\n scalar delete <enlistment>\n \n DESCRIPTION\n@@ -139,6 +140,17 @@ reconfigure the enlistment.\n With the `--all` option, all enlistments currently registered with Scalar\n will be reconfigured. Use this option after each Scalar upgrade.\n \n+Diagnose\n+~~~~~~~~\n+\n+diagnose [<enlistment>]::\n+    When reporting issues with Scalar, it is often helpful to provide the\n+    information gathered by this command, including logs and certain\n+    statistics describing the data shape of the current enlistment.\n++\n+The output of this command is a `.zip` file that is written into\n+a directory adjacent to the worktree in the `src` directory.\n+\n Delete\n ~~~~~~\n \ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex bb42354a8b..fbb1df2049 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -98,4 +98,18 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n \tgrep \"cloned. does not exist\" err\n '\n \n+SQ=\"'\"\n+test_expect_success UNZIP 'scalar diagnose' '\n+\tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tscalar diagnose cloned >out 2>err &&\n+\tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n+\tzip_path=$(cat zip_path) &&\n+\ttest -n \"$zip_path\" &&\n+\tunzip -v \"$zip_path\" &&\n+\tfolder=${zip_path%.zip} &&\n+\ttest_path_is_missing \"$folder\" &&\n+\tunzip -p \"$zip_path\" diagnostics.log >out &&\n+\ttest_file_not_empty out\n+'\n+\n test_done\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456352","messageId":"20220528231118.3504387-7-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 6/7] scalar: teach `diagnose` to gather packfile info","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:17Z","receivedAt":"2022-05-28T23:11:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nIt's helpful to see if there are other crud files in the pack\ndirectory. Let's teach the `scalar diagnose` command to gather\nfile size information about pack files.\n\nWhile at it, also enumerate the pack files in the alternate\nobject directories, if any are registered.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n contrib/scalar/scalar.c          | 30 ++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  6 +++++-\n 2 files changed, 35 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex f06a2f3576..f745519038 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -12,6 +12,7 @@\n #include \"packfile.h\"\n #include \"help.h\"\n #include \"archive.h\"\n+#include \"object-store.h\"\n \n /*\n  * Remove the deepest subdirectory in the provided path string. Path must not\n@@ -594,6 +595,29 @@ static int cmd_clone(int argc, const char **argv)\n \treturn res;\n }\n \n+static void dir_file_stats_objects(const char *full_path, size_t full_path_len,\n+\t\t\t\t   const char *file_name, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\tstruct stat st;\n+\n+\tif (!stat(full_path, &st))\n+\t\tstrbuf_addf(buf, \"%-70s %16\" PRIuMAX \"\\n\", file_name,\n+\t\t\t    (uintmax_t)st.st_size);\n+}\n+\n+static int dir_file_stats(struct object_directory *object_dir, void *data)\n+{\n+\tstruct strbuf *buf = data;\n+\n+\tstrbuf_addf(buf, \"Contents of %s:\\n\", object_dir->path);\n+\n+\tfor_each_file_in_pack_dir(object_dir->path, dir_file_stats_objects,\n+\t\t\t\t  data);\n+\n+\treturn 0;\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -656,6 +680,12 @@ static int cmd_diagnose(int argc, const char **argv)\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\n \t\t     (int)buf.len, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=packs-local.txt:\");\n+\tdir_file_stats(the_repository->objects->odb, &buf);\n+\tforeach_alt_odb(dir_file_stats, &buf);\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 6e52088919..2603e2278f 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -101,6 +101,8 @@ test_expect_success '`scalar [...] <dir>` errors out when dir is missing' '\n SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n+\tgit repack &&\n+\techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -110,7 +112,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tfolder=${zip_path%.zip} &&\n \ttest_path_is_missing \"$folder\" &&\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n-\ttest_file_not_empty out\n+\ttest_file_not_empty out &&\n+\tunzip -p \"$zip_path\" packs-local.txt >out &&\n+\tgrep \"$(pwd)/.git/objects\" out\n '\n \n test_done\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456353","messageId":"20220528231118.3504387-8-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 7/7] scalar: teach `diagnose` to gather loose objects information","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:18Z","receivedAt":"2022-05-28T23:11:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Matthew John Cheetham <mjcheetham@outlook.com>\n\nWhen operating at the scale that Scalar wants to support, certain data\nshapes are more likely to cause undesirable performance issues, such as\nlarge numbers of loose objects.\n\nBy including statistics about this, `scalar diagnose` now makes it\neasier to identify such scenarios.\n\nSigned-off-by: Matthew John Cheetham <mjcheetham@outlook.com>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n contrib/scalar/scalar.c          | 59 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  5 ++-\n 2 files changed, 63 insertions(+), 1 deletion(-)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex f745519038..28176914e5 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -618,6 +618,60 @@ static int dir_file_stats(struct object_directory *object_dir, void *data)\n \treturn 0;\n }\n \n+static int count_files(char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count = 0;\n+\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) && e->d_type == DT_REG)\n+\t\t\tcount++;\n+\n+\tclosedir(dir);\n+\treturn count;\n+}\n+\n+static void loose_objs_stats(struct strbuf *buf, const char *path)\n+{\n+\tDIR *dir = opendir(path);\n+\tstruct dirent *e;\n+\tint count;\n+\tint total = 0;\n+\tunsigned char c;\n+\tstruct strbuf count_path = STRBUF_INIT;\n+\tsize_t base_path_len;\n+\n+\tif (!dir)\n+\t\treturn;\n+\n+\tstrbuf_addstr(buf, \"Object directory stats for \");\n+\tstrbuf_add_absolute_path(buf, path);\n+\tstrbuf_addstr(buf, \":\\n\");\n+\n+\tstrbuf_add_absolute_path(&count_path, path);\n+\tstrbuf_addch(&count_path, '/');\n+\tbase_path_len = count_path.len;\n+\n+\twhile ((e = readdir(dir)) != NULL)\n+\t\tif (!is_dot_or_dotdot(e->d_name) &&\n+\t\t    e->d_type == DT_DIR && strlen(e->d_name) == 2 &&\n+\t\t    !hex_to_bytes(&c, e->d_name, 1)) {\n+\t\t\tstrbuf_setlen(&count_path, base_path_len);\n+\t\t\tstrbuf_addstr(&count_path, e->d_name);\n+\t\t\ttotal += (count = count_files(count_path.buf));\n+\t\t\tstrbuf_addf(buf, \"%s : %7d files\\n\", e->d_name, count);\n+\t\t}\n+\n+\tstrbuf_addf(buf, \"Total: %d loose objects\", total);\n+\n+\tstrbuf_release(&count_path);\n+\tclosedir(dir);\n+}\n+\n static int cmd_diagnose(int argc, const char **argv)\n {\n \tstruct option options[] = {\n@@ -686,6 +740,11 @@ static int cmd_diagnose(int argc, const char **argv)\n \tforeach_alt_odb(dir_file_stats, &buf);\n \tstrvec_push(&archiver_args, buf.buf);\n \n+\tstrbuf_reset(&buf);\n+\tstrbuf_addstr(&buf, \"--add-virtual-file=objects-local.txt:\");\n+\tloose_objs_stats(&buf, \".git/objects\");\n+\tstrvec_push(&archiver_args, buf.buf);\n+\n \tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n \t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex 2603e2278f..10b1172a8a 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -103,6 +103,7 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tgit repack &&\n \techo \"$(pwd)/.git/objects/\" >>cloned/src/.git/objects/info/alternates &&\n+\ttest_commit -C cloned/src loose &&\n \tscalar diagnose cloned >out 2>err &&\n \tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n@@ -114,7 +115,9 @@ test_expect_success UNZIP 'scalar diagnose' '\n \tunzip -p \"$zip_path\" diagnostics.log >out &&\n \ttest_file_not_empty out &&\n \tunzip -p \"$zip_path\" packs-local.txt >out &&\n-\tgrep \"$(pwd)/.git/objects\" out\n+\tgrep \"$(pwd)/.git/objects\" out &&\n+\tunzip -p \"$zip_path\" objects-local.txt >out &&\n+\tgrep \"^Total: [1-9]\" out\n '\n \n test_done\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456354","messageId":"20220528231118.3504387-6-gitster@pobox.com","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"[PATCH v6+ 5/7] scalar diagnose: include disk space information","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-28T23:11:16Z","receivedAt":"2022-05-28T23:11:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen analyzing problems with large worktrees/repositories, it is useful\nto know how close to a \"full disk\" situation Scalar/Git operates. Let's\ninclude this information.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n contrib/scalar/scalar.c          | 53 ++++++++++++++++++++++++++++++++\n contrib/scalar/t/t9099-scalar.sh |  1 +\n 2 files changed, 54 insertions(+)\n\ndiff --git a/contrib/scalar/scalar.c b/contrib/scalar/scalar.c\nindex a1e05a2146..f06a2f3576 100644\n--- a/contrib/scalar/scalar.c\n+++ b/contrib/scalar/scalar.c\n@@ -302,6 +302,58 @@ static int add_directory_to_archiver(struct strvec *archiver_args,\n \treturn res;\n }\n \n+#ifndef WIN32\n+#include <sys/statvfs.h>\n+#endif\n+\n+static int get_disk_info(struct strbuf *out)\n+{\n+#ifdef WIN32\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar volume_name[MAX_PATH], fs_name[MAX_PATH];\n+\tDWORD serial_number, component_length, flags;\n+\tULARGE_INTEGER avail2caller, total, avail;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (!GetDiskFreeSpaceExA(buf.buf, &avail2caller, &total, &avail)) {\n+\t\terror(_(\"could not determine free disk size for '%s'\"),\n+\t\t      buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_setlen(&buf, offset_1st_component(buf.buf));\n+\tif (!GetVolumeInformationA(buf.buf, volume_name, sizeof(volume_name),\n+\t\t\t\t   &serial_number, &component_length, &flags,\n+\t\t\t\t   fs_name, sizeof(fs_name))) {\n+\t\terror(_(\"could not get info for '%s'\"), buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, avail2caller.QuadPart);\n+\tstrbuf_addch(out, '\\n');\n+\tstrbuf_release(&buf);\n+#else\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct statvfs stat;\n+\n+\tstrbuf_realpath(&buf, \".\", 1);\n+\tif (statvfs(buf.buf, &stat) < 0) {\n+\t\terror_errno(_(\"could not determine free disk size for '%s'\"),\n+\t\t\t    buf.buf);\n+\t\tstrbuf_release(&buf);\n+\t\treturn -1;\n+\t}\n+\n+\tstrbuf_addf(out, \"Available space on '%s': \", buf.buf);\n+\tstrbuf_humanise_bytes(out, st_mult(stat.f_bsize, stat.f_bavail));\n+\tstrbuf_addf(out, \" (mount flags 0x%lx)\\n\", stat.f_flag);\n+\tstrbuf_release(&buf);\n+#endif\n+\treturn 0;\n+}\n+\n /* printf-style interface, expects `<key>=<value>` argument */\n static int set_config(const char *fmt, ...)\n {\n@@ -598,6 +650,7 @@ static int cmd_diagnose(int argc, const char **argv)\n \tget_version_info(&buf, 1);\n \n \tstrbuf_addf(&buf, \"Enlistment root: %s\\n\", the_repository->worktree);\n+\tget_disk_info(&buf);\n \twrite_or_die(stdout_fd, buf.buf, buf.len);\n \tstrvec_pushf(&archiver_args,\n \t\t     \"--add-virtual-file=diagnostics.log:%.*s\",\ndiff --git a/contrib/scalar/t/t9099-scalar.sh b/contrib/scalar/t/t9099-scalar.sh\nindex fbb1df2049..6e52088919 100755\n--- a/contrib/scalar/t/t9099-scalar.sh\n+++ b/contrib/scalar/t/t9099-scalar.sh\n@@ -102,6 +102,7 @@ SQ=\"'\"\n test_expect_success UNZIP 'scalar diagnose' '\n \tscalar clone \"file://$(pwd)\" cloned --single-branch &&\n \tscalar diagnose cloned >out 2>err &&\n+\tgrep \"Available space\" out &&\n \tsed -n \"s/.*$SQ\\\\(.*\\\\.zip\\\\)$SQ.*/\\\\1/p\" <err >zip_path &&\n \tzip_path=$(cat zip_path) &&\n \ttest -n \"$zip_path\" &&\n-- \n2.36.1-385-g60203f3fdb\n\n"},{"id":"456366","messageId":"nycvar.QRO.7.76.6.2205301205450.349@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"20220528231118.3504387-1-gitster@pobox.com","subject":"Re: [PATCH v6+ 0/7] js/scalar-diagnose rebased","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-05-30T10:12:46Z","receivedAt":"2022-05-30T10:13:05Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Sat, 28 May 2022, Junio C Hamano wrote:\n\n> Recent document clarification on the \"--prefix\" option of the \"git\n> archive\" command from René serves as a good basis for the\n> documentation of the \"--add-virtual-file\" option added by this\n> series, so here is my attempt to rebase js/scalar-diagnose topic\n> on it to hopefully help reduce Dscho's workload ;-)\n\nI usually frown upon sending patches on other people's behalf without\nobtaining their consent first [*1*], but in this case I have to admit that\nI appreciate your help very much.\n\nThe range-diff looks good.\n\nThank you,\nDscho\n\nFootnote *1*: In case it was unclear, I consider submitting PRs at\nhttps://github.com/git-for-windows/git as an implicit request to shepherd\nthe patches onto the Git mailing list, i.e. as consent to have me send\nthose patches on the original contributors' behalf.\n"},{"id":"456375","messageId":"xmqqwne2x6oo.fsf@gitster.g","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2205301205450.349@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v6+ 0/7] js/scalar-diagnose rebased","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-30T17:37:43Z","receivedAt":"2022-05-30T17:37:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> Recent document clarification on the \"--prefix\" option of the \"git\n>> archive\" command from René serves as a good basis for the\n>> documentation of the \"--add-virtual-file\" option added by this\n>> series, so here is my attempt to rebase js/scalar-diagnose topic\n>> on it to hopefully help reduce Dscho's workload ;-)\n>\n> I usually frown upon sending patches on other people's behalf without\n> obtaining their consent first [*1*], but in this case I have to admit that\n> I appreciate your help very much.\n\nI understand what you mean.\n\nConsider this as an extended form of the usual notes I send to a\nthread to say \"ok, based on the discussion I saw on the list, I'll\ntweak OP's patch <this way> while queuing; thank you all for\ncontributing.\"  The way I try to convey <this way> can range from\nwords (e.g. when a reviewer points out a typo) to a fixup patch\n(e.g. when the necessary update is a bit more involved), and this\ntime it took a full series with interdiff form.  Of course I do not\nhave to do any of the above and just leave it up to the OP to pick\nup ideas from the discussion while sending updates, but sometimes\nit is quicker to skip round-trips.\n\nI do not say \"Please holler if I misunderstood the discussion and\ncorrect me, and the OP can always update/override with a rerolled\nseries.\" when I send out such a \"here is how the version queued\nwould be different from the original\" notice, but I always mean\nthat, this time included ;-).\n\nYour \"frowning upon\" is understandable in that it can become a\nhostile behaviour towards others, including the maintainer who is\nforced to ignore or pick.  It is never fun to be in the position to\nalways exclude half of the patches posted to the list by\ncontributors who are competing instead of cooperating, and resending\na tweaked patch to show \"here is how I would imagine is a better\nversion of your series\" needs to be done with care.\n\nThanks.\n"},{"id":"457006","messageId":"220610.86ilp9s1x7.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"20220528231118.3504387-5-gitster@pobox.com","subject":"Re: [PATCH v6+ 4/7] scalar: implement `scalar diagnose`","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-10T02:08:34Z","receivedAt":"2022-06-10T02:11:24Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, May 28 2022, Junio C Hamano wrote:\n\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> [...]\n> The `diagnose` command is the culmination of this hard-won knowledge: it\n> gathers the installed hooks, the config, a couple statistics describing\n> the data shape, among other pieces of information, and then wraps\n> everything up in a tidy, neat `.zip` archive.\n> [...]\n> +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n> +\t\tgoto diagnose_cleanup;\n\nNoticed on top of some local changes I have to not add a .git/hooks (the\n--no-template topic), but this fails to diagnose any repo that doesn't\nhave these paths, which are optional, either because a user could have manually removed them, or used --template=.\n\nalthough I don't think there's a way to create that sort of repo with\nthe scalar tooling, it doesn't seem to forward that option, but I didn't\nlook deeply.\n\nSo, no big deal, but it would be nice to have that fixed. Is there a\nreason for why this mere addition of various stuff for diagnosis goes\nstraight to an opendir() and error on failure, as opposed to doing an\nlstat() etc. first?\n"},{"id":"457032","messageId":"xmqqpmjgfoxm.fsf@gitster.g","threadId":"57313","inReplyTo":"220610.86ilp9s1x7.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v6+ 4/7] scalar: implement `scalar diagnose`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-10T16:44:53Z","receivedAt":"2022-06-10T16:44:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> On Sat, May 28 2022, Junio C Hamano wrote:\n>\n>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>> [...]\n>> The `diagnose` command is the culmination of this hard-won knowledge: it\n>> gathers the installed hooks, the config, a couple statistics describing\n>> the data shape, among other pieces of information, and then wraps\n>> everything up in a tidy, neat `.zip` archive.\n>> [...]\n>> +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n>> +\t\tgoto diagnose_cleanup;\n>\n> Noticed on top of some local changes I have to not add a\n> .git/hooks (the --no-template topic), but this fails to diagnose\n> any repo that doesn't have these paths, which are optional, either\n> because a user could have manually removed them, or used\n> --template=.\n\nQuite honestly, if it lacks any directory that we traditionally\ncreated upon \"git init\", with our standard templates, we can and\nshould call such a repository \"broken\" and move on.\n\n\n"},{"id":"457038","messageId":"220610.86a6aks9j1.gmgdl@evledraar.gmail.com","threadId":"57313","inReplyTo":"xmqqpmjgfoxm.fsf@gitster.g","subject":"Re: [PATCH v6+ 4/7] scalar: implement `scalar diagnose`","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-10T17:35:58Z","receivedAt":"2022-06-10T17:39:20Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Jun 10 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>> On Sat, May 28 2022, Junio C Hamano wrote:\n>>\n>>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>>> [...]\n>>> The `diagnose` command is the culmination of this hard-won knowledge: it\n>>> gathers the installed hooks, the config, a couple statistics describing\n>>> the data shape, among other pieces of information, and then wraps\n>>> everything up in a tidy, neat `.zip` archive.\n>>> [...]\n>>> +\tif ((res = add_directory_to_archiver(&archiver_args, \".git\", 0)) ||\n>>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/hooks\", 0)) ||\n>>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/info\", 0)) ||\n>>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/logs\", 1)) ||\n>>> +\t    (res = add_directory_to_archiver(&archiver_args, \".git/objects/info\", 0)))\n>>> +\t\tgoto diagnose_cleanup;\n>>\n>> Noticed on top of some local changes I have to not add a\n>> .git/hooks (the --no-template topic), but this fails to diagnose\n>> any repo that doesn't have these paths, which are optional, either\n>> because a user could have manually removed them, or used\n>> --template=.\n>\n> Quite honestly, if it lacks any directory that we traditionally\n> created upon \"git init\", with our standard templates, we can and\n> should call such a repository \"broken\" and move on.\n\nIn our own test suite we do e.g. (and did more of that until some recent\nchanges of mine):\n\n    git mv .git/hooks .git/hooks.disabled\n\nWe've never documented in \"git init\" or the like that these very\noptional directories in .git/ were some sort of hard requirenment, and\ne.g. core.hooksPath and gitrepository-layout(5) explicitly seem to\nsuggest otherwise.\n\nIn any case, there's the golden rule about being strict in what you emit\nand loose in what you accept, which we've taken with repository\ncompatibility. Having a tool that's designed to aid bugreporting be\npicky about what sort of repository it supports seems to go against the\npoint of such a tool.\n\nParticularly in this case, where it seems easy to just guard it with a\nstat() check, or not error out if we fail to add this to the *.zip file,\nno?\n"},{"id":"457300","messageId":"20220615181641.vltm3qtbsckp5s56@lucy.dinwoodie.org","threadId":"57313","inReplyTo":"20220528231118.3504387-3-gitster@pobox.com","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Adam Dinwoodie","fromEmail":"adam@dinwoodie.org","sentAt":"2022-06-15T18:16:41Z","receivedAt":"2022-06-15T18:16:47Z","isPatch":true,"sender":{"key":"adam@dinwoodie.org","avatar":"https://avatars.githubusercontent.com/u/1397507?v=4"},"body":"On Sat, May 28, 2022 at 04:11:13PM -0700, Junio C Hamano wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> By allowing the path to be enclosed in double-quotes, we can avoid\n> the limitation that paths cannot contain colons.\n> \n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  * Tightened shell variable quoting\n> \n> <snip>\n>\n> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> index d6027189e2..3992d08158 100755\n> --- a/t/t5003-archive-zip.sh\n> +++ b/t/t5003-archive-zip.sh\n> @@ -207,13 +207,21 @@ check_zip with_untracked\n>  check_added with_untracked untracked untracked\n>  \n>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> +\tif test_have_prereq FUNNYNAMES\n> +\tthen\n> +\t\tPATHNAME=\"pathname with : colon\"\n> +\telse\n> +\t\tPATHNAME=\"pathname without colon\"\n> +\tfi &&\n>  \tgit archive --format=zip >with_file_with_content.zip \\\n> +\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n>  \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n>  \tmkdir tmp-unpack && (\n>  \t\tcd tmp-unpack &&\n>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>  \t\ttest_path_is_file hello &&\n> +\t\ttest_path_is_file \"$PATHNAME\" &&\n>  \t\ttest world = $(cat hello)\n>  \t)\n>  '\n\nThis test is currently failing on Cygwin: it looks like it's exposing a\nbug in Cygwin that means files with colons in their name aren't\ncorrectly extracted from zip archives.  I'm going to report that to the\nCygwin mailing list, but I wanted to note it for the record here, too.\n\nAdam\n"},{"id":"457309","messageId":"xmqqpmj9zohk.fsf@gitster.g","threadId":"57313","inReplyTo":"20220615181641.vltm3qtbsckp5s56@lucy.dinwoodie.org","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-15T20:00:07Z","receivedAt":"2022-06-15T20:00:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Dinwoodie <adam@dinwoodie.org> writes:\n\n>> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n>> index d6027189e2..3992d08158 100755\n>> --- a/t/t5003-archive-zip.sh\n>> +++ b/t/t5003-archive-zip.sh\n>> @@ -207,13 +207,21 @@ check_zip with_untracked\n>>  check_added with_untracked untracked untracked\n>>  \n>>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n>> +\tif test_have_prereq FUNNYNAMES\n>> +\tthen\n>> +\t\tPATHNAME=\"pathname with : colon\"\n>> +\telse\n>> +\t\tPATHNAME=\"pathname without colon\"\n>> +\tfi &&\n>>  \tgit archive --format=zip >with_file_with_content.zip \\\n>> +\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n>>  \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n>>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n>>  \tmkdir tmp-unpack && (\n>>  \t\tcd tmp-unpack &&\n>>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n>>  \t\ttest_path_is_file hello &&\n>> +\t\ttest_path_is_file \"$PATHNAME\" &&\n>>  \t\ttest world = $(cat hello)\n>>  \t)\n>>  '\n>\n> This test is currently failing on Cygwin: it looks like it's exposing a\n> bug in Cygwin that means files with colons in their name aren't\n> correctly extracted from zip archives.  I'm going to report that to the\n> Cygwin mailing list, but I wanted to note it for the record here, too.\n\nDoes this mean that our code to set FUNNYNAMES prerequiste is\nslightly broken?  IOW, should we check with a path with a colon in\nit, as well as whatever we use currently for FUNNYNAMES?\n\nSomething like the attached patch?  \n\nOr does Cygwin otherwise work perfectly well with a path with a\ncolon in it, but only $GIT_UNZIP command has problem with it?  If\nthat is the case, then please disregard the attached.\n\nThanks.\n\n t/test-lib.sh | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git i/t/test-lib.sh w/t/test-lib.sh\nindex 55857af601..5dce7d95c9 100644\n--- i/t/test-lib.sh\n+++ w/t/test-lib.sh\n@@ -1620,6 +1620,7 @@ test_lazy_prereq FUNNYNAMES '\n \ttouch -- \\\n \t\t\"FUNNYNAMES tab\tembedded\" \\\n \t\t\"FUNNYNAMES \\\"quote embedded\\\"\" \\\n+\t\t\"FUNNYNAMES colon : embedded\" \\\n \t\t\"FUNNYNAMES newline\n embedded\" 2>/dev/null &&\n \trm -- \\\n"},{"id":"457321","messageId":"20220615213656.zp36wdwbcz7yevac@lucy.dinwoodie.org","threadId":"57313","inReplyTo":"xmqqpmj9zohk.fsf@gitster.g","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Adam Dinwoodie","fromEmail":"adam@dinwoodie.org","sentAt":"2022-06-15T21:36:56Z","receivedAt":"2022-06-15T21:37:05Z","isPatch":true,"sender":{"key":"adam@dinwoodie.org","avatar":"https://avatars.githubusercontent.com/u/1397507?v=4"},"body":"On Wed, Jun 15, 2022 at 01:00:07PM -0700, Junio C Hamano wrote:\n> Adam Dinwoodie <adam@dinwoodie.org> writes:\n> \n> >> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> >> index d6027189e2..3992d08158 100755\n> >> --- a/t/t5003-archive-zip.sh\n> >> +++ b/t/t5003-archive-zip.sh\n> >> @@ -207,13 +207,21 @@ check_zip with_untracked\n> >>  check_added with_untracked untracked untracked\n> >>  \n> >>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> >> +\tif test_have_prereq FUNNYNAMES\n> >> +\tthen\n> >> +\t\tPATHNAME=\"pathname with : colon\"\n> >> +\telse\n> >> +\t\tPATHNAME=\"pathname without colon\"\n> >> +\tfi &&\n> >>  \tgit archive --format=zip >with_file_with_content.zip \\\n> >> +\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n> >>  \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n> >>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n> >>  \tmkdir tmp-unpack && (\n> >>  \t\tcd tmp-unpack &&\n> >>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> >>  \t\ttest_path_is_file hello &&\n> >> +\t\ttest_path_is_file \"$PATHNAME\" &&\n> >>  \t\ttest world = $(cat hello)\n> >>  \t)\n> >>  '\n> >\n> > This test is currently failing on Cygwin: it looks like it's exposing a\n> > bug in Cygwin that means files with colons in their name aren't\n> > correctly extracted from zip archives.  I'm going to report that to the\n> > Cygwin mailing list, but I wanted to note it for the record here, too.\n> \n> Does this mean that our code to set FUNNYNAMES prerequiste is\n> slightly broken?  IOW, should we check with a path with a colon in\n> it, as well as whatever we use currently for FUNNYNAMES?\n> \n> Something like the attached patch?  \n> \n> Or does Cygwin otherwise work perfectly well with a path with a\n> colon in it, but only $GIT_UNZIP command has problem with it?  If\n> that is the case, then please disregard the attached.\n\nThe latter: Cygwin works perfectly with paths containing colons, except\nthat Cygwin's `unzip` is seemingly buggy and doesn't work.  The file\nsystems Cygwin runs on don't support colons in paths, but Cygwin hides\nthat problem by rewriting ASCII colons to some high Unicode code point\non the filesystem, meaning Cygwin-native applications see a regular\ncolon, while Windows-native applications see an unusual but perfectly\nvalid Unicode character.\n\nI tested the same patch to FUNNYNAMES myself before reporting, and the\ntest fails exactly the same way.  If we wanted to catch this, I think\nwe'd need a test that explicitly attempted to unzip an archive\ncontaining a path with a colon.\n\n(The code to set FUNNYNAMES *is* slightly broken, per the discussions\naround 6d340dfaef (\"t9902: split test to run on appropriate systems\",\n2022-04-08), and my to-do list still features tidying up and\nresubmitting the patch Ævar wrote in that discussion thread.  But it\nwouldn't help here because this issue is specific to Cygwin's `unzip`,\nrather than a general limitation of running on Cygwin.)\n"},{"id":"457503","messageId":"nycvar.QRO.7.76.6.2206182213290.349@tvgsbejvaqbjf.bet","threadId":"57313","inReplyTo":"20220615213656.zp36wdwbcz7yevac@lucy.dinwoodie.org","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-06-18T20:19:28Z","receivedAt":"2022-06-18T20:21:28Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Adam,\n\nOn Wed, 15 Jun 2022, Adam Dinwoodie wrote:\n\n> On Wed, Jun 15, 2022 at 01:00:07PM -0700, Junio C Hamano wrote:\n> > Adam Dinwoodie <adam@dinwoodie.org> writes:\n> >\n> > >> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> > >> index d6027189e2..3992d08158 100755\n> > >> --- a/t/t5003-archive-zip.sh\n> > >> +++ b/t/t5003-archive-zip.sh\n> > >> @@ -207,13 +207,21 @@ check_zip with_untracked\n> > >>  check_added with_untracked untracked untracked\n> > >>\n> > >>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> > >> +\tif test_have_prereq FUNNYNAMES\n> > >> +\tthen\n> > >> +\t\tPATHNAME=\"pathname with : colon\"\n> > >> +\telse\n> > >> +\t\tPATHNAME=\"pathname without colon\"\n> > >> +\tfi &&\n> > >>  \tgit archive --format=zip >with_file_with_content.zip \\\n> > >> +\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n> > >>  \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n> > >>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n> > >>  \tmkdir tmp-unpack && (\n> > >>  \t\tcd tmp-unpack &&\n> > >>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> > >>  \t\ttest_path_is_file hello &&\n> > >> +\t\ttest_path_is_file \"$PATHNAME\" &&\n> > >>  \t\ttest world = $(cat hello)\n> > >>  \t)\n> > >>  '\n> > >\n> > > This test is currently failing on Cygwin: it looks like it's exposing a\n> > > bug in Cygwin that means files with colons in their name aren't\n> > > correctly extracted from zip archives.  I'm going to report that to the\n> > > Cygwin mailing list, but I wanted to note it for the record here, too.\n> >\n> > Does this mean that our code to set FUNNYNAMES prerequiste is\n> > slightly broken?  IOW, should we check with a path with a colon in\n> > it, as well as whatever we use currently for FUNNYNAMES?\n> >\n> > Something like the attached patch?\n> >\n> > Or does Cygwin otherwise work perfectly well with a path with a\n> > colon in it, but only $GIT_UNZIP command has problem with it?  If\n> > that is the case, then please disregard the attached.\n>\n> The latter: Cygwin works perfectly with paths containing colons, except\n> that Cygwin's `unzip` is seemingly buggy and doesn't work.  The file\n> systems Cygwin runs on don't support colons in paths, but Cygwin hides\n> that problem by rewriting ASCII colons to some high Unicode code point\n> on the filesystem,\n\nLet me throw in a bit more detail: The forbidden characters are mapped\ninto the Unicode page U+f0XX, which is supposed to be used \"for private\npurposes\". Even more detail can be found here:\nhttps://github.com/cygwin/cygwin/blob/cygwin-3_3_5-release/winsup/cygwin/strfuncs.cc#L19-L23\n\n> meaning Cygwin-native applications see a regular colon, while\n> Windows-native applications see an unusual but perfectly valid Unicode\n> character.\n\nNow, I have two questions:\n\n- Why does `unzip` not use Cygwin's regular functions (which should all be\n  aware of that U+f0XX <-> U+00XX mapping)?\n\n- Even more importantly: would the test case pass if we simply used\n  another forbidden character, such as `?` or `*`?\n\n> I tested the same patch to FUNNYNAMES myself before reporting, and the\n> test fails exactly the same way.  If we wanted to catch this, I think\n> we'd need a test that explicitly attempted to unzip an archive\n> containing a path with a colon.\n>\n> (The code to set FUNNYNAMES *is* slightly broken, per the discussions\n> around 6d340dfaef (\"t9902: split test to run on appropriate systems\",\n> 2022-04-08), and my to-do list still features tidying up and\n> resubmitting the patch Ævar wrote in that discussion thread.  But it\n> wouldn't help here because this issue is specific to Cygwin's `unzip`,\n> rather than a general limitation of running on Cygwin.)\n\nI'd rather avoid changing FUNNYNAMES at this stage, if we can help it.\n\nThanks,\nDscho\n"},{"id":"457507","messageId":"xmqqlettljad.fsf@gitster.g","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2206182213290.349@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-18T22:05:14Z","receivedAt":"2022-06-18T22:05:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> I'd rather avoid changing FUNNYNAMES at this stage, if we can help it.\n\nI wonder if it is sufficient to ask \"unzip -l\" the names of the\nfiles in the archive, without having to materialize these files on\nthe filesystem.  Would that bypass the whole FUNNYNAMES business, or\nis \"unzip\" paranoid enough to reject an archive, even when it is not\nextracting into the local filesystem, with a path that it would not\nbe able to extract if it were asked to?\n\nI do not know how standardized different implementations of \"unzip\"\nis, and how similar output \"unzip -l\" implementations produce are,\nbut the following seems to pass for me locally.\n\n t/t5003-archive-zip.sh | 18 ++++--------------\n 1 file changed, 4 insertions(+), 14 deletions(-)\n\ndiff --git c/t/t5003-archive-zip.sh w/t/t5003-archive-zip.sh\nindex 3992d08158..f2fdf2c235 100755\n--- c/t/t5003-archive-zip.sh\n+++ w/t/t5003-archive-zip.sh\n@@ -207,23 +207,13 @@ check_zip with_untracked\n check_added with_untracked untracked untracked\n \n test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n-\tif test_have_prereq FUNNYNAMES\n-\tthen\n-\t\tPATHNAME=\"pathname with : colon\"\n-\telse\n-\t\tPATHNAME=\"pathname without colon\"\n-\tfi &&\n+\tPATHNAME=\"pathname with : colon\" &&\n \tgit archive --format=zip >with_file_with_content.zip \\\n \t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n-\ttest_when_finished \"rm -rf tmp-unpack\" &&\n-\tmkdir tmp-unpack && (\n-\t\tcd tmp-unpack &&\n-\t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n-\t\ttest_path_is_file hello &&\n-\t\ttest_path_is_file \"$PATHNAME\" &&\n-\t\ttest world = $(cat hello)\n-\t)\n+\t\"$GIT_UNZIP\" -l with_file_with_content.zip >toc &&\n+\tgrep -e \" $PATHNAME\\$\" toc &&\n+\tgrep -e \" hello\\$\" toc\n '\n \n test_expect_success 'git archive --format=zip --add-file twice' '\n"},{"id":"457546","messageId":"20220620094141.uwvofvefoq26xxdu@lucy.dinwoodie.org","threadId":"57313","inReplyTo":"nycvar.QRO.7.76.6.2206182213290.349@tvgsbejvaqbjf.bet","subject":"Re: [PATCH v6+ 2/7] archive --add-virtual-file: allow paths containing colons","fromName":"Adam Dinwoodie","fromEmail":"adam@dinwoodie.org","sentAt":"2022-06-20T09:41:41Z","receivedAt":"2022-06-20T09:42:30Z","isPatch":true,"sender":{"key":"adam@dinwoodie.org","avatar":"https://avatars.githubusercontent.com/u/1397507?v=4"},"body":"On Sat, Jun 18, 2022 at 10:19:28PM +0200, Johannes Schindelin wrote:\n> Hi Adam,\n> \n> On Wed, 15 Jun 2022, Adam Dinwoodie wrote:\n> \n> > On Wed, Jun 15, 2022 at 01:00:07PM -0700, Junio C Hamano wrote:\n> > > Adam Dinwoodie <adam@dinwoodie.org> writes:\n> > >\n> > > >> diff --git a/t/t5003-archive-zip.sh b/t/t5003-archive-zip.sh\n> > > >> index d6027189e2..3992d08158 100755\n> > > >> --- a/t/t5003-archive-zip.sh\n> > > >> +++ b/t/t5003-archive-zip.sh\n> > > >> @@ -207,13 +207,21 @@ check_zip with_untracked\n> > > >>  check_added with_untracked untracked untracked\n> > > >>\n> > > >>  test_expect_success UNZIP 'git archive --format=zip --add-virtual-file' '\n> > > >> +\tif test_have_prereq FUNNYNAMES\n> > > >> +\tthen\n> > > >> +\t\tPATHNAME=\"pathname with : colon\"\n> > > >> +\telse\n> > > >> +\t\tPATHNAME=\"pathname without colon\"\n> > > >> +\tfi &&\n> > > >>  \tgit archive --format=zip >with_file_with_content.zip \\\n> > > >> +\t\t--add-virtual-file=\\\"\"$PATHNAME\"\\\": \\\n> > > >>  \t\t--add-virtual-file=hello:world $EMPTY_TREE &&\n> > > >>  \ttest_when_finished \"rm -rf tmp-unpack\" &&\n> > > >>  \tmkdir tmp-unpack && (\n> > > >>  \t\tcd tmp-unpack &&\n> > > >>  \t\t\"$GIT_UNZIP\" ../with_file_with_content.zip &&\n> > > >>  \t\ttest_path_is_file hello &&\n> > > >> +\t\ttest_path_is_file \"$PATHNAME\" &&\n> > > >>  \t\ttest world = $(cat hello)\n> > > >>  \t)\n> > > >>  '\n> > > >\n> > > > This test is currently failing on Cygwin: it looks like it's exposing a\n> > > > bug in Cygwin that means files with colons in their name aren't\n> > > > correctly extracted from zip archives.  I'm going to report that to the\n> > > > Cygwin mailing list, but I wanted to note it for the record here, too.\n> > >\n> > > Does this mean that our code to set FUNNYNAMES prerequiste is\n> > > slightly broken?  IOW, should we check with a path with a colon in\n> > > it, as well as whatever we use currently for FUNNYNAMES?\n> > >\n> > > Something like the attached patch?\n> > >\n> > > Or does Cygwin otherwise work perfectly well with a path with a\n> > > colon in it, but only $GIT_UNZIP command has problem with it?  If\n> > > that is the case, then please disregard the attached.\n> >\n> > The latter: Cygwin works perfectly with paths containing colons, except\n> > that Cygwin's `unzip` is seemingly buggy and doesn't work.  The file\n> > systems Cygwin runs on don't support colons in paths, but Cygwin hides\n> > that problem by rewriting ASCII colons to some high Unicode code point\n> > on the filesystem,\n> \n> Let me throw in a bit more detail: The forbidden characters are mapped\n> into the Unicode page U+f0XX, which is supposed to be used \"for private\n> purposes\". Even more detail can be found here:\n> https://github.com/cygwin/cygwin/blob/cygwin-3_3_5-release/winsup/cygwin/strfuncs.cc#L19-L23\n> \n> > meaning Cygwin-native applications see a regular colon, while\n> > Windows-native applications see an unusual but perfectly valid Unicode\n> > character.\n> \n> Now, I have two questions:\n> \n> - Why does `unzip` not use Cygwin's regular functions (which should all be\n>   aware of that U+f0XX <-> U+00XX mapping)?\n\nThat is an excellent question!  This behaviour came from an `#ifdef\n__CYGWIN__` in the upstream unzip package; with that #ifdef removed,\neverything works as expected.  The folk on the Cygwin mailing list had\nno idea *why* that #ifdef was there, given it's evidently unnecessary;\nmy best guess is that it was added a long time ago before Cygwin could\nhandle those characters in the general case.\n\nSince my report, the Cygwin package has picked up a new maintainer who\nhas released a version of the unzip package with that #ifdef removed, so\nthis test is now passing.\n\n> - Even more importantly: would the test case pass if we simply used\n>   another forbidden character, such as `?` or `*`?\n\nThe set of characters that had special handling in unzip was \"*:?|<> all\nof which are handled appropriately by Cygwin applications in general,\nand all of which had this unnecessary handling in `unzip`\n\n> > I tested the same patch to FUNNYNAMES myself before reporting, and the\n> > test fails exactly the same way.  If we wanted to catch this, I think\n> > we'd need a test that explicitly attempted to unzip an archive\n> > containing a path with a colon.\n> >\n> > (The code to set FUNNYNAMES *is* slightly broken, per the discussions\n> > around 6d340dfaef (\"t9902: split test to run on appropriate systems\",\n> > 2022-04-08), and my to-do list still features tidying up and\n> > resubmitting the patch Ævar wrote in that discussion thread.  But it\n> > wouldn't help here because this issue is specific to Cygwin's `unzip`,\n> > rather than a general limitation of running on Cygwin.)\n> \n> I'd rather avoid changing FUNNYNAMES at this stage, if we can help it.\n\nOh yes, I definitely wasn't proposing changing things for 2.37.0!  I\njust wanted to acknowledge that there is a known issue here that has\nbeen discussed on this list previously, that we (I) would hopefully get\naround to fixing at some point.\n\nAdam\n"}]}