{"thread":{"id":"42968","subject":"[PATCH v3 00/10] Git filter protocol","startedAt":"2016-07-29T23:38:12Z","lastAt":"2016-08-08T16:26:50Z","messageCount":100,"participants":["larsxschneider@gmail.com","Johannes Sixt","Jakub Narębski","Lars Schneider","Torstem Bögershausen","Torsten Bögershausen","Junio C Hamano","Jeff King"],"isPatch":true,"patchVersion":3,"patchTotal":10},"messages":[{"id":"292573","messageId":"20160729233801.82844-1-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160727000605.49982-1-larsxschneider%40gmail.com/","subject":"[PATCH v3 00/10] Git filter protocol","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:51Z","receivedAt":"2016-07-29T23:38:12Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nHi,\n\nthanks a lot to Jakub, Peff, Torsten, Eric, and Junio for comments\nand reviews.\n\nHere is what has changed since v2:\n\n* replace `/dev/urandom` with `test-genrandom` (Torsten, Peff)\n* improve commit message \"Git filter driver command with spaces\" (Jakub)\n* use proper types for memory and disk (Peff)\n* create packet read buffer with overflow check (Peff)\n* change capabilities format: \"capabilities clean smudge\" (Jakub)\n* replace \"%zu\" (Eric)\n* remove &= error handling (Eric, Peff)\n* initialize *argv[] with  { cmd, NULL } (Jakub)\n* reorder multi_packet_read() parameters to match read(2) (Eric)\n* do not continue if fstat fails (Eric)\n* filter: add reject response\n* add functions to pkt-line.h/c that: (Jakub, Peff)\n    - can write a packet without creating a new buffer\n    - do not die in case of a failure\n* add function to pkt-line.h/c that writes a pkt-line flush and does not die on error\n* add filter stream capability\n* add filter shutdown capability\n* docs: fix LARGE_PACKET_MAX documentaion\n    see http://public-inbox.org/git/20160726134257.GB19277%40sigill.intra.peff.net/\n* docs: fix s/seperated/separated/ (Jakub)\n* docs: \"mis-configured one-shot filters would hang\" (Jakub)\n* docs: filter protocol filename absolute (Jakub)\n* docs: state that Git can use more than one packet (Jabub)\n* docs: add \"\\n\" to lines (Jakub)\n* docs: filter precedence (Jakub)\n\nCheers,\nLars\n\nPS: If you prefer checkout the code from a Git repo instead then you can find\nit here: https://github.com/larsxschneider/git/tree/protocol-filter/v3\n\n\nLars Schneider (10):\n  pkt-line: extract set_packet_header()\n  pkt-line: add direct_packet_write() and direct_packet_write_data()\n  pkt-line: add packet_flush_gentle()\n  pkt-line: call packet_trace() only if a packet is actually send\n  pack-protocol: fix maximum pkt-line size\n  run-command: add clean_on_exit_handler\n  convert: quote filter names in error messages\n  convert: modernize tests\n  convert: generate large test files only once\n  convert: add filter.<driver>.process option\n\n Documentation/gitattributes.txt             |  84 ++++-\n Documentation/technical/protocol-common.txt |   6 +-\n convert.c                                   | 412 +++++++++++++++++++++-\n pkt-line.c                                  |  53 ++-\n pkt-line.h                                  |   6 +\n run-command.c                               |  12 +-\n run-command.h                               |   1 +\n t/t0021-conversion.sh                       | 515 +++++++++++++++++++++++++---\n t/t0021/rot13-filter.pl                     | 177 ++++++++++\n 9 files changed, 1193 insertions(+), 73 deletions(-)\n create mode 100755 t/t0021/rot13-filter.pl\n\n--\n2.9.0\n\n"},{"id":"292574","messageId":"20160729233801.82844-2-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 01/10] pkt-line: extract set_packet_header()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:52Z","receivedAt":"2016-07-29T23:38:16Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nset_packet_header() converts an integer to a 4 byte hex string. Make\nthis function locally available so that other pkt-line functions can\nuse it.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 15 ++++++++++-----\n 1 file changed, 10 insertions(+), 5 deletions(-)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 62fdb37..445b8e1 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -98,9 +98,17 @@ void packet_buf_flush(struct strbuf *buf)\n }\n \n #define hex(a) (hexchar[(a) & 15])\n-static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n+static void set_packet_header(char *buf, const int size)\n {\n \tstatic char hexchar[] = \"0123456789abcdef\";\n+\tbuf[0] = hex(size >> 12);\n+\tbuf[1] = hex(size >> 8);\n+\tbuf[2] = hex(size >> 4);\n+\tbuf[3] = hex(size);\n+}\n+\n+static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n+{\n \tsize_t orig_len, n;\n \n \torig_len = out->len;\n@@ -111,10 +119,7 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n \tif (n > LARGE_PACKET_MAX)\n \t\tdie(\"protocol error: impossibly long line\");\n \n-\tout->buf[orig_len + 0] = hex(n >> 12);\n-\tout->buf[orig_len + 1] = hex(n >> 8);\n-\tout->buf[orig_len + 2] = hex(n >> 4);\n-\tout->buf[orig_len + 3] = hex(n);\n+\tset_packet_header(&out->buf[orig_len], n);\n \tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n }\n \n-- \n2.9.0\n\n"},{"id":"292575","messageId":"20160729233801.82844-3-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 02/10] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:53Z","receivedAt":"2016-07-29T23:38:20Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nSometimes pkt-line data is already available in a buffer and it would\nbe a waste of resources to write the packet using packet_write() which\nwould copy the existing buffer into a strbuf before writing it.\n\nIf the caller has control over the buffer creation then the\nPKTLINE_DATA_START macro can be used to skip the header and write\ndirectly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\nwould be the maximum). direct_packet_write() would take this buffer,\nadjust the pkt-line header and write it.\n\nIf the caller has no control over the buffer creation then\ndirect_packet_write_data() can be used. This function creates a pkt-line\nheader. Afterwards the header and the data buffer are written using two\nconsecutive write calls.\n\nBoth functions have a gentle parameter that indicates if Git should die\nin case of a write error (gentle set to 0) or return with a error (gentle\nset to 1).\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 30 ++++++++++++++++++++++++++++++\n pkt-line.h |  5 +++++\n 2 files changed, 35 insertions(+)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 445b8e1..6fae508 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -135,6 +135,36 @@ void packet_write(int fd, const char *fmt, ...)\n \twrite_or_die(fd, buf.buf, buf.len);\n }\n \n+int direct_packet_write(int fd, char *buf, size_t size, int gentle)\n+{\n+\tint ret = 0;\n+\tpacket_trace(buf + 4, size - 4, 1);\n+\tset_packet_header(buf, size);\n+\tif (gentle)\n+\t\tret = !write_or_whine_pipe(fd, buf, size, \"pkt-line\");\n+\telse\n+\t\twrite_or_die(fd, buf, size);\n+\treturn ret;\n+}\n+\n+int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle)\n+{\n+\tint ret = 0;\n+\tchar hdr[4];\n+\tset_packet_header(hdr, sizeof(hdr) + size);\n+\tpacket_trace(buf, size, 1);\n+\tif (gentle) {\n+\t\tret = (\n+\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n+\t\t\t!write_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n+\t\t);\n+\t} else {\n+\t\twrite_or_die(fd, hdr, sizeof(hdr));\n+\t\twrite_or_die(fd, buf, size);\n+\t}\n+\treturn ret;\n+}\n+\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...)\n {\n \tva_list args;\ndiff --git a/pkt-line.h b/pkt-line.h\nindex 3cb9d91..02dcced 100644\n--- a/pkt-line.h\n+++ b/pkt-line.h\n@@ -23,6 +23,8 @@ void packet_flush(int fd);\n void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n void packet_buf_flush(struct strbuf *buf);\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n+int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n+int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n \n /*\n  * Read a packetized line into the buffer, which must be at least size bytes\n@@ -77,6 +79,9 @@ char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);\n \n #define DEFAULT_PACKET_MAX 1000\n #define LARGE_PACKET_MAX 65520\n+#define PKTLINE_HEADER_LEN 4\n+#define PKTLINE_DATA_START(pkt) ((pkt) + PKTLINE_HEADER_LEN)\n+#define PKTLINE_DATA_LEN (LARGE_PACKET_MAX - PKTLINE_HEADER_LEN)\n extern char packet_buffer[LARGE_PACKET_MAX];\n \n #endif\n-- \n2.9.0\n\n"},{"id":"292576","messageId":"20160729233801.82844-4-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:54Z","receivedAt":"2016-07-29T23:38:23Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\npacket_flush() would die in case of a write error even though for some callers\nan error would be acceptable. Add packet_flush_gentle() which writes a pkt-line\nflush packet and returns `0` for success and `1` for failure.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 6 ++++++\n pkt-line.h | 1 +\n 2 files changed, 7 insertions(+)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 6fae508..1728690 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -91,6 +91,12 @@ void packet_flush(int fd)\n \twrite_or_die(fd, \"0000\", 4);\n }\n \n+int packet_flush_gentle(int fd)\n+{\n+\tpacket_trace(\"0000\", 4, 1);\n+\treturn !write_or_whine_pipe(fd, \"0000\", 4, \"flush packet\");\n+}\n+\n void packet_buf_flush(struct strbuf *buf)\n {\n \tpacket_trace(\"0000\", 4, 1);\ndiff --git a/pkt-line.h b/pkt-line.h\nindex 02dcced..3953c98 100644\n--- a/pkt-line.h\n+++ b/pkt-line.h\n@@ -23,6 +23,7 @@ void packet_flush(int fd);\n void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n void packet_buf_flush(struct strbuf *buf);\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n+int packet_flush_gentle(int fd);\n int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n \n-- \n2.9.0\n\n"},{"id":"292577","messageId":"20160729233801.82844-5-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 04/10] pkt-line: call packet_trace() only if a packet is actually send","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:55Z","receivedAt":"2016-07-29T23:38:26Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nThe packet_trace() call is not ideal in format_packet() as we would print\na trace when a packet is formatted and (potentially) when the packet is\nactually send. This was no problem up until now because format_packet()\nwas only used by one function. Fix it by moving the trace call into the\nfunction that actally sends the packet.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 1728690..32c0a34 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -126,7 +126,6 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n \t\tdie(\"protocol error: impossibly long line\");\n \n \tset_packet_header(&out->buf[orig_len], n);\n-\tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n }\n \n void packet_write(int fd, const char *fmt, ...)\n@@ -138,6 +137,7 @@ void packet_write(int fd, const char *fmt, ...)\n \tva_start(args, fmt);\n \tformat_packet(&buf, fmt, args);\n \tva_end(args);\n+\tpacket_trace(buf.buf + 4, buf.len - 4, 1);\n \twrite_or_die(fd, buf.buf, buf.len);\n }\n \n-- \n2.9.0\n\n"},{"id":"292578","messageId":"20160729233801.82844-6-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 05/10] pack-protocol: fix maximum pkt-line size","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:56Z","receivedAt":"2016-07-29T23:38:29Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nAccording to LARGE_PACKET_MAX in pkt-line.h the maximal lenght of a\npkt-line packet is 65520 bytes. The pkt-line header takes 4 bytes and\ntherefore the pkt-line data component must not exceed 65516 bytes.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n Documentation/technical/protocol-common.txt | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/technical/protocol-common.txt b/Documentation/technical/protocol-common.txt\nindex bf30167..ecedb34 100644\n--- a/Documentation/technical/protocol-common.txt\n+++ b/Documentation/technical/protocol-common.txt\n@@ -67,9 +67,9 @@ with non-binary data the same whether or not they contain the trailing\n LF (stripping the LF if present, and not complaining when it is\n missing).\n \n-The maximum length of a pkt-line's data component is 65520 bytes.\n-Implementations MUST NOT send pkt-line whose length exceeds 65524\n-(65520 bytes of payload + 4 bytes of length data).\n+The maximum length of a pkt-line's data component is 65516 bytes.\n+Implementations MUST NOT send pkt-line whose length exceeds 65520\n+(65516 bytes of payload + 4 bytes of length data).\n \n Implementations SHOULD NOT send an empty pkt-line (\"0004\").\n \n-- \n2.9.0\n\n"},{"id":"292579","messageId":"20160729233801.82844-8-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 07/10] convert: quote filter names in error messages","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:58Z","receivedAt":"2016-07-29T23:38:32Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGit filter driver commands with spaces (e.g. `filter.sh foo`) are hard to\nread in error messages. Quote them to improve the readability.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n convert.c | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/convert.c b/convert.c\nindex b1614bf..522e2c5 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -397,7 +397,7 @@ static int filter_buffer_or_fd(int in, int out, void *data)\n \tchild_process.out = out;\n \n \tif (start_command(&child_process))\n-\t\treturn error(\"cannot fork to run external filter %s\", params->cmd);\n+\t\treturn error(\"cannot fork to run external filter '%s'\", params->cmd);\n \n \tsigchain_push(SIGPIPE, SIG_IGN);\n \n@@ -415,13 +415,13 @@ static int filter_buffer_or_fd(int in, int out, void *data)\n \tif (close(child_process.in))\n \t\twrite_err = 1;\n \tif (write_err)\n-\t\terror(\"cannot feed the input to external filter %s\", params->cmd);\n+\t\terror(\"cannot feed the input to external filter '%s'\", params->cmd);\n \n \tsigchain_pop(SIGPIPE);\n \n \tstatus = finish_command(&child_process);\n \tif (status)\n-\t\terror(\"external filter %s failed %d\", params->cmd, status);\n+\t\terror(\"external filter '%s' failed %d\", params->cmd, status);\n \n \tstrbuf_release(&cmd);\n \treturn (write_err || status);\n@@ -462,15 +462,15 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n \t\treturn 0;\t/* error was already reported */\n \n \tif (strbuf_read(&nbuf, async.out, len) < 0) {\n-\t\terror(\"read from external filter %s failed\", cmd);\n+\t\terror(\"read from external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \tif (close(async.out)) {\n-\t\terror(\"read from external filter %s failed\", cmd);\n+\t\terror(\"read from external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \tif (finish_async(&async)) {\n-\t\terror(\"external filter %s failed\", cmd);\n+\t\terror(\"external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \n-- \n2.9.0\n\n"},{"id":"292580","messageId":"20160729233801.82844-7-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 06/10] run-command: add clean_on_exit_handler","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:57Z","receivedAt":"2016-07-29T23:38:35Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nSome commands might need to perform cleanup tasks on exit. Let's give\nthem an interface for doing this.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n run-command.c | 12 ++++++++----\n run-command.h |  1 +\n 2 files changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/run-command.c b/run-command.c\nindex 33bc63a..197b534 100644\n--- a/run-command.c\n+++ b/run-command.c\n@@ -21,6 +21,7 @@ void child_process_clear(struct child_process *child)\n \n struct child_to_clean {\n \tpid_t pid;\n+\tvoid (*clean_on_exit_handler)(pid_t);\n \tstruct child_to_clean *next;\n };\n static struct child_to_clean *children_to_clean;\n@@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n {\n \twhile (children_to_clean) {\n \t\tstruct child_to_clean *p = children_to_clean;\n+\t\tif (p->clean_on_exit_handler)\n+\t\t\tp->clean_on_exit_handler(p->pid);\n \t\tchildren_to_clean = p->next;\n \t\tkill(p->pid, sig);\n \t\tif (!in_signal)\n@@ -49,10 +52,11 @@ static void cleanup_children_on_exit(void)\n \tcleanup_children(SIGTERM, 0);\n }\n \n-static void mark_child_for_cleanup(pid_t pid)\n+static void mark_child_for_cleanup(pid_t pid, void (*clean_on_exit_handler)(pid_t))\n {\n \tstruct child_to_clean *p = xmalloc(sizeof(*p));\n \tp->pid = pid;\n+\tp->clean_on_exit_handler = clean_on_exit_handler;\n \tp->next = children_to_clean;\n \tchildren_to_clean = p;\n \n@@ -422,7 +426,7 @@ int start_command(struct child_process *cmd)\n \tif (cmd->pid < 0)\n \t\terror_errno(\"cannot fork() for %s\", cmd->argv[0]);\n \telse if (cmd->clean_on_exit)\n-\t\tmark_child_for_cleanup(cmd->pid);\n+\t\tmark_child_for_cleanup(cmd->pid, cmd->clean_on_exit_handler);\n \n \t/*\n \t * Wait for child's execvp. If the execvp succeeds (or if fork()\n@@ -483,7 +487,7 @@ int start_command(struct child_process *cmd)\n \tif (cmd->pid < 0 && (!cmd->silent_exec_failure || errno != ENOENT))\n \t\terror_errno(\"cannot spawn %s\", cmd->argv[0]);\n \tif (cmd->clean_on_exit && cmd->pid >= 0)\n-\t\tmark_child_for_cleanup(cmd->pid);\n+\t\tmark_child_for_cleanup(cmd->pid, cmd->clean_on_exit_handler);\n \n \targv_array_clear(&nargv);\n \tcmd->argv = sargv;\n@@ -752,7 +756,7 @@ int start_async(struct async *async)\n \t\texit(!!async->proc(proc_in, proc_out, async->data));\n \t}\n \n-\tmark_child_for_cleanup(async->pid);\n+\tmark_child_for_cleanup(async->pid, NULL);\n \n \tif (need_in)\n \t\tclose(fdin[0]);\ndiff --git a/run-command.h b/run-command.h\nindex 5066649..59d21ea 100644\n--- a/run-command.h\n+++ b/run-command.h\n@@ -43,6 +43,7 @@ struct child_process {\n \tunsigned stdout_to_stderr:1;\n \tunsigned use_shell:1;\n \tunsigned clean_on_exit:1;\n+\tvoid (*clean_on_exit_handler)(pid_t);\n };\n \n #define CHILD_PROCESS_INIT { NULL, ARGV_ARRAY_INIT, ARGV_ARRAY_INIT }\n-- \n2.9.0\n\n"},{"id":"292581","messageId":"20160729233801.82844-9-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 08/10] convert: modernize tests","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:37:59Z","receivedAt":"2016-07-29T23:38:36Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nUse `test_config` to set the config, check that files are empty with\n`test_must_be_empty`, compare files with `test_cmp`, and remove spaces\nafter \">\" and \"<\".\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh | 62 +++++++++++++++++++++++++--------------------------\n 1 file changed, 31 insertions(+), 31 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 7bac2bc..7b45136 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -13,8 +13,8 @@ EOF\n chmod +x rot13.sh\n \n test_expect_success setup '\n-\tgit config filter.rot13.smudge ./rot13.sh &&\n-\tgit config filter.rot13.clean ./rot13.sh &&\n+\ttest_config filter.rot13.smudge ./rot13.sh &&\n+\ttest_config filter.rot13.clean ./rot13.sh &&\n \n \t{\n \t    echo \"*.t filter=rot13\"\n@@ -38,8 +38,8 @@ script='s/^\\$Id: \\([0-9a-f]*\\) \\$/\\1/p'\n \n test_expect_success check '\n \n-\tcmp test.o test &&\n-\tcmp test.o test.t &&\n+\ttest_cmp test.o test &&\n+\ttest_cmp test.o test.t &&\n \n \t# ident should be stripped in the repository\n \tgit diff --raw --exit-code :test :test.i &&\n@@ -47,10 +47,10 @@ test_expect_success check '\n \tembedded=$(sed -ne \"$script\" test.i) &&\n \ttest \"z$id\" = \"z$embedded\" &&\n \n-\tgit cat-file blob :test.t > test.r &&\n+\tgit cat-file blob :test.t >test.r &&\n \n-\t./rot13.sh < test.o > test.t &&\n-\tcmp test.r test.t\n+\t./rot13.sh <test.o >test.t &&\n+\ttest_cmp test.r test.t\n '\n \n # If an expanded ident ever gets into the repository, we want to make sure that\n@@ -130,7 +130,7 @@ test_expect_success 'filter shell-escaped filenames' '\n \n \t# delete the files and check them out again, using a smudge filter\n \t# that will count the args and echo the command-line back to us\n-\tgit config filter.argc.smudge \"sh ./argc.sh %f\" &&\n+\ttest_config filter.argc.smudge \"sh ./argc.sh %f\" &&\n \trm \"$normal\" \"$special\" &&\n \tgit checkout -- \"$normal\" \"$special\" &&\n \n@@ -141,7 +141,7 @@ test_expect_success 'filter shell-escaped filenames' '\n \ttest_cmp expect \"$special\" &&\n \n \t# do the same thing, but with more args in the filter expression\n-\tgit config filter.argc.smudge \"sh ./argc.sh %f --my-extra-arg\" &&\n+\ttest_config filter.argc.smudge \"sh ./argc.sh %f --my-extra-arg\" &&\n \trm \"$normal\" \"$special\" &&\n \tgit checkout -- \"$normal\" \"$special\" &&\n \n@@ -154,9 +154,9 @@ test_expect_success 'filter shell-escaped filenames' '\n '\n \n test_expect_success 'required filter should filter data' '\n-\tgit config filter.required.smudge ./rot13.sh &&\n-\tgit config filter.required.clean ./rot13.sh &&\n-\tgit config filter.required.required true &&\n+\ttest_config filter.required.smudge ./rot13.sh &&\n+\ttest_config filter.required.clean ./rot13.sh &&\n+\ttest_config filter.required.required true &&\n \n \techo \"*.r filter=required\" >.gitattributes &&\n \n@@ -165,17 +165,17 @@ test_expect_success 'required filter should filter data' '\n \n \trm -f test.r &&\n \tgit checkout -- test.r &&\n-\tcmp test.o test.r &&\n+\ttest_cmp test.o test.r &&\n \n \t./rot13.sh <test.o >expected &&\n \tgit cat-file blob :test.r >actual &&\n-\tcmp expected actual\n+\ttest_cmp expected actual\n '\n \n test_expect_success 'required filter smudge failure' '\n-\tgit config filter.failsmudge.smudge false &&\n-\tgit config filter.failsmudge.clean cat &&\n-\tgit config filter.failsmudge.required true &&\n+\ttest_config filter.failsmudge.smudge false &&\n+\ttest_config filter.failsmudge.clean cat &&\n+\ttest_config filter.failsmudge.required true &&\n \n \techo \"*.fs filter=failsmudge\" >.gitattributes &&\n \n@@ -186,9 +186,9 @@ test_expect_success 'required filter smudge failure' '\n '\n \n test_expect_success 'required filter clean failure' '\n-\tgit config filter.failclean.smudge cat &&\n-\tgit config filter.failclean.clean false &&\n-\tgit config filter.failclean.required true &&\n+\ttest_config filter.failclean.smudge cat &&\n+\ttest_config filter.failclean.clean false &&\n+\ttest_config filter.failclean.required true &&\n \n \techo \"*.fc filter=failclean\" >.gitattributes &&\n \n@@ -197,8 +197,8 @@ test_expect_success 'required filter clean failure' '\n '\n \n test_expect_success 'filtering large input to small output should use little memory' '\n-\tgit config filter.devnull.clean \"cat >/dev/null\" &&\n-\tgit config filter.devnull.required true &&\n+\ttest_config filter.devnull.clean \"cat >/dev/null\" &&\n+\ttest_config filter.devnull.required true &&\n \tfor i in $(test_seq 1 30); do printf \"%1048576d\" 1; done >30MB &&\n \techo \"30MB filter=devnull\" >.gitattributes &&\n \tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add 30MB\n@@ -207,7 +207,7 @@ test_expect_success 'filtering large input to small output should use little mem\n test_expect_success 'filter that does not read is fine' '\n \ttest-genrandom foo $((128 * 1024 + 1)) >big &&\n \techo \"big filter=epipe\" >.gitattributes &&\n-\tgit config filter.epipe.clean \"echo xyzzy\" &&\n+\ttest_config filter.epipe.clean \"echo xyzzy\" &&\n \tgit add big &&\n \tgit cat-file blob :big >actual &&\n \techo xyzzy >expect &&\n@@ -215,20 +215,20 @@ test_expect_success 'filter that does not read is fine' '\n '\n \n test_expect_success EXPENSIVE 'filter large file' '\n-\tgit config filter.largefile.smudge cat &&\n-\tgit config filter.largefile.clean cat &&\n+\ttest_config filter.largefile.smudge cat &&\n+\ttest_config filter.largefile.clean cat &&\n \tfor i in $(test_seq 1 2048); do printf \"%1048576d\" 1; done >2GB &&\n \techo \"2GB filter=largefile\" >.gitattributes &&\n \tgit add 2GB 2>err &&\n-\t! test -s err &&\n+\ttest_must_be_empty err &&\n \trm -f 2GB &&\n \tgit checkout -- 2GB 2>err &&\n-\t! test -s err\n+\ttest_must_be_empty err\n '\n \n test_expect_success \"filter: clean empty file\" '\n-\tgit config filter.in-repo-header.clean  \"echo cleaned && cat\" &&\n-\tgit config filter.in-repo-header.smudge \"sed 1d\" &&\n+\ttest_config filter.in-repo-header.clean  \"echo cleaned && cat\" &&\n+\ttest_config filter.in-repo-header.smudge \"sed 1d\" &&\n \n \techo \"empty-in-worktree    filter=in-repo-header\" >>.gitattributes &&\n \t>empty-in-worktree &&\n@@ -240,8 +240,8 @@ test_expect_success \"filter: clean empty file\" '\n '\n \n test_expect_success \"filter: smudge empty file\" '\n-\tgit config filter.empty-in-repo.clean \"cat >/dev/null\" &&\n-\tgit config filter.empty-in-repo.smudge \"echo smudged && cat\" &&\n+\ttest_config filter.empty-in-repo.clean \"cat >/dev/null\" &&\n+\ttest_config filter.empty-in-repo.smudge \"echo smudged && cat\" &&\n \n \techo \"empty-in-repo filter=empty-in-repo\" >>.gitattributes &&\n \techo dead data walking >empty-in-repo &&\n-- \n2.9.0\n\n"},{"id":"292582","messageId":"20160729233801.82844-10-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 09/10] convert: generate large test files only once","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:38:00Z","receivedAt":"2016-07-29T23:38:40Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGenerate more interesting large test files with pseudo random characters\nin between and reuse these test files in multiple tests. Run tests formerly\nmarked as EXPENSIVE every time but with a smaller data set.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh | 48 ++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 38 insertions(+), 10 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 7b45136..34c8eb9 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -4,6 +4,15 @@ test_description='blob conversion via gitattributes'\n \n . ./test-lib.sh\n \n+if test_have_prereq EXPENSIVE\n+then\n+\tT0021_LARGE_FILE_SIZE=2048\n+\tT0021_LARGISH_FILE_SIZE=100\n+else\n+\tT0021_LARGE_FILE_SIZE=30\n+\tT0021_LARGISH_FILE_SIZE=2\n+fi\n+\n cat <<EOF >rot13.sh\n #!$SHELL_PATH\n tr \\\n@@ -31,7 +40,26 @@ test_expect_success setup '\n \tcat test >test.i &&\n \tgit add test test.t test.i &&\n \trm -f test test.t test.i &&\n-\tgit checkout -- test test.t test.i\n+\tgit checkout -- test test.t test.i &&\n+\n+\tmkdir generated-test-data &&\n+\tfor i in $(test_seq 1 $T0021_LARGE_FILE_SIZE)\n+\tdo\n+\t\tRANDOM_STRING=\"$(test-genrandom end $i | tr -dc \"A-Za-z0-9\" )\"\n+\t\tROT_RANDOM_STRING=\"$(echo $RANDOM_STRING | ./rot13.sh )\"\n+\t\t# Generate 1MB of empty data and 100 bytes of random characters\n+\t\t# printf \"$(test-genrandom start $i)\"\n+\t\tprintf \"%1048576d\" 1 >>generated-test-data/large.file &&\n+\t\tprintf \"$RANDOM_STRING\" >>generated-test-data/large.file &&\n+\t\tprintf \"%1048576d\" 1 >>generated-test-data/large.file.rot13 &&\n+\t\tprintf \"$ROT_RANDOM_STRING\" >>generated-test-data/large.file.rot13 &&\n+\n+\t\tif test $i = $T0021_LARGISH_FILE_SIZE\n+\t\tthen\n+\t\t\tcat generated-test-data/large.file >generated-test-data/largish.file &&\n+\t\t\tcat generated-test-data/large.file.rot13 >generated-test-data/largish.file.rot13\n+\t\tfi\n+\tdone\n '\n \n script='s/^\\$Id: \\([0-9a-f]*\\) \\$/\\1/p'\n@@ -199,9 +227,9 @@ test_expect_success 'required filter clean failure' '\n test_expect_success 'filtering large input to small output should use little memory' '\n \ttest_config filter.devnull.clean \"cat >/dev/null\" &&\n \ttest_config filter.devnull.required true &&\n-\tfor i in $(test_seq 1 30); do printf \"%1048576d\" 1; done >30MB &&\n-\techo \"30MB filter=devnull\" >.gitattributes &&\n-\tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add 30MB\n+\tcp generated-test-data/large.file large.file &&\n+\techo \"large.file filter=devnull\" >.gitattributes &&\n+\tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add large.file\n '\n \n test_expect_success 'filter that does not read is fine' '\n@@ -214,15 +242,15 @@ test_expect_success 'filter that does not read is fine' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success EXPENSIVE 'filter large file' '\n+test_expect_success 'filter large file' '\n \ttest_config filter.largefile.smudge cat &&\n \ttest_config filter.largefile.clean cat &&\n-\tfor i in $(test_seq 1 2048); do printf \"%1048576d\" 1; done >2GB &&\n-\techo \"2GB filter=largefile\" >.gitattributes &&\n-\tgit add 2GB 2>err &&\n+\techo \"large.file filter=largefile\" >.gitattributes &&\n+\tcp generated-test-data/large.file large.file &&\n+\tgit add large.file 2>err &&\n \ttest_must_be_empty err &&\n-\trm -f 2GB &&\n-\tgit checkout -- 2GB 2>err &&\n+\trm -f large.file &&\n+\tgit checkout -- large.file 2>err &&\n \ttest_must_be_empty err\n '\n \n-- \n2.9.0\n\n"},{"id":"292583","messageId":"20160729233801.82844-11-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-29T23:38:01Z","receivedAt":"2016-07-29T23:38:44Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGit's clean/smudge mechanism invokes an external filter process for every\nsingle blob that is affected by a filter. If Git filters a lot of blobs\nthen the startup time of the external filter processes can become a\nsignificant part of the overall Git execution time.\n\nThis patch adds the filter.<driver>.process string option which, if used,\nkeeps the external filter process running and processes all blobs with\nthe following packet format (pkt-line) based protocol over standard input\nand standard output.\n\nGit starts the filter on first usage and expects a welcome\nmessage, protocol version number, and filter capabilities\nseparated by spaces:\n------------------------\npacket:          git< git-filter-protocol\\n\npacket:          git< version 2\\n\npacket:          git< capabilities clean smudge\\n\n------------------------\nSupported filter capabilities are \"clean\", \"smudge\", \"stream\",\nand \"shutdown\".\n\nAfterwards Git sends a command (based on the supported\ncapabilities), the filename including its path\nrelative to the repository root, the content size as ASCII number\nin bytes, the content split in zero or many pkt-line packets,\nand a flush packet at the end:\n------------------------\npacket:          git> smudge\\n\npacket:          git> filename=path/testfile.dat\\n\npacket:          git> size=7\\n\npacket:          git> CONTENT\npacket:          git> 0000\n------------------------\n\nThe filter is expected to respond with the result content size as\nASCII number in bytes. If the capability \"stream\" is defined then\nthe filter must not send the content size. Afterwards the result\ncontent in send in zero or many pkt-line packets and a flush packet\nat the end. Finally a \"success\" packet is send to indicate that\neverything went well.\n------------------------\npacket:          git< size=57\\n   (omitted with capability \"stream\")\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< success\\n\n------------------------\n\nIn case the filter cannot process the content, it is expected\nto respond with the result content size 0 (only if \"stream\" is\nnot defined) and a \"reject\" packet.\n------------------------\npacket:          git< size=0\\n    (omitted with capability \"stream\")\npacket:          git< reject\\n\n------------------------\n\nAfter the filter has processed a blob it is expected to wait for\nthe next command. A demo implementation can be found in\n`t/t0021/rot13-filter.pl` located in the Git core repository.\n\nIf the filter supports the \"shutdown\" capability then Git will\nsend the \"shutdown\" command and wait until the filter answers\nwith \"done\". This gives the filter the opportunity to perform\ncleanup tasks. Afterwards the filter is expected to exit.\n------------------------\npacket:          git> shutdown\\n\npacket:          git< done\\n\n------------------------\n\nIf a filter.<driver>.clean or filter.<driver>.smudge command\nis configured then these commands always take precedence over\na configured filter.<driver>.process command.\n\nPlease note that you cannot use an existing filter.<driver>.clean\nor filter.<driver>.smudge command as filter.<driver>.process\ncommand. As soon as Git would detect a file that needs to be\nprocessed by this filter, it would stop responding.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\nHelped-by: Martin-Louis Bright <mlbright@gmail.com>\n---\n Documentation/gitattributes.txt |  84 ++++++++-\n convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--\n t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++\n t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++\n 4 files changed, 1053 insertions(+), 13 deletions(-)\n create mode 100755 t/t0021/rot13-filter.pl\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 8882a3e..e3fbcc2 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -300,7 +300,11 @@ checkout, when the `smudge` command is specified, the command is\n fed the blob object from its standard input, and its standard\n output is used to update the worktree file.  Similarly, the\n `clean` command is used to convert the contents of worktree file\n-upon checkin.\n+upon checkin. By default these commands process only a single\n+blob and terminate. If a long running filter process (see section\n+below) is used then Git can process all blobs with a single filter\n+invocation for the entire life of a single Git command (e.g.\n+`git add .`).\n \n One use of the content filtering is to massage the content into a shape\n that is more convenient for the platform, filesystem, and the user to use.\n@@ -375,6 +379,84 @@ substitution.  For example:\n ------------------------\n \n \n+Long Running Filter Process\n+^^^^^^^^^^^^^^^^^^^^^^^^^^^\n+\n+If the filter command (string value) is defined via\n+filter.<driver>.process then Git can process all blobs with a\n+single filter invocation for the entire life of a single Git\n+command. This is achieved by using the following packet\n+format (pkt-line, see protocol-common.txt) based protocol over\n+standard input and standard output.\n+\n+Git starts the filter on first usage and expects a welcome\n+message, protocol version number, and filter capabilities\n+separated by spaces:\n+------------------------\n+packet:          git< git-filter-protocol\\n\n+packet:          git< version 2\\n\n+packet:          git< capabilities clean smudge\\n\n+------------------------\n+Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n+and \"shutdown\".\n+\n+Afterwards Git sends a command (based on the supported\n+capabilities), the filename including its path\n+relative to the repository root, the content size as ASCII number\n+in bytes, the content split in zero or many pkt-line packets,\n+and a flush packet at the end:\n+------------------------\n+packet:          git> smudge\\n\n+packet:          git> filename=path/testfile.dat\\n\n+packet:          git> size=7\\n\n+packet:          git> CONTENT\n+packet:          git> 0000\n+------------------------\n+\n+The filter is expected to respond with the result content size as\n+ASCII number in bytes. If the capability \"stream\" is defined then\n+the filter must not send the content size. Afterwards the result\n+content in send in zero or many pkt-line packets and a flush packet\n+at the end. Finally a \"success\" packet is send to indicate that\n+everything went well.\n+------------------------\n+packet:          git< size=57\\n   (omitted with capability \"stream\")\n+packet:          git< SMUDGED_CONTENT\n+packet:          git< 0000\n+packet:          git< success\\n\n+------------------------\n+\n+In case the filter cannot process the content, it is expected\n+to respond with the result content size 0 (only if \"stream\" is\n+not defined) and a \"reject\" packet.\n+------------------------\n+packet:          git< size=0\\n    (omitted with capability \"stream\")\n+packet:          git< reject\\n\n+------------------------\n+\n+After the filter has processed a blob it is expected to wait for\n+the next command. A demo implementation can be found in\n+`t/t0021/rot13-filter.pl` located in the Git core repository.\n+\n+If the filter supports the \"shutdown\" capability then Git will\n+send the \"shutdown\" command and wait until the filter answers\n+with \"done\". This gives the filter the opportunity to perform\n+cleanup tasks. Afterwards the filter is expected to exit.\n+------------------------\n+packet:          git> shutdown\\n\n+packet:          git< done\\n\n+------------------------\n+\n+If a filter.<driver>.clean or filter.<driver>.smudge command\n+is configured then these commands always take precedence over\n+a configured filter.<driver>.process command.\n+\n+Please note that you cannot use an existing filter.<driver>.clean\n+or filter.<driver>.smudge command as filter.<driver>.process\n+command. As soon as Git would detect a file that needs to be\n+processed by this filter, it would stop responding.\n+\n+\n Interaction between checkin/checkout attributes\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \ndiff --git a/convert.c b/convert.c\nindex 522e2c5..be6405c 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -3,6 +3,7 @@\n #include \"run-command.h\"\n #include \"quote.h\"\n #include \"sigchain.h\"\n+#include \"pkt-line.h\"\n \n /*\n  * convert.c - convert a file when checking it out and checking it in.\n@@ -481,11 +482,355 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n \treturn ret;\n }\n \n+static int multi_packet_read(int fd_in, struct strbuf *sb, size_t expected_bytes, int is_stream)\n+{\n+\tint bytes_read;\n+\tsize_t total_bytes_read = 0;\n+\tif (expected_bytes == 0 && !is_stream)\n+\t\treturn 0;\n+\n+\tif (is_stream)\n+\t\tstrbuf_grow(sb, LARGE_PACKET_MAX);           // allocate space for at least one packet\n+\telse\n+\t\tstrbuf_grow(sb, st_add(expected_bytes, 1));  // add one extra byte for the packet flush\n+\n+\tdo {\n+\t\tbytes_read = packet_read(\n+\t\t\tfd_in, NULL, NULL,\n+\t\t\tsb->buf + total_bytes_read, sb->len - total_bytes_read - 1,\n+\t\t\tPACKET_READ_GENTLE_ON_EOF\n+\t\t);\n+\t\tif (bytes_read < 0)\n+\t\t\treturn 1;  // unexpected EOF\n+\n+\t\tif (is_stream &&\n+\t\t\tbytes_read > 0 &&\n+\t\t\tsb->len - total_bytes_read - 1 <= 0)\n+\t\t\tstrbuf_grow(sb, st_add(sb->len, LARGE_PACKET_MAX));\n+\t\ttotal_bytes_read += bytes_read;\n+\t}\n+\twhile (\n+\t\tbytes_read > 0 &&                   // the last packet was no flush\n+\t\tsb->len - total_bytes_read - 1 > 0  // we still have space left in the buffer\n+\t);\n+\tstrbuf_setlen(sb, total_bytes_read);\n+\treturn (is_stream ? 0 : expected_bytes != total_bytes_read);\n+}\n+\n+static int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n+{\n+\tint did_fail = 0;\n+\tssize_t bytes_to_write;\n+\twhile (!did_fail) {\n+\t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_LEN);\n+\t\tif (bytes_to_write < 0)\n+\t\t\treturn 1;\n+\t\tif (bytes_to_write == 0)\n+\t\t\tbreak;\n+\t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n+\t}\n+\tif (!did_fail)\n+\t\tdid_fail = packet_flush_gentle(fd_out);\n+\treturn did_fail;\n+}\n+\n+static int multi_packet_write_from_buf(const char *src, size_t len, int fd_out)\n+{\n+\tint did_fail = 0;\n+\tsize_t bytes_written = 0;\n+\tsize_t bytes_to_write;\n+\twhile (!did_fail) {\n+\t\tif ((len - bytes_written) > PKTLINE_DATA_LEN)\n+\t\t\tbytes_to_write = PKTLINE_DATA_LEN;\n+\t\telse\n+\t\t\tbytes_to_write = len - bytes_written;\n+\t\tif (bytes_to_write == 0)\n+\t\t\tbreak;\n+\t\tdid_fail |= direct_packet_write_data(fd_out, src + bytes_written, bytes_to_write, 1);\n+\t\tbytes_written += bytes_to_write;\n+\t}\n+\tif (!did_fail)\n+\t\tdid_fail = packet_flush_gentle(fd_out);\n+\treturn did_fail;\n+}\n+\n+#define FILTER_CAPABILITIES_STREAM   0x1\n+#define FILTER_CAPABILITIES_CLEAN    0x2\n+#define FILTER_CAPABILITIES_SMUDGE   0x4\n+#define FILTER_CAPABILITIES_SHUTDOWN 0x8\n+#define FILTER_SUPPORTS_STREAM(type) ((type) & FILTER_CAPABILITIES_STREAM)\n+#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n+#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n+#define FILTER_SUPPORTS_SHUTDOWN(type) ((type) & FILTER_CAPABILITIES_SHUTDOWN)\n+\n+struct cmd2process {\n+\tstruct hashmap_entry ent; /* must be the first member! */\n+\tconst char *cmd;\n+\tint supported_capabilities;\n+\tstruct child_process process;\n+};\n+\n+static int cmd_process_map_initialized = 0;\n+static struct hashmap cmd_process_map;\n+\n+static int cmd2process_cmp(const struct cmd2process *e1,\n+\t\t\t\t\t\t\tconst struct cmd2process *e2,\n+\t\t\t\t\t\t\tconst void *unused)\n+{\n+\treturn strcmp(e1->cmd, e2->cmd);\n+}\n+\n+static struct cmd2process *find_protocol2_filter_entry(struct hashmap *hashmap, const char *cmd)\n+{\n+\tstruct cmd2process k;\n+\thashmap_entry_init(&k, strhash(cmd));\n+\tk.cmd = cmd;\n+\treturn hashmap_get(hashmap, &k, NULL);\n+}\n+\n+static void kill_protocol2_filter(struct hashmap *hashmap, struct cmd2process *entry) {\n+\tif (!entry)\n+\t\treturn;\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\tclose(entry->process.in);\n+\tclose(entry->process.out);\n+\tsigchain_pop(SIGPIPE);\n+\tfinish_command(&entry->process);\n+\tchild_process_clear(&entry->process);\n+\thashmap_remove(hashmap, entry, NULL);\n+\tfree(entry);\n+}\n+\n+void shutdown_protocol2_filter(pid_t pid)\n+{\n+\tint did_fail;\n+\tstruct cmd2process *entry;\n+\tstruct hashmap_iter iter;\n+\tstatic const char shutdown[] = \"shutdown\\n\";\n+\tchar *result = NULL;\n+\n+\tif (!cmd_process_map_initialized)\n+\t\treturn;\n+\n+    hashmap_iter_init(&cmd_process_map, &iter);\n+\twhile ((entry = hashmap_iter_next(&iter))) {\n+\t\tif (entry->process.pid == pid &&\n+\t\t\tFILTER_SUPPORTS_SHUTDOWN(entry->supported_capabilities)\n+\t\t) {\n+\t\t\tsigchain_push(SIGPIPE, SIG_IGN);\n+\t\t\tdid_fail = direct_packet_write_data(\n+\t\t\t\tentry->process.in, shutdown, strlen(shutdown), 1);\n+\t\t\tif (!did_fail)\n+\t\t\t\tresult = packet_read_line(entry->process.out, NULL);\n+\t\t\tclose(entry->process.in);\n+\t\t\tclose(entry->process.out);\n+\t\t\tsigchain_pop(SIGPIPE);\n+\n+\t\t\tif (did_fail || !result || strcmp(result, \"done\"))\n+\t\t\t\terror(\"shutdown of external filter '%s' failed\", entry->cmd);\n+\t\t}\n+\t}\n+}\n+\n+static struct cmd2process *start_protocol2_filter(struct hashmap *hashmap, const char *cmd)\n+{\n+\tint did_fail;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tconst char *argv[] = { cmd, NULL };\n+\tstruct string_list capabilities = STRING_LIST_INIT_NODUP;\n+\tchar *capabilities_buffer;\n+\tint i;\n+\n+\tentry = xmalloc(sizeof(*entry));\n+\thashmap_entry_init(entry, strhash(cmd));\n+\tentry->cmd = cmd;\n+\tentry->supported_capabilities = 0;\n+\tprocess = &entry->process;\n+\n+\tchild_process_init(process);\n+\tprocess->argv = argv;\n+\tprocess->use_shell = 1;\n+\tprocess->in = -1;\n+\tprocess->out = -1;\n+\tprocess->clean_on_exit = 1;\n+\tprocess->clean_on_exit_handler = shutdown_protocol2_filter;\n+\n+\tif (start_command(process)) {\n+\t\terror(\"cannot fork to run external filter '%s'\", cmd);\n+\t\tkill_protocol2_filter(hashmap, entry);\n+\t\treturn NULL;\n+\t}\n+\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\tdid_fail = strcmp(packet_read_line(process->out, NULL), \"git-filter-protocol\");\n+\tif (!did_fail)\n+\t\tdid_fail |= strcmp(packet_read_line(process->out, NULL), \"version 2\");\n+\tif (!did_fail)\n+\t\tcapabilities_buffer = packet_read_line(process->out, NULL);\n+\telse\n+\t\tcapabilities_buffer = NULL;\n+\tsigchain_pop(SIGPIPE);\n+\n+\tif (!did_fail && capabilities_buffer) {\n+\t\tstring_list_split_in_place(&capabilities, capabilities_buffer, ' ', -1);\n+\t\tif (capabilities.nr > 1 &&\n+\t\t\t!strcmp(capabilities.items[0].string, \"capabilities\")) {\n+\t\t\tfor (i = 1; i < capabilities.nr; i++) {\n+\t\t\t\tconst char *requested = capabilities.items[i].string;\n+\t\t\t\tif (!strcmp(requested, \"stream\")) {\n+\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_STREAM;\n+\t\t\t\t} else if (!strcmp(requested, \"clean\")) {\n+\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_CLEAN;\n+\t\t\t\t} else if (!strcmp(requested, \"smudge\")) {\n+\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SMUDGE;\n+\t\t\t\t} else if (!strcmp(requested, \"shutdown\")) {\n+\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SHUTDOWN;\n+\t\t\t\t} else {\n+\t\t\t\t\twarning(\n+\t\t\t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n+\t\t\t\t\t\tcmd, requested\n+\t\t\t\t\t);\n+\t\t\t\t}\n+\t\t\t}\n+\t\t} else {\n+\t\t\terror(\"filter capabilities not found\");\n+\t\t\tdid_fail = 1;\n+\t\t}\n+\t\tstring_list_clear(&capabilities, 0);\n+\t}\n+\n+\tif (did_fail) {\n+\t\terror(\"initialization for external filter '%s' failed\", cmd);\n+\t\tkill_protocol2_filter(hashmap, entry);\n+\t\treturn NULL;\n+\t}\n+\n+\thashmap_add(hashmap, entry);\n+\treturn entry;\n+}\n+\n+static int apply_protocol2_filter(const char *path, const char *src, size_t len,\n+\t\t\t\t\t\tint fd, struct strbuf *dst, const char *cmd,\n+\t\t\t\t\t\tconst int wanted_capability)\n+{\n+\tint ret = 1;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tstruct stat file_stat;\n+\tstruct strbuf nbuf = STRBUF_INIT;\n+\tsize_t expected_bytes = 0;\n+\tchar *strtol_end;\n+\tchar *strbuf;\n+\tchar *filter_type;\n+\tchar *filter_result = NULL;\n+\n+\tif (!cmd || !*cmd)\n+\t\treturn 0;\n+\n+\tif (!dst)\n+\t\treturn 1;\n+\n+\tif (!cmd_process_map_initialized) {\n+\t\tcmd_process_map_initialized = 1;\n+\t\thashmap_init(&cmd_process_map, (hashmap_cmp_fn) cmd2process_cmp, 0);\n+\t\tentry = NULL;\n+\t} else {\n+\t\tentry = find_protocol2_filter_entry(&cmd_process_map, cmd);\n+\t}\n+\n+\tfflush(NULL);\n+\n+\tif (!entry) {\n+\t\tentry = start_protocol2_filter(&cmd_process_map, cmd);\n+\t\tif (!entry) {\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\tprocess = &entry->process;\n+\n+\tif (!(wanted_capability & entry->supported_capabilities))\n+\t\treturn 1;  // it is OK if the wanted capability is not supported\n+\n+\tif FILTER_SUPPORTS_CLEAN(wanted_capability)\n+\t\tfilter_type = \"clean\";\n+\telse if FILTER_SUPPORTS_SMUDGE(wanted_capability)\n+\t\tfilter_type = \"smudge\";\n+\telse\n+\t\tdie(\"unexpected filter type\");\n+\n+\tif (fd >= 0 && !src) {\n+\t\tif (fstat(fd, &file_stat) == -1)\n+\t\t\treturn 0;\n+\t\tlen = file_stat.st_size;\n+\t}\n+\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\n+\tpacket_buf_write(&nbuf, \"%s\\n\", filter_type);\n+\tret &= !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n+\n+\tif (ret) {\n+\t\tstrbuf_reset(&nbuf);\n+\t\tpacket_buf_write(&nbuf, \"filename=%s\\n\", path);\n+\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n+\t}\n+\n+\tif (ret) {\n+\t\tstrbuf_reset(&nbuf);\n+\t\tpacket_buf_write(&nbuf, \"size=%\"PRIuMAX\"\\n\", (uintmax_t)len);\n+\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n+\t}\n+\n+\tif (ret) {\n+\t\tif (fd >= 0)\n+\t\t\tret = !multi_packet_write_from_fd(fd, process->in);\n+\t\telse\n+\t\t\tret = !multi_packet_write_from_buf(src, len, process->in);\n+\t}\n+\n+\tif (ret && !FILTER_SUPPORTS_STREAM(entry->supported_capabilities)) {\n+\t\tstrbuf = packet_read_line(process->out, NULL);\n+\t\tif (strlen(strbuf) > 5 && !strncmp(\"size=\", strbuf, 5)) {\n+\t\t\texpected_bytes = (off_t)strtol(strbuf + 5, &strtol_end, 10);\n+\t\t\tret = (strtol_end != strbuf && errno != ERANGE);\n+\t\t} else {\n+\t\t\tret = 0;\n+\t\t}\n+\t}\n+\n+\tif (ret) {\n+\t\tstrbuf_reset(&nbuf);\n+\t\tret = !multi_packet_read(process->out, &nbuf, expected_bytes,\n+\t\t\tFILTER_SUPPORTS_STREAM(entry->supported_capabilities));\n+\t}\n+\n+\tif (ret) {\n+\t\tfilter_result = packet_read_line(process->out, NULL);\n+\t\tret = !strcmp(filter_result, \"success\");\n+\t}\n+\n+\tsigchain_pop(SIGPIPE);\n+\n+\tif (ret) {\n+\t\tstrbuf_swap(dst, &nbuf);\n+\t} else {\n+\t\tif (!filter_result || strcmp(filter_result, \"reject\")) {\n+\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n+\t\t\terror(\"external filter '%s' failed\", cmd);\n+\t\t\tkill_protocol2_filter(&cmd_process_map, entry);\n+\t\t}\n+\t}\n+\tstrbuf_release(&nbuf);\n+\treturn ret;\n+}\n+\n static struct convert_driver {\n \tconst char *name;\n \tstruct convert_driver *next;\n \tconst char *smudge;\n \tconst char *clean;\n+\tconst char *process;\n \tint required;\n } *user_convert, **user_convert_tail;\n \n@@ -526,6 +871,10 @@ static int read_convert_config(const char *var, const char *value, void *cb)\n \tif (!strcmp(\"clean\", key))\n \t\treturn git_config_string(&drv->clean, var, value);\n \n+\tif (!strcmp(\"process\", key)) {\n+\t\treturn git_config_string(&drv->process, var, value);\n+\t}\n+\n \tif (!strcmp(\"required\", key)) {\n \t\tdrv->required = git_config_bool(var, value);\n \t\treturn 0;\n@@ -823,7 +1172,12 @@ int would_convert_to_git_filter_fd(const char *path)\n \tif (!ca.drv->required)\n \t\treturn 0;\n \n-\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n+\tif (!ca.drv->clean && ca.drv->process)\n+\t\treturn apply_protocol2_filter(\n+\t\t\tpath, NULL, 0, -1, NULL, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n+\t\t);\n+\telse\n+\t\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n }\n \n const char *get_convert_attr_ascii(const char *path)\n@@ -856,17 +1210,24 @@ int convert_to_git(const char *path, const char *src, size_t len,\n                    struct strbuf *dst, enum safe_crlf checksafe)\n {\n \tint ret = 0;\n-\tconst char *filter = NULL;\n+\tconst char *clean_filter = NULL;\n+\tconst char *process_filter = NULL;\n \tint required = 0;\n \tstruct conv_attrs ca;\n \n \tconvert_attrs(&ca, path);\n \tif (ca.drv) {\n-\t\tfilter = ca.drv->clean;\n+\t\tclean_filter = ca.drv->clean;\n+\t\tprocess_filter = ca.drv->process;\n \t\trequired = ca.drv->required;\n \t}\n \n-\tret |= apply_filter(path, src, len, -1, dst, filter);\n+\tif (!clean_filter && process_filter)\n+\t\tret |= apply_protocol2_filter(\n+\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_CLEAN\n+\t\t);\n+\telse\n+\t\tret |= apply_filter(path, src, len, -1, dst, clean_filter);\n \tif (!ret && required)\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n@@ -885,13 +1246,21 @@ int convert_to_git(const char *path, const char *src, size_t len,\n void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n \t\t\t      enum safe_crlf checksafe)\n {\n+\tint ret = 0;\n \tstruct conv_attrs ca;\n \tconvert_attrs(&ca, path);\n \n \tassert(ca.drv);\n-\tassert(ca.drv->clean);\n+\tassert(ca.drv->clean || ca.drv->process);\n+\n+\tif (!ca.drv->clean && ca.drv->process)\n+\t\tret = apply_protocol2_filter(\n+\t\t\tpath, NULL, 0, fd, dst, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n+\t\t);\n+\telse\n+\t\tret = apply_filter(path, NULL, 0, fd, dst, ca.drv->clean);\n \n-\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv->clean))\n+\tif (!ret)\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n@@ -902,14 +1271,16 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t\t\t\t    size_t len, struct strbuf *dst,\n \t\t\t\t\t    int normalizing)\n {\n-\tint ret = 0, ret_filter = 0;\n-\tconst char *filter = NULL;\n+\tint ret = 0, ret_filter;\n+\tconst char *smudge_filter = NULL;\n+\tconst char *process_filter = NULL;\n \tint required = 0;\n \tstruct conv_attrs ca;\n \n \tconvert_attrs(&ca, path);\n \tif (ca.drv) {\n-\t\tfilter = ca.drv->smudge;\n+\t\tprocess_filter = ca.drv->process;\n+\t\tsmudge_filter = ca.drv->smudge;\n \t\trequired = ca.drv->required;\n \t}\n \n@@ -922,7 +1293,7 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t * CRLF conversion can be skipped if normalizing, unless there\n \t * is a smudge filter.  The filter might expect CRLFs.\n \t */\n-\tif (filter || !normalizing) {\n+\tif (smudge_filter || process_filter || !normalizing) {\n \t\tret |= crlf_to_worktree(path, src, len, dst, ca.crlf_action);\n \t\tif (ret) {\n \t\t\tsrc = dst->buf;\n@@ -930,7 +1301,12 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t}\n \t}\n \n-\tret_filter = apply_filter(path, src, len, -1, dst, filter);\n+\tif (!smudge_filter && process_filter)\n+\t\tret_filter = apply_protocol2_filter(\n+\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_SMUDGE\n+\t\t);\n+\telse\n+\t\tret_filter = apply_filter(path, src, len, -1, dst, smudge_filter);\n \tif (!ret_filter && required)\n \t\tdie(\"%s: smudge filter %s failed\", path, ca.drv->name);\n \n@@ -1383,7 +1759,7 @@ struct stream_filter *get_stream_filter(const char *path, const unsigned char *s\n \tstruct stream_filter *filter = NULL;\n \n \tconvert_attrs(&ca, path);\n-\tif (ca.drv && (ca.drv->smudge || ca.drv->clean))\n+\tif (ca.drv && (ca.drv->process || ca.drv->smudge || ca.drv->clean))\n \t\treturn NULL;\n \n \tif (ca.crlf_action == CRLF_AUTO || ca.crlf_action == CRLF_AUTO_CRLF)\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 34c8eb9..e8a7703 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -296,4 +296,409 @@ test_expect_success 'disable filter with empty override' '\n \ttest_must_be_empty err\n '\n \n+test_expect_success PERL 'required process filter should filter data' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge shutdown\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"test22\" >test2.r &&\n+\t\tmkdir testsubdir &&\n+\t\techo \"test333\" >testsubdir/test3.r &&\n+\n+\t\trm -f rot13-filter.log &&\n+\t\tgit add . &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_add.log <<-\\EOF &&\n+\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t1 IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\t1 IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\t\t1 IN: shutdown -- [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_add.log uniq-rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n+\t\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_commit.log <<-\\EOF &&\n+\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\tx IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\tx IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\t\t1 IN: shutdown -- [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_commit.log uniq-rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\trm -f test?.r testsubdir/test3.r &&\n+\t\tgit checkout . &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\t\tIN: shutdown -- [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n+\n+\t\tgit checkout empty &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit checkout master &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout_master.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\t\tIN: shutdown -- [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n+\n+\t\t./../rot13.sh <test.r >expected &&\n+\t\tgit cat-file blob :test.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <test2.r >expected &&\n+\t\tgit cat-file blob :test2.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n+\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n+\t\ttest_cmp expected actual\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should filter data stream' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl stream clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"test22\" >test2.r &&\n+\t\tmkdir testsubdir &&\n+\t\techo \"test333\" >testsubdir/test3.r &&\n+\n+\t\trm -f rot13-filter.log &&\n+\t\tgit add . &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_add.log <<-\\EOF &&\n+\t\t\t1 IN: clean test.r 57 [OK] -- OUT: STREAM [OK]\n+\t\t\t1 IN: clean test2.r 7 [OK] -- OUT: STREAM [OK]\n+\t\t\t1 IN: clean testsubdir/test3.r 8 [OK] -- OUT: STREAM [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_add.log uniq-rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n+\t\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_commit.log <<-\\EOF &&\n+\t\t\tx IN: clean test.r 57 [OK] -- OUT: STREAM [OK]\n+\t\t\tx IN: clean test2.r 7 [OK] -- OUT: STREAM [OK]\n+\t\t\tx IN: clean testsubdir/test3.r 8 [OK] -- OUT: STREAM [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_commit.log uniq-rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\trm -f test?.r testsubdir/test3.r &&\n+\t\tgit checkout . &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: STREAM [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: STREAM [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n+\n+\t\tgit checkout empty &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit checkout master &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout_master.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test.r 57 [OK] -- OUT: STREAM [OK]\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: STREAM [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: STREAM [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n+\n+\t\t./../rot13.sh <test.r >expected &&\n+\t\tgit cat-file blob :test.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <test2.r >expected &&\n+\t\tgit cat-file blob :test2.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n+\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n+\t\ttest_cmp expected actual\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should filter smudge data and one-shot filter should clean' '\n+\ttest_config_global filter.protocol.clean ./../rot13.sh &&\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"test22\" >test2.r &&\n+\t\tmkdir testsubdir &&\n+\t\techo \"test333\" >testsubdir/test3.r &&\n+\n+\t\trm -f rot13-filter.log &&\n+\t\tgit add . &&\n+\t\ttest_must_be_empty rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\ttest_must_be_empty rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\trm -f test?.r testsubdir/test3.r &&\n+\t\tgit checkout . &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n+\n+\t\tgit checkout empty &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit checkout master &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout_master.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n+\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n+\n+\t\t./../rot13.sh <test.r >expected &&\n+\t\tgit cat-file blob :test.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <test2.r >expected &&\n+\t\tgit cat-file blob :test2.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n+\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n+\t\ttest_cmp expected actual\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should clean only' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\n+\t\trm -f rot13-filter.log &&\n+\t\tgit add . &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_add.log <<-\\EOF &&\n+\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_add.log uniq-rot13-filter.log &&\n+\n+\t\t>rot13-filter.log &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n+\t\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >uniq-rot13-filter.log &&\n+\t\tcat >expected_commit.log <<-\\EOF &&\n+\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t1 start\n+\t\t\t1 wrote filter header\n+\t\tEOF\n+\t\ttest_cmp expected_commit.log uniq-rot13-filter.log\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should process files larger LARGE_PACKET_MAX' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.file filter=protocol\" >.gitattributes &&\n+\t\tcat ../generated-test-data/largish.file.rot13 >large.rot13 &&\n+\t\tcat ../generated-test-data/largish.file >large.file &&\n+\t\tcat large.file >large.original &&\n+\n+\t\tgit add large.file .gitattributes &&\n+\t\tgit commit . -m \"test commit\" &&\n+\n+\t\trm -f large.file &&\n+\t\tgit checkout -- large.file &&\n+\t\tgit cat-file blob :large.file >actual &&\n+\t\ttest_cmp large.rot13 actual\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should with clean error should fail' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"this is going to fail\" >clean-write-fail.r &&\n+\t\techo \"test333\" >test3.r &&\n+\n+\t\t# Note: There are three clean paths in convert.c we just test one here.\n+\t\ttest_must_fail git add .\n+\t)\n+'\n+\n+test_expect_success PERL 'process filter should restart after unexpected write failure' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"1234567\" >test2.o &&\n+\t\tcat test2.o >test2.r &&\n+\t\techo \"this is going to fail\" >smudge-write-fail.o &&\n+\t\tcat smudge-write-fail.o >smudge-write-fail.r &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\trm -f *.r &&\n+\n+\t\tprintf \"\" >rot13-filter.log &&\n+\t\tgit checkout . &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout_master.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge smudge-write-fail.r 22 [OK] -- OUT: 22 [WRITE FAIL]\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\tIN: smudge test2.r 8 [OK] -- OUT: 8 [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n+\n+\t\ttest_cmp ../test.o test.r &&\n+\t\t./../rot13.sh <../test.o >expected &&\n+\t\tgit cat-file blob :test.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\ttest_cmp test2.o test2.r &&\n+\t\t./../rot13.sh <test2.o >expected &&\n+\t\tgit cat-file blob :test2.r >actual &&\n+\t\ttest_cmp expected actual &&\n+\n+\t\t! test_cmp smudge-write-fail.o smudge-write-fail.r && # Smudge failed!\n+\t\t./../rot13.sh <smudge-write-fail.o >expected &&\n+\t\tgit cat-file blob :smudge-write-fail.r >actual &&\n+\t\ttest_cmp expected actual\t\t\t\t\t\t\t  # Clean worked!\n+\t)\n+'\n+\n+test_expect_success PERL 'process filter should not restart after intentionally rejected file' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"1234567\" >test2.o &&\n+\t\tcat test2.o >test2.r &&\n+\t\techo \"this is going to fail\" >reject.o &&\n+\t\tcat reject.o >reject.r &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\trm -f *.r &&\n+\n+\t\tprintf \"\" >rot13-filter.log &&\n+\t\tgit checkout . &&\n+\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n+\t\tcat >expected_checkout_master.log <<-\\EOF &&\n+\t\t\tstart\n+\t\t\twrote filter header\n+\t\t\tIN: smudge reject.r 22 [OK] -- OUT: 0 [REJECT]\n+\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\tIN: smudge test2.r 8 [OK] -- OUT: 8 [OK]\n+\t\tEOF\n+\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log\n+\t)\n+'\n test_done\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nnew file mode 100755\nindex 0000000..cb0925d\n--- /dev/null\n+++ b/t/t0021/rot13-filter.pl\n@@ -0,0 +1,177 @@\n+#!/usr/bin/perl\n+#\n+# Example implementation for the Git filter protocol version 2\n+# See Documentation/gitattributes.txt, section \"Filter Protocol\"\n+#\n+# The script takes the list of supported protocol capabilities as\n+# arguments (\"stream\", \"clean\", and \"smudge\" are supported).\n+#\n+# This implementation supports three special test cases:\n+# (1) If data with the filename \"clean-write-fail.r\" is processed with\n+#     a \"clean\" operation then the write operation will die.\n+# (2) If data with the filename \"smudge-write-fail.r\" is processed with\n+#     a \"smudge\" operation then the write operation will die.\n+# (3) If data with the filename \"failure.r\" is processed with any\n+#     operation then the filter signals that the operation was not\n+#     successful.\n+#\n+\n+use strict;\n+use warnings;\n+\n+my $MAX_PACKET_CONTENT_SIZE = 65516;\n+my @capabilities            = @ARGV;\n+\n+sub rot13 {\n+    my ($str) = @_;\n+    $str =~ y/A-Za-z/N-ZA-Mn-za-m/;\n+    return $str;\n+}\n+\n+sub packet_read {\n+    my $buffer;\n+    my $bytes_read = read STDIN, $buffer, 4;\n+    if ( $bytes_read == 0 ) {\n+        return;\n+    }\n+    elsif ( $bytes_read != 4 ) {\n+        die \"invalid packet size '$bytes_read' field\";\n+    }\n+    my $pkt_size = hex($buffer);\n+    if ( $pkt_size == 0 ) {\n+        return ( 1, \"\" );\n+    }\n+    elsif ( $pkt_size > 4 ) {\n+        my $content_size = $pkt_size - 4;\n+        $bytes_read = read STDIN, $buffer, $content_size;\n+        if ( $bytes_read != $content_size ) {\n+            die \"invalid packet\";\n+        }\n+        return ( 0, $buffer );\n+    }\n+    else {\n+        die \"invalid packet size\";\n+    }\n+}\n+\n+sub packet_write {\n+    my ($packet) = @_;\n+    print STDOUT sprintf( \"%04x\", length($packet) + 4 );\n+    print STDOUT $packet;\n+    STDOUT->flush();\n+}\n+\n+sub packet_flush {\n+    print STDOUT sprintf( \"%04x\", 0 );\n+    STDOUT->flush();\n+}\n+\n+open my $debug, \">>\", \"rot13-filter.log\";\n+print $debug \"start\\n\";\n+$debug->flush();\n+\n+packet_write(\"git-filter-protocol\\n\");\n+packet_write(\"version 2\\n\");\n+packet_write( \"capabilities \" . join( ' ', @capabilities ) . \"\\n\" );\n+print $debug \"wrote filter header\\n\";\n+$debug->flush();\n+\n+while (1) {\n+    my $command = packet_read();\n+    unless ( defined($command) ) {\n+        exit();\n+    }\n+    chomp $command;\n+    print $debug \"IN: $command\";\n+    $debug->flush();\n+\n+    if ( $command eq \"shutdown\" ) {\n+        print $debug \" -- [OK]\";\n+        $debug->flush();\n+        packet_write(\"done\\n\");\n+        exit();\n+    }\n+\n+    my ($filename) = packet_read() =~ /filename=([^=]+)\\n/;\n+    print $debug \" $filename\";\n+    $debug->flush();\n+    my ($filelen) = packet_read() =~ /size=([^=]+)\\n/;\n+    chomp $filelen;\n+    print $debug \" $filelen\";\n+    $debug->flush();\n+\n+    $filelen =~ /\\A\\d+\\z/ or die \"bad filelen: $filelen\";\n+    my $output;\n+\n+    if ( $filelen > 0 ) {\n+        my $input = \"\";\n+        {\n+            binmode(STDIN);\n+            my $buffer;\n+            my $done = 0;\n+            while ( !$done ) {\n+                ( $done, $buffer ) = packet_read();\n+                $input .= $buffer;\n+            }\n+            print $debug \" [OK] -- \";\n+            $debug->flush();\n+        }\n+\n+        if ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n+            $output = rot13($input);\n+        }\n+        elsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n+            $output = rot13($input);\n+        }\n+        else {\n+            die \"bad command $command\";\n+        }\n+    }\n+\n+    my $output_len = length($output);\n+    if ( $filename eq \"reject.r\" ) {\n+        $output_len = 0;\n+    }\n+\n+    if ( grep( /^stream$/, @capabilities ) ) {\n+        print $debug \"OUT: STREAM \";\n+    }\n+    else {\n+        packet_write(\"size=$output_len\\n\");\n+        print $debug \"OUT: $output_len \";\n+    }\n+    $debug->flush();\n+\n+    if ( $filename eq \"reject.r\" ) {\n+        packet_write(\"reject\\n\");\n+        print $debug \"[REJECT]\\n\";    # Could also be an error\n+        $debug->flush();\n+    }\n+\n+    if ( $output_len > 0 ) {\n+        if (( $command eq \"clean\" and $filename eq \"clean-write-fail.r\" )\n+            or\n+            ( $command eq \"smudge\" and $filename eq \"smudge-write-fail.r\" ))\n+        {\n+            print $debug \"[WRITE FAIL]\\n\";\n+            $debug->flush();\n+            die \"write error\";\n+        }\n+        else {\n+            while ( length($output) > 0 ) {\n+                my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n+                packet_write($packet);\n+                if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n+                    $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n+                }\n+                else {\n+                    $output = \"\";\n+                }\n+            }\n+            packet_flush();\n+            packet_write(\"success\\n\");\n+            print $debug \"[OK]\\n\";\n+            $debug->flush();\n+        }\n+    }\n+}\n-- \n2.9.0\n\n"},{"id":"292596","messageId":"ef6c6152-a720-6bd5-22bb-6ebf375ca919@kdbg.org","threadId":"42968","inReplyTo":"20160729233801.82844-7-larsxschneider@gmail.com","subject":"Re: [PATCH v3 06/10] run-command: add clean_on_exit_handler","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2016-07-30T09:50:12Z","receivedAt":"2016-07-30T09:50:28Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 30.07.2016 um 01:37 schrieb larsxschneider@gmail.com:\n> Some commands might need to perform cleanup tasks on exit. Let's give\n> them an interface for doing this.\n>\n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n>  run-command.c | 12 ++++++++----\n>  run-command.h |  1 +\n>  2 files changed, 9 insertions(+), 4 deletions(-)\n>\n> diff --git a/run-command.c b/run-command.c\n> index 33bc63a..197b534 100644\n> --- a/run-command.c\n> +++ b/run-command.c\n> @@ -21,6 +21,7 @@ void child_process_clear(struct child_process *child)\n>\n>  struct child_to_clean {\n>  \tpid_t pid;\n> +\tvoid (*clean_on_exit_handler)(pid_t);\n>  \tstruct child_to_clean *next;\n>  };\n>  static struct child_to_clean *children_to_clean;\n> @@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n>  {\n>  \twhile (children_to_clean) {\n>  \t\tstruct child_to_clean *p = children_to_clean;\n> +\t\tif (p->clean_on_exit_handler)\n> +\t\t\tp->clean_on_exit_handler(p->pid);\n\nThis summons demons. cleanup_children() is invoked from a signal \nhandler. In this case, it can call only async-signal-safe functions. It \ndoes not look like the handler that you are going to install later will \ntake note of this caveat!\n\n>  \t\tchildren_to_clean = p->next;\n>  \t\tkill(p->pid, sig);\n>  \t\tif (!in_signal)\n\nThe condition that we see here in the context protects free(p) (which is \nnot async-signal-safe). Perhaps the invocation of the new callback \nshould be skipped in the same manner when this is called from a signal \nhandler? 507d7804 (pager: don't use unsafe functions in signal handlers) \nmay be worth a look.\n\n-- Hannes\n\n"},{"id":"292597","messageId":"4081bc44-d964-79ec-165f-f49f33823c17@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-2-larsxschneider@gmail.com","subject":"Re: [PATCH v3 01/10] pkt-line: extract set_packet_header()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T10:30:02Z","receivedAt":"2016-07-30T10:33:33Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> set_packet_header() converts an integer to a 4 byte hex string. Make\n> this function locally available so that other pkt-line functions can\n> use it.\n\nThis description is not that clear that set_packet_header() is a new\nfunction.  Perhaps something like the following\n\n  Extract the part of format_packet() that converts an integer to a 4 byte\n  hex string into set_packet_header().  Make this new function ...\n\nI also wonder if the part \"Make this [new] function locally available...\"\nis needed; we need to justify exports, but I think we don't need to\njustify limiting it to a module.  If you want to justify that it is\n\"static\", perhaps it would be better to say why not to export it.\n\nAnyway, I think it is worthy refactoring (and compiler should be\nable to inline it, so there are no nano-performance considerations).\n\nGood work!\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n>  pkt-line.c | 15 ++++++++++-----\n>  1 file changed, 10 insertions(+), 5 deletions(-)\n> \n> diff --git a/pkt-line.c b/pkt-line.c\n> index 62fdb37..445b8e1 100644\n> --- a/pkt-line.c\n> +++ b/pkt-line.c\n> @@ -98,9 +98,17 @@ void packet_buf_flush(struct strbuf *buf)\n>  }\n>  \n>  #define hex(a) (hexchar[(a) & 15])\n\nI guess that this is inherited from the original, but this preprocessor\nmacro is local to the format_header() / set_packet_header() function,\nand would not work outside it.  Therefore I think we should #undef it\nafter set_packet_header(), just in case somebody mistakes it for\na generic hex() function.  Perhaps even put it inside set_packet_header(),\ntogether with #undef.\n\nBut I might be mistaken... let's check... no, it isn't used outside it.\n\n> -static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n> +static void set_packet_header(char *buf, const int size)\n>  {\n>  \tstatic char hexchar[] = \"0123456789abcdef\";\n> +\tbuf[0] = hex(size >> 12);\n> +\tbuf[1] = hex(size >> 8);\n> +\tbuf[2] = hex(size >> 4);\n> +\tbuf[3] = hex(size);\n> +}\n> +\n> +static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n\nIt is strange how 'git diff' chosen to represent this patch...\n\n> +{\n>  \tsize_t orig_len, n;\n>  \n>  \torig_len = out->len;\n> @@ -111,10 +119,7 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n>  \tif (n > LARGE_PACKET_MAX)\n>  \t\tdie(\"protocol error: impossibly long line\");\n>  \n> -\tout->buf[orig_len + 0] = hex(n >> 12);\n> -\tout->buf[orig_len + 1] = hex(n >> 8);\n> -\tout->buf[orig_len + 2] = hex(n >> 4);\n> -\tout->buf[orig_len + 3] = hex(n);\n> +\tset_packet_header(&out->buf[orig_len], n);\n>  \tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n>  }\n>  \n> \n\n"},{"id":"292598","messageId":"58e4737b-6e0e-565c-2468-05c705dea426@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-3-larsxschneider@gmail.com","subject":"Re: [PATCH v3 02/10] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T10:49:03Z","receivedAt":"2016-07-30T10:49:27Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> Sometimes pkt-line data is already available in a buffer and it would\n> be a waste of resources to write the packet using packet_write() which\n> would copy the existing buffer into a strbuf before writing it.\n> \n> If the caller has control over the buffer creation then the\n> PKTLINE_DATA_START macro can be used to skip the header and write\n> directly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\n> would be the maximum). direct_packet_write() would take this buffer,\n> adjust the pkt-line header and write it.\n> \n> If the caller has no control over the buffer creation then\n> direct_packet_write_data() can be used. This function creates a pkt-line\n> header. Afterwards the header and the data buffer are written using two\n> consecutive write calls.\n\nI don't quite understand what do you mean by \"caller has control\nover the buffer creation\".  Do you mean that caller either can write\nover the buffer, or cannot overwrite the buffer?  Or do you mean that\ncaller either can allocate buffer to hold header, or is getting\nonly the data?\n\n> \n> Both functions have a gentle parameter that indicates if Git should die\n> in case of a write error (gentle set to 0) or return with a error (gentle\n> set to 1).\n\nSo they are *_maybe_gently(), isn't it ;-)?  Are there any existing\nfunctions in Git codebase that take 'gently' / 'strict' / 'die_on_error'\nparameter?\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n>  pkt-line.c | 30 ++++++++++++++++++++++++++++++\n>  pkt-line.h |  5 +++++\n>  2 files changed, 35 insertions(+)\n> \n> diff --git a/pkt-line.c b/pkt-line.c\n> index 445b8e1..6fae508 100644\n> --- a/pkt-line.c\n> +++ b/pkt-line.c\n> @@ -135,6 +135,36 @@ void packet_write(int fd, const char *fmt, ...)\n>  \twrite_or_die(fd, buf.buf, buf.len);\n>  }\n>  \n> +int direct_packet_write(int fd, char *buf, size_t size, int gentle)\n> +{\n> +\tint ret = 0;\n> +\tpacket_trace(buf + 4, size - 4, 1);\n> +\tset_packet_header(buf, size);\n> +\tif (gentle)\n> +\t\tret = !write_or_whine_pipe(fd, buf, size, \"pkt-line\");\n> +\telse\n> +\t\twrite_or_die(fd, buf, size);\n\nHmmm... in gently case we get the information in the warning that\nit is about \"pkt-line\", which is missing from !gently case.  But\nit is probably not important.\n\n> +\treturn ret;\n> +}\n\nNice clean function, thanks to extracting set_packet_header().\n\n> +\n> +int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle)\n\nI would name the parameter 'data', rather than 'buf'; IMVHO it\nbetter describes it.\n\n> +{\n> +\tint ret = 0;\n> +\tchar hdr[4];\n> +\tset_packet_header(hdr, sizeof(hdr) + size);\n> +\tpacket_trace(buf, size, 1);\n> +\tif (gentle) {\n> +\t\tret = (\n> +\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n\nYou can write '4' here, no need for sizeof(hdr)... though compiler would\noptimize it away.\n\n> +\t\t\t!write_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n> +\t\t);\n\nDo we want to try to write \"pkt-line data\" if \"pkt-line header\" failed?\nIf not, perhaps De Morgan-ize it\n\n  +\t\tret = !(\n  +\t\t\twrite_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") &&\n  +\t\t\twrite_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n  +\t\t);\n\n\n> +\t} else {\n> +\t\twrite_or_die(fd, hdr, sizeof(hdr));\n> +\t\twrite_or_die(fd, buf, size);\n\nI guess these two writes (here and in 'gently' case) are unavoidable...\n\n> +\t}\n> +\treturn ret;\n> +}\n> +\n>  void packet_buf_write(struct strbuf *buf, const char *fmt, ...)\n>  {\n>  \tva_list args;\n> diff --git a/pkt-line.h b/pkt-line.h\n> index 3cb9d91..02dcced 100644\n> --- a/pkt-line.h\n> +++ b/pkt-line.h\n> @@ -23,6 +23,8 @@ void packet_flush(int fd);\n>  void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n>  void packet_buf_flush(struct strbuf *buf);\n>  void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n> +int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n> +int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n>  \n>  /*\n>   * Read a packetized line into the buffer, which must be at least size bytes\n> @@ -77,6 +79,9 @@ char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);\n>  \n>  #define DEFAULT_PACKET_MAX 1000\n>  #define LARGE_PACKET_MAX 65520\n> +#define PKTLINE_HEADER_LEN 4\n> +#define PKTLINE_DATA_START(pkt) ((pkt) + PKTLINE_HEADER_LEN)\n> +#define PKTLINE_DATA_LEN (LARGE_PACKET_MAX - PKTLINE_HEADER_LEN)\n\nThose are not used in direct_packet_write() and direct_packet_write_data();\nbut they would make them more verbose and less readable.\n\n>  extern char packet_buffer[LARGE_PACKET_MAX];\n>  \n>  #endif\n> \n\n"},{"id":"292599","messageId":"41184531-d3c2-43c0-d3b8-23cc913dbf86@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-4-larsxschneider@gmail.com","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T12:04:55Z","receivedAt":"2016-07-30T12:05:22Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> packet_flush() would die in case of a write error even though for some callers\n> an error would be acceptable. Add packet_flush_gentle() which writes a pkt-line\n> flush packet and returns `0` for success and `1` for failure.\n\nI think it should be packet_flush_gently(), as in \"to flush gently\",\nbut this is only my opinion; I have not checked the naming rules and\npractices for the rest of Git codebase.\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n\n"},{"id":"292600","messageId":"786f0b8e-29f0-3dd3-7bb4-5f6558f8ec84@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-5-larsxschneider@gmail.com","subject":"Re: [PATCH v3 04/10] pkt-line: call packet_trace() only if a packet is actually send","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T12:29:17Z","receivedAt":"2016-07-30T12:29:49Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> The packet_trace() call is not ideal in format_packet() as we would print\n\nStyle; I think the following is more readable:\n\n  The packet_trace() call in format_packet() is not ideal, as we would...\n\n> a trace when a packet is formatted and (potentially) when the packet is\n> actually send. This was no problem up until now because format_packet()\n> was only used by one function. Fix it by moving the trace call into the\n> function that actally sends the packet.\n\ns/actally/actually/\n\nI don't buy this explanation.  If you want to trace packets, you might\ndo it on input (when formatting packet), or on output (when writing\npacket).  It's when there are more than one formatting function, but\none writing function, then placing trace call in write function means\nless code duplication; and of course the reverse.\n\nAnother issue is that something may happen between formatting packet\nand sending it, and we probably want to packet_trace() when packet\nis actually send.\n\nNeither of those is visible in commit message.\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n>  pkt-line.c | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n> \n> diff --git a/pkt-line.c b/pkt-line.c\n> index 1728690..32c0a34 100644\n> --- a/pkt-line.c\n> +++ b/pkt-line.c\n> @@ -126,7 +126,6 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n>  \t\tdie(\"protocol error: impossibly long line\");\n>  \n>  \tset_packet_header(&out->buf[orig_len], n);\n> -\tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n>  }\n>  \n>  void packet_write(int fd, const char *fmt, ...)\n> @@ -138,6 +137,7 @@ void packet_write(int fd, const char *fmt, ...)\n>  \tva_start(args, fmt);\n>  \tformat_packet(&buf, fmt, args);\n>  \tva_end(args);\n> +\tpacket_trace(buf.buf + 4, buf.len - 4, 1);\n>  \twrite_or_die(fd, buf.buf, buf.len);\n>  }\n>  \n> \n\n"},{"id":"292603","messageId":"cb6721b8-2a6a-deb1-2fc7-59399d118cec@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-6-larsxschneider@gmail.com","subject":"Re: [PATCH v3 05/10] pack-protocol: fix maximum pkt-line size","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T13:58:36Z","receivedAt":"2016-07-30T13:59:00Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> According to LARGE_PACKET_MAX in pkt-line.h the maximal lenght of a\n> pkt-line packet is 65520 bytes. The pkt-line header takes 4 bytes and\n> therefore the pkt-line data component must not exceed 65516 bytes.\n\ns/lenght/length/\n\nIs it maximum length of pkt-line packet, or maximum length of data\nthat can be send in a packet?\n\nWith 4 hex digits, maximal length if pkt-line packet (together\nwith length) is ffff_16, that is 2^16-1 = 65535.  Where does the\nnumber 65520 comes from?\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n>  Documentation/technical/protocol-common.txt | 6 +++---\n>  1 file changed, 3 insertions(+), 3 deletions(-)\n> \n> diff --git a/Documentation/technical/protocol-common.txt b/Documentation/technical/protocol-common.txt\n> index bf30167..ecedb34 100644\n> --- a/Documentation/technical/protocol-common.txt\n> +++ b/Documentation/technical/protocol-common.txt\n> @@ -67,9 +67,9 @@ with non-binary data the same whether or not they contain the trailing\n>  LF (stripping the LF if present, and not complaining when it is\n>  missing).\n>  \n> -The maximum length of a pkt-line's data component is 65520 bytes.\n> -Implementations MUST NOT send pkt-line whose length exceeds 65524\n> -(65520 bytes of payload + 4 bytes of length data).\n> +The maximum length of a pkt-line's data component is 65516 bytes.\n> +Implementations MUST NOT send pkt-line whose length exceeds 65520\n> +(65516 bytes of payload + 4 bytes of length data).\n>  \n>  Implementations SHOULD NOT send an empty pkt-line (\"0004\").\n>  \n> \n\n"},{"id":"292665","messageId":"b4c9ac5d-bd6b-141b-5b85-ab4aa719ccb0@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-11-larsxschneider@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-30T22:05:31Z","receivedAt":"2016-07-30T22:06:07Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> Git's clean/smudge mechanism invokes an external filter process for every\n> single blob that is affected by a filter. If Git filters a lot of blobs\n> then the startup time of the external filter processes can become a\n> significant part of the overall Git execution time.\n> \n> This patch adds the filter.<driver>.process string option which, if used,\n> keeps the external filter process running and processes all blobs with\n> the following packet format (pkt-line) based protocol over standard input\n> and standard output.\n\nI think it would be nice to have here at least summary of the benchmarks\nyou did in https://github.com/github/git-lfs/pull/1382\n\n> \n> Git starts the filter on first usage and expects a welcome\n> message, protocol version number, and filter capabilities\n> separated by spaces:\n> ------------------------\n> packet:          git< git-filter-protocol\\n\n> packet:          git< version 2\\n\n> packet:          git< capabilities clean smudge\\n\n\nSorry for going back and forth, but now I think that 'capabilities' are\nnot really needed here, though they are in line with \"version\" in\nthe second packet / line, namely \"version 2\".  If it does not make\nparsing more difficult...\n\n> ------------------------\n> Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n> and \"shutdown\".\n\nI'd rather put \"stream\" and \"shutdown\" capabilities into separate\npatches, for easier review.\n\n> \n> Afterwards Git sends a command (based on the supported\n> capabilities), the filename including its path\n> relative to the repository root, the content size as ASCII number\n> in bytes, the content split in zero or many pkt-line packets,\n> and a flush packet at the end:\n\nI guess the following is the most basic example, with mode detailed\ndescription left for the documentation.\n\n> ------------------------\n> packet:          git> smudge\\n\n> packet:          git> filename=path/testfile.dat\\n\n> packet:          git> size=7\\n\n\nSo I see you went with \"<variable>=<value>\" idea, rather than \"<value>\"\n(with <variable> defined by position in a sequence of 'header' packets),\nor \"<variable> <value>...\" that introductory header uses.\n\n> packet:          git> CONTENT\n> packet:          git> 0000\n> ------------------------\n> \n> The filter is expected to respond with the result content size as\n> ASCII number in bytes. If the capability \"stream\" is defined then\n> the filter must not send the content size. Afterwards the result\n> content in send in zero or many pkt-line packets and a flush packet\n> at the end. \n\nIf it does not cost filter anything, it could send size upfront\n(based on size of original, or based on external data), even if\nit is prepared for streaming.\n\nIn the opposite case, where filter cannot stream because it requires\nwhole contents upfront (e.g. to calculate hash of the contents, or\nto do operation that needs whole file like sorting or reversing lines),\nit should always be able to calculate the size... or not.  For\nexample 'sort | uniq' filter needs whole input upfront for sort,\nbut it does not know how many lines will be in output without doing\nthe 'uniq' part.\n\nSo I think the ability of filter to provide size (or size hint) of\nits output should be decoupled from streaming support.\n\n>             Finally a \"success\" packet is send to indicate that\n> everything went well.\n\nThat's a nice addition, and probably a necessary one, to the stream\nprotocol.  Git must know and consume it - we wouldn't be able to\nretrofit it later.\n\n> ------------------------\n> packet:          git< size=57\\n   (omitted with capability \"stream\")\n\nI was thinking about having possible responses to receiving file\ncontents (or starting receiving in the streaming case) to be:\n\n  packet:          git< ok size=7\\n    (or \"ok 7\\n\", if size is known)\n\nor\n\n  packet:          git< ok\\n           (if filter does not know size upfront)\n\nor\n\n  packet:          git< fail <msg>\\n   (or just \"fail\" + packet with msg)\n\nThe last would be when filter knows upfront that it cannot perform\nthe operation.  Though sending an empty file with non-\"success\" final\nwould work as well.\n\nFor example LFS filter (that is configured as not required) may refuse\nto store files which are smaller than some pre-defined constant threshold.\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< success\\n\n> ------------------------\n> \n> In case the filter cannot process the content, it is expected\n> to respond with the result content size 0 (only if \"stream\" is\n> not defined) and a \"reject\" packet.\n> ------------------------\n> packet:          git< size=0\\n    (omitted with capability \"stream\")\n> packet:          git< reject\\n\n> ------------------------\n\nThis is *wrong* idea!  Empty file, with size=0, can be a perfectly\nlegitimate response.  \n\nFor example rot13 filter should respond to an empty file on input\nwith an empty file on output.  LFS-like filters and encryption\nmechanism should return empty file on fetch / decryption\nif such empty file was stored / encrypted.\n\nA strange LFS could even use filenames (with files being empty\nthemselves) as a lookup key for artifactory.  For example a kind\nof CDN for common libraries, with version embedded in filename,\nlike 'libs/jquery-1.9.0.min.js', etc.\n\n> \n> After the filter has processed a blob it is expected to wait for\n> the next command. A demo implementation can be found in\n> `t/t0021/rot13-filter.pl` located in the Git core repository.\n\nIf filter does not support \"shutdown\" capability (or if said\ncapability is postponed for later patch), it should behave sanely\nwhen Git command reaps it (SIGTERM + wait + SIGKILL?, SIGCHLD?).\n\n> \n> If the filter supports the \"shutdown\" capability then Git will\n> send the \"shutdown\" command and wait until the filter answers\n> with \"done\". This gives the filter the opportunity to perform\n> cleanup tasks. Afterwards the filter is expected to exit.\n> ------------------------\n> packet:          git> shutdown\\n\n> packet:          git< done\\n\n> ------------------------\n\nI guess there is no timeout mechanism: if filter hangs on shutdown,\nthen git command would also hang waiting for signal to exit.\n\n> \n> If a filter.<driver>.clean or filter.<driver>.smudge command\n> is configured then these commands always take precedence over\n> a configured filter.<driver>.process command.\n\nNote: the value of `clean`, `smudge` and `process` is a command,\nnot just a string.\n\nI wonder if it would be worth it to explain the reasoning behind\nthis solution and show alternate ones.\n\n * Using a separate variable to signal that filters are invoked\n   per-command rather than per-file, and use pkt-line interface,\n   like boolean-valued `useProtocol`, or `protocolVersion` set\n   to '2' or 'v2', or `persistence` set to 'per-command', there\n   is high risk of user's trying to use exiting one-shot per-file\n   filters... and Git hanging.\n\n * Using new variables for each capability, e.g. `processSmudge`\n   and `processClean` would lead to explosion of variable names;\n   I think.\n\n * Current solution of using `process` in addition to `clean`\n   and `smudge` clearly says that you need to use different\n   command for per-file (`clean` and `smudge`), and per-command\n   filter, while allowing to use them together.\n\n   The possible disadvantage is Git command starting `process`\n   filter, only to see that it doesn't offer required capability,\n   for example offering only \"clean\" but not \"smudge\".  There\n   is simple workaround - set `smudge` variable (same as not\n   present capability) to empty string.\n\n> \n> Please note that you cannot use an existing filter.<driver>.clean\n> or filter.<driver>.smudge command as filter.<driver>.process\n> command. As soon as Git would detect a file that needs to be\n> processed by this filter, it would stop responding.\n\nI think this needs to be in the documentation (I have not checked\nyet if it is), but is not needed in the already long commit message.\n\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> Helped-by: Martin-Louis Bright <mlbright@gmail.com>\n> ---\n>  Documentation/gitattributes.txt |  84 ++++++++-\n>  convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--\n>  t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++\n>  t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++\n>  4 files changed, 1053 insertions(+), 13 deletions(-)\n>  create mode 100755 t/t0021/rot13-filter.pl\n> \n> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\n> index 8882a3e..e3fbcc2 100644\n> --- a/Documentation/gitattributes.txt\n> +++ b/Documentation/gitattributes.txt\n> @@ -300,7 +300,11 @@ checkout, when the `smudge` command is specified, the command is\n>  fed the blob object from its standard input, and its standard\n>  output is used to update the worktree file.  Similarly, the\n>  `clean` command is used to convert the contents of worktree file\n> -upon checkin.\n> +upon checkin. By default these commands process only a single\n> +blob and terminate. If a long running filter process (see section\n> +below) is used then Git can process all blobs with a single filter\n> +invocation for the entire life of a single Git command (e.g.\n> +`git add .`).\n\nProposed improvement:\n\n                       If a long running `process` filter is used\n   in place of `clean` and/or `smudge` filters, then Git can process\n   all blobs with a single filter command invocation for the entire\n   life of a single Git command, for example `git add --all`.  See\n   section below for the description of the protocol used to\n   communicate with a `process` filter.\n\n>  \n>  One use of the content filtering is to massage the content into a shape\n>  that is more convenient for the platform, filesystem, and the user to use.\n> @@ -375,6 +379,84 @@ substitution.  For example:\n>  ------------------------\n>  \n>  \n> +Long Running Filter Process\n> +^^^^^^^^^^^^^^^^^^^^^^^^^^^\n> +\n> +If the filter command (string value) is defined via\n\nThis is no mere string value, this is command invocation (with its\nown rules, e.g. splitting parameters on whitespace, etc.).  Though\nI'm not sure how to say it succintly.  Maybe skip \"(string value)\"?\nBut it is there for a reason...\n\n> +filter.<driver>.process then Git can process all blobs with a\n\nShouldn't it be `filter.<driver>.process`?\n\n> +single filter invocation for the entire life of a single Git\n> +command. This is achieved by using the following packet\n> +format (pkt-line, see protocol-common.txt) based protocol over\n\nCan we linkgit-it (to technical documentation)?\n\n> +standard input and standard output.\n> +\n> +Git starts the filter on first usage and expects a welcome\n\nIs \"usage\" here correct?  Perhaps it would be more readable\nto say that Git starts filter when encountering first file\nthat needs cleaning or smudgeing.\n\n> +message, protocol version number, and filter capabilities\n> +separated by spaces:\n> +------------------------\n> +packet:          git< git-filter-protocol\\n\n> +packet:          git< version 2\\n\n> +packet:          git< capabilities clean smudge\\n\n> +------------------------\n> +Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n> +and \"shutdown\".\n\nFilter should include at least one of \"clean\" and \"smudge\"\ncapabilities (currently), otherwise it wouldn't do anything.\n\nI don't know if it is a good place to say that because of pkt-line\nrecommendations about text-content packets, each of those should\nterminate in endline, with \"\\n\" included in pkt-line length.\n\n> +\n> +Afterwards Git sends a command (based on the supported\n> +capabilities),\n\nI think it should be something like the following:\n\n   If among filter `process` capabilities there is capability\n   that corresponds to the operation performed by a Git command\n   (that is, either \"clean\" or \"smudge\"), then Git would send,\n   in separate packets, a command (based on supported capabilites),\n\nthough it feels too \"chatty\" (and the sentence gets quite long).\n\n>                the filename including its path\n> +relative to the repository root, \n\nErrr... \"the filename including its path\"? Wouldn't be it simpler\nto just say:\n\n  the pathname of a file relative to the repository root,\n\nAlso, isn't it now \"filename=<pathname>\\n\"?\n\n>                                   the content size as ASCII number\n> +in bytes, \n\nCould Git not give the size, for example if fstat() fails? Do\nwe reserve space for other information here?\n\nAlso, isn't it now \"size=<bytes>\\n\"?\n\n>             the content split in zero or many pkt-line packets,\n\ns/zero or many/zero or more/\n\n> +and a flush packet at the end:\n\nI wonder if instead of long sentence, it would be more readable\nto use enumeration (ordered list) or itemize (unordered list).\n\n> +------------------------\n> +packet:          git> smudge\\n\n> +packet:          git> filename=path/testfile.dat\\n\n> +packet:          git> size=7\\n\n> +packet:          git> CONTENT\n> +packet:          git> 0000\n> +------------------------\n> +\n> +The filter is expected to respond with the result content size as\n> +ASCII number in bytes. If the capability \"stream\" is defined then\n> +the filter must not send the content size.\n\nAs I wrote earlier, I think sending or not the size of the output\nshould be decoupled from the \"stream\" capability.\n\nStreaming is IMVHO rather a capability of starting to send parts\nof response before the whole contents of input arrives.  I think\nper-file filters support that and that's what start_async() there\nis about.\n\n>                                             Afterwards the result\n> +content in send in zero or many pkt-line packets and a flush packet\n> +at the end. Finally a \"success\" packet is send to indicate that\n> +everything went well.\n\nI guess it is \"success\" packet if everything went well, and place\nfor informing about errors in the future - filter is assumed to die\nif there are errors in filtering, isn't it?\n\nThat is, not \"send to indicate\", but \"send if\".\n\n> +------------------------\n> +packet:          git< size=57\\n   (omitted with capability \"stream\")\n> +packet:          git< SMUDGED_CONTENT\n> +packet:          git< 0000\n> +packet:          git< success\\n\n> +------------------------\n> +\n> +In case the filter cannot process the content, it is expected\n> +to respond with the result content size 0 (only if \"stream\" is\n> +not defined) and a \"reject\" packet.\n> +------------------------\n> +packet:          git< size=0\\n    (omitted with capability \"stream\")\n> +packet:          git< reject\\n\n> +------------------------\n\nI would assume that we have two error conditions.  \n\nFirst situation is when the filter knows upfront (after receiving name\nand size of file, and after receiving contents for not-streaming filters)\nthat it cannot process the file (like e.g. LFS filter with artifactory\nreplica/shard being a bit behind master, and not including contents of\nthe file being filtered).\n\nMy proposal is to reply with \"fail\" _in place of_ size of reply:\n\n   packet:         git< fail\\n       (any case: size known or not, stream or not)\n\nIt could be \"reject\", or \"error\" instead of \"fail\".\n\n\nAnother situation is if filter encounters error during output,\neither with streaming filter (or non-stream, but not storing whole\ninput upfront) realizing in the middle of output that there is something\nwrong with input (e.g. converting between encoding, and encountering\ncharacter that cannot be represented in output encoding), or e.g. filter\nprocess being killed, or network connection dropping with LFS filter, etc.\nThe filter has send some packets with output already.  In this case\nfilter should flush, and send \"reject\" or \"error\" packet.\n\n   <error condition>\n   packet:         git< \"0000\"       (flush packet)\n   packet:         git< reject\\n\n\nShould there be a place for an error message, or would standard error\n(stderr) be used for this?\n\n> +\n> +After the filter has processed a blob it is expected to wait for\n> +the next command. A demo implementation can be found in\n> +`t/t0021/rot13-filter.pl` located in the Git core repository.\n\nIt is actually in Git sources.  Is it the best way to refer to\nsuch files?\n\n> +\n> +If the filter supports the \"shutdown\" capability then Git will\n> +send the \"shutdown\" command and wait until the filter answers\n> +with \"done\". This gives the filter the opportunity to perform\n> +cleanup tasks. Afterwards the filter is expected to exit.\n> +------------------------\n> +packet:          git> shutdown\\n\n> +packet:          git< done\\n\n> +------------------------\n> +\n> +If a filter.<driver>.clean or filter.<driver>.smudge command\n> +is configured then these commands always take precedence over\n> +a configured filter.<driver>.process command.\n\nAll right; this is quite clear.\n\n> +\n> +Please note that you cannot use an existing filter.<driver>.clean\n> +or filter.<driver>.smudge command as filter.<driver>.process\n> +command. As soon as Git would detect a file that needs to be\n> +processed by this filter, it would stop responding.\n\nThis isn't.\n\n\nP.S. I will comment about the implementation part in the next email.\n-- \nJakub Narębski\n\n"},{"id":"292677","messageId":"69988611-06ec-048d-12e7-7b87882ddc6a@gmail.com","threadId":"42968","inReplyTo":"b4c9ac5d-bd6b-141b-5b85-ab4aa719ccb0@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-31T09:42:11Z","receivedAt":"2016-07-31T09:42:38Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Excuse me replying to myself, but there are a few things I forgot,\n or realized only later]\n\nW dniu 31.07.2016 o 00:05, Jakub Narębski pisze:\n> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>>\n>> Git's clean/smudge mechanism invokes an external filter process for every\n>> single blob that is affected by a filter. If Git filters a lot of blobs\n>> then the startup time of the external filter processes can become a\n>> significant part of the overall Git execution time.\n>>\n>> This patch adds the filter.<driver>.process string option which, if used,\n>> keeps the external filter process running and processes all blobs with\n>> the following packet format (pkt-line) based protocol over standard input\n>> and standard output.\n> \n> I think it would be nice to have here at least summary of the benchmarks\n> you did in https://github.com/github/git-lfs/pull/1382\n\nNote that this feature is especially useful if startup time is long,\nthat is if you are using an operating system with costly fork / new process\nstartup time like MS Windows (which you have mentioned), or writing\nfilter in a programming language with large startup time like Java\nor Python (the latter may have changed since).\n\n  https://gnustavo.wordpress.com/2012/06/28/programming-languages-start-up-times/\n\n[...]\n> I was thinking about having possible responses to receiving file\n> contents (or starting receiving in the streaming case) to be:\n> \n>   packet:          git< ok size=7\\n    (or \"ok 7\\n\", if size is known)\n> \n> or\n> \n>   packet:          git< ok\\n           (if filter does not know size upfront)\n> \n> or\n> \n>   packet:          git< fail <msg>\\n   (or just \"fail\" + packet with msg)\n> \n> The last would be when filter knows upfront that it cannot perform\n> the operation.  Though sending an empty file with non-\"success\" final\n> would work as well.\n\n[...]\n\n>> In case the filter cannot process the content, it is expected\n>> to respond with the result content size 0 (only if \"stream\" is\n>> not defined) and a \"reject\" packet.\n>> ------------------------\n>> packet:          git< size=0\\n    (omitted with capability \"stream\")\n>> packet:          git< reject\\n\n>> ------------------------\n> \n> This is *wrong* idea!  Empty file, with size=0, can be a perfectly\n> legitimate response.  \n\nActually, I think I have misunderstood your intent.  If you want to have\nsimpler protocol, with only one place to signal errors, that is after\nsending a response, then proper way of signaling the error condition\nwould be to send empty file and then \"reject\" instead of \"success\":\n\n   packet:          git< size=0\\n    (omitted with capability \"stream\")\n   packet:          git< 0000        (we need this flush packet)\n   packet:          git< reject\\n\n\nOtherwise in the case without size upfront (capability \"stream\")\nfile with contents \"reject\" would be mistaken for the \"reject\" packet.\n\nSee below for proposal with two places to signal errors: before sending\nfirst byte, and after.\n\n\nNOTE: there is a bit of mixed and possibly confusing notation, that\nis 0000 is flush packet, not packet with 0000 as content.  Perhaps\nwrite pkt-line in full?\n\n\n[...]\n>> ---\n>>  Documentation/gitattributes.txt |  84 ++++++++-\n>>  convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--\n>>  t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++\n>>  t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++\n>>  4 files changed, 1053 insertions(+), 13 deletions(-)\n>>  create mode 100755 t/t0021/rot13-filter.pl\n\nWouldn't it be better for easier review to split it into separate patches?\nPerhaps at least the new test...\n\n[...]\n> I would assume that we have two error conditions.  \n> \n> First situation is when the filter knows upfront (after receiving name\n> and size of file, and after receiving contents for not-streaming filters)\n> that it cannot process the file (like e.g. LFS filter with artifactory\n> replica/shard being a bit behind master, and not including contents of\n> the file being filtered).\n> \n> My proposal is to reply with \"fail\" _in place of_ size of reply:\n> \n>    packet:         git< fail\\n       (any case: size known or not, stream or not)\n> \n> It could be \"reject\", or \"error\" instead of \"fail\".\n> \n> \n> Another situation is if filter encounters error during output,\n> either with streaming filter (or non-stream, but not storing whole\n> input upfront) realizing in the middle of output that there is something\n> wrong with input (e.g. converting between encoding, and encountering\n> character that cannot be represented in output encoding), or e.g. filter\n> process being killed, or network connection dropping with LFS filter, etc.\n> The filter has send some packets with output already.  In this case\n> filter should flush, and send \"reject\" or \"error\" packet.\n> \n>    <error condition>\n>    packet:         git< \"0000\"       (flush packet)\n>    packet:         git< reject\\n\n> \n> Should there be a place for an error message, or would standard error\n> (stderr) be used for this?\n\n"},{"id":"292690","messageId":"6765D972-876A-4F94-A170-468002498296@gmail.com","threadId":"42968","inReplyTo":"69988611-06ec-048d-12e7-7b87882ddc6a@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-31T19:49:04Z","receivedAt":"2016-07-31T19:49:14Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 31 Jul 2016, at 11:42, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> [Excuse me replying to myself, but there are a few things I forgot,\n> or realized only later]\n\nNo worries :)\n\n> \n> W dniu 31.07.2016 o 00:05, Jakub Narębski pisze:\n>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n>>> From: Lars Schneider <larsxschneider@gmail.com>\n>>> \n>>> Git's clean/smudge mechanism invokes an external filter process for every\n>>> single blob that is affected by a filter. If Git filters a lot of blobs\n>>> then the startup time of the external filter processes can become a\n>>> significant part of the overall Git execution time.\n>>> \n>>> This patch adds the filter.<driver>.process string option which, if used,\n>>> keeps the external filter process running and processes all blobs with\n>>> the following packet format (pkt-line) based protocol over standard input\n>>> and standard output.\n>> \n>> I think it would be nice to have here at least summary of the benchmarks\n>> you did in https://github.com/github/git-lfs/pull/1382\n> \n> Note that this feature is especially useful if startup time is long,\n> that is if you are using an operating system with costly fork / new process\n> startup time like MS Windows (which you have mentioned), or writing\n> filter in a programming language with large startup time like Java\n> or Python (the latter may have changed since).\n> \n>  https://gnustavo.wordpress.com/2012/06/28/programming-languages-start-up-times/\n\nOK, I will add this. Is it OK to add the link to the commit message?\n(since I don't know how long the link will be available).\n\n\n> [...]\n>> I was thinking about having possible responses to receiving file\n>> contents (or starting receiving in the streaming case) to be:\n>> \n>>  packet:          git< ok size=7\\n    (or \"ok 7\\n\", if size is known)\n>> \n>> or\n>> \n>>  packet:          git< ok\\n           (if filter does not know size upfront)\n>> \n>> or\n>> \n>>  packet:          git< fail <msg>\\n   (or just \"fail\" + packet with msg)\n>> \n>> The last would be when filter knows upfront that it cannot perform\n>> the operation.  Though sending an empty file with non-\"success\" final\n>> would work as well.\n> \n> [...]\n> \n>>> In case the filter cannot process the content, it is expected\n>>> to respond with the result content size 0 (only if \"stream\" is\n>>> not defined) and a \"reject\" packet.\n>>> ------------------------\n>>> packet:          git< size=0\\n    (omitted with capability \"stream\")\n>>> packet:          git< reject\\n\n>>> ------------------------\n>> \n>> This is *wrong* idea!  Empty file, with size=0, can be a perfectly\n>> legitimate response.  \n> \n> Actually, I think I have misunderstood your intent.  If you want to have\n> simpler protocol, with only one place to signal errors, that is after\n> sending a response, then proper way of signaling the error condition\n> would be to send empty file and then \"reject\" instead of \"success\":\n> \n>   packet:          git< size=0\\n    (omitted with capability \"stream\")\n>   packet:          git< 0000        (we need this flush packet)\n>   packet:          git< reject\\n\n> \n> Otherwise in the case without size upfront (capability \"stream\")\n> file with contents \"reject\" would be mistaken for the \"reject\" packet.\n> \n> See below for proposal with two places to signal errors: before sending\n> first byte, and after.\n\nRight now the protocol is implemented covering the following cases:\n\n## CASE 1 - no stream success\n\npacket:          git< size=57\\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< success\\n\n\n\n## CASE 2 - no stream success but 0 byte response\n\npacket:          git< size=0\\n\npacket:          git< success\\n\n\n\n## CASE 3 - no stream filter; filter doesn't want to process the file\n\npacket:          git< size=0\\n\npacket:          git< reject\\n\n\n\n## CASE 4 - no stream filter; filter error\n\npacket:          git< size=57\\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< error\\n\n\nCASE 4 is not explicitly checked. If a final message is neither\n\"success\" nor \"reject\" then it is interpreted as error. If that\nhappens then Git will shutdown and restart the filter process\nif there is another file to filter. \n\nAlternatively a filter process can shutdown itself, too, to signal\nan error.\n\nThe corresponding stream filter look like this:\n\n## CASE 1 - stream success\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< success\\n\n\n\n## CASE 2 - stream success but 0 byte response\n\npacket:          git< 0000\npacket:          git< success\\n\n\n\n## CASE 3 - stream filter; filter doesn't want to process the file\n\npacket:          git< 0000\npacket:          git< reject\\n\n\n\n## CASE 4 - stream filter; filter error\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< error\\n\n\n--\n\nI just realized that the size 0 case is a bit inconsistent\nin the no stream case as it has no flush packet. Maybe I \nshould indeed remove the flush packet in the no stream case\ncompletely?!\n\nDo the cases above make sense to you?\n\nRegarding error handling. I would prefer it if the filter prints\nall errors to STDERR by itself. I think that is the safest\noption to communicate errors to the users because if the communication\ngot into a bad state then Git might not be able to read the errors\nproperly.\n\nSee Peff's response on the topic, too:\nhttp://public-inbox.org/git/20160729165018.GA6553%40sigill.intra.peff.net/\n\n\n> NOTE: there is a bit of mixed and possibly confusing notation, that\n> is 0000 is flush packet, not packet with 0000 as content.  Perhaps\n> write pkt-line in full?\n\nI am not sure I understand what you mean (maybe it's too late for me...).\nCan you try to rephrase or give an example?\n\nThank you,\nLars\n\n\n\n> \n> \n> [...]\n>>> ---\n>>> Documentation/gitattributes.txt |  84 ++++++++-\n>>> convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--\n>>> t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++\n>>> t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++\n>>> 4 files changed, 1053 insertions(+), 13 deletions(-)\n>>> create mode 100755 t/t0021/rot13-filter.pl\n> \n> Wouldn't it be better for easier review to split it into separate patches?\n> Perhaps at least the new test...\n> \n> [...]\n>> I would assume that we have two error conditions.  \n>> \n>> First situation is when the filter knows upfront (after receiving name\n>> and size of file, and after receiving contents for not-streaming filters)\n>> that it cannot process the file (like e.g. LFS filter with artifactory\n>> replica/shard being a bit behind master, and not including contents of\n>> the file being filtered).\n>> \n>> My proposal is to reply with \"fail\" _in place of_ size of reply:\n>> \n>>   packet:         git< fail\\n       (any case: size known or not, stream or not)\n>> \n>> It could be \"reject\", or \"error\" instead of \"fail\".\n>> \n>> \n>> Another situation is if filter encounters error during output,\n>> either with streaming filter (or non-stream, but not storing whole\n>> input upfront) realizing in the middle of output that there is something\n>> wrong with input (e.g. converting between encoding, and encountering\n>> character that cannot be represented in output encoding), or e.g. filter\n>> process being killed, or network connection dropping with LFS filter, etc.\n>> The filter has send some packets with output already.  In this case\n>> filter should flush, and send \"reject\" or \"error\" packet.\n>> \n>>   <error condition>\n>>   packet:         git< \"0000\"       (flush packet)\n>>   packet:         git< reject\\n\n>> \n>> Should there be a place for an error message, or would standard error\n>> (stderr) be used for this?\n> \n\n"},{"id":"292692","messageId":"63231F5B-959F-4A9D-89B9-E4A42AF34AB1@web.de","threadId":"42968","inReplyTo":"20160729233801.82844-4-larsxschneider@gmail.com","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Torstem Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2016-07-31T20:36:22Z","receivedAt":"2016-07-31T20:42:04Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"\n\n> Am 29.07.2016 um 20:37 schrieb larsxschneider@gmail.com:\n> \n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> packet_flush() would die in case of a write error even though for some callers\n> an error would be acceptable.\nWhat happens if there is a write error ?\nBasically the protocol is out of synch.\nLenght information is mixed up with payload, or the other way\naround.\nIt may be, that the consequences of a write error are acceptable,\nbecause a filter is allowed to fail.\nWhat is not acceptable is a \"broken\" protocol.\nThe consequence schould be to close the fd and tear down all\nresources. connected to it.\nIn our case to terminate the external filter daemon in some way,\nand to never use this instance again.\n\n\n> Add packet_flush_gentle() which writes a pkt-line\n> flush packet and returns `0` for success and `1` for failure.\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n> pkt-line.c | 6 ++++++\n> pkt-line.h | 1 +\n> 2 files changed, 7 insertions(+)\n> \n> diff --git a/pkt-line.c b/pkt-line.c\n> index 6fae508..1728690 100644\n> --- a/pkt-line.c\n> +++ b/pkt-line.c\n> @@ -91,6 +91,12 @@ void packet_flush(int fd)\n>  write_or_die(fd, \"0000\", 4);\n> }\n> \n> +int packet_flush_gentle(int fd)\n> +{\n> +    packet_trace(\"0000\", 4, 1);\n> +    return !write_or_whine_pipe(fd, \"0000\", 4, \"flush packet\");\n> +}\n> +\n> void packet_buf_flush(struct strbuf *buf)\n> {\n>  packet_trace(\"0000\", 4, 1);\n> diff --git a/pkt-line.h b/pkt-line.h\n> index 02dcced..3953c98 100644\n> --- a/pkt-line.h\n> +++ b/pkt-line.h\n> @@ -23,6 +23,7 @@ void packet_flush(int fd);\n> void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n> void packet_buf_flush(struct strbuf *buf);\n> void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n> +int packet_flush_gentle(int fd);\n> int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n> int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n> \n> -- \n> 2.9.0\n> \n"},{"id":"292693","messageId":"8FC2D283-AF8D-4643-834E-3D1927C558C0@gmail.com","threadId":"42968","inReplyTo":"63231F5B-959F-4A9D-89B9-E4A42AF34AB1@web.de","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-07-31T21:45:08Z","receivedAt":"2016-07-31T21:45:18Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 31 Jul 2016, at 22:36, Torstem Bögershausen <tboegi@web.de> wrote:\n> \n> \n> \n>> Am 29.07.2016 um 20:37 schrieb larsxschneider@gmail.com:\n>> \n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> packet_flush() would die in case of a write error even though for some callers\n>> an error would be acceptable.\n> What happens if there is a write error ?\n> Basically the protocol is out of synch.\n> Lenght information is mixed up with payload, or the other way\n> around.\n> It may be, that the consequences of a write error are acceptable,\n> because a filter is allowed to fail.\n> What is not acceptable is a \"broken\" protocol.\n> The consequence schould be to close the fd and tear down all\n> resources. connected to it.\n> In our case to terminate the external filter daemon in some way,\n> and to never use this instance again.\n\nCorrect! That is exactly what is happening in kill_protocol2_filter()\nhere:\n\n\n+static int apply_protocol2_filter(const char *path, const char *src, size_t len,\n+\t\t\t\t\t\tint fd, struct strbuf *dst, const char *cmd,\n+\t\t\t\t\t\tconst int wanted_capability)\n+{\n...\n+\tif (ret) {\n+\t\tstrbuf_swap(dst, &nbuf);\n+\t} else {\n+\t\tif (!filter_result || strcmp(filter_result, \"reject\")) {\n+\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n+\t\t\terror(\"external filter '%s' failed\", cmd);\n+\t\t\tkill_protocol2_filter(&cmd_process_map, entry);\n+\t\t}\n+\t}\n+\tstrbuf_release(&nbuf);\n+\treturn ret;\n+}\n\nMore context:\nhttps://github.com/larsxschneider/git/blob/e128326070847ac596e8bb21adebc8abab2003fc/convert.c#L821\n\n- Lars\n\n\n> \n> \n>> Add packet_flush_gentle() which writes a pkt-line\n>> flush packet and returns `0` for success and `1` for failure.\n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> pkt-line.c | 6 ++++++\n>> pkt-line.h | 1 +\n>> 2 files changed, 7 insertions(+)\n>> \n>> diff --git a/pkt-line.c b/pkt-line.c\n>> index 6fae508..1728690 100644\n>> --- a/pkt-line.c\n>> +++ b/pkt-line.c\n>> @@ -91,6 +91,12 @@ void packet_flush(int fd)\n>> write_or_die(fd, \"0000\", 4);\n>> }\n>> \n>> +int packet_flush_gentle(int fd)\n>> +{\n>> +    packet_trace(\"0000\", 4, 1);\n>> +    return !write_or_whine_pipe(fd, \"0000\", 4, \"flush packet\");\n>> +}\n>> +\n>> void packet_buf_flush(struct strbuf *buf)\n>> {\n>> packet_trace(\"0000\", 4, 1);\n>> diff --git a/pkt-line.h b/pkt-line.h\n>> index 02dcced..3953c98 100644\n>> --- a/pkt-line.h\n>> +++ b/pkt-line.h\n>> @@ -23,6 +23,7 @@ void packet_flush(int fd);\n>> void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n>> void packet_buf_flush(struct strbuf *buf);\n>> void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n>> +int packet_flush_gentle(int fd);\n>> int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n>> int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n>> \n>> -- \n>> 2.9.0\n>> \n\n"},{"id":"292694","messageId":"2f4743d1-3c93-406d-8b44-da0eb075e65c@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-11-larsxschneider@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-31T22:19:43Z","receivedAt":"2016-07-31T22:21:11Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n[...]\n> +Please note that you cannot use an existing filter.<driver>.clean\n> +or filter.<driver>.smudge command as filter.<driver>.process\n> +command.\n\nI think it would be more readable and easier to understand to write:\n\n  ... you cannot use an existing ... command with\n  filter.<driver>.process\n\nAbout the style: wouldn't `filter.<driver>.process` be better?\n\n>              As soon as Git would detect a file that needs to be\n> +processed by this filter, it would stop responding.\n\nThis is quite convoluted, and hard to understand.  I would say\nthat because `clean` and `smudge` filters are expected to read\nfirst, while Git expects `process` filter to say first, using\n`clean` or `smudge` filter without changes as `process` filter\nwould lead to git command deadlocking / hanging / stopping\nresponding.\n\n> +\n> +\n>  Interaction between checkin/checkout attributes\n>  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n>  \n> diff --git a/convert.c b/convert.c\n> index 522e2c5..be6405c 100644\n> --- a/convert.c\n> +++ b/convert.c\n> @@ -3,6 +3,7 @@\n>  #include \"run-command.h\"\n>  #include \"quote.h\"\n>  #include \"sigchain.h\"\n> +#include \"pkt-line.h\"\n>  \n>  /*\n>   * convert.c - convert a file when checking it out and checking it in.\n> @@ -481,11 +482,355 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n>  \treturn ret;\n>  }\n>  \n> +static int multi_packet_read(int fd_in, struct strbuf *sb, size_t expected_bytes, int is_stream)\n\nAbout name of this function: `multi_packet_read` is fine, though I wonder\nif `packet_read_in_full` with nearly the same parameters as `packet_read`,\nor `packet_read_till_flush`, or `read_in_full_packetized` would be better.\n\nAlso, the problem is that while we know that what packet_read() stores\nwould fit in memory (in size_t), it is not true for reading whole file,\nwhich might be very large - for example huge graphical assets like raw\nimages or raw videos, or virtual machine images.  Isn't that the goal\nof git-LFS solutions, which need this feature?  Shouldn't we have then\nboth `multi_packet_read_to_fd` and `multi_packet_read_to_buf`,\nor whatever?\n\nAlso, if we have `fd_in`, then perhaps `sb_out`?\n\nI am also unsure if `expected_bytes` (or `expected_size`) should not be\njust a size hint, leaving handing mismatch between expected size and\nreal size of output to the caller; then the `is_stream` would be not\nneeded.\n\n> +{\n> +\tint bytes_read;\n> +\tsize_t total_bytes_read = 0;\n\nWhy `bytes_read` is int, while `total_bytes_read` is size_t? Ah, I see\nthat packet_read() returns an int.  It should be ssize_t, just like\nread(), isn't it?  But we know that packet size is limited, and would\nfit in an int (or would it?).\n\nAlso, total_bytes_read could overflow size_t, but then we would have\nproblems storing the result in strbuf.\n\n> +\tif (expected_bytes == 0 && !is_stream)\n> +\t\treturn 0;\n\nSo in all cases *except* size = 0 we expect flush packet after the\ncontents, but size = 0 is a corner case without flush packet?\n\n> +\n> +\tif (is_stream)\n> +\t\tstrbuf_grow(sb, LARGE_PACKET_MAX);           // allocate space for at least one packet\n> +\telse\n> +\t\tstrbuf_grow(sb, st_add(expected_bytes, 1));  // add one extra byte for the packet flush\n> +\n> +\tdo {\n> +\t\tbytes_read = packet_read(\n> +\t\t\tfd_in, NULL, NULL,\n> +\t\t\tsb->buf + total_bytes_read, sb->len - total_bytes_read - 1,\n> +\t\t\tPACKET_READ_GENTLE_ON_EOF\n> +\t\t);\n> +\t\tif (bytes_read < 0)\n> +\t\t\treturn 1;  // unexpected EOF\n\nDon't we usually return negative numbers on error?  Ah, I see that the\nreturn is a bool, which allows to use boolean expression with 'return'.\nBut I am still unsure if it is good API, this return value.\n\nIf we move handling of size mismatch to the caller, then the function\ncan simply return the size of data read (probably off_t or uint64_t).\nThen the caller can check if it is what it expected, and react accordingly.\n\n> +\n> +\t\tif (is_stream &&\n> +\t\t\tbytes_read > 0 &&\n> +\t\t\tsb->len - total_bytes_read - 1 <= 0)\n> +\t\t\tstrbuf_grow(sb, st_add(sb->len, LARGE_PACKET_MAX));\n> +\t\ttotal_bytes_read += bytes_read;\n> +\t}\n> +\twhile (\n> +\t\tbytes_read > 0 &&                   // the last packet was no flush\n> +\t\tsb->len - total_bytes_read - 1 > 0  // we still have space left in the buffer\n\nAh, so buffer is resized only in the 'is_stream' case.  Perhaps then\nuse an \"int options\" instead of 'is_stream', and have one of flags\ntell if we should resize or not, that is if size parameter is hint\nor a strict limit.\n\n> +\t);\n> +\tstrbuf_setlen(sb, total_bytes_read);\n> +\treturn (is_stream ? 0 : expected_bytes != total_bytes_read);\n> +}\n> +\n> +static int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n\nIs it equivalent of copy_fd() function, but where destination uses pkt-line\nand we need to pack data into pkt-lines?\n\n> +{\n> +\tint did_fail = 0;\n> +\tssize_t bytes_to_write;\n> +\twhile (!did_fail) {\n> +\t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_LEN);\n\nUsing global variable packet_buffer makes this code thread-unsafe, isn't it?\nBut perhaps that is not a problem, because other functions are also\nusing this global variable.\n\nIt is more of PKTLINE_DATA_MAXLEN, isn't it?\n\n> +\t\tif (bytes_to_write < 0)\n> +\t\t\treturn 1;\n> +\t\tif (bytes_to_write == 0)\n> +\t\t\tbreak;\n> +\t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n> +\t}\n> +\tif (!did_fail)\n> +\t\tdid_fail = packet_flush_gentle(fd_out);\n\nShouldn't we try to flush even if there was an error?  Or is it\nthat if there is an error writing, then there is some problem\nsuch that we know that flush would not work?\n\n> +\treturn did_fail;\n\nReturn true on fail?  Shouldn't we follow example of copy_fd()\nfrom copy.c, and return COPY_READ_ERROR, or COPY_WRITE_ERROR,\nor PKTLINE_WRITE_ERROR?\n\n\n> +}\n> +\n> +static int multi_packet_write_from_buf(const char *src, size_t len, int fd_out)\n\nIt is equivalent of write_in_full(), with different order of parameters,\nbut where destination file descriptor expects pkt-line and we need to pack\ndata into pkt-lines?\n\nNOTE: function description comments?\n\n> +{\n> +\tint did_fail = 0;\n> +\tsize_t bytes_written = 0;\n> +\tsize_t bytes_to_write;\n\nNote to self: bytes_to_write should fit in size_t, as it is limited to\nPKTLINE_DATA_LEN.  bytes_written should fit in size_t, as it is at most\nlen, which is of type size_t.\n\n> +\twhile (!did_fail) {\n> +\t\tif ((len - bytes_written) > PKTLINE_DATA_LEN)\n> +\t\t\tbytes_to_write = PKTLINE_DATA_LEN;\n> +\t\telse\n> +\t\t\tbytes_to_write = len - bytes_written;\n> +\t\tif (bytes_to_write == 0)\n> +\t\t\tbreak;\n> +\t\tdid_fail |= direct_packet_write_data(fd_out, src + bytes_written, bytes_to_write, 1);\n> +\t\tbytes_written += bytes_to_write;\n\nAh, I see now why we need both direct_packet_write() and\ndirect_packet_write_data().  Nice abstraction, makes for\nclear code.\n\nThe last parameter of '1' means 'gently', isn't it?\n\n> +\t}\n> +\tif (!did_fail)\n> +\t\tdid_fail = packet_flush_gentle(fd_out);\n> +\treturn did_fail;\n> +}\n\nI think all three/four of those functions should be added in a separate\ncommit, separate patch in patch series.  Namely:\n\n - for git -> filter:\n    * read from fd,      write pkt-line to fd  (off_t)\n    * read from str+len, write pkt-line to fd  (size_t, ssize_t)\n - for filter -> git:\n    * read pkt-line from fd, write to fd       (off_t)\n    * read pkt-line from fd, write to str+len  (size_t, ssize_t)\n\nPerhaps some of those can be in one overloaded function, perhaps it would\nbe easier to keep them separate.\n\nAlso, I do wonder how the fetch / push code spools pack file received\nover pkt-lines to disk.  Can we reuse that code?  Or maybe that code\ncould use those new functions?\n\n\n> +\n> +#define FILTER_CAPABILITIES_STREAM   0x1\n> +#define FILTER_CAPABILITIES_CLEAN    0x2\n> +#define FILTER_CAPABILITIES_SMUDGE   0x4\n> +#define FILTER_CAPABILITIES_SHUTDOWN 0x8\n> +#define FILTER_SUPPORTS_STREAM(type) ((type) & FILTER_CAPABILITIES_STREAM)\n> +#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n> +#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n> +#define FILTER_SUPPORTS_SHUTDOWN(type) ((type) & FILTER_CAPABILITIES_SHUTDOWN)\n> +\n> +struct cmd2process {\n> +\tstruct hashmap_entry ent; /* must be the first member! */\n> +\tconst char *cmd;\n> +\tint supported_capabilities;\n\nI wonder if switching from int (perhaps with field width of 1 to denote\nthat it is boolean-like flag) to mask makes it more readable, or less.\nBut I think it is.\n\n\nReading Documentation/technical/api-hashmap.txt I found the following\nrecommendation:\n\n  `struct hashmap_entry`::\n\n        An opaque structure representing an entry in the hash table, which must\n        be used as first member of user data structures. Ideally it should be\n        followed by an int-sized member to prevent unused memory on 64-bit\n        systems due to alignment.\n\nTherefore it \"int supported_capabilities\" should precede\n\"const char *cmd\", I think.  Though it is not strictly necessary; it\nis not as if this hash table were large (maximum size is limited by\nthe number of filter drivers configured), so we don't waste much space\ndue to internal padding / due to alignment.\n\n> +\tstruct child_process process;\n> +};\n> +\n> +static int cmd_process_map_initialized = 0;\n> +static struct hashmap cmd_process_map;\n\nReading Documentation/technical/api-hashmap.txt I see that:\n\n  `tablesize` is the allocated size of the hash table. A non-0 value indicates\n  that the hashmap is initialized.\n\nSo cmd_process_map_initialized is not really needed, is it?\n\n> +\n> +static int cmd2process_cmp(const struct cmd2process *e1,\n> +\t\t\t\t\t\t\tconst struct cmd2process *e2,\n> +\t\t\t\t\t\t\tconst void *unused)\n> +{\n> +\treturn strcmp(e1->cmd, e2->cmd);\n> +}\n\nWell, to be exact (which is decidely not needed!) two commands might\nbe equivalent not being identical as strings (e.g. extra space between\nparameters).  But it is something the user should care about, not Git.\n\n> +\n> +static struct cmd2process *find_protocol2_filter_entry(struct hashmap *hashmap, const char *cmd)\n\nI'm not sure if *_protocol2_* is needed; those functions are static,\nlocal to convert.c.\n\n> +{\n> +\tstruct cmd2process k;\n\nDoes this name of variable 'k' follow established convention?\n'key' would be more descriptive, but it's not as if this function\nwas long; so 'k' is all right, I think.\n\n> +\thashmap_entry_init(&k, strhash(cmd));\n> +\tk.cmd = cmd;\n> +\treturn hashmap_get(hashmap, &k, NULL);\n> +}\n> +\n> +static void kill_protocol2_filter(struct hashmap *hashmap, struct cmd2process *entry) {\n\nProgramming style: the opening brace should be on separate line,\nthat is:\n\n  +static void kill_protocol2_filter(struct hashmap *hashmap, struct cmd2process *entry)\n  +{\n\n> +\tif (!entry)\n> +\t\treturn;\n> +\tsigchain_push(SIGPIPE, SIG_IGN);\n> +\tclose(entry->process.in);\n> +\tclose(entry->process.out);\n> +\tsigchain_pop(SIGPIPE);\n> +\tfinish_command(&entry->process);\n> +\tchild_process_clear(&entry->process);\n> +\thashmap_remove(hashmap, entry, NULL);\n> +\tfree(entry);\n> +}\n\nAll those, from #define FILTER_CAPABILITIES_ to here could be put\nin a separate patch, to reduce size of this one.  But I am less\nsure that it is worth it for this case.\n\n> +\n> +void shutdown_protocol2_filter(pid_t pid)\n> +{\n[...]\n\nIn my opinion this should be postponed to a separate commit.\n\n> +}\n> +\n> +static struct cmd2process *start_protocol2_filter(struct hashmap *hashmap, const char *cmd)\n\nThis has some parts in common with existing filter_buffer_or_fd().\nI wonder if it would be worth to extract those common parts.\n\nBut perhaps it would be better to leave such refactoring for later.\n\n> +{\n> +\tint did_fail;\n> +\tstruct cmd2process *entry;\n> +\tstruct child_process *process;\n> +\tconst char *argv[] = { cmd, NULL };\n> +\tstruct string_list capabilities = STRING_LIST_INIT_NODUP;\n> +\tchar *capabilities_buffer;\n> +\tint i;\n> +\n> +\tentry = xmalloc(sizeof(*entry));\n> +\thashmap_entry_init(entry, strhash(cmd));\n> +\tentry->cmd = cmd;\n> +\tentry->supported_capabilities = 0;\n> +\tprocess = &entry->process;\n> +\n> +\tchild_process_init(process);\n\nfilter_buffer_or_fd() uses instead\n\n  struct child_process child_process = CHILD_PROCESS_INIT;\n\nBut I see that you need to access &entry->process anyway, so you\nneed to have it here, and in this case child_process_init() is\nequivalent.\n\nI wonder if it would be worth it to use strbuf for cmd.\n\n> +\tprocess->argv = argv;\n> +\tprocess->use_shell = 1;\n> +\tprocess->in = -1;\n> +\tprocess->out = -1;\n> +\tprocess->clean_on_exit = 1;\n> +\tprocess->clean_on_exit_handler = shutdown_protocol2_filter;\n\nThese two lines are new, and related to the \"shutdown\" capability, isn't it?\n\n> +\n> +\tif (start_command(process)) {\n> +\t\terror(\"cannot fork to run external filter '%s'\", cmd);\n> +\t\tkill_protocol2_filter(hashmap, entry);\n\nI guess the alternative solution of adding filter to the hashmap only\nafter starting the process would be racy?\n\nAh, disregard that. I see that this pattern is a common way to error\nout in this function (for process-related errors).\n\n> +\t\treturn NULL;\n> +\t}\n> +\n> +\tsigchain_push(SIGPIPE, SIG_IGN);\n> +\tdid_fail = strcmp(packet_read_line(process->out, NULL), \"git-filter-protocol\");\n> +\tif (!did_fail)\n> +\t\tdid_fail |= strcmp(packet_read_line(process->out, NULL), \"version 2\");\n> +\tif (!did_fail)\n> +\t\tcapabilities_buffer = packet_read_line(process->out, NULL);\n> +\telse\n> +\t\tcapabilities_buffer = NULL;\n> +\tsigchain_pop(SIGPIPE);\n> +\n> +\tif (!did_fail && capabilities_buffer) {\n> +\t\tstring_list_split_in_place(&capabilities, capabilities_buffer, ' ', -1);\n> +\t\tif (capabilities.nr > 1 &&\n> +\t\t\t!strcmp(capabilities.items[0].string, \"capabilities\")) {\n> +\t\t\tfor (i = 1; i < capabilities.nr; i++) {\n> +\t\t\t\tconst char *requested = capabilities.items[i].string;\n> +\t\t\t\tif (!strcmp(requested, \"stream\")) {\n> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_STREAM;\n> +\t\t\t\t} else if (!strcmp(requested, \"clean\")) {\n> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_CLEAN;\n> +\t\t\t\t} else if (!strcmp(requested, \"smudge\")) {\n> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SMUDGE;\n> +\t\t\t\t} else if (!strcmp(requested, \"shutdown\")) {\n> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SHUTDOWN;\n> +\t\t\t\t} else {\n> +\t\t\t\t\twarning(\n> +\t\t\t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n> +\t\t\t\t\t\tcmd, requested\n> +\t\t\t\t\t);\n> +\t\t\t\t}\n> +\t\t\t}\n> +\t\t} else {\n> +\t\t\terror(\"filter capabilities not found\");\n> +\t\t\tdid_fail = 1;\n> +\t\t}\n> +\t\tstring_list_clear(&capabilities, 0);\n> +\t}\n\nI wonder if the above conditional wouldn't be better to be put in\na separate function, parse_filter_capabilities(capabilities_buffer),\nreturning a mask, or having mask as an out parameter, and returning\nan error condition.\n\n> +\n> +\tif (did_fail) {\n> +\t\terror(\"initialization for external filter '%s' failed\", cmd);\n\nMore detailed information not needed, because one can use GIT_PACKET_TRACE.\nWould it be worth add this information as a kind of advice, or put it\nin the documentation of the `process` option?\n\n> +\t\tkill_protocol2_filter(hashmap, entry);\n> +\t\treturn NULL;\n> +\t}\n> +\n> +\thashmap_add(hashmap, entry);\n> +\treturn entry;\n> +}\n> +\n> +static int apply_protocol2_filter(const char *path, const char *src, size_t len,\n> +\t\t\t\t\t\tint fd, struct strbuf *dst, const char *cmd,\n> +\t\t\t\t\t\tconst int wanted_capability)\n\napply_protocol2_filter, or apply_process_filter?  Or rather,\ns/_protocol2_/_process_/g ?\n\nThis is equivalent to\n\n   static int apply_filter(const char *path, const char *src, size_t len, int fd,\n                           struct strbuf *dst, const char *cmd)\n\nCould we have extended that one instead?\n\n> +{\n> +\tint ret = 1;\n> +\tstruct cmd2process *entry;\n> +\tstruct child_process *process;\n> +\tstruct stat file_stat;\n> +\tstruct strbuf nbuf = STRBUF_INIT;\n> +\tsize_t expected_bytes = 0;\n> +\tchar *strtol_end;\n> +\tchar *strbuf;\n> +\tchar *filter_type;\n> +\tchar *filter_result = NULL;\n> +\n\n> +\tif (!cmd || !*cmd)\n> +\t\treturn 0;\n> +\n> +\tif (!dst)\n> +\t\treturn 1;\n\nThis is the same as in apply_filter().\n\n> +\n> +\tif (!cmd_process_map_initialized) {\n> +\t\tcmd_process_map_initialized = 1;\n> +\t\thashmap_init(&cmd_process_map, (hashmap_cmp_fn) cmd2process_cmp, 0);\n> +\t\tentry = NULL;\n> +\t} else {\n> +\t\tentry = find_protocol2_filter_entry(&cmd_process_map, cmd);\n> +\t}\n\nHere we try to find existing process, rather than starting new\nas in apply_filter()\n\n> +\n> +\tfflush(NULL);\n\nThis is the same as in apply_filter(), but I wonder what it is for.\n\n> +\n> +\tif (!entry) {\n> +\t\tentry = start_protocol2_filter(&cmd_process_map, cmd);\n> +\t\tif (!entry) {\n> +\t\t\treturn 0;\n> +\t\t}\n\nStyle; we prefer:\n\n  +\t\tif (!entry)\n  +\t\t\treturn 0;\n\nThis is very similar to apply_filter(), but the latter uses start_async()\nfrom \"run-command.h\", with filter_buffer_or_fd() as asynchronous process,\nwhich gets passed command to run in struct filter_params.  In this\nfunction start_protocol2_filter() runs start_command(), synchronous API.\n\nWhy the difference?\n\n> +\t}\n> +\tprocess = &entry->process;\n> +\n> +\tif (!(wanted_capability & entry->supported_capabilities))\n> +\t\treturn 1;  // it is OK if the wanted capability is not supported\n> +\n> +\tif FILTER_SUPPORTS_CLEAN(wanted_capability)\n> +\t\tfilter_type = \"clean\";\n> +\telse if FILTER_SUPPORTS_SMUDGE(wanted_capability)\n> +\t\tfilter_type = \"smudge\";\n> +\telse\n> +\t\tdie(\"unexpected filter type\");\n\nStyle: it should be\n\n  +\tif (FILTER_SUPPORTS_CLEAN(wanted_capability))\n  +\t\tfilter_type = \"clean\";\n  +\telse if (FILTER_SUPPORTS_SMUDGE(wanted_capability))\n  +\t\tfilter_type = \"smudge\";\n  +\telse\n  +\t\tdie(\"unexpected filter type\");\n\neven though by accident the macro provides the parentheses to \"if\".\n\nCan we make an error/die message more detailed?  Maybe it is\nnot possible...\n\n> +\n> +\tif (fd >= 0 && !src) {\n> +\t\tif (fstat(fd, &file_stat) == -1)\n> +\t\t\treturn 0;\n> +\t\tlen = file_stat.st_size;\n> +\t}\n\nAll right, when fstat() can fail?  Could we then send contents without\nsize upfront, or is it better to require size to make it more consistent\nfor filter drivers scripts?\n\nCould this whole \"send single file\" be put in a separate function?\nOr is it not worth it?\n\n> +\n> +\tsigchain_push(SIGPIPE, SIG_IGN);\n\nHmmm... ignoring SIGPIPE was good for one-shot filters.  Is it still\nO.K. for per-command persistent ones?\n\n> +\n> +\tpacket_buf_write(&nbuf, \"%s\\n\", filter_type);\n> +\tret &= !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n> +\n> +\tif (ret) {\n> +\t\tstrbuf_reset(&nbuf);\n> +\t\tpacket_buf_write(&nbuf, \"filename=%s\\n\", path);\n> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n> +\t}\n\nPerhaps a better solution would be\n\n        if (err)\n        \tgoto fin_error;\n\nrather than this.\n\n> +\n> +\tif (ret) {\n> +\t\tstrbuf_reset(&nbuf);\n> +\t\tpacket_buf_write(&nbuf, \"size=%\"PRIuMAX\"\\n\", (uintmax_t)len);\n> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n> +\t}\n\nOr maybe extract writing the header for a file into a separate function?\nThis one gets a bit long...\n\n> +\n> +\tif (ret) {\n> +\t\tif (fd >= 0)\n> +\t\t\tret = !multi_packet_write_from_fd(fd, process->in);\n> +\t\telse\n> +\t\t\tret = !multi_packet_write_from_buf(src, len, process->in);\n> +\t}\n\nThis is not streaming.  The above sends whole file, or whole string to\nthe filter process, without draining filter output.  If the filter were\nto read some, then write some, it might deadlock on full buffers, isn't it?\nOr am I mistaken?\n\n> +\n> +\tif (ret && !FILTER_SUPPORTS_STREAM(entry->supported_capabilities)) {\n> +\t\tstrbuf = packet_read_line(process->out, NULL);\n> +\t\tif (strlen(strbuf) > 5 && !strncmp(\"size=\", strbuf, 5)) {\n> +\t\t\texpected_bytes = (off_t)strtol(strbuf + 5, &strtol_end, 10);\n> +\t\t\tret = (strtol_end != strbuf && errno != ERANGE);\n> +\t\t} else {\n> +\t\t\tret = 0;\n> +\t\t}\n> +\t}\n> +\n> +\tif (ret) {\n> +\t\tstrbuf_reset(&nbuf);\n> +\t\tret = !multi_packet_read(process->out, &nbuf, expected_bytes,\n> +\t\t\tFILTER_SUPPORTS_STREAM(entry->supported_capabilities));\n> +\t}\n\nWhat happens if the output of filter does not fit in size_t?  I see that\n(I think) this problem is inherited from the original implementation.\n\n> +\n> +\tif (ret) {\n> +\t\tfilter_result = packet_read_line(process->out, NULL);\n> +\t\tret = !strcmp(filter_result, \"success\");\n> +\t}\n> +\n> +\tsigchain_pop(SIGPIPE);\n> +\n> +\tif (ret) {\n> +\t\tstrbuf_swap(dst, &nbuf);\n> +\t} else {\n> +\t\tif (!filter_result || strcmp(filter_result, \"reject\")) {\n> +\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n> +\t\t\terror(\"external filter '%s' failed\", cmd);\n> +\t\t\tkill_protocol2_filter(&cmd_process_map, entry);\n> +\t\t}\n> +\t}\n\nSo if Git gets finish signal \"success\" from filter, it accepts the output.\nIf Git gets finish signal \"reject\" from filter, it restarts filter (and\nreject the output - user can retry the command himself / herself).\nIf Git gets any other finish signal, for example \"error\" (but this is not\nstandarized), then it rejects the output, keeping the unfiltered result,\nbut keeps filtering.\n\nI think it is not described in this detail in the documentation of the\nnew protocol.\n\n> +\tstrbuf_release(&nbuf);\n> +\treturn ret;\n> +}\n\nI wonder if this point might be start of the new patch... but then you\nwould have no way to test what you wrote.\n\n> +\n>  static struct convert_driver {\n>  \tconst char *name;\n>  \tstruct convert_driver *next;\n>  \tconst char *smudge;\n>  \tconst char *clean;\n> +\tconst char *process;\n>  \tint required;\n>  } *user_convert, **user_convert_tail;\n\nAll right.\n\n>  \n> @@ -526,6 +871,10 @@ static int read_convert_config(const char *var, const char *value, void *cb)\n>  \tif (!strcmp(\"clean\", key))\n>  \t\treturn git_config_string(&drv->clean, var, value);\n>  \n> +\tif (!strcmp(\"process\", key)) {\n> +\t\treturn git_config_string(&drv->process, var, value);\n> +\t}\n> +\n\nAll right.\n\n>  \tif (!strcmp(\"required\", key)) {\n>  \t\tdrv->required = git_config_bool(var, value);\n>  \t\treturn 0;\n> @@ -823,7 +1172,12 @@ int would_convert_to_git_filter_fd(const char *path)\n>  \tif (!ca.drv->required)\n>  \t\treturn 0;\n>  \n> -\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n> +\tif (!ca.drv->clean && ca.drv->process)\n> +\t\treturn apply_protocol2_filter(\n> +\t\t\tpath, NULL, 0, -1, NULL, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n> +\t\t);\n> +\telse\n> +\t\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n\nCould we augment apply_filter() instead, so that the invocation is\n\n        return apply_filter(path, NULL, 0, -1, NULL, ca.drv, FILTER_CLEAN);\n\nThough I am not sure if moving this conditional to apply_filter would\nbe a good idea; maybe wrapper around augmented apply_filter_do()?\n\n>  }\n>  \n>  const char *get_convert_attr_ascii(const char *path)\n> @@ -856,17 +1210,24 @@ int convert_to_git(const char *path, const char *src, size_t len,\n>                     struct strbuf *dst, enum safe_crlf checksafe)\n>  {\n>  \tint ret = 0;\n> -\tconst char *filter = NULL;\n> +\tconst char *clean_filter = NULL;\n> +\tconst char *process_filter = NULL;\n>  \tint required = 0;\n>  \tstruct conv_attrs ca;\n>  \n>  \tconvert_attrs(&ca, path);\n>  \tif (ca.drv) {\n> -\t\tfilter = ca.drv->clean;\n> +\t\tclean_filter = ca.drv->clean;\n> +\t\tprocess_filter = ca.drv->process;\n>  \t\trequired = ca.drv->required;\n>  \t}\n\nAll right (assuming un-augmented apply_filter()).\n\n>  \n> -\tret |= apply_filter(path, src, len, -1, dst, filter);\n> +\tif (!clean_filter && process_filter)\n> +\t\tret |= apply_protocol2_filter(\n> +\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_CLEAN\n> +\t\t);\n> +\telse\n> +\t\tret |= apply_filter(path, src, len, -1, dst, clean_filter);\n\nI wonder if it would be more readable to write it like this\n(and of course elsewhere too):\n\n  +\tif (!clean_filter && process_filter)\n  +\t\tret |= apply_protocol2_filter(\n  +\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_CLEAN\n  +\t\t);\n  +\telse\n  +\t\tret |= apply_filter(\n  +\t\t\tpath, src, len, -1, dst, clean_filter);\n  +\t\t);\n\n\nThough it would screw up \"git blame -C -C -w\"\n\n>  \tif (!ret && required)\n>  \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n>  \n> @@ -885,13 +1246,21 @@ int convert_to_git(const char *path, const char *src, size_t len,\n>  void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n>  \t\t\t      enum safe_crlf checksafe)\n>  {\n> +\tint ret = 0;\n\nRight, 'ret' is needed because we now have two possibilities:\n`clean` filter and `process` filter.\n\n>  \tstruct conv_attrs ca;\n>  \tconvert_attrs(&ca, path);\n>  \n>  \tassert(ca.drv);\n> -\tassert(ca.drv->clean);\n> +\tassert(ca.drv->clean || ca.drv->process);\n> +\n> +\tif (!ca.drv->clean && ca.drv->process)\n> +\t\tret = apply_protocol2_filter(\n> +\t\t\tpath, NULL, 0, fd, dst, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n> +\t\t);\n> +\telse\n> +\t\tret = apply_filter(path, NULL, 0, fd, dst, ca.drv->clean);\n>  \n> -\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv->clean))\n> +\tif (!ret)\n>  \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n>  \n>  \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n> @@ -902,14 +1271,16 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n>  \t\t\t\t\t    size_t len, struct strbuf *dst,\n>  \t\t\t\t\t    int normalizing)\n>  {\n> -\tint ret = 0, ret_filter = 0;\n> -\tconst char *filter = NULL;\n> +\tint ret = 0, ret_filter;\n\nWhy the change:\n\n  -\tint ret = 0, ret_filter = 0;\n  +\tint ret = 0, ret_filter;\n\n> +\tconst char *smudge_filter = NULL;\n> +\tconst char *process_filter = NULL;\n>  \tint required = 0;\n>  \tstruct conv_attrs ca;\n>  \n>  \tconvert_attrs(&ca, path);\n>  \tif (ca.drv) {\n> -\t\tfilter = ca.drv->smudge;\n> +\t\tprocess_filter = ca.drv->process;\n> +\t\tsmudge_filter = ca.drv->smudge;\n>  \t\trequired = ca.drv->required;\n>  \t}\n\nAll right, the same.\n\n[...]\n> diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\n> index 34c8eb9..e8a7703 100755\n> --- a/t/t0021-conversion.sh\n> +++ b/t/t0021-conversion.sh\n> @@ -296,4 +296,409 @@ test_expect_success 'disable filter with empty override' '\n>  \ttest_must_be_empty err\n>  '\n>  \n> +test_expect_success PERL 'required process filter should filter data' '\n> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge shutdown\" &&\n> +\ttest_config_global filter.protocol.required true &&\n> +\trm -rf repo &&\n> +\tmkdir repo &&\n> +\t(\n> +\t\tcd repo &&\n> +\t\tgit init &&\n> +\n> +\t\techo \"*.r filter=protocol\" >.gitattributes &&\n> +\t\tgit add . &&\n> +\t\tgit commit . -m \"test commit\" &&\n\nThis is more of \"Initial commit\", not that it matters\n\n> +\t\tgit branch empty &&\n> +\n> +\t\tcat ../test.o >test.r &&\n\nErr, the above is just copying file, isn't it?\nMaybe it was copied from other tests, I have not checked.\n\n> +\t\techo \"test22\" >test2.r &&\n> +\t\tmkdir testsubdir &&\n> +\t\techo \"test333\" >testsubdir/test3.r &&\n\nAll right, we test text file, we test binary file (I assume), we test\nfile in a subdirectory.  What about testing empty file?  Or large file\nwhich would not fit in the stdin/stdout buffer (as EXPENSIVE test)?\n\n> +\n> +\t\trm -f rot13-filter.log &&\n> +\t\tgit add . &&\n\nSo this runs \"clean\" filter, storing cleaned contents in the index.\n\n> +\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >uniq-rot13-filter.log &&\n> +\t\tcat >expected_add.log <<-\\EOF &&\n> +\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n> +\t\t\t1 IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n> +\t\t\t1 IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n\nAnd we check the \"know size upfront\" case (mistakenly called non-\"stream\").\n\n> +\t\t\t1 IN: shutdown -- [OK]\n\nAnd test \"shutdown\" capability (not as separate test).\n\n> +\t\t\t1 start\n> +\t\t\t1 wrote filter header\n> +\t\tEOF\n\nAnd we are required to keep the expected_add.log file sorted by hand???\n\n> +\t\ttest_cmp expected_add.log uniq-rot13-filter.log &&\n> +\n> +\t\t>rot13-filter.log &&\n\nTruncate log. Still in the same test.\n\n> +\t\tgit commit . -m \"test commit\" &&\n\nThis is test commit with files undergoing \"clean\" part of filter.\n\n> +\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n> +\t\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >uniq-rot13-filter.log &&\n\nThere is known performance regression, in that filter is run more\nthan once on given file.\n\nActually... why it does not use cleaned-up contents from the index?\n\n> +\t\tcat >expected_commit.log <<-\\EOF &&\n> +\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n> +\t\t\tx IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n> +\t\t\tx IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n> +\t\t\t1 IN: shutdown -- [OK]\n> +\t\t\t1 start\n> +\t\t\t1 wrote filter header\n\nRight, this is the goal of the patch series: for filter to be started\nonly once per git command invocation.\n\n> +\t\tEOF\n> +\t\ttest_cmp expected_commit.log uniq-rot13-filter.log &&\n> +\n\nStill in the same test, even though we would be testing \"smudge\"\ncapability now.  \n\nIt's a pity that t/test-lib.sh does not support subtests from\nthe TAP specification (Test Anything Protocol that Git testsuite\nuses).\n\n> +\t\t>rot13-filter.log &&\n> +\t\trm -f test?.r testsubdir/test3.r &&\n> +\t\tgit checkout . &&\n\nAll right, we removed some files so that \"git checkout .\" could\nrestore them to life.\n\n> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n\nUseless use of cat\n\n  +\t\tgrep -v \"IN: clean\"  rot13-filter.log  >smudge-rot13-filter.log &&\n\nAlso: why 'git checkout <path>' would run \"clean\" filter?\nIs it existing strange behaviour?\n\n> +\t\tcat >expected_checkout.log <<-\\EOF &&\n> +\t\t\tstart\n> +\t\t\twrote filter header\n> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n> +\t\t\tIN: shutdown -- [OK]\n> +\t\tEOF\n\nThis time without 'sort | uniq -c'.  Is it really needed for the\n\"good\" case, or is it there for two cases to look similar?\n\n> +\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n> +\n> +\t\tgit checkout empty &&\n\nShouldn't we check that switching to branch 'empty' does not run\nfilters, or is it covered by other tests?  Or perhaps this simply\ndoes not matter here, is it?\n\n> +\n> +\t\t>rot13-filter.log &&\n> +\t\tgit checkout master &&\n\nDoes it test different callpath than 'git checkout .'?  Well, the\nset of files is different...\n\n> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n> +\t\tcat >expected_checkout_master.log <<-\\EOF &&\n> +\t\t\tstart\n> +\t\t\twrote filter header\n> +\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n> +\t\t\tIN: shutdown -- [OK]\n> +\t\tEOF\n> +\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n> +\n\nAnd here we start checking that the filter did filter,\nthat is the content in the repository is \"clean\"ed-up.\nStill the same test.\n\n> +\t\t./../rot13.sh <test.r >expected &&\n> +\t\tgit cat-file blob :test.r >actual &&\n> +\t\ttest_cmp expected actual &&\n> +\n> +\t\t./../rot13.sh <test2.r >expected &&\n> +\t\tgit cat-file blob :test2.r >actual &&\n> +\t\ttest_cmp expected actual &&\n> +\n> +\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n> +\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n> +\t\ttest_cmp expected actual\n> +\t)\n> +'\n> +\n> +test_expect_success PERL 'required process filter should filter data stream' '\n> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl stream clean smudge\" &&\n> +\ttest_config_global filter.protocol.required true &&\n\nErrr... I don't see how it is different from the previous test.\n[...]\n\n> +\n> +test_expect_success PERL 'required process filter should filter smudge data and one-shot filter should clean' '\n\nAll right, so this tests the precedence... well, it doesn't.\n\nIt tests that `process` filter with \"smudge\" capability only works well\nwith one-shot `clean` filter.\n\n> +\ttest_config_global filter.protocol.clean ./../rot13.sh &&\n> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl smudge\" &&\n\nWhy the difference in pathnames (the directory part) between those two?\n\n> +\ttest_config_global filter.protocol.required true &&\n> +\trm -rf repo &&\n> +\tmkdir repo &&\n> +\t(\n> +\t\tcd repo &&\n> +\t\tgit init &&\n> +\n> +\t\techo \"*.r filter=protocol\" >.gitattributes &&\n> +\t\tgit add . &&\n> +\t\tgit commit . -m \"test commit\" &&\n> +\t\tgit branch empty &&\n> +\n> +\t\tcat ../test.o >test.r &&\n> +\t\techo \"test22\" >test2.r &&\n> +\t\tmkdir testsubdir &&\n> +\t\techo \"test333\" >testsubdir/test3.r &&\n> +\n> +\t\trm -f rot13-filter.log &&\n> +\t\tgit add . &&\n> +\t\ttest_must_be_empty rot13-filter.log &&\n> +\n> +\t\t>rot13-filter.log &&\n> +\t\tgit commit . -m \"test commit\" &&\n> +\t\ttest_must_be_empty rot13-filter.log &&\n\nAll right, these tests that `process` filter is not ran.  But we don't\nknow if it is because it lacks capability, or because it is overriden\nby one-shot filter (well, that comes later).\n\n> +\n> +\t\t>rot13-filter.log &&\n> +\t\trm -f test?.r testsubdir/test3.r &&\n> +\t\tgit checkout . &&\n> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n> +\t\tcat >expected_checkout.log <<-\\EOF &&\n> +\t\t\tstart\n> +\t\t\twrote filter header\n> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n> +\t\tEOF\n> +\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n\nThis part is repeated many, many times.  Maybe add some helper\nshell function for this?\n\n[...]\n> +\t\t./../rot13.sh <test.r >expected &&\n> +\t\tgit cat-file blob :test.r >actual &&\n> +\t\ttest_cmp expected actual &&\n> +\n> +\t\t./../rot13.sh <test2.r >expected &&\n> +\t\tgit cat-file blob :test2.r >actual &&\n> +\t\ttest_cmp expected actual &&\n> +\n> +\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n> +\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n> +\t\ttest_cmp expected actual\n\nHere we test that equivalent one-shot cleanup filter was run.\nHere also we have repeated contents; maybe some helper function\nwould make it shorter?\n\n> +\t)\n> +'\n\nHere I am stopping examining tests in detail.\n\n> +test_expect_success PERL 'required process filter should clean only' '\n> +test_expect_success PERL 'required process filter should process files larger LARGE_PACKET_MAX' '\n\nThose two tests do not depend on being required or not; it is only\nthat without required they would fail softly in case of latter test\n(which we can detect too).\n\n> +test_expect_success PERL 'required process filter should with clean error should fail' '\n> +test_expect_success PERL 'process filter should restart after unexpected write failure' '\n\nSo these two are sort of complimentary.  When `process` is required,\nthen it should fail if it cannot filter some file.  If it is not,\nit should keep processing other files.\n\n> +test_expect_success PERL 'process filter should not restart after intentionally rejected file' '\n\nUh... all right, so \"reject\" means that filter cannot continue?\nStrange meaning for 'reject', though ;-)\n\n>  test_done\n> diff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\n> new file mode 100755\n> index 0000000..cb0925d\n> --- /dev/null\n> +++ b/t/t0021/rot13-filter.pl\n> @@ -0,0 +1,177 @@\n> +#!/usr/bin/perl\n> +#\n> +# Example implementation for the Git filter protocol version 2\n> +# See Documentation/gitattributes.txt, section \"Filter Protocol\"\n> +#\n> +# The script takes the list of supported protocol capabilities as\n> +# arguments (\"stream\", \"clean\", and \"smudge\" are supported).\n\nWhat about \"shutdown\"?\n\n> +#\n> +# This implementation supports three special test cases:\n> +# (1) If data with the filename \"clean-write-fail.r\" is processed with\n> +#     a \"clean\" operation then the write operation will die.\n> +# (2) If data with the filename \"smudge-write-fail.r\" is processed with\n> +#     a \"smudge\" operation then the write operation will die.\n\nAll right, so it is hard failure with filter script dying.\n\n> +# (3) If data with the filename \"failure.r\" is processed with any\n> +#     operation then the filter signals that the operation was not\n> +#     successful.\n\nAll right, so it is failure detected by filter script and signalled to Git.\n\n> +#\n> +\n> +use strict;\n> +use warnings;\n\nSo no more \"use autodie\", because of compatibility with old Perls.\n\n> +\n> +my $MAX_PACKET_CONTENT_SIZE = 65516;\n> +my @capabilities            = @ARGV;\n\nNo autoflush this time?\n\n> +\n> +sub rot13 {\n> +    my ($str) = @_;\n> +    $str =~ y/A-Za-z/N-ZA-Mn-za-m/;\n> +    return $str;\n> +}\n> +\n> +sub packet_read {\n> +    my $buffer;\n> +    my $bytes_read = read STDIN, $buffer, 4;\n> +    if ( $bytes_read == 0 ) {\n> +        return;\n> +    }\n> +    elsif ( $bytes_read != 4 ) {\n> +        die \"invalid packet size '$bytes_read' field\";\n> +    }\n> +    my $pkt_size = hex($buffer);\n> +    if ( $pkt_size == 0 ) {\n> +        return ( 1, \"\" );\n\nUnusual return convention.  Though it is a test script, so\nit doesn't matter much.\n\n> +    }\n> +    elsif ( $pkt_size > 4 ) {\n> +        my $content_size = $pkt_size - 4;\n> +        $bytes_read = read STDIN, $buffer, $content_size;\n> +        if ( $bytes_read != $content_size ) {\n> +            die \"invalid packet\";\n\nMore detailed error message, maybe?\n\n> +        }\n> +        return ( 0, $buffer );\n> +    }\n> +    else {\n> +        die \"invalid packet size\";\n> +    }\n> +}\n> +\n> +sub packet_write {\n> +    my ($packet) = @_;\n> +    print STDOUT sprintf( \"%04x\", length($packet) + 4 );\n> +    print STDOUT $packet;\n> +    STDOUT->flush();\n> +}\n> +\n> +sub packet_flush {\n> +    print STDOUT sprintf( \"%04x\", 0 );\n> +    STDOUT->flush();\n> +}\n> +\n> +open my $debug, \">>\", \"rot13-filter.log\";\n> +print $debug \"start\\n\";\n> +$debug->flush();\n> +\n> +packet_write(\"git-filter-protocol\\n\");\n> +packet_write(\"version 2\\n\");\n> +packet_write( \"capabilities \" . join( ' ', @capabilities ) . \"\\n\" );\n> +print $debug \"wrote filter header\\n\";\n> +$debug->flush();\n> +\n> +while (1) {\n> +    my $command = packet_read();\n> +    unless ( defined($command) ) {\n> +        exit();\n> +    }\n> +    chomp $command;\n> +    print $debug \"IN: $command\";\n> +    $debug->flush();\n> +\n> +    if ( $command eq \"shutdown\" ) {\n> +        print $debug \" -- [OK]\";\n> +        $debug->flush();\n> +        packet_write(\"done\\n\");\n> +        exit();\n> +    }\n> +\n> +    my ($filename) = packet_read() =~ /filename=([^=]+)\\n/;\n> +    print $debug \" $filename\";\n> +    $debug->flush();\n> +    my ($filelen) = packet_read() =~ /size=([^=]+)\\n/;\n> +    chomp $filelen;\n\nI think this chomp is not needed, as \"\\n\" is not included.\nThough the regexp should probably be anchored.\n\n> +    print $debug \" $filelen\";\n> +    $debug->flush();\n> +\n> +    $filelen =~ /\\A\\d+\\z/ or die \"bad filelen: $filelen\";\n> +    my $output;\n> +\n> +    if ( $filelen > 0 ) {\n\nSo here is a special case for $filelen = 0.\nNegative $filelen is not allowed, via regexp.\n\n> +        my $input = \"\";\n> +        {\n> +            binmode(STDIN);\n> +            my $buffer;\n> +            my $done = 0;\n> +            while ( !$done ) {\n> +                ( $done, $buffer ) = packet_read();\n> +                $input .= $buffer;\n> +            }\n> +            print $debug \" [OK] -- \";\n> +            $debug->flush();\n> +        }\n> +\n> +        if ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n> +            $output = rot13($input);\n> +        }\n> +        elsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n> +            $output = rot13($input);\n> +        }\n\nThese two conditionals could be shortened, but then they would be less\nreadable.  Or not:\n\n           if ( grep { $_ eq $command } @capabilities ) {\n           \t$output = rot13($input);\n           }\n\n> +        else {\n> +            die \"bad command $command\";\n> +        }\n> +    }\n> +\n> +    my $output_len = length($output);\n> +    if ( $filename eq \"reject.r\" ) {\n> +        $output_len = 0;\n> +    }\n> +\n> +    if ( grep( /^stream$/, @capabilities ) ) {\n> +        print $debug \"OUT: STREAM \";\n> +    }\n> +    else {\n> +        packet_write(\"size=$output_len\\n\");\n> +        print $debug \"OUT: $output_len \";\n> +    }\n> +    $debug->flush();\n> +\n> +    if ( $filename eq \"reject.r\" ) {\n> +        packet_write(\"reject\\n\");\n> +        print $debug \"[REJECT]\\n\";    # Could also be an error\n\nHow if could be an error?\n\n> +        $debug->flush();\n> +    }\n> +\n> +    if ( $output_len > 0 ) {\n> +        if (( $command eq \"clean\" and $filename eq \"clean-write-fail.r\" )\n> +            or\n> +            ( $command eq \"smudge\" and $filename eq \"smudge-write-fail.r\" ))\n\nPerhaps simply:\n\n  +        if ( $filename eq \"${command}-write-fail.r\" ) {\n\n> +        {\n> +            print $debug \"[WRITE FAIL]\\n\";\n> +            $debug->flush();\n> +            die \"write error\";\n> +        }\n> +        else {\n> +            while ( length($output) > 0 ) {\n> +                my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n> +                packet_write($packet);\n> +                if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n> +                    $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n> +                }\n> +                else {\n> +                    $output = \"\";\n> +                }\n> +            }\n> +            packet_flush();\n> +            packet_write(\"success\\n\");\n> +            print $debug \"[OK]\\n\";\n> +            $debug->flush();\n> +        }\n> +    }\n> +}\n> \n\n"},{"id":"292695","messageId":"7255ef06-a9a0-91b7-b6da-a90322de926b@gmail.com","threadId":"42968","inReplyTo":"6765D972-876A-4F94-A170-468002498296@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-07-31T22:59:31Z","receivedAt":"2016-07-31T23:00:06Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 31.07.2016 o 21:49, Lars Schneider pisze: \n> On 31 Jul 2016, at 11:42, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 31.07.2016 o 00:05, Jakub Narębski pisze:\n>>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n[...]\n>>> I think it would be nice to have here at least summary of the benchmarks\n>>> you did in https://github.com/github/git-lfs/pull/1382\n\nThis would be nice to have in the commit message: real benchmarks.\n\n>>\n>> Note that this feature is especially useful if startup time is long,\n>> that is if you are using an operating system with costly fork / new process\n>> startup time like MS Windows (which you have mentioned), or writing\n>> filter in a programming language with large startup time like Java\n>> or Python (the latter may have changed since).\n>>\n>>  https://gnustavo.wordpress.com/2012/06/28/programming-languages-start-up-times/\n> \n> OK, I will add this. Is it OK to add the link to the commit message?\n> (since I don't know how long the link will be available).\n\nI don't think it is needed.  Perhaps only a sentence or half to notify\nwhere you could get most from this feature, but even then it is not\nnecessary.\n\nI'm sorry for the confusion.\n\n>> See below for proposal with two places to signal errors: before sending\n>> first byte, and after.\n> \n> Right now the protocol is implemented covering the following cases:\n> \n> ## CASE 1 - no stream success\n\nIt is less \"stream\", more \"size unknown\".  Real streaming is interleaving\nreading and writing, which is currently not supported due to lack of\nstart_async() - I think.\n\n> \n> packet:          git< size=57\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< success\\n\n\nRight.  What happens if either length(SMUDGED_CONTENT) < size,\nor length(SMUDGED_CONTENT) > size?  It could conceivably happen,\ne.g. due to an error in size calculation.\n\nNOTE that without using flush packet to signal end of contents,\nwe would be not able to signal a situation when filter encounters\nan error (per-file, or long temporary) when it have written some\ncontent already.  For example this may happen for git-LFS filter,\nif the server hosting artifactory (or even whole network) gets\ndown during cleanup / smudging.\n\nWell, unless we would use other special packets:\n - empty packet, that is \"0004\" pkt-line\n - invalid packet, that is \"0001\", \"0002\", \"0003\" pkt-line\nto signal premature end of SMUDGED_CONTENT.\n\n> \n> \n> ## CASE 2 - no stream success but 0 byte response\n> \n> packet:          git< size=0\\n\n> packet:          git< success\\n\n\nWhy there is need to special case 0 byte (empty file) response?\n\n  packet:          git< size=0\\n\n  packet:          git< 0000\n  packet:          git< success\\n\n\nis perfectly fine.\n  \n> ## CASE 3 - no stream filter; filter doesn't want to process the file\n> \n> packet:          git< size=0\\n\n> packet:          git< reject\\n\n\nWhy not simply\n \n  packet:          git< reject\\n\n\nOr, if we are going success/reject/whatever route\n\n  packet:          git< size=0\\n\n  packet:          git< 0000\n  packet:          git< reject\\n\n\n> ## CASE 4 - no stream filter; filter error\n> \n> packet:          git< size=57\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< error\\n\n> \n> CASE 4 is not explicitly checked. If a final message is neither\n> \"success\" nor \"reject\" then it is interpreted as error. If that\n> happens then Git will shutdown and restart the filter process\n> if there is another file to filter. \n\nThis should be documented.\n\n> \n> Alternatively a filter process can shutdown itself, too, to signal\n> an error.\n> \n> The corresponding stream filter look like this:\n> \n> ## CASE 1 - stream success\n> \n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< success\\n\n> \n> \n> ## CASE 2 - stream success but 0 byte response\n> \n> packet:          git< 0000\n> packet:          git< success\\n\n> \n> \n> ## CASE 3 - stream filter; filter doesn't want to process the file\n> \n> packet:          git< 0000\n> packet:          git< reject\\n\n> \n> \n> ## CASE 4 - stream filter; filter error\n> \n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< error\\n\n> \n> --\n> \n> I just realized that the size 0 case is a bit inconsistent\n> in the no stream case as it has no flush packet. Maybe I \n> should indeed remove the flush packet in the no stream case\n> completely?!\n\nThat's what I wrote about SPOT (single point of truth), of using\neither size or flush packet, but not both.  But...\n\nAs I wrote, you need some mechanism to signal premature end\nof contents, and start of an error description.\n\n> \n> Do the cases above make sense to you?\n\nExcept for the inconsistency of the size 0 case.  This what\nI meant to say.\n\n> \n> Regarding error handling. I would prefer it if the filter prints\n> all errors to STDERR by itself. I think that is the safest\n> option to communicate errors to the users because if the communication\n> got into a bad state then Git might not be able to read the errors\n> properly.\n> \n> See Peff's response on the topic, too:\n> http://public-inbox.org/git/20160729165018.GA6553%40sigill.intra.peff.net/\n\nActually it looks like Peff is slightly against using stderr.\n\nJK> Git-LFS sends to stderr because there's no other option. I wonder if it\nJK> would be nicer to make it Git's responsibility to talk to the user,\nJK> because then it could respect things like \"--quiet\". I guess error\nJK> messages are generally printed regardless of verbosity, though, so\nJK> printing them unconditionally is OK.\n\nI think it should be O.K., and it makes writing filter drivers\nsimpler if we don't have multiplex channels.\n\n>> NOTE: there is a bit of mixed and possibly confusing notation, that\n>> is 0000 is flush packet, not packet with 0000 as content.  Perhaps\n>> write pkt-line in full?\n> \n> I am not sure I understand what you mean (maybe it's too late for me...).\n> Can you try to rephrase or give an example?\n\nCompare\n\n  packet:          git< 0000\n\nwith\n\n  packet:          git< success\\n\n\nThe former as pkt-line is\n\n  git< 0000\n\nthe latter is\n\n  git< 000csuccess\\n\n       ^^^^\n           \\-- packet header\n\n-- \nJakub Narębski\n\n"},{"id":"292720","messageId":"73AED3FA-C666-4C49-90B2-387E410F7D52@gmail.com","threadId":"42968","inReplyTo":"4081bc44-d964-79ec-165f-f49f33823c17@gmail.com","subject":"Re: [PATCH v3 01/10] pkt-line: extract set_packet_header()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T11:33:32Z","receivedAt":"2016-08-01T11:33:49Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 12:30, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> set_packet_header() converts an integer to a 4 byte hex string. Make\n>> this function locally available so that other pkt-line functions can\n>> use it.\n> \n> This description is not that clear that set_packet_header() is a new\n> function.  Perhaps something like the following\n> \n>  Extract the part of format_packet() that converts an integer to a 4 byte\n>  hex string into set_packet_header().  Make this new function ...\n> \n> I also wonder if the part \"Make this [new] function locally available...\"\n> is needed; we need to justify exports, but I think we don't need to\n> justify limiting it to a module.  If you want to justify that it is\n> \"static\", perhaps it would be better to say why not to export it.\n> \n> Anyway, I think it is worthy refactoring (and compiler should be\n> able to inline it, so there are no nano-performance considerations).\n> \n> Good work!\n\nThank you! I would go with this then:\n\nExtract the part of format_packet() that converts an integer to a 4 byte\nhex string into set_packet_header().\n\nOK?\n\n\n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> pkt-line.c | 15 ++++++++++-----\n>> 1 file changed, 10 insertions(+), 5 deletions(-)\n>> \n>> diff --git a/pkt-line.c b/pkt-line.c\n>> index 62fdb37..445b8e1 100644\n>> --- a/pkt-line.c\n>> +++ b/pkt-line.c\n>> @@ -98,9 +98,17 @@ void packet_buf_flush(struct strbuf *buf)\n>> }\n>> \n>> #define hex(a) (hexchar[(a) & 15])\n> \n> I guess that this is inherited from the original, but this preprocessor\n> macro is local to the format_header() / set_packet_header() function,\n> and would not work outside it.  Therefore I think we should #undef it\n> after set_packet_header(), just in case somebody mistakes it for\n> a generic hex() function.  Perhaps even put it inside set_packet_header(),\n> together with #undef.\n> \n> But I might be mistaken... let's check... no, it isn't used outside it.\n\nAgreed. Would that be OK?\n\nstatic void set_packet_header(char *buf, const int size)\n{\n\tstatic char hexchar[] = \"0123456789abcdef\";\n\t#define hex(a) (hexchar[(a) & 15])\n\tbuf[0] = hex(size >> 12);\n\tbuf[1] = hex(size >> 8);\n\tbuf[2] = hex(size >> 4);\n\tbuf[3] = hex(size);\n\t#undef hex\n}\n\n- Lars"},{"id":"292738","messageId":"EBBE9E5E-1A39-4124-AB0D-D74EE01FA0DA@gmail.com","threadId":"42968","inReplyTo":"ef6c6152-a720-6bd5-22bb-6ebf375ca919@kdbg.org","subject":"Re: [PATCH v3 06/10] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T11:14:06Z","receivedAt":"2016-08-01T12:05:34Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 11:50, Johannes Sixt <j6t@kdbg.org> wrote:\n> \n> Am 30.07.2016 um 01:37 schrieb larsxschneider@gmail.com:\n>> Some commands might need to perform cleanup tasks on exit. Let's give\n>> them an interface for doing this.\n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> run-command.c | 12 ++++++++----\n>> run-command.h |  1 +\n>> 2 files changed, 9 insertions(+), 4 deletions(-)\n>> \n>> diff --git a/run-command.c b/run-command.c\n>> index 33bc63a..197b534 100644\n>> --- a/run-command.c\n>> +++ b/run-command.c\n>> @@ -21,6 +21,7 @@ void child_process_clear(struct child_process *child)\n>> \n>> struct child_to_clean {\n>> \tpid_t pid;\n>> +\tvoid (*clean_on_exit_handler)(pid_t);\n>> \tstruct child_to_clean *next;\n>> };\n>> static struct child_to_clean *children_to_clean;\n>> @@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n>> {\n>> \twhile (children_to_clean) {\n>> \t\tstruct child_to_clean *p = children_to_clean;\n>> +\t\tif (p->clean_on_exit_handler)\n>> +\t\t\tp->clean_on_exit_handler(p->pid);\n> \n> This summons demons. cleanup_children() is invoked from a signal handler. In this case, it can call only async-signal-safe functions. It does not look like the handler that you are going to install later will take note of this caveat!\n> \n>> \t\tchildren_to_clean = p->next;\n>> \t\tkill(p->pid, sig);\n>> \t\tif (!in_signal)\n> \n> The condition that we see here in the context protects free(p) (which is not async-signal-safe). Perhaps the invocation of the new callback should be skipped in the same manner when this is called from a signal handler? 507d7804 (pager: don't use unsafe functions in signal handlers) may be worth a look.\n\nThanks a lot of pointing this out to me!\n\nDo I get it right that after the signal \"SIGTERM\" I can do a cleanup and don't \nneed to worry about any function calls but if I get any other signal then I can \nonly perform async-signal-safe calls?\n\nIf this is correct, then the following solution would work great:\n\n\t\tif (!in_signal && p->clean_on_exit_handler)\n\t\t\tp->clean_on_exit_handler(p->pid);\n\nThanks,\nLars"},{"id":"292739","messageId":"64783AA5-D579-4783-88E7-E0B3BDE5FDEB@gmail.com","threadId":"42968","inReplyTo":"58e4737b-6e0e-565c-2468-05c705dea426@gmail.com","subject":"Re: [PATCH v3 02/10] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T12:00:53Z","receivedAt":"2016-08-01T12:14:13Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 12:49, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> Sometimes pkt-line data is already available in a buffer and it would\n>> be a waste of resources to write the packet using packet_write() which\n>> would copy the existing buffer into a strbuf before writing it.\n>> \n>> If the caller has control over the buffer creation then the\n>> PKTLINE_DATA_START macro can be used to skip the header and write\n>> directly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\n>> would be the maximum). direct_packet_write() would take this buffer,\n>> adjust the pkt-line header and write it.\n>> \n>> If the caller has no control over the buffer creation then\n>> direct_packet_write_data() can be used. This function creates a pkt-line\n>> header. Afterwards the header and the data buffer are written using two\n>> consecutive write calls.\n> \n> I don't quite understand what do you mean by \"caller has control\n> over the buffer creation\".  Do you mean that caller either can write\n> over the buffer, or cannot overwrite the buffer?  Or do you mean that\n> caller either can allocate buffer to hold header, or is getting\n> only the data?\n\nHow about this:\n\n[...]\n\nIf the caller creates the buffer then a proper pkt-line buffer with header\nand data section can be created. The PKTLINE_DATA_START macro can be used \nto skip the header section and write directly to the data section (PKTLINE_DATA_LEN \nbytes would be the maximum). direct_packet_write() would take this buffer, \nfill the pkt-line header section with the appropriate data length value and \nwrite the entire buffer.\n\nIf the caller does not create the buffer, and consequently cannot leave room\nfor the pkt-line header, then direct_packet_write_data() can be used. This \nfunction creates an extra buffer for the pkt-line header and afterwards writes\nthe header buffer and the data buffer with two consecutive write calls.\n\n---\nIs that more clear?\n\n> \n>> \n>> Both functions have a gentle parameter that indicates if Git should die\n>> in case of a write error (gentle set to 0) or return with a error (gentle\n>> set to 1).\n> \n> So they are *_maybe_gently(), isn't it ;-)?  Are there any existing\n> functions in Git codebase that take 'gently' / 'strict' / 'die_on_error'\n> parameter?\n\nYes, git grep \"gentle\" reveals:\n\nwrapper.c:static int memory_limit_check(size_t size, int gentle)\nobject.c:int type_from_string_gently(const char *str, ssize_t len, int gentle)\n\n\n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> pkt-line.c | 30 ++++++++++++++++++++++++++++++\n>> pkt-line.h |  5 +++++\n>> 2 files changed, 35 insertions(+)\n>> \n>> diff --git a/pkt-line.c b/pkt-line.c\n>> index 445b8e1..6fae508 100644\n>> --- a/pkt-line.c\n>> +++ b/pkt-line.c\n>> @@ -135,6 +135,36 @@ void packet_write(int fd, const char *fmt, ...)\n>> \twrite_or_die(fd, buf.buf, buf.len);\n>> }\n>> \n>> +int direct_packet_write(int fd, char *buf, size_t size, int gentle)\n>> +{\n>> +\tint ret = 0;\n>> +\tpacket_trace(buf + 4, size - 4, 1);\n>> +\tset_packet_header(buf, size);\n>> +\tif (gentle)\n>> +\t\tret = !write_or_whine_pipe(fd, buf, size, \"pkt-line\");\n>> +\telse\n>> +\t\twrite_or_die(fd, buf, size);\n> \n> Hmmm... in gently case we get the information in the warning that\n> it is about \"pkt-line\", which is missing from !gently case.  But\n> it is probably not important.\n> \n>> +\treturn ret;\n>> +}\n> \n> Nice clean function, thanks to extracting set_packet_header().\n> \n>> +\n>> +int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle)\n> \n> I would name the parameter 'data', rather than 'buf'; IMVHO it\n> better describes it.\n\nAgreed!\n\n> \n>> +{\n>> +\tint ret = 0;\n>> +\tchar hdr[4];\n>> +\tset_packet_header(hdr, sizeof(hdr) + size);\n>> +\tpacket_trace(buf, size, 1);\n>> +\tif (gentle) {\n>> +\t\tret = (\n>> +\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n> \n> You can write '4' here, no need for sizeof(hdr)... though compiler would\n> optimize it away.\n\nRight, it would be optimized. However, I don't like the 4 there either. OK to use a macro\ninstead? PKTLINE_HEADER_LEN ?\n\n\n>> +\t\t\t!write_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>> +\t\t);\n> \n> Do we want to try to write \"pkt-line data\" if \"pkt-line header\" failed?\n> If not, perhaps De Morgan-ize it\n> \n>  +\t\tret = !(\n>  +\t\t\twrite_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") &&\n>  +\t\t\twrite_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>  +\t\t);\n\n\nOriginal:\n\t\tret = (\n\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n\t\t\t!write_or_whine_pipe(fd, data, size, \"pkt-line data\")\n\t\t);\n\nWell, if the first write call fails (return == 0), then it is negated and evaluates to true.\nI would think the second call is not evaluated, then?!\n\nCPP reference:\n\"For the built-in logical OR operator, the result is true if either the first or the second \noperand (or both) is true. If the firstoperand is true, the second operand is not evaluated.\"\nhttp://en.cppreference.com/w/cpp/language/operator_logical\n\nShould I make this more explicit with a if clause?\n \n\n>> +\t} else {\n>> +\t\twrite_or_die(fd, hdr, sizeof(hdr));\n>> +\t\twrite_or_die(fd, buf, size);\n> \n> I guess these two writes (here and in 'gently' case) are unavoidable...\n\nI think so, too.\n\n\n> \n>> +\t}\n>> +\treturn ret;\n>> +}\n>> +\n>> void packet_buf_write(struct strbuf *buf, const char *fmt, ...)\n>> {\n>> \tva_list args;\n>> diff --git a/pkt-line.h b/pkt-line.h\n>> index 3cb9d91..02dcced 100644\n>> --- a/pkt-line.h\n>> +++ b/pkt-line.h\n>> @@ -23,6 +23,8 @@ void packet_flush(int fd);\n>> void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n>> void packet_buf_flush(struct strbuf *buf);\n>> void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n>> +int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n>> +int direct_packet_write_data(int fd, const char *buf, size_t size, int gentle);\n>> \n>> /*\n>>  * Read a packetized line into the buffer, which must be at least size bytes\n>> @@ -77,6 +79,9 @@ char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);\n>> \n>> #define DEFAULT_PACKET_MAX 1000\n>> #define LARGE_PACKET_MAX 65520\n>> +#define PKTLINE_HEADER_LEN 4\n>> +#define PKTLINE_DATA_START(pkt) ((pkt) + PKTLINE_HEADER_LEN)\n>> +#define PKTLINE_DATA_LEN (LARGE_PACKET_MAX - PKTLINE_HEADER_LEN)\n> \n> Those are not used in direct_packet_write() and direct_packet_write_data();\n> but they would make them more verbose and less readable.\n\nGood point, I should use them to check for the maximal packet length!\n\nThanks,\nLars\n\n> \n>> extern char packet_buffer[LARGE_PACKET_MAX];\n>> \n>> #endif\n>> \n> \n\n"},{"id":"292740","messageId":"52A5703D-AF9E-405D-A4F8-E30A7B923400@gmail.com","threadId":"42968","inReplyTo":"786f0b8e-29f0-3dd3-7bb4-5f6558f8ec84@gmail.com","subject":"Re: [PATCH v3 04/10] pkt-line: call packet_trace() only if a packet is actually send","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T12:18:36Z","receivedAt":"2016-08-01T12:18:51Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 14:29, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> The packet_trace() call is not ideal in format_packet() as we would print\n> \n> Style; I think the following is more readable:\n> \n>  The packet_trace() call in format_packet() is not ideal, as we would...\n\nAgreed!\n\n\n>> a trace when a packet is formatted and (potentially) when the packet is\n>> actually send. This was no problem up until now because format_packet()\n>> was only used by one function. Fix it by moving the trace call into the\n>> function that actally sends the packet.\n> \n> s/actally/actually/\n\nThanks!\n\n\n> I don't buy this explanation.  If you want to trace packets, you might\n> do it on input (when formatting packet), or on output (when writing\n> packet).  It's when there are more than one formatting function, but\n> one writing function, then placing trace call in write function means\n> less code duplication; and of course the reverse.\n> \n> Another issue is that something may happen between formatting packet\n> and sending it, and we probably want to packet_trace() when packet\n> is actually send.\n> \n> Neither of those is visible in commit message.\n\nThe packet_trace() call in format_packet() is not ideal, as we would print\na trace when a packet is formatted and (potentially) when the same packet is\nactually written. This was no problem up until now because packet_write(),\nthe function that uses format_packet() and writes the formatted packet,\ndid not trace the packet.\n\nThis developer believes that trace calls should only happen when a packet\nis actually written as the packet could be modified between formatting\nand writing. Therefore the trace call was moved from format_packet() to \npacket_write().\n\n--\n\nBetter?\n\n> \n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> pkt-line.c | 2 +-\n>> 1 file changed, 1 insertion(+), 1 deletion(-)\n>> \n>> diff --git a/pkt-line.c b/pkt-line.c\n>> index 1728690..32c0a34 100644\n>> --- a/pkt-line.c\n>> +++ b/pkt-line.c\n>> @@ -126,7 +126,6 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n>> \t\tdie(\"protocol error: impossibly long line\");\n>> \n>> \tset_packet_header(&out->buf[orig_len], n);\n>> -\tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n>> }\n>> \n>> void packet_write(int fd, const char *fmt, ...)\n>> @@ -138,6 +137,7 @@ void packet_write(int fd, const char *fmt, ...)\n>> \tva_start(args, fmt);\n>> \tformat_packet(&buf, fmt, args);\n>> \tva_end(args);\n>> +\tpacket_trace(buf.buf + 4, buf.len - 4, 1);\n>> \twrite_or_die(fd, buf.buf, buf.len);\n>> }\n>> \n>> \n> \n\n"},{"id":"292741","messageId":"B8F75E2C-2E85-4E91-B51F-90C3BF238072@gmail.com","threadId":"42968","inReplyTo":"cb6721b8-2a6a-deb1-2fc7-59399d118cec@gmail.com","subject":"Re: [PATCH v3 05/10] pack-protocol: fix maximum pkt-line size","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T12:23:19Z","receivedAt":"2016-08-01T12:23:37Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 15:58, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> According to LARGE_PACKET_MAX in pkt-line.h the maximal lenght of a\n>> pkt-line packet is 65520 bytes. The pkt-line header takes 4 bytes and\n>> therefore the pkt-line data component must not exceed 65516 bytes.\n> \n> s/lenght/length/\n\nThanks!\n\n\n> Is it maximum length of pkt-line packet, or maximum length of data\n> that can be send in a packet?\n\n65520 is the maximum length of a pkt-line.\n\n\n> With 4 hex digits, maximal length if pkt-line packet (together\n> with length) is ffff_16, that is 2^16-1 = 65535.  Where does the\n> number 65520 comes from?\n\nHistoric reasons, I guess? However, it won't be changed. See response\nfrom Peff here:\nhttp://public-inbox.org/git/20160726134257.GB19277%40sigill.intra.peff.net/\n\n\n> \n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n>> Documentation/technical/protocol-common.txt | 6 +++---\n>> 1 file changed, 3 insertions(+), 3 deletions(-)\n>> \n>> diff --git a/Documentation/technical/protocol-common.txt b/Documentation/technical/protocol-common.txt\n>> index bf30167..ecedb34 100644\n>> --- a/Documentation/technical/protocol-common.txt\n>> +++ b/Documentation/technical/protocol-common.txt\n>> @@ -67,9 +67,9 @@ with non-binary data the same whether or not they contain the trailing\n>> LF (stripping the LF if present, and not complaining when it is\n>> missing).\n>> \n>> -The maximum length of a pkt-line's data component is 65520 bytes.\n>> -Implementations MUST NOT send pkt-line whose length exceeds 65524\n>> -(65520 bytes of payload + 4 bytes of length data).\n>> +The maximum length of a pkt-line's data component is 65516 bytes.\n>> +Implementations MUST NOT send pkt-line whose length exceeds 65520\n>> +(65516 bytes of payload + 4 bytes of length data).\n>> \n>> Implementations SHOULD NOT send an empty pkt-line (\"0004\").\n>> \n>> \n> \n\n"},{"id":"292742","messageId":"4CDC1706-33FD-4525-8B32-800005F5855B@gmail.com","threadId":"42968","inReplyTo":"41184531-d3c2-43c0-d3b8-23cc913dbf86@gmail.com","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T12:28:35Z","receivedAt":"2016-08-01T12:30:01Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 30 Jul 2016, at 14:04, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> packet_flush() would die in case of a write error even though for some callers\n>> an error would be acceptable. Add packet_flush_gentle() which writes a pkt-line\n>> flush packet and returns `0` for success and `1` for failure.\n> \n> I think it should be packet_flush_gently(), as in \"to flush gently\",\n> but this is only my opinion; I have not checked the naming rules and\n> practices for the rest of Git codebase.\n\nAgreed. This would match:\n\nobject.c:int type_from_string_gently(const char *str, ssize_t len, int gentle)\n\nThanks,\nLars\n\n> \n>> \n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> ---\n> \n\n"},{"id":"292743","messageId":"ABE7D2DB-C45F-4F29-8CC2-8D873FD6C36A@gmail.com","threadId":"42968","inReplyTo":"b4c9ac5d-bd6b-141b-5b85-ab4aa719ccb0@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T13:32:13Z","receivedAt":"2016-08-01T13:32:40Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 31 Jul 2016, at 00:05, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> Git's clean/smudge mechanism invokes an external filter process for every\n>> single blob that is affected by a filter. If Git filters a lot of blobs\n>> then the startup time of the external filter processes can become a\n>> significant part of the overall Git execution time.\n>> \n>> This patch adds the filter.<driver>.process string option which, if used,\n>> keeps the external filter process running and processes all blobs with\n>> the following packet format (pkt-line) based protocol over standard input\n>> and standard output.\n> \n> I think it would be nice to have here at least summary of the benchmarks\n> you did in https://github.com/github/git-lfs/pull/1382\n\nOK.\n\n\n>> Git starts the filter on first usage and expects a welcome\n>> message, protocol version number, and filter capabilities\n>> separated by spaces:\n>> ------------------------\n>> packet:          git< git-filter-protocol\\n\n>> packet:          git< version 2\\n\n>> packet:          git< capabilities clean smudge\\n\n> \n> Sorry for going back and forth, but now I think that 'capabilities' are\n> not really needed here, though they are in line with \"version\" in\n> the second packet / line, namely \"version 2\".  If it does not make\n> parsing more difficult...\n\nI don't understand what you mean with \"they are not really needed\"?\nThe field is necessary to understand the protocol, no?\n\nIn the last roll I added the \"key=value\" format to the protocol upon\nyours and Peff's suggestion. Would it be OK to change the startup\nsequence accordingly?\n\npacket:          git< version=2\\n\npacket:          git< capabilities=clean smudge\\n\n\n\n>> ------------------------\n>> Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n>> and \"shutdown\".\n> \n> I'd rather put \"stream\" and \"shutdown\" capabilities into separate\n> patches, for easier review.\n\nI agree with \"shutdown\". I think I would like to remove the \"stream\"\noption and make it the default for the following reasons:\n\n(1) As you mentioned elsewhere, \"stream\" is not really streaming at this\npoint because we don't read/write in parallel.\n\n(2) Junio and you pointed out that if we transmit size and flush packet\nthen we have redundancy in the protocol.\n\n(3) With the newly introduced \"success\"/\"reject\"/\"failure\" packet at the \nend of a filter operation, a filter process has a way to signal Git that\nsomething went wrong. Initially I had the idea that a filter process just\nstops writing and Git would detect the mismatch between expected bytes\nand received bytes. But the final status packet is a much clearer solution.\n\n(4) Maintaining two slightly different protocols is a waste of resources \nand only increases the size of this (already large) patch.\n\nMy only argument for the size packet was that this allows efficient buffer\nallocation. However, in non of my benchmarks this was actually a problem.\nTherefore this is probably a epsilon optimization and should be removed.\n\nOK with everyone?\n\n\n>> Afterwards Git sends a command (based on the supported\n>> capabilities), the filename including its path\n>> relative to the repository root, the content size as ASCII number\n>> in bytes, the content split in zero or many pkt-line packets,\n>> and a flush packet at the end:\n> \n> I guess the following is the most basic example, with mode detailed\n> description left for the documentation.\n> \n>> ------------------------\n>> packet:          git> smudge\\n\n>> packet:          git> filename=path/testfile.dat\\n\n>> packet:          git> size=7\\n\n> \n> So I see you went with \"<variable>=<value>\" idea, rather than \"<value>\"\n> (with <variable> defined by position in a sequence of 'header' packets),\n> or \"<variable> <value>...\" that introductory header uses.\n\nThe implementation still requires the exact sequence of the packets.\nHowever, we could make this more flexible in a later patch with the\n\"key=value\" formatting.\n\n> \n>> packet:          git> CONTENT\n>> packet:          git> 0000\n>> ------------------------\n>> \n>> The filter is expected to respond with the result content size as\n>> ASCII number in bytes. If the capability \"stream\" is defined then\n>> the filter must not send the content size. Afterwards the result\n>> content in send in zero or many pkt-line packets and a flush packet\n>> at the end. \n> \n> If it does not cost filter anything, it could send size upfront\n> (based on size of original, or based on external data), even if\n> it is prepared for streaming.\n> \n> In the opposite case, where filter cannot stream because it requires\n> whole contents upfront (e.g. to calculate hash of the contents, or\n> to do operation that needs whole file like sorting or reversing lines),\n> it should always be able to calculate the size... or not.  For\n> example 'sort | uniq' filter needs whole input upfront for sort,\n> but it does not know how many lines will be in output without doing\n> the 'uniq' part.\n> \n> So I think the ability of filter to provide size (or size hint) of\n> its output should be decoupled from streaming support.\n\nAS mentioned above, I would like to remove the size packet completely\nto simplify this patch. If there is really a need for such a packet\nthen we could add it later (given the flexible \"key=value\" format of the\nprotocol).\n\n\n>>            Finally a \"success\" packet is send to indicate that\n>> everything went well.\n> \n> That's a nice addition, and probably a necessary one, to the stream\n> protocol.  Git must know and consume it - we wouldn't be able to\n> retrofit it later.\n> \n>> ------------------------\n>> packet:          git< size=57\\n   (omitted with capability \"stream\")\n> \n> I was thinking about having possible responses to receiving file\n> contents (or starting receiving in the streaming case) to be:\n> \n>  packet:          git< ok size=7\\n    (or \"ok 7\\n\", if size is known)\n> \n> or\n> \n>  packet:          git< ok\\n           (if filter does not know size upfront)\n> \n> or\n> \n>  packet:          git< fail <msg>\\n   (or just \"fail\" + packet with msg)\n> \n> The last would be when filter knows upfront that it cannot perform\n> the operation.  Though sending an empty file with non-\"success\" final\n> would work as well.\n> \n> For example LFS filter (that is configured as not required) may refuse\n> to store files which are smaller than some pre-defined constant threshold.\n\nDiscussed in http://public-inbox.org/git/7255ef06-a9a0-91b7-b6da-a90322de926b%40gmail.com/\n\n> \n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n>> packet:          git< success\\n\n>> ------------------------\n>> \n>> In case the filter cannot process the content, it is expected\n>> to respond with the result content size 0 (only if \"stream\" is\n>> not defined) and a \"reject\" packet.\n>> ------------------------\n>> packet:          git< size=0\\n    (omitted with capability \"stream\")\n>> packet:          git< reject\\n\n>> ------------------------\n> \n> This is *wrong* idea!  Empty file, with size=0, can be a perfectly\n> legitimate response.\n\nDiscussed in http://public-inbox.org/git/7255ef06-a9a0-91b7-b6da-a90322de926b%40gmail.com/\n\n\n> For example rot13 filter should respond to an empty file on input\n> with an empty file on output.  LFS-like filters and encryption\n> mechanism should return empty file on fetch / decryption\n> if such empty file was stored / encrypted.\n> \n> A strange LFS could even use filenames (with files being empty\n> themselves) as a lookup key for artifactory.  For example a kind\n> of CDN for common libraries, with version embedded in filename,\n> like 'libs/jquery-1.9.0.min.js', etc.\n\nRight, that would be possible with the current implementation. I will\nadd an empty file to the test case to prove it.\n\n\n>> After the filter has processed a blob it is expected to wait for\n>> the next command. A demo implementation can be found in\n>> `t/t0021/rot13-filter.pl` located in the Git core repository.\n> \n> If filter does not support \"shutdown\" capability (or if said\n> capability is postponed for later patch), it should behave sanely\n> when Git command reaps it (SIGTERM + wait + SIGKILL?, SIGCHLD?).\n\nHow would you do this? Don't you think the current solution is\ngood enough for processes that don't need a proper shutdown?\n\n\n>> \n>> If the filter supports the \"shutdown\" capability then Git will\n>> send the \"shutdown\" command and wait until the filter answers\n>> with \"done\". This gives the filter the opportunity to perform\n>> cleanup tasks. Afterwards the filter is expected to exit.\n>> ------------------------\n>> packet:          git> shutdown\\n\n>> packet:          git< done\\n\n>> ------------------------\n> \n> I guess there is no timeout mechanism: if filter hangs on shutdown,\n> then git command would also hang waiting for signal to exit.\n\nCorrect. Even if we implement a timeout - what time would we wait?\nI think this is still the best option for now.\n\n\n>> If a filter.<driver>.clean or filter.<driver>.smudge command\n>> is configured then these commands always take precedence over\n>> a configured filter.<driver>.process command.\n> \n> Note: the value of `clean`, `smudge` and `process` is a command,\n> not just a string.\n\nOK\n\n> I wonder if it would be worth it to explain the reasoning behind\n> this solution and show alternate ones.\n> \n> * Using a separate variable to signal that filters are invoked\n>   per-command rather than per-file, and use pkt-line interface,\n>   like boolean-valued `useProtocol`, or `protocolVersion` set\n>   to '2' or 'v2', or `persistence` set to 'per-command', there\n>   is high risk of user's trying to use exiting one-shot per-file\n>   filters... and Git hanging.\n> \n> * Using new variables for each capability, e.g. `processSmudge`\n>   and `processClean` would lead to explosion of variable names;\n>   I think.\n> \n> * Current solution of using `process` in addition to `clean`\n>   and `smudge` clearly says that you need to use different\n>   command for per-file (`clean` and `smudge`), and per-command\n>   filter, while allowing to use them together.\n> \n>   The possible disadvantage is Git command starting `process`\n>   filter, only to see that it doesn't offer required capability,\n>   for example offering only \"clean\" but not \"smudge\".  There\n>   is simple workaround - set `smudge` variable (same as not\n>   present capability) to empty string.\n\nIf you think it is necessary to have this discussion in the\ncommit message, then I will add it.\n\n\n>> Please note that you cannot use an existing filter.<driver>.clean\n>> or filter.<driver>.smudge command as filter.<driver>.process\n>> command. As soon as Git would detect a file that needs to be\n>> processed by this filter, it would stop responding.\n> \n> I think this needs to be in the documentation (I have not checked\n> yet if it is), but is not needed in the already long commit message.\n\nOK\n\n\n>> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n>> Helped-by: Martin-Louis Bright <mlbright@gmail.com>\n>> ---\n>> Documentation/gitattributes.txt |  84 ++++++++-\n>> convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--\n>> t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++\n>> t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++\n>> 4 files changed, 1053 insertions(+), 13 deletions(-)\n>> create mode 100755 t/t0021/rot13-filter.pl\n>> \n>> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\n>> index 8882a3e..e3fbcc2 100644\n>> --- a/Documentation/gitattributes.txt\n>> +++ b/Documentation/gitattributes.txt\n>> @@ -300,7 +300,11 @@ checkout, when the `smudge` command is specified, the command is\n>> fed the blob object from its standard input, and its standard\n>> output is used to update the worktree file.  Similarly, the\n>> `clean` command is used to convert the contents of worktree file\n>> -upon checkin.\n>> +upon checkin. By default these commands process only a single\n>> +blob and terminate. If a long running filter process (see section\n>> +below) is used then Git can process all blobs with a single filter\n>> +invocation for the entire life of a single Git command (e.g.\n>> +`git add .`).\n> \n> Proposed improvement:\n> \n>                       If a long running `process` filter is used\n>   in place of `clean` and/or `smudge` filters, then Git can process\n>   all blobs with a single filter command invocation for the entire\n>   life of a single Git command, for example `git add --all`.  See\n>   section below for the description of the protocol used to\n>   communicate with a `process` filter.\n\nSounds good. I will use this!\n\n\n>> One use of the content filtering is to massage the content into a shape\n>> that is more convenient for the platform, filesystem, and the user to use.\n>> @@ -375,6 +379,84 @@ substitution.  For example:\n>> ------------------------\n>> \n>> \n>> +Long Running Filter Process\n>> +^^^^^^^^^^^^^^^^^^^^^^^^^^^\n>> +\n>> +If the filter command (string value) is defined via\n> \n> This is no mere string value, this is command invocation (with its\n> own rules, e.g. splitting parameters on whitespace, etc.).  Though\n> I'm not sure how to say it succintly.  Maybe skip \"(string value)\"?\n> But it is there for a reason...\n\nHow about: \"If the filter command as string value is defined via ...\"?\n\n\n>> +filter.<driver>.process then Git can process all blobs with a\n> \n> Shouldn't it be `filter.<driver>.process`?\n\nOK, I will change it.\n\n\n>> +single filter invocation for the entire life of a single Git\n>> +command. This is achieved by using the following packet\n>> +format (pkt-line, see protocol-common.txt) based protocol over\n> \n> Can we linkgit-it (to technical documentation)?\n\nI don't think that is possible because it was never done. See:\ngit grep \"linkgit:tech\"\n\n\n>> +standard input and standard output.\n>> +\n>> +Git starts the filter on first usage and expects a welcome\n> \n> Is \"usage\" here correct?  Perhaps it would be more readable\n> to say that Git starts filter when encountering first file\n> that needs cleaning or smudgeing.\n\nOK. How about this:\n\nGit starts the filter when it encounters the first file\nthat needs to be cleaned or smudged. After the filter started\nGit expects a welcome message, protocol version number, and \nfilter capabilities separated by spaces:\n\n\n>> +message, protocol version number, and filter capabilities\n>> +separated by spaces:\n>> +------------------------\n>> +packet:          git< git-filter-protocol\\n\n>> +packet:          git< version 2\\n\n>> +packet:          git< capabilities clean smudge\\n\n>> +------------------------\n>> +Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n>> +and \"shutdown\".\n> \n> Filter should include at least one of \"clean\" and \"smudge\"\n> capabilities (currently), otherwise it wouldn't do anything.\n\nWell, I think that should be clear to the reader, no?\n\n\n> I don't know if it is a good place to say that because of pkt-line\n> recommendations about text-content packets, each of those should\n> terminate in endline, with \"\\n\" included in pkt-line length.\n> \n>> +\n>> +Afterwards Git sends a command (based on the supported\n>> +capabilities),\n> \n> I think it should be something like the following:\n> \n>   If among filter `process` capabilities there is capability\n>   that corresponds to the operation performed by a Git command\n>   (that is, either \"clean\" or \"smudge\"), then Git would send,\n>   in separate packets, a command (based on supported capabilites),\n> \n> though it feels too \"chatty\" (and the sentence gets quite long).\n> \n>>               the filename including its path\n>> +relative to the repository root, \n> \n> Errr... \"the filename including its path\"? Wouldn't be it simpler\n> to just say:\n> \n>  the pathname of a file relative to the repository root,\n\nAgreed\n\n\n> Also, isn't it now \"filename=<pathname>\\n\"?\n\nYou mean I should change filename to pathname? Agreed!\n\n\n>>                                  the content size as ASCII number\n>> +in bytes, \n> \n> Could Git not give the size, for example if fstat() fails? Do\n> we reserve space for other information here?\n> \n> Also, isn't it now \"size=<bytes>\\n\"?\n\nSize will go away as discussed in the beginning of this email.\n\n> \n>>            the content split in zero or many pkt-line packets,\n> \n> s/zero or many/zero or more/\n\nAgreed!\n\n\n>> +and a flush packet at the end:\n> \n> I wonder if instead of long sentence, it would be more readable\n> to use enumeration (ordered list) or itemize (unordered list).\n\nMaybe, but I think that is good enough for now.\n\n\n>> +------------------------\n>> +packet:          git> smudge\\n\n>> +packet:          git> filename=path/testfile.dat\\n\n>> +packet:          git> size=7\\n\n>> +packet:          git> CONTENT\n>> +packet:          git> 0000\n>> +------------------------\n>> +\n>> +The filter is expected to respond with the result content size as\n>> +ASCII number in bytes. If the capability \"stream\" is defined then\n>> +the filter must not send the content size.\n> \n> As I wrote earlier, I think sending or not the size of the output\n> should be decoupled from the \"stream\" capability.\n> \n> Streaming is IMVHO rather a capability of starting to send parts\n> of response before the whole contents of input arrives.  I think\n> per-file filters support that and that's what start_async() there\n> is about.\n\nCorrect. However, size will go away as discussed in the beginning of \nthis email. \n\n\n>>                                            Afterwards the result\n>> +content in send in zero or many pkt-line packets and a flush packet\n>> +at the end. Finally a \"success\" packet is send to indicate that\n>> +everything went well.\n> \n> I guess it is \"success\" packet if everything went well, and place\n> for informing about errors in the future - filter is assumed to die\n> if there are errors in filtering, isn't it?\n\nCorrect.\n\n\n> That is, not \"send to indicate\", but \"send if\".\n\nAgreed!\n\n\n> \n>> +------------------------\n>> +packet:          git< size=57\\n   (omitted with capability \"stream\")\n>> +packet:          git< SMUDGED_CONTENT\n>> +packet:          git< 0000\n>> +packet:          git< success\\n\n>> +------------------------\n>> +\n>> +In case the filter cannot process the content, it is expected\n>> +to respond with the result content size 0 (only if \"stream\" is\n>> +not defined) and a \"reject\" packet.\n>> +------------------------\n>> +packet:          git< size=0\\n    (omitted with capability \"stream\")\n>> +packet:          git< reject\\n\n>> +------------------------\n> \n> I would assume that we have two error conditions.  \n> \n> First situation is when the filter knows upfront (after receiving name\n> and size of file, and after receiving contents for not-streaming filters)\n> that it cannot process the file (like e.g. LFS filter with artifactory\n> replica/shard being a bit behind master, and not including contents of\n> the file being filtered).\n> \n> My proposal is to reply with \"fail\" _in place of_ size of reply:\n> \n>   packet:         git< fail\\n       (any case: size known or not, stream or not)\n> \n> It could be \"reject\", or \"error\" instead of \"fail\".\n> \n> \n> Another situation is if filter encounters error during output,\n> either with streaming filter (or non-stream, but not storing whole\n> input upfront) realizing in the middle of output that there is something\n> wrong with input (e.g. converting between encoding, and encountering\n> character that cannot be represented in output encoding), or e.g. filter\n> process being killed, or network connection dropping with LFS filter, etc.\n> The filter has send some packets with output already.  In this case\n> filter should flush, and send \"reject\" or \"error\" packet.\n> \n>   <error condition>\n>   packet:         git< \"0000\"       (flush packet)\n>   packet:         git< reject\\n\n> \n> Should there be a place for an error message, or would standard error\n> (stderr) be used for this?\n\nAlready discussed in http://public-inbox.org/git/6765D972-876A-4F94-A170-468002498296%40gmail.com/\n\nI will add an example for the error case, too.\n\n\n>> +\n>> +After the filter has processed a blob it is expected to wait for\n>> +the next command. A demo implementation can be found in\n>> +`t/t0021/rot13-filter.pl` located in the Git core repository.\n> \n> It is actually in Git sources.  Is it the best way to refer to\n> such files?\n\nWell, I could add a github.com link but I don't think everyone\nwould like that. What would you suggest?\n\n\n>> +\n>> +If the filter supports the \"shutdown\" capability then Git will\n>> +send the \"shutdown\" command and wait until the filter answers\n>> +with \"done\". This gives the filter the opportunity to perform\n>> +cleanup tasks. Afterwards the filter is expected to exit.\n>> +------------------------\n>> +packet:          git> shutdown\\n\n>> +packet:          git< done\\n\n>> +------------------------\n>> +\n>> +If a filter.<driver>.clean or filter.<driver>.smudge command\n>> +is configured then these commands always take precedence over\n>> +a configured filter.<driver>.process command.\n> \n> All right; this is quite clear.\n> \n>> +\n>> +Please note that you cannot use an existing filter.<driver>.clean\n>> +or filter.<driver>.smudge command as filter.<driver>.process\n>> +command. As soon as Git would detect a file that needs to be\n>> +processed by this filter, it would stop responding.\n> \n> This isn't.\n\nWould that be better?\n\n\nPlease note that you cannot use an existing `filter.<driver>.clean`\nor `filter.<driver>.smudge` command as `filter.<driver>.process`\ncommand because the former two use a different inter process \ncommunication protocol than the latter one. As soon as Git would detect \na file that needs to be processed by such an invalid \"process\" filter, \nit would wait for a proper protocol handshake and appear \"hanging\".\n\n\n> \n> P.S. I will comment about the implementation part in the next email.\n\nSure! Thanks again for the extensive review,\nLars\n\n"},{"id":"292769","messageId":"5180D54D-92C4-4875-AEB3-801663D70A8B@gmail.com","threadId":"42968","inReplyTo":"2f4743d1-3c93-406d-8b44-da0eb075e65c@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-01T17:55:46Z","receivedAt":"2016-08-01T18:14:47Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 01 Aug 2016, at 00:19, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n> [...]\n>> +Please note that you cannot use an existing filter.<driver>.clean\n>> +or filter.<driver>.smudge command as filter.<driver>.process\n>> +command.\n> \n> I think it would be more readable and easier to understand to write:\n> \n>  ... you cannot use an existing ... command with\n>  filter.<driver>.process\n> \n> About the style: wouldn't `filter.<driver>.process` be better?\n\nOK, changed it!\n\n\n>>             As soon as Git would detect a file that needs to be\n>> +processed by this filter, it would stop responding.\n> \n> This is quite convoluted, and hard to understand.  I would say\n> that because `clean` and `smudge` filters are expected to read\n> first, while Git expects `process` filter to say first, using\n> `clean` or `smudge` filter without changes as `process` filter\n> would lead to git command deadlocking / hanging / stopping\n> responding.\n\nHow about this:\n\nPlease note that you cannot use an existing `filter.<driver>.clean`\nor `filter.<driver>.smudge` command with `filter.<driver>.process`\nbecause the former two use a different inter process communication\nprotocol than the latter one. As soon as Git would detect a file\nthat needs to be processed by such an invalid \"process\" filter, \nit would wait for a proper protocol handshake and appear \"hanging\".\n\n\n>> +\n>> +\n>> Interaction between checkin/checkout attributes\n>> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n>> \n>> diff --git a/convert.c b/convert.c\n>> index 522e2c5..be6405c 100644\n>> --- a/convert.c\n>> +++ b/convert.c\n>> @@ -3,6 +3,7 @@\n>> #include \"run-command.h\"\n>> #include \"quote.h\"\n>> #include \"sigchain.h\"\n>> +#include \"pkt-line.h\"\n>> \n>> /*\n>>  * convert.c - convert a file when checking it out and checking it in.\n>> @@ -481,11 +482,355 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n>> \treturn ret;\n>> }\n>> \n>> +static int multi_packet_read(int fd_in, struct strbuf *sb, size_t expected_bytes, int is_stream)\n> \n> About name of this function: `multi_packet_read` is fine, though I wonder\n> if `packet_read_in_full` with nearly the same parameters as `packet_read`,\n> or `packet_read_till_flush`, or `read_in_full_packetized` would be better.\n\nI like `multi_packet_read` and will rename!\n\n\n> Also, the problem is that while we know that what packet_read() stores\n> would fit in memory (in size_t), it is not true for reading whole file,\n> which might be very large - for example huge graphical assets like raw\n> images or raw videos, or virtual machine images.  Isn't that the goal\n> of git-LFS solutions, which need this feature?  Shouldn't we have then\n> both `multi_packet_read_to_fd` and `multi_packet_read_to_buf`,\n> or whatever?\n\nGit LFS works well with the current clean/smudge mechanism that uses the\nsame on in memory buffers. I understand your concern but I think this\nimprovement is out of scope for this patch series.\n\n\n> Also, if we have `fd_in`, then perhaps `sb_out`?\n\nAgreed!\n\n\n> I am also unsure if `expected_bytes` (or `expected_size`) should not be\n> just a size hint, leaving handing mismatch between expected size and\n> real size of output to the caller; then the `is_stream` would be not\n> needed.\n\nAs mentioned in a previous email... I will drop the \"size\" support in\nthis patch series as it is not really needed.\n\n\n>> +{\n>> +\tint bytes_read;\n>> +\tsize_t total_bytes_read = 0;\n> \n> Why `bytes_read` is int, while `total_bytes_read` is size_t? Ah, I see\n> that packet_read() returns an int.  It should be ssize_t, just like\n> read(), isn't it?  But we know that packet size is limited, and would\n> fit in an int (or would it?).\n\nYes, it is limited but I agree on ssize_t!\n\n\n> Also, total_bytes_read could overflow size_t, but then we would have\n> problems storing the result in strbuf.\n\nWould that check be ok?\n\n\t\tif (total_bytes_read > SIZE_MAX - bytes_read)\n\t\t\treturn 1;  // `total_bytes_read` would overflow and is not representable\n\n\n>> +\tif (expected_bytes == 0 && !is_stream)\n>> +\t\treturn 0;\n> \n> So in all cases *except* size = 0 we expect flush packet after the\n> contents, but size = 0 is a corner case without flush packet?\n\nI agree that is inconsistent... I will change it!\n\n\n>> +\n>> +\tif (is_stream)\n>> +\t\tstrbuf_grow(sb, LARGE_PACKET_MAX);           // allocate space for at least one packet\n>> +\telse\n>> +\t\tstrbuf_grow(sb, st_add(expected_bytes, 1));  // add one extra byte for the packet flush\n>> +\n>> +\tdo {\n>> +\t\tbytes_read = packet_read(\n>> +\t\t\tfd_in, NULL, NULL,\n>> +\t\t\tsb->buf + total_bytes_read, sb->len - total_bytes_read - 1,\n>> +\t\t\tPACKET_READ_GENTLE_ON_EOF\n>> +\t\t);\n>> +\t\tif (bytes_read < 0)\n>> +\t\t\treturn 1;  // unexpected EOF\n> \n> Don't we usually return negative numbers on error?  Ah, I see that the\n> return is a bool, which allows to use boolean expression with 'return'.\n> But I am still unsure if it is good API, this return value.\n\nAccording to Peff zero for success is the usual style:\nhttp://public-inbox.org/git/20160728133523.GB21311%40sigill.intra.peff.net/\n\n\n> If we move handling of size mismatch to the caller, then the function\n> can simply return the size of data read (probably off_t or uint64_t).\n> Then the caller can check if it is what it expected, and react accordingly.\n\nTrue, but as discussed previously I will remove the size.\n\n\n>> +\n>> +\t\tif (is_stream &&\n>> +\t\t\tbytes_read > 0 &&\n>> +\t\t\tsb->len - total_bytes_read - 1 <= 0)\n>> +\t\t\tstrbuf_grow(sb, st_add(sb->len, LARGE_PACKET_MAX));\n>> +\t\ttotal_bytes_read += bytes_read;\n>> +\t}\n>> +\twhile (\n>> +\t\tbytes_read > 0 &&                   // the last packet was no flush\n>> +\t\tsb->len - total_bytes_read - 1 > 0  // we still have space left in the buffer\n> \n> Ah, so buffer is resized only in the 'is_stream' case.  Perhaps then\n> use an \"int options\" instead of 'is_stream', and have one of flags\n> tell if we should resize or not, that is if size parameter is hint\n> or a strict limit.\n\nObsolete\n\n\n>> +\t);\n>> +\tstrbuf_setlen(sb, total_bytes_read);\n>> +\treturn (is_stream ? 0 : expected_bytes != total_bytes_read);\n>> +}\n>> +\n>> +static int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n> \n> Is it equivalent of copy_fd() function, but where destination uses pkt-line\n> and we need to pack data into pkt-lines?\n\nCorrect!\n\n\n>> +{\n>> +\tint did_fail = 0;\n>> +\tssize_t bytes_to_write;\n>> +\twhile (!did_fail) {\n>> +\t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_LEN);\n> \n> Using global variable packet_buffer makes this code thread-unsafe, isn't it?\n> But perhaps that is not a problem, because other functions are also\n> using this global variable.\n\nCorrect!\n\n\n> It is more of PKTLINE_DATA_MAXLEN, isn't it?\n\nAgreed, will change!\n\n\n> \n>> +\t\tif (bytes_to_write < 0)\n>> +\t\t\treturn 1;\n>> +\t\tif (bytes_to_write == 0)\n>> +\t\t\tbreak;\n>> +\t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n>> +\t}\n>> +\tif (!did_fail)\n>> +\t\tdid_fail = packet_flush_gentle(fd_out);\n> \n> Shouldn't we try to flush even if there was an error?  Or is it\n> that if there is an error writing, then there is some problem\n> such that we know that flush would not work?\n\nRight, that's what I though.\n\n\n>> +\treturn did_fail;\n> \n> Return true on fail?  Shouldn't we follow example of copy_fd()\n> from copy.c, and return COPY_READ_ERROR, or COPY_WRITE_ERROR,\n> or PKTLINE_WRITE_ERROR?\n\nOK. How about this?\n\nstatic int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n{\n\tint did_fail = 0;\n\tssize_t bytes_to_write;\n\twhile (!did_fail) {\n\t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_MAXLEN);\n\t\tif (bytes_to_write < 0)\n\t\t\treturn COPY_READ_ERROR;\n\t\tif (bytes_to_write == 0)\n\t\t\tbreak;\n\t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n\t}\n\tif (!did_fail)\n\t\tdid_fail = packet_flush_gently(fd_out);\n\treturn (did_fail ? COPY_WRITE_ERROR : 0);\n}\n\n\n>> +}\n>> +\n>> +static int multi_packet_write_from_buf(const char *src, size_t len, int fd_out)\n> \n> It is equivalent of write_in_full(), with different order of parameters,\n> but where destination file descriptor expects pkt-line and we need to pack\n> data into pkt-lines?\n\nTrue. Do you suggest to reorder parameters? I also would like to rename `src` to `src_in`, OK?\n\n> \n> NOTE: function description comments?\n\nWhat do you mean here?\n\n\n>> +{\n>> +\tint did_fail = 0;\n>> +\tsize_t bytes_written = 0;\n>> +\tsize_t bytes_to_write;\n> \n> Note to self: bytes_to_write should fit in size_t, as it is limited to\n> PKTLINE_DATA_LEN.  bytes_written should fit in size_t, as it is at most\n> len, which is of type size_t.\n> \n>> +\twhile (!did_fail) {\n>> +\t\tif ((len - bytes_written) > PKTLINE_DATA_LEN)\n>> +\t\t\tbytes_to_write = PKTLINE_DATA_LEN;\n>> +\t\telse\n>> +\t\t\tbytes_to_write = len - bytes_written;\n>> +\t\tif (bytes_to_write == 0)\n>> +\t\t\tbreak;\n>> +\t\tdid_fail |= direct_packet_write_data(fd_out, src + bytes_written, bytes_to_write, 1);\n>> +\t\tbytes_written += bytes_to_write;\n> \n> Ah, I see now why we need both direct_packet_write() and\n> direct_packet_write_data().  Nice abstraction, makes for\n> clear code.\n> \n> The last parameter of '1' means 'gently', isn't it?\n\nCorrect. Thanks :)\n\n\n>> +\t}\n>> +\tif (!did_fail)\n>> +\t\tdid_fail = packet_flush_gentle(fd_out);\n>> +\treturn did_fail;\n>> +}\n> \n> I think all three/four of those functions should be added in a separate\n> commit, separate patch in patch series.\n\nOK\n\n>  Namely:\n> \n> - for git -> filter:\n>    * read from fd,      write pkt-line to fd  (off_t)\n>    * read from str+len, write pkt-line to fd  (size_t, ssize_t)\n> - for filter -> git:\n>    * read pkt-line from fd, write to fd       (off_t)\n\nThis one does not exist.\n\n\n>    * read pkt-line from fd, write to str+len  (size_t, ssize_t)\n> \n> Perhaps some of those can be in one overloaded function, perhaps it would\n> be easier to keep them separate.\n\nI would like to keep them separate as it is easier to comprehend.\n\n> \n> Also, I do wonder how the fetch / push code spools pack file received\n> over pkt-lines to disk.  Can we reuse that code?\n\nI haven't found any.\n\n\n>  Or maybe that code\n> could use those new functions?\n\nI think so, but this would be out of scope for this series :)\n\n\n>> +\n>> +#define FILTER_CAPABILITIES_STREAM   0x1\n>> +#define FILTER_CAPABILITIES_CLEAN    0x2\n>> +#define FILTER_CAPABILITIES_SMUDGE   0x4\n>> +#define FILTER_CAPABILITIES_SHUTDOWN 0x8\n>> +#define FILTER_SUPPORTS_STREAM(type) ((type) & FILTER_CAPABILITIES_STREAM)\n>> +#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n>> +#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n>> +#define FILTER_SUPPORTS_SHUTDOWN(type) ((type) & FILTER_CAPABILITIES_SHUTDOWN)\n>> +\n>> +struct cmd2process {\n>> +\tstruct hashmap_entry ent; /* must be the first member! */\n>> +\tconst char *cmd;\n>> +\tint supported_capabilities;\n> \n> I wonder if switching from int (perhaps with field width of 1 to denote\n> that it is boolean-like flag) to mask makes it more readable, or less.\n> But I think it is.\n> \n> \n> Reading Documentation/technical/api-hashmap.txt I found the following\n> recommendation:\n> \n>  `struct hashmap_entry`::\n> \n>        An opaque structure representing an entry in the hash table, which must\n>        be used as first member of user data structures. Ideally it should be\n>        followed by an int-sized member to prevent unused memory on 64-bit\n>        systems due to alignment.\n> \n> Therefore it \"int supported_capabilities\" should precede\n> \"const char *cmd\", I think.  Though it is not strictly necessary; it\n> is not as if this hash table were large (maximum size is limited by\n> the number of filter drivers configured), so we don't waste much space\n> due to internal padding / due to alignment.\n\nThanks! I will change it to your suggestion anyway!\n\n\n> \n>> +\tstruct child_process process;\n>> +};\n>> +\n>> +static int cmd_process_map_initialized = 0;\n>> +static struct hashmap cmd_process_map;\n> \n> Reading Documentation/technical/api-hashmap.txt I see that:\n> \n>  `tablesize` is the allocated size of the hash table. A non-0 value indicates\n>  that the hashmap is initialized.\n> \n> So cmd_process_map_initialized is not really needed, is it?\n\nI copied that from config.c:\nhttps://github.com/git/git/blob/f8f7adce9fc50a11a764d57815602dcb818d1816/config.c#L1425-L1428\n\n`git grep \"tablesize\"` reveals that the check for `tablesize` is only used\nin hashmap.c ... so what approach should we use?\n\n\n>> +\n>> +static int cmd2process_cmp(const struct cmd2process *e1,\n>> +\t\t\t\t\t\t\tconst struct cmd2process *e2,\n>> +\t\t\t\t\t\t\tconst void *unused)\n>> +{\n>> +\treturn strcmp(e1->cmd, e2->cmd);\n>> +}\n> \n> Well, to be exact (which is decidely not needed!) two commands might\n> be equivalent not being identical as strings (e.g. extra space between\n> parameters).  But it is something the user should care about, not Git.\n> \n>> +\n>> +static struct cmd2process *find_protocol2_filter_entry(struct hashmap *hashmap, const char *cmd)\n> \n> I'm not sure if *_protocol2_* is needed; those functions are static,\n> local to convert.c.\n\nI want to make sure that the reader understands that these functions are\nrelated to the filter protocol version 2. Not OK?\n\n\n>> +{\n>> +\tstruct cmd2process k;\n> \n> Does this name of variable 'k' follow established convention?\n> 'key' would be more descriptive, but it's not as if this function\n> was long; so 'k' is all right, I think.\n\nI agree on \"key\".\n\n\n> \n>> +\thashmap_entry_init(&k, strhash(cmd));\n>> +\tk.cmd = cmd;\n>> +\treturn hashmap_get(hashmap, &k, NULL);\n>> +}\n>> +\n>> +static void kill_protocol2_filter(struct hashmap *hashmap, struct cmd2process *entry) {\n> \n> Programming style: the opening brace should be on separate line,\n> that is:\n> \n>  +static void kill_protocol2_filter(struct hashmap *hashmap, struct cmd2process *entry)\n>  +{\n\nAgreed!\n\n\n>> +\tif (!entry)\n>> +\t\treturn;\n>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n>> +\tclose(entry->process.in);\n>> +\tclose(entry->process.out);\n>> +\tsigchain_pop(SIGPIPE);\n>> +\tfinish_command(&entry->process);\n>> +\tchild_process_clear(&entry->process);\n>> +\thashmap_remove(hashmap, entry, NULL);\n>> +\tfree(entry);\n>> +}\n> \n> All those, from #define FILTER_CAPABILITIES_ to here could be put\n> in a separate patch, to reduce size of this one.  But I am less\n> sure that it is worth it for this case.\n> \n>> +\n>> +void shutdown_protocol2_filter(pid_t pid)\n>> +{\n> [...]\n> \n> In my opinion this should be postponed to a separate commit.\n\nAgreed!\n\n> \n>> +}\n>> +\n>> +static struct cmd2process *start_protocol2_filter(struct hashmap *hashmap, const char *cmd)\n> \n> This has some parts in common with existing filter_buffer_or_fd().\n> I wonder if it would be worth to extract those common parts.\n> \n> But perhaps it would be better to leave such refactoring for later.\n> \n>> +{\n>> +\tint did_fail;\n>> +\tstruct cmd2process *entry;\n>> +\tstruct child_process *process;\n>> +\tconst char *argv[] = { cmd, NULL };\n>> +\tstruct string_list capabilities = STRING_LIST_INIT_NODUP;\n>> +\tchar *capabilities_buffer;\n>> +\tint i;\n>> +\n>> +\tentry = xmalloc(sizeof(*entry));\n>> +\thashmap_entry_init(entry, strhash(cmd));\n>> +\tentry->cmd = cmd;\n>> +\tentry->supported_capabilities = 0;\n>> +\tprocess = &entry->process;\n>> +\n>> +\tchild_process_init(process);\n> \n> filter_buffer_or_fd() uses instead\n> \n>  struct child_process child_process = CHILD_PROCESS_INIT;\n> \n> But I see that you need to access &entry->process anyway, so you\n> need to have it here, and in this case child_process_init() is\n> equivalent.\n> \n> I wonder if it would be worth it to use strbuf for cmd.\n\nWhat do you mean by \"worth it to use strbuf for cmd\"? Why would\nwe need a strbuf?\n\n\n>> +\tprocess->argv = argv;\n>> +\tprocess->use_shell = 1;\n>> +\tprocess->in = -1;\n>> +\tprocess->out = -1;\n>> +\tprocess->clean_on_exit = 1;\n>> +\tprocess->clean_on_exit_handler = shutdown_protocol2_filter;\n> \n> These two lines are new, and related to the \"shutdown\" capability, isn't it?\n\nYes.\n\n\n> \n>> +\n>> +\tif (start_command(process)) {\n>> +\t\terror(\"cannot fork to run external filter '%s'\", cmd);\n>> +\t\tkill_protocol2_filter(hashmap, entry);\n> \n> I guess the alternative solution of adding filter to the hashmap only\n> after starting the process would be racy?\n> \n> Ah, disregard that. I see that this pattern is a common way to error\n> out in this function (for process-related errors).\n> \n>> +\t\treturn NULL;\n>> +\t}\n>> +\n>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n>> +\tdid_fail = strcmp(packet_read_line(process->out, NULL), \"git-filter-protocol\");\n>> +\tif (!did_fail)\n>> +\t\tdid_fail |= strcmp(packet_read_line(process->out, NULL), \"version 2\");\n>> +\tif (!did_fail)\n>> +\t\tcapabilities_buffer = packet_read_line(process->out, NULL);\n>> +\telse\n>> +\t\tcapabilities_buffer = NULL;\n>> +\tsigchain_pop(SIGPIPE);\n>> +\n>> +\tif (!did_fail && capabilities_buffer) {\n>> +\t\tstring_list_split_in_place(&capabilities, capabilities_buffer, ' ', -1);\n>> +\t\tif (capabilities.nr > 1 &&\n>> +\t\t\t!strcmp(capabilities.items[0].string, \"capabilities\")) {\n>> +\t\t\tfor (i = 1; i < capabilities.nr; i++) {\n>> +\t\t\t\tconst char *requested = capabilities.items[i].string;\n>> +\t\t\t\tif (!strcmp(requested, \"stream\")) {\n>> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_STREAM;\n>> +\t\t\t\t} else if (!strcmp(requested, \"clean\")) {\n>> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_CLEAN;\n>> +\t\t\t\t} else if (!strcmp(requested, \"smudge\")) {\n>> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SMUDGE;\n>> +\t\t\t\t} else if (!strcmp(requested, \"shutdown\")) {\n>> +\t\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SHUTDOWN;\n>> +\t\t\t\t} else {\n>> +\t\t\t\t\twarning(\n>> +\t\t\t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n>> +\t\t\t\t\t\tcmd, requested\n>> +\t\t\t\t\t);\n>> +\t\t\t\t}\n>> +\t\t\t}\n>> +\t\t} else {\n>> +\t\t\terror(\"filter capabilities not found\");\n>> +\t\t\tdid_fail = 1;\n>> +\t\t}\n>> +\t\tstring_list_clear(&capabilities, 0);\n>> +\t}\n> \n> I wonder if the above conditional wouldn't be better to be put in\n> a separate function, parse_filter_capabilities(capabilities_buffer),\n> returning a mask, or having mask as an out parameter, and returning\n> an error condition.\n\nAgreed.\n\n\n>> +\n>> +\tif (did_fail) {\n>> +\t\terror(\"initialization for external filter '%s' failed\", cmd);\n> \n> More detailed information not needed, because one can use GIT_PACKET_TRACE.\n> Would it be worth add this information as a kind of advice, or put it\n> in the documentation of the `process` option?\n\nI will put it into the docs.\n\n\n> \n>> +\t\tkill_protocol2_filter(hashmap, entry);\n>> +\t\treturn NULL;\n>> +\t}\n>> +\n>> +\thashmap_add(hashmap, entry);\n>> +\treturn entry;\n>> +}\n>> +\n>> +static int apply_protocol2_filter(const char *path, const char *src, size_t len,\n>> +\t\t\t\t\t\tint fd, struct strbuf *dst, const char *cmd,\n>> +\t\t\t\t\t\tconst int wanted_capability)\n> \n> apply_protocol2_filter, or apply_process_filter?  Or rather,\n> s/_protocol2_/_process_/g ?\n\nMh. I wanted to convey that this functions is protocol V2 related...\n\n> \n> This is equivalent to\n> \n>   static int apply_filter(const char *path, const char *src, size_t len, int fd,\n>                           struct strbuf *dst, const char *cmd)\n> \n> Could we have extended that one instead?\n\nInitially I had one function but that got kind of long ... I prefer two for now.\n\n\n>> +{\n>> +\tint ret = 1;\n>> +\tstruct cmd2process *entry;\n>> +\tstruct child_process *process;\n>> +\tstruct stat file_stat;\n>> +\tstruct strbuf nbuf = STRBUF_INIT;\n>> +\tsize_t expected_bytes = 0;\n>> +\tchar *strtol_end;\n>> +\tchar *strbuf;\n>> +\tchar *filter_type;\n>> +\tchar *filter_result = NULL;\n>> +\n> \n>> +\tif (!cmd || !*cmd)\n>> +\t\treturn 0;\n>> +\n>> +\tif (!dst)\n>> +\t\treturn 1;\n> \n> This is the same as in apply_filter().\n> \n>> +\n>> +\tif (!cmd_process_map_initialized) {\n>> +\t\tcmd_process_map_initialized = 1;\n>> +\t\thashmap_init(&cmd_process_map, (hashmap_cmp_fn) cmd2process_cmp, 0);\n>> +\t\tentry = NULL;\n>> +\t} else {\n>> +\t\tentry = find_protocol2_filter_entry(&cmd_process_map, cmd);\n>> +\t}\n> \n> Here we try to find existing process, rather than starting new\n> as in apply_filter()\n> \n>> +\n>> +\tfflush(NULL);\n> \n> This is the same as in apply_filter(), but I wonder what it is for.\n\n\"If the stream argument is NULL, fflush() flushes all\n open output streams.\"\n\nhttp://man7.org/linux/man-pages/man3/fflush.3.html\n\n> \n>> +\n>> +\tif (!entry) {\n>> +\t\tentry = start_protocol2_filter(&cmd_process_map, cmd);\n>> +\t\tif (!entry) {\n>> +\t\t\treturn 0;\n>> +\t\t}\n> \n> Style; we prefer:\n> \n>  +\t\tif (!entry)\n>  +\t\t\treturn 0;\n\nAgreed.\n\n\n> This is very similar to apply_filter(), but the latter uses start_async()\n> from \"run-command.h\", with filter_buffer_or_fd() as asynchronous process,\n> which gets passed command to run in struct filter_params.  In this\n> function start_protocol2_filter() runs start_command(), synchronous API.\n> \n> Why the difference?\n\nThe protocol V2 requires a sequential processing of the packets. See\ndiscussion with Junio here:\nhttp://public-inbox.org/git/xmqqbn1th5qn.fsf%40gitster.mtv.corp.google.com/\n\n[LONG SNIP]\n\nI will answer the second half in a separate email.\n\nThanks for the review,\nLars\n\n"},{"id":"292818","messageId":"c2a8149b-8a61-81d8-f07b-31b80e565fda@kdbg.org","threadId":"42968","inReplyTo":"EBBE9E5E-1A39-4124-AB0D-D74EE01FA0DA@gmail.com","subject":"Re: [PATCH v3 06/10] run-command: add clean_on_exit_handler","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2016-08-02T05:53:31Z","receivedAt":"2016-08-02T06:17:58Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 01.08.2016 um 13:14 schrieb Lars Schneider:\n >> On 30 Jul 2016, at 11:50, Johannes Sixt <j6t@kdbg.org> wrote:\n >> Am 30.07.2016 um 01:37 schrieb larsxschneider@gmail.com:\n >>> static struct child_to_clean *children_to_clean;\n >>> @@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n >>> {\n >>> \twhile (children_to_clean) {\n >>> \t\tstruct child_to_clean *p = children_to_clean;\n >>> +\t\tif (p->clean_on_exit_handler)\n >>> +\t\t\tp->clean_on_exit_handler(p->pid);\n >>\n >> This summons demons. cleanup_children() is invoked from a signal\n >> handler. In this case, it can call only async-signal-safe functions.\n >> It does not look like the handler that you are going to install\n >> later will take note of this caveat!\n >>\n >>> \t\tchildren_to_clean = p->next;\n >>> \t\tkill(p->pid, sig);\n >>> \t\tif (!in_signal)\n >>\n >> The condition that we see here in the context protects free(p)\n >> (which is not async-signal-safe). Perhaps the invocation of the new\n >> callback should be skipped in the same manner when this is called\n >> from a signal handler? 507d7804 (pager: don't use unsafe functions\n >> in signal handlers) may be worth a look.\n >\n > Thanks a lot of pointing this out to me!\n >\n > Do I get it right that after the signal \"SIGTERM\" I can do a cleanup\n > and don't need to worry about any function calls but if I get any\n > other signal then I can only perform async-signal-safe calls?\n\nNo. SIGTERM is not special.\n\nPerhaps you were misled by the SIGTERM mentioned in \ncleanup_children_on_exit()? This function is invoked on regular exit \n(not from a signal). SIGTERM is used in this case to terminate children \nthat are still lingering around.\n\n > If this is correct, then the following solution would work great:\n >\n > \t\tif (!in_signal && p->clean_on_exit_handler)\n > \t\t\tp->clean_on_exit_handler(p->pid);\n\nThis should work nevertheless because in_signal is set when the function \nis invoked from a signal handler (of any signal that is caught) via \ncleanup_children_on_signal().\n\n-- Hannes\n\n"},{"id":"292819","messageId":"1C660676-D19E-4F73-BDD9-5F18CC0245EA@gmail.com","threadId":"42968","inReplyTo":"c2a8149b-8a61-81d8-f07b-31b80e565fda@kdbg.org","subject":"Re: [PATCH v3 06/10] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-02T07:41:11Z","receivedAt":"2016-08-02T07:59:09Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 02 Aug 2016, at 07:53, Johannes Sixt <j6t@kdbg.org> wrote:\n> \n> Am 01.08.2016 um 13:14 schrieb Lars Schneider:\n> >> On 30 Jul 2016, at 11:50, Johannes Sixt <j6t@kdbg.org> wrote:\n> >> Am 30.07.2016 um 01:37 schrieb larsxschneider@gmail.com:\n> >>> static struct child_to_clean *children_to_clean;\n> >>> @@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n> >>> {\n> >>> \twhile (children_to_clean) {\n> >>> \t\tstruct child_to_clean *p = children_to_clean;\n> >>> +\t\tif (p->clean_on_exit_handler)\n> >>> +\t\t\tp->clean_on_exit_handler(p->pid);\n> >>\n> >> This summons demons. cleanup_children() is invoked from a signal\n> >> handler. In this case, it can call only async-signal-safe functions.\n> >> It does not look like the handler that you are going to install\n> >> later will take note of this caveat!\n> >>\n> >>> \t\tchildren_to_clean = p->next;\n> >>> \t\tkill(p->pid, sig);\n> >>> \t\tif (!in_signal)\n> >>\n> >> The condition that we see here in the context protects free(p)\n> >> (which is not async-signal-safe). Perhaps the invocation of the new\n> >> callback should be skipped in the same manner when this is called\n> >> from a signal handler? 507d7804 (pager: don't use unsafe functions\n> >> in signal handlers) may be worth a look.\n> >\n> > Thanks a lot of pointing this out to me!\n> >\n> > Do I get it right that after the signal \"SIGTERM\" I can do a cleanup\n> > and don't need to worry about any function calls but if I get any\n> > other signal then I can only perform async-signal-safe calls?\n> \n> No. SIGTERM is not special.\n> \n> Perhaps you were misled by the SIGTERM mentioned in cleanup_children_on_exit()? This function is invoked on regular exit (not from a signal). SIGTERM is used in this case to terminate children that are still lingering around.\n\nYes, that was my source of confusion. Thanks for the clarification!\n\n> \n> > If this is correct, then the following solution would work great:\n> >\n> > \t\tif (!in_signal && p->clean_on_exit_handler)\n> > \t\t\tp->clean_on_exit_handler(p->pid);\n> \n> This should work nevertheless because in_signal is set when the function is invoked from a signal handler (of any signal that is caught) via cleanup_children_on_signal().\n\nRight. Thank you!\n\n- Lars"},{"id":"292862","messageId":"20160802195546.GA2660@atze2.lan","threadId":"42968","inReplyTo":"8FC2D283-AF8D-4643-834E-3D1927C558C0@gmail.com","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2016-08-02T19:56:38Z","receivedAt":"2016-08-02T19:58:01Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Sun, Jul 31, 2016 at 11:45:08PM +0200, Lars Schneider wrote:\n> \n> > On 31 Jul 2016, at 22:36, Torstem Bögershausen <tboegi@web.de> wrote:\n> > \n> > \n> > \n> >> Am 29.07.2016 um 20:37 schrieb larsxschneider@gmail.com:\n> >> \n> >> From: Lars Schneider <larsxschneider@gmail.com>\n> >> \n> >> packet_flush() would die in case of a write error even though for some callers\n> >> an error would be acceptable.\n> > What happens if there is a write error ?\n> > Basically the protocol is out of synch.\n> > Lenght information is mixed up with payload, or the other way\n> > around.\n> > It may be, that the consequences of a write error are acceptable,\n> > because a filter is allowed to fail.\n> > What is not acceptable is a \"broken\" protocol.\n> > The consequence schould be to close the fd and tear down all\n> > resources. connected to it.\n> > In our case to terminate the external filter daemon in some way,\n> > and to never use this instance again.\n> \n> Correct! That is exactly what is happening in kill_protocol2_filter()\n> here:\n\nWait a second.\nIs kill the same as shutdown ?\nI would expect that\nThe process terminates itself as soon as it detects EOF.\nAs there is nothing more read.\n\nThen the next question: The combination of kill & protocol in kill_protocol(),\nwhat does it mean ?\nIs it more like a graceful shutdown_protocol() ?\n\n"},{"id":"292895","messageId":"9DDA993E-2AFD-4C69-8E22-58601EEC8A40@gmail.com","threadId":"42968","inReplyTo":"2f4743d1-3c93-406d-8b44-da0eb075e65c@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T13:10:19Z","receivedAt":"2016-08-03T13:11:12Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 01 Aug 2016, at 00:19, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n> \n\n[LONG SNIP]\n\nFirst part answered here:\nhttp://public-inbox.org/git/5180D54D-92C4-4875-AEB3-801663D70A8B%40gmail.com/\n\n> \n>> +\t}\n>> +\tprocess = &entry->process;\n>> +\n>> +\tif (!(wanted_capability & entry->supported_capabilities))\n>> +\t\treturn 1;  // it is OK if the wanted capability is not supported\n>> +\n>> +\tif FILTER_SUPPORTS_CLEAN(wanted_capability)\n>> +\t\tfilter_type = \"clean\";\n>> +\telse if FILTER_SUPPORTS_SMUDGE(wanted_capability)\n>> +\t\tfilter_type = \"smudge\";\n>> +\telse\n>> +\t\tdie(\"unexpected filter type\");\n> \n> Style: it should be\n> \n>  +\tif (FILTER_SUPPORTS_CLEAN(wanted_capability))\n>  +\t\tfilter_type = \"clean\";\n>  +\telse if (FILTER_SUPPORTS_SMUDGE(wanted_capability))\n>  +\t\tfilter_type = \"smudge\";\n>  +\telse\n>  +\t\tdie(\"unexpected filter type\");\n> \n> even though by accident the macro provides the parentheses to \"if\".\n\nAgreed.\n\n\n> Can we make an error/die message more detailed?  Maybe it is\n> not possible...\n\nYeah, I don't see an easy way...\n\n> \n>> +\n>> +\tif (fd >= 0 && !src) {\n>> +\t\tif (fstat(fd, &file_stat) == -1)\n>> +\t\t\treturn 0;\n>> +\t\tlen = file_stat.st_size;\n>> +\t}\n> \n> All right, when fstat() can fail?  Could we then send contents without\n> size upfront, or is it better to require size to make it more consistent\n> for filter drivers scripts?\n\nIf fstat() fails then there is clearly something wrong and the filter\nshould fail.\n\n\n> Could this whole \"send single file\" be put in a separate function?\n> Or is it not worth it?\n\nThis function would have almost the same signature as apply_protocol2_filter\nand therefore I would say it's not worth it since the function is not\ncrazy long.\n\n\n> \n>> +\n>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n> \n> Hmmm... ignoring SIGPIPE was good for one-shot filters.  Is it still\n> O.K. for per-command persistent ones?\n\nVery good question. You are right... we don't want to ignore any errors\nduring the protocol... I will remove it.\n\n\n> \n>> +\n>> +\tpacket_buf_write(&nbuf, \"%s\\n\", filter_type);\n>> +\tret &= !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>> +\n>> +\tif (ret) {\n>> +\t\tstrbuf_reset(&nbuf);\n>> +\t\tpacket_buf_write(&nbuf, \"filename=%s\\n\", path);\n>> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>> +\t}\n> \n> Perhaps a better solution would be\n> \n>        if (err)\n>        \tgoto fin_error;\n> \n> rather than this.\n\nOK, I change it to goto error handling style.\n\n> \n>> +\n>> +\tif (ret) {\n>> +\t\tstrbuf_reset(&nbuf);\n>> +\t\tpacket_buf_write(&nbuf, \"size=%\"PRIuMAX\"\\n\", (uintmax_t)len);\n>> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>> +\t}\n> \n> Or maybe extract writing the header for a file into a separate function?\n> This one gets a bit long...\n\nMaybe... but I think that would make it harder to understand the protocol. I\nthink I would prefer to have all the communication in one function layer.\n\n\n>> +\n>> +\tif (ret) {\n>> +\t\tif (fd >= 0)\n>> +\t\t\tret = !multi_packet_write_from_fd(fd, process->in);\n>> +\t\telse\n>> +\t\t\tret = !multi_packet_write_from_buf(src, len, process->in);\n>> +\t}\n> \n> This is not streaming.  The above sends whole file, or whole string to\n> the filter process, without draining filter output.  If the filter were\n> to read some, then write some, it might deadlock on full buffers, isn't it?\n> Or am I mistaken?\n\nCorrect.\n\n\n>> +\n>> +\tif (ret && !FILTER_SUPPORTS_STREAM(entry->supported_capabilities)) {\n>> +\t\tstrbuf = packet_read_line(process->out, NULL);\n>> +\t\tif (strlen(strbuf) > 5 && !strncmp(\"size=\", strbuf, 5)) {\n>> +\t\t\texpected_bytes = (off_t)strtol(strbuf + 5, &strtol_end, 10);\n>> +\t\t\tret = (strtol_end != strbuf && errno != ERANGE);\n>> +\t\t} else {\n>> +\t\t\tret = 0;\n>> +\t\t}\n>> +\t}\n>> +\n>> +\tif (ret) {\n>> +\t\tstrbuf_reset(&nbuf);\n>> +\t\tret = !multi_packet_read(process->out, &nbuf, expected_bytes,\n>> +\t\t\tFILTER_SUPPORTS_STREAM(entry->supported_capabilities));\n>> +\t}\n> \n> What happens if the output of filter does not fit in size_t?  I see that\n> (I think) this problem is inherited from the original implementation.\n\nCorrect. And therefore I would prefer not to change this in this series.\n\n\n>> +\n>> +\tif (ret) {\n>> +\t\tfilter_result = packet_read_line(process->out, NULL);\n>> +\t\tret = !strcmp(filter_result, \"success\");\n>> +\t}\n>> +\n>> +\tsigchain_pop(SIGPIPE);\n>> +\n>> +\tif (ret) {\n>> +\t\tstrbuf_swap(dst, &nbuf);\n>> +\t} else {\n>> +\t\tif (!filter_result || strcmp(filter_result, \"reject\")) {\n>> +\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n>> +\t\t\terror(\"external filter '%s' failed\", cmd);\n>> +\t\t\tkill_protocol2_filter(&cmd_process_map, entry);\n>> +\t\t}\n>> +\t}\n> \n> So if Git gets finish signal \"success\" from filter, it accepts the output.\n> If Git gets finish signal \"reject\" from filter, it restarts filter (and\n> reject the output - user can retry the command himself / herself).\n> If Git gets any other finish signal, for example \"error\" (but this is not\n> standarized), then it rejects the output, keeping the unfiltered result,\n> but keeps filtering.\n> \n> I think it is not described in this detail in the documentation of the\n> new protocol.\n\nAgreed, will add!\n\n> \n>> +\tstrbuf_release(&nbuf);\n>> +\treturn ret;\n>> +}\n> \n> I wonder if this point might be start of the new patch... but then you\n> would have no way to test what you wrote.\n> \n>> +\n>> static struct convert_driver {\n>> \tconst char *name;\n>> \tstruct convert_driver *next;\n>> \tconst char *smudge;\n>> \tconst char *clean;\n>> +\tconst char *process;\n>> \tint required;\n>> } *user_convert, **user_convert_tail;\n> \n> All right.\n> \n>> \n>> @@ -526,6 +871,10 @@ static int read_convert_config(const char *var, const char *value, void *cb)\n>> \tif (!strcmp(\"clean\", key))\n>> \t\treturn git_config_string(&drv->clean, var, value);\n>> \n>> +\tif (!strcmp(\"process\", key)) {\n>> +\t\treturn git_config_string(&drv->process, var, value);\n>> +\t}\n>> +\n> \n> All right.\n> \n>> \tif (!strcmp(\"required\", key)) {\n>> \t\tdrv->required = git_config_bool(var, value);\n>> \t\treturn 0;\n>> @@ -823,7 +1172,12 @@ int would_convert_to_git_filter_fd(const char *path)\n>> \tif (!ca.drv->required)\n>> \t\treturn 0;\n>> \n>> -\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n>> +\tif (!ca.drv->clean && ca.drv->process)\n>> +\t\treturn apply_protocol2_filter(\n>> +\t\t\tpath, NULL, 0, -1, NULL, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n>> +\t\t);\n>> +\telse\n>> +\t\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n> \n> Could we augment apply_filter() instead, so that the invocation is\n> \n>        return apply_filter(path, NULL, 0, -1, NULL, ca.drv, FILTER_CLEAN);\n> \n> Though I am not sure if moving this conditional to apply_filter would\n> be a good idea; maybe wrapper around augmented apply_filter_do()?\n\nYes, a wrapper makes it way cleaner!\n\n\n>> }\n>> \n>> const char *get_convert_attr_ascii(const char *path)\n>> @@ -856,17 +1210,24 @@ int convert_to_git(const char *path, const char *src, size_t len,\n>>                    struct strbuf *dst, enum safe_crlf checksafe)\n>> {\n>> \tint ret = 0;\n>> -\tconst char *filter = NULL;\n>> +\tconst char *clean_filter = NULL;\n>> +\tconst char *process_filter = NULL;\n>> \tint required = 0;\n>> \tstruct conv_attrs ca;\n>> \n>> \tconvert_attrs(&ca, path);\n>> \tif (ca.drv) {\n>> -\t\tfilter = ca.drv->clean;\n>> +\t\tclean_filter = ca.drv->clean;\n>> +\t\tprocess_filter = ca.drv->process;\n>> \t\trequired = ca.drv->required;\n>> \t}\n> \n> All right (assuming un-augmented apply_filter()).\n> \n>> \n>> -\tret |= apply_filter(path, src, len, -1, dst, filter);\n>> +\tif (!clean_filter && process_filter)\n>> +\t\tret |= apply_protocol2_filter(\n>> +\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_CLEAN\n>> +\t\t);\n>> +\telse\n>> +\t\tret |= apply_filter(path, src, len, -1, dst, clean_filter);\n> \n> I wonder if it would be more readable to write it like this\n> (and of course elsewhere too):\n> \n>  +\tif (!clean_filter && process_filter)\n>  +\t\tret |= apply_protocol2_filter(\n>  +\t\t\tpath, src, len, -1, dst, process_filter, FILTER_CAPABILITIES_CLEAN\n>  +\t\t);\n>  +\telse\n>  +\t\tret |= apply_filter(\n>  +\t\t\tpath, src, len, -1, dst, clean_filter);\n>  +\t\t);\n> \n> \n> Though it would screw up \"git blame -C -C -w\"\n\nObsolete with the wrapper mentioned above.\n\n\n>> \tif (!ret && required)\n>> \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n>> \n>> @@ -885,13 +1246,21 @@ int convert_to_git(const char *path, const char *src, size_t len,\n>> void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n>> \t\t\t      enum safe_crlf checksafe)\n>> {\n>> +\tint ret = 0;\n> \n> Right, 'ret' is needed because we now have two possibilities:\n> `clean` filter and `process` filter.\n> \n>> \tstruct conv_attrs ca;\n>> \tconvert_attrs(&ca, path);\n>> \n>> \tassert(ca.drv);\n>> -\tassert(ca.drv->clean);\n>> +\tassert(ca.drv->clean || ca.drv->process);\n>> +\n>> +\tif (!ca.drv->clean && ca.drv->process)\n>> +\t\tret = apply_protocol2_filter(\n>> +\t\t\tpath, NULL, 0, fd, dst, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n>> +\t\t);\n>> +\telse\n>> +\t\tret = apply_filter(path, NULL, 0, fd, dst, ca.drv->clean);\n>> \n>> -\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv->clean))\n>> +\tif (!ret)\n>> \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n>> \n>> \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n>> @@ -902,14 +1271,16 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n>> \t\t\t\t\t    size_t len, struct strbuf *dst,\n>> \t\t\t\t\t    int normalizing)\n>> {\n>> -\tint ret = 0, ret_filter = 0;\n>> -\tconst char *filter = NULL;\n>> +\tint ret = 0, ret_filter;\n> \n> Why the change:\n> \n>  -\tint ret = 0, ret_filter = 0;\n>  +\tint ret = 0, ret_filter;\n\nReverted with the wrapper.\n\n\n>> +\tconst char *smudge_filter = NULL;\n>> +\tconst char *process_filter = NULL;\n>> \tint required = 0;\n>> \tstruct conv_attrs ca;\n>> \n>> \tconvert_attrs(&ca, path);\n>> \tif (ca.drv) {\n>> -\t\tfilter = ca.drv->smudge;\n>> +\t\tprocess_filter = ca.drv->process;\n>> +\t\tsmudge_filter = ca.drv->smudge;\n>> \t\trequired = ca.drv->required;\n>> \t}\n> \n> All right, the same.\n> \n> [...]\n>> diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\n>> index 34c8eb9..e8a7703 100755\n>> --- a/t/t0021-conversion.sh\n>> +++ b/t/t0021-conversion.sh\n>> @@ -296,4 +296,409 @@ test_expect_success 'disable filter with empty override' '\n>> \ttest_must_be_empty err\n>> '\n>> \n>> +test_expect_success PERL 'required process filter should filter data' '\n>> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge shutdown\" &&\n>> +\ttest_config_global filter.protocol.required true &&\n>> +\trm -rf repo &&\n>> +\tmkdir repo &&\n>> +\t(\n>> +\t\tcd repo &&\n>> +\t\tgit init &&\n>> +\n>> +\t\techo \"*.r filter=protocol\" >.gitattributes &&\n>> +\t\tgit add . &&\n>> +\t\tgit commit . -m \"test commit\" &&\n> \n> This is more of \"Initial commit\", not that it matters\n> \n>> +\t\tgit branch empty &&\n>> +\n>> +\t\tcat ../test.o >test.r &&\n> \n> Err, the above is just copying file, isn't it?\n> Maybe it was copied from other tests, I have not checked.\n\nIt was created in the \"setup\" test.\n\n\n>> +\t\techo \"test22\" >test2.r &&\n>> +\t\tmkdir testsubdir &&\n>> +\t\techo \"test333\" >testsubdir/test3.r &&\n> \n> All right, we test text file, we test binary file (I assume), we test\n> file in a subdirectory.  What about testing empty file?  Or large file\n> which would not fit in the stdin/stdout buffer (as EXPENSIVE test)?\n\nNo binary file. The main reason for this test is to check multiple files.\nI'll add a empty file. A large file is tested in the next test.\n\n> \n>> +\n>> +\t\trm -f rot13-filter.log &&\n>> +\t\tgit add . &&\n> \n> So this runs \"clean\" filter, storing cleaned contents in the index.\n\nCorrect.\n\n\n>> +\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >uniq-rot13-filter.log &&\n>> +\t\tcat >expected_add.log <<-\\EOF &&\n>> +\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n>> +\t\t\t1 IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n>> +\t\t\t1 IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n> \n> And we check the \"know size upfront\" case (mistakenly called non-\"stream\").\n\nCorrect - however, I removed non-stream\n\n\n>> +\t\t\t1 IN: shutdown -- [OK]\n> \n> And test \"shutdown\" capability (not as separate test).\n\nFixed.\n\n\n>> +\t\t\t1 start\n>> +\t\t\t1 wrote filter header\n>> +\t\tEOF\n> \n> And we are required to keep the expected_add.log file sorted by hand???\n\nWell, the clean invocations (and therefore their order of appearance)\nare not deterministic. See my discussion with Junio here:\nhttp://public-inbox.org/git/xmqqshv18i8i.fsf%40gitster.mtv.corp.google.com/\n\n> \n>> +\t\ttest_cmp expected_add.log uniq-rot13-filter.log &&\n>> +\n>> +\t\t>rot13-filter.log &&\n> \n> Truncate log. Still in the same test.\n> \n>> +\t\tgit commit . -m \"test commit\" &&\n> \n> This is test commit with files undergoing \"clean\" part of filter.\n> \n>> +\t\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n>> +\t\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >uniq-rot13-filter.log &&\n> \n> There is known performance regression, in that filter is run more\n> than once on given file.\n> \n> Actually... why it does not use cleaned-up contents from the index?\n\nSee discussion here: http://public-inbox.org/git/20160722152753.GA6859%40sigill.intra.peff.net/\n\n\n>> +\t\tcat >expected_commit.log <<-\\EOF &&\n>> +\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n>> +\t\t\tx IN: clean test2.r 7 [OK] -- OUT: 7 [OK]\n>> +\t\t\tx IN: clean testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n>> +\t\t\t1 IN: shutdown -- [OK]\n>> +\t\t\t1 start\n>> +\t\t\t1 wrote filter header\n> \n> Right, this is the goal of the patch series: for filter to be started\n> only once per git command invocation.\n> \n>> +\t\tEOF\n>> +\t\ttest_cmp expected_commit.log uniq-rot13-filter.log &&\n>> +\n> \n> Still in the same test, even though we would be testing \"smudge\"\n> capability now.  \n> \n> It's a pity that t/test-lib.sh does not support subtests from\n> the TAP specification (Test Anything Protocol that Git testsuite\n> uses).\n> \n>> +\t\t>rot13-filter.log &&\n>> +\t\trm -f test?.r testsubdir/test3.r &&\n>> +\t\tgit checkout . &&\n> \n> All right, we removed some files so that \"git checkout .\" could\n> restore them to life.\n> \n>> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n> \n> Useless use of cat\n> \n>  +\t\tgrep -v \"IN: clean\"  rot13-filter.log  >smudge-rot13-filter.log &&\n\nFixed, thanks!\n\n\n> Also: why 'git checkout <path>' would run \"clean\" filter?\n> Is it existing strange behaviour?\n\nAFAIK, that's existing behavior.\n\n\n>> +\t\tcat >expected_checkout.log <<-\\EOF &&\n>> +\t\t\tstart\n>> +\t\t\twrote filter header\n>> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n>> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n>> +\t\t\tIN: shutdown -- [OK]\n>> +\t\tEOF\n> \n> This time without 'sort | uniq -c'.\n\nYes, because the smudge calls are deterministic!\n\n\n>  Is it really needed for the\n> \"good\" case, or is it there for two cases to look similar?\n\nI am not sure what you mean?!\n\n\n>> +\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n>> +\n>> +\t\tgit checkout empty &&\n> \n> Shouldn't we check that switching to branch 'empty' does not run\n> filters, or is it covered by other tests?  Or perhaps this simply\n> does not matter here, is it?\n\nEasy enough to check. I will add this.\n\n> \n>> +\n>> +\t\t>rot13-filter.log &&\n>> +\t\tgit checkout master &&\n> \n> Does it test different callpath than 'git checkout .'?  Well, the\n> set of files is different...\n> \n>> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n>> +\t\tcat >expected_checkout_master.log <<-\\EOF &&\n>> +\t\t\tstart\n>> +\t\t\twrote filter header\n>> +\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n>> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n>> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n>> +\t\t\tIN: shutdown -- [OK]\n>> +\t\tEOF\n>> +\t\ttest_cmp expected_checkout_master.log smudge-rot13-filter.log &&\n>> +\n> \n> And here we start checking that the filter did filter,\n> that is the content in the repository is \"clean\"ed-up.\n> Still the same test.\n> \n>> +\t\t./../rot13.sh <test.r >expected &&\n>> +\t\tgit cat-file blob :test.r >actual &&\n>> +\t\ttest_cmp expected actual &&\n>> +\n>> +\t\t./../rot13.sh <test2.r >expected &&\n>> +\t\tgit cat-file blob :test2.r >actual &&\n>> +\t\ttest_cmp expected actual &&\n>> +\n>> +\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n>> +\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n>> +\t\ttest_cmp expected actual\n>> +\t)\n>> +'\n>> +\n>> +test_expect_success PERL 'required process filter should filter data stream' '\n>> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl stream clean smudge\" &&\n>> +\ttest_config_global filter.protocol.required true &&\n> \n> Errr... I don't see how it is different from the previous test.\n> [...]\n\nstream/non-stream ... but this is obsolete in the next roll. fixed!\n\n> \n>> +\n>> +test_expect_success PERL 'required process filter should filter smudge data and one-shot filter should clean' '\n> \n> All right, so this tests the precedence... well, it doesn't.\n> \n> It tests that `process` filter with \"smudge\" capability only works well\n> with one-shot `clean` filter.\n\nTrue. Isn't that what the test description indicates?\n\n\n>> +\ttest_config_global filter.protocol.clean ./../rot13.sh &&\n>> +\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl smudge\" &&\n> \n> Why the difference in pathnames (the directory part) between those two?\n\nrot13.sh is generated in the header of the file.\nrot13-filter.pl is part of the test suite\n\n\n>> +\ttest_config_global filter.protocol.required true &&\n>> +\trm -rf repo &&\n>> +\tmkdir repo &&\n>> +\t(\n>> +\t\tcd repo &&\n>> +\t\tgit init &&\n>> +\n>> +\t\techo \"*.r filter=protocol\" >.gitattributes &&\n>> +\t\tgit add . &&\n>> +\t\tgit commit . -m \"test commit\" &&\n>> +\t\tgit branch empty &&\n>> +\n>> +\t\tcat ../test.o >test.r &&\n>> +\t\techo \"test22\" >test2.r &&\n>> +\t\tmkdir testsubdir &&\n>> +\t\techo \"test333\" >testsubdir/test3.r &&\n>> +\n>> +\t\trm -f rot13-filter.log &&\n>> +\t\tgit add . &&\n>> +\t\ttest_must_be_empty rot13-filter.log &&\n>> +\n>> +\t\t>rot13-filter.log &&\n>> +\t\tgit commit . -m \"test commit\" &&\n>> +\t\ttest_must_be_empty rot13-filter.log &&\n> \n> All right, these tests that `process` filter is not ran.  But we don't\n> know if it is because it lacks capability, or because it is overriden\n> by one-shot filter (well, that comes later).\n\nOnly the clean one shot filter is configured. Therefore that shouldn't be a\nproblem, right?\n\n\n>> +\n>> +\t\t>rot13-filter.log &&\n>> +\t\trm -f test?.r testsubdir/test3.r &&\n>> +\t\tgit checkout . &&\n>> +\t\tcat rot13-filter.log | grep -v \"IN: clean\" >smudge-rot13-filter.log &&\n>> +\t\tcat >expected_checkout.log <<-\\EOF &&\n>> +\t\t\tstart\n>> +\t\t\twrote filter header\n>> +\t\t\tIN: smudge test2.r 7 [OK] -- OUT: 7 [OK]\n>> +\t\t\tIN: smudge testsubdir/test3.r 8 [OK] -- OUT: 8 [OK]\n>> +\t\tEOF\n>> +\t\ttest_cmp expected_checkout.log smudge-rot13-filter.log &&\n> \n> This part is repeated many, many times.  Maybe add some helper\n> shell function for this?\n\nGood idea! Will add!\n\n\n> [...]\n>> +\t\t./../rot13.sh <test.r >expected &&\n>> +\t\tgit cat-file blob :test.r >actual &&\n>> +\t\ttest_cmp expected actual &&\n>> +\n>> +\t\t./../rot13.sh <test2.r >expected &&\n>> +\t\tgit cat-file blob :test2.r >actual &&\n>> +\t\ttest_cmp expected actual &&\n>> +\n>> +\t\t./../rot13.sh <testsubdir/test3.r >expected &&\n>> +\t\tgit cat-file blob :testsubdir/test3.r >actual &&\n>> +\t\ttest_cmp expected actual\n> \n> Here we test that equivalent one-shot cleanup filter was run.\n> Here also we have repeated contents; maybe some helper function\n> would make it shorter?\n\nAgreed!\n\n\n>> +\t)\n>> +'\n> \n> Here I am stopping examining tests in detail.\n> \n>> +test_expect_success PERL 'required process filter should clean only' '\n>> +test_expect_success PERL 'required process filter should process files larger LARGE_PACKET_MAX' '\n> \n> Those two tests do not depend on being required or not; it is only\n> that without required they would fail softly in case of latter test\n> (which we can detect too).\n\nTrue, but since they fail hard it is easier to check.\n\n\n>> +test_expect_success PERL 'required process filter should with clean error should fail' '\n>> +test_expect_success PERL 'process filter should restart after unexpected write failure' '\n> \n> So these two are sort of complimentary.  When `process` is required,\n> then it should fail if it cannot filter some file.  If it is not,\n> it should keep processing other files.\n\nTrue.\n\n\n>> +test_expect_success PERL 'process filter should not restart after intentionally rejected file' '\n> \n> Uh... all right, so \"reject\" means that filter cannot continue?\n> Strange meaning for 'reject', though ;-)\n\nNo, with reject a filter can say \"I don't want to process that file\". This is a legitimate\nresponse and I don't Git to restart the filter in that case.\n\n\n>> test_done\n>> diff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\n>> new file mode 100755\n>> index 0000000..cb0925d\n>> --- /dev/null\n>> +++ b/t/t0021/rot13-filter.pl\n>> @@ -0,0 +1,177 @@\n>> +#!/usr/bin/perl\n>> +#\n>> +# Example implementation for the Git filter protocol version 2\n>> +# See Documentation/gitattributes.txt, section \"Filter Protocol\"\n>> +#\n>> +# The script takes the list of supported protocol capabilities as\n>> +# arguments (\"stream\", \"clean\", and \"smudge\" are supported).\n> \n> What about \"shutdown\"?\n\nWill fix.\n\n\n>> +#\n>> +# This implementation supports three special test cases:\n>> +# (1) If data with the filename \"clean-write-fail.r\" is processed with\n>> +#     a \"clean\" operation then the write operation will die.\n>> +# (2) If data with the filename \"smudge-write-fail.r\" is processed with\n>> +#     a \"smudge\" operation then the write operation will die.\n> \n> All right, so it is hard failure with filter script dying.\n\nCorrect.\n\n> \n>> +# (3) If data with the filename \"failure.r\" is processed with any\n>> +#     operation then the filter signals that the operation was not\n>> +#     successful.\n> \n> All right, so it is failure detected by filter script and signalled to Git.\n> \n>> +#\n>> +\n>> +use strict;\n>> +use warnings;\n> \n> So no more \"use autodie\", because of compatibility with old Perls.\n> \n>> +\n>> +my $MAX_PACKET_CONTENT_SIZE = 65516;\n>> +my @capabilities            = @ARGV;\n> \n> No autoflush this time?\n\nEric recommended to disable it:\nhttp://public-inbox.org/git/20160723072721.GA20875%40starla/\n\n\n>> +\n>> +sub rot13 {\n>> +    my ($str) = @_;\n>> +    $str =~ y/A-Za-z/N-ZA-Mn-za-m/;\n>> +    return $str;\n>> +}\n>> +\n>> +sub packet_read {\n>> +    my $buffer;\n>> +    my $bytes_read = read STDIN, $buffer, 4;\n>> +    if ( $bytes_read == 0 ) {\n>> +        return;\n>> +    }\n>> +    elsif ( $bytes_read != 4 ) {\n>> +        die \"invalid packet size '$bytes_read' field\";\n>> +    }\n>> +    my $pkt_size = hex($buffer);\n>> +    if ( $pkt_size == 0 ) {\n>> +        return ( 1, \"\" );\n> \n> Unusual return convention.  Though it is a test script, so\n> it doesn't matter much.\n> \n>> +    }\n>> +    elsif ( $pkt_size > 4 ) {\n>> +        my $content_size = $pkt_size - 4;\n>> +        $bytes_read = read STDIN, $buffer, $content_size;\n>> +        if ( $bytes_read != $content_size ) {\n>> +            die \"invalid packet\";\n> \n> More detailed error message, maybe?\n\nOK\n\n\n>> +        }\n>> +        return ( 0, $buffer );\n>> +    }\n>> +    else {\n>> +        die \"invalid packet size\";\n>> +    }\n>> +}\n>> +\n>> +sub packet_write {\n>> +    my ($packet) = @_;\n>> +    print STDOUT sprintf( \"%04x\", length($packet) + 4 );\n>> +    print STDOUT $packet;\n>> +    STDOUT->flush();\n>> +}\n>> +\n>> +sub packet_flush {\n>> +    print STDOUT sprintf( \"%04x\", 0 );\n>> +    STDOUT->flush();\n>> +}\n>> +\n>> +open my $debug, \">>\", \"rot13-filter.log\";\n>> +print $debug \"start\\n\";\n>> +$debug->flush();\n>> +\n>> +packet_write(\"git-filter-protocol\\n\");\n>> +packet_write(\"version 2\\n\");\n>> +packet_write( \"capabilities \" . join( ' ', @capabilities ) . \"\\n\" );\n>> +print $debug \"wrote filter header\\n\";\n>> +$debug->flush();\n>> +\n>> +while (1) {\n>> +    my $command = packet_read();\n>> +    unless ( defined($command) ) {\n>> +        exit();\n>> +    }\n>> +    chomp $command;\n>> +    print $debug \"IN: $command\";\n>> +    $debug->flush();\n>> +\n>> +    if ( $command eq \"shutdown\" ) {\n>> +        print $debug \" -- [OK]\";\n>> +        $debug->flush();\n>> +        packet_write(\"done\\n\");\n>> +        exit();\n>> +    }\n>> +\n>> +    my ($filename) = packet_read() =~ /filename=([^=]+)\\n/;\n>> +    print $debug \" $filename\";\n>> +    $debug->flush();\n>> +    my ($filelen) = packet_read() =~ /size=([^=]+)\\n/;\n>> +    chomp $filelen;\n> \n> I think this chomp is not needed, as \"\\n\" is not included.\n> Though the regexp should probably be anchored.\n\nAgreed.\n\n\n>> +    print $debug \" $filelen\";\n>> +    $debug->flush();\n>> +\n>> +    $filelen =~ /\\A\\d+\\z/ or die \"bad filelen: $filelen\";\n>> +    my $output;\n>> +\n>> +    if ( $filelen > 0 ) {\n> \n> So here is a special case for $filelen = 0.\n> Negative $filelen is not allowed, via regexp.\n\nObsolete in v4.\n\n\n>> +        my $input = \"\";\n>> +        {\n>> +            binmode(STDIN);\n>> +            my $buffer;\n>> +            my $done = 0;\n>> +            while ( !$done ) {\n>> +                ( $done, $buffer ) = packet_read();\n>> +                $input .= $buffer;\n>> +            }\n>> +            print $debug \" [OK] -- \";\n>> +            $debug->flush();\n>> +        }\n>> +\n>> +        if ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n>> +            $output = rot13($input);\n>> +        }\n>> +        elsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n>> +            $output = rot13($input);\n>> +        }\n> \n> These two conditionals could be shortened, but then they would be less\n> readable.  Or not:\n> \n>           if ( grep { $_ eq $command } @capabilities ) {\n>           \t$output = rot13($input);\n>           }\n\nI would like to keep it that way for readability since\nthe test script also serves as example implementation.\n\n\n>> +        else {\n>> +            die \"bad command $command\";\n>> +        }\n>> +    }\n>> +\n>> +    my $output_len = length($output);\n>> +    if ( $filename eq \"reject.r\" ) {\n>> +        $output_len = 0;\n>> +    }\n>> +\n>> +    if ( grep( /^stream$/, @capabilities ) ) {\n>> +        print $debug \"OUT: STREAM \";\n>> +    }\n>> +    else {\n>> +        packet_write(\"size=$output_len\\n\");\n>> +        print $debug \"OUT: $output_len \";\n>> +    }\n>> +    $debug->flush();\n>> +\n>> +    if ( $filename eq \"reject.r\" ) {\n>> +        packet_write(\"reject\\n\");\n>> +        print $debug \"[REJECT]\\n\";    # Could also be an error\n> \n> How if could be an error?\n\nRemoved.\n\n\n> \n>> +        $debug->flush();\n>> +    }\n>> +\n>> +    if ( $output_len > 0 ) {\n>> +        if (( $command eq \"clean\" and $filename eq \"clean-write-fail.r\" )\n>> +            or\n>> +            ( $command eq \"smudge\" and $filename eq \"smudge-write-fail.r\" ))\n> \n> Perhaps simply:\n> \n>  +        if ( $filename eq \"${command}-write-fail.r\" ) {\n\nNice! Will fix!\n\n\n>> +        {\n>> +            print $debug \"[WRITE FAIL]\\n\";\n>> +            $debug->flush();\n>> +            die \"write error\";\n>> +        }\n>> +        else {\n>> +            while ( length($output) > 0 ) {\n>> +                my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n>> +                packet_write($packet);\n>> +                if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n>> +                    $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n>> +                }\n>> +                else {\n>> +                    $output = \"\";\n>> +                }\n>> +            }\n>> +            packet_flush();\n>> +            packet_write(\"success\\n\");\n>> +            print $debug \"[OK]\\n\";\n>> +            $debug->flush();\n>> +        }\n>> +    }\n>> +}\n>> \n> \n\n\nThank you very much (again!) for your extensive review,\nLars\n\n\n"},{"id":"292907","messageId":"20160803164225.46355-2-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:14Z","receivedAt":"2016-08-03T17:04:52Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nset_packet_header() converts an integer to a 4 byte hex string. Make\nthis function locally available so that other pkt-line functions can\nuse it.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 18 ++++++++++++------\n 1 file changed, 12 insertions(+), 6 deletions(-)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 62fdb37..177dc73 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -97,10 +97,19 @@ void packet_buf_flush(struct strbuf *buf)\n \tstrbuf_add(buf, \"0000\", 4);\n }\n \n-#define hex(a) (hexchar[(a) & 15])\n-static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n+static void set_packet_header(char *buf, const int size)\n {\n \tstatic char hexchar[] = \"0123456789abcdef\";\n+\t#define hex(a) (hexchar[(a) & 15])\n+\tbuf[0] = hex(size >> 12);\n+\tbuf[1] = hex(size >> 8);\n+\tbuf[2] = hex(size >> 4);\n+\tbuf[3] = hex(size);\n+\t#undef hex\n+}\n+\n+static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n+{\n \tsize_t orig_len, n;\n \n \torig_len = out->len;\n@@ -111,10 +120,7 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n \tif (n > LARGE_PACKET_MAX)\n \t\tdie(\"protocol error: impossibly long line\");\n \n-\tout->buf[orig_len + 0] = hex(n >> 12);\n-\tout->buf[orig_len + 1] = hex(n >> 8);\n-\tout->buf[orig_len + 2] = hex(n >> 4);\n-\tout->buf[orig_len + 3] = hex(n);\n+\tset_packet_header(&out->buf[orig_len], n);\n \tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n }\n \n-- \n2.9.0\n\n"},{"id":"292909","messageId":"20160803164225.46355-8-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:20Z","receivedAt":"2016-08-03T17:04:54Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nSome commands might need to perform cleanup tasks on exit. Let's give\nthem an interface for doing this.\n\nPlease note, that the cleanup callback is not executed if Git dies of a\nsignal. The reason is that only \"async-signal-safe\" functions would be\nallowed to be call in that case. Since we cannot control what functions\nthe callback will use, we will not support the case. See 507d7804 for\nmore details.\n\nHelped-by: Johannes Sixt <j6t@kdbg.org>\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n run-command.c | 12 ++++++++----\n run-command.h |  1 +\n 2 files changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/run-command.c b/run-command.c\nindex 33bc63a..6ca75f3 100644\n--- a/run-command.c\n+++ b/run-command.c\n@@ -21,6 +21,7 @@ void child_process_clear(struct child_process *child)\n \n struct child_to_clean {\n \tpid_t pid;\n+\tvoid (*clean_on_exit_handler)(pid_t);\n \tstruct child_to_clean *next;\n };\n static struct child_to_clean *children_to_clean;\n@@ -30,6 +31,8 @@ static void cleanup_children(int sig, int in_signal)\n {\n \twhile (children_to_clean) {\n \t\tstruct child_to_clean *p = children_to_clean;\n+\t\tif (!in_signal && p->clean_on_exit_handler)\n+\t\t\tp->clean_on_exit_handler(p->pid);\n \t\tchildren_to_clean = p->next;\n \t\tkill(p->pid, sig);\n \t\tif (!in_signal)\n@@ -49,10 +52,11 @@ static void cleanup_children_on_exit(void)\n \tcleanup_children(SIGTERM, 0);\n }\n \n-static void mark_child_for_cleanup(pid_t pid)\n+static void mark_child_for_cleanup(pid_t pid, void (*clean_on_exit_handler)(pid_t))\n {\n \tstruct child_to_clean *p = xmalloc(sizeof(*p));\n \tp->pid = pid;\n+\tp->clean_on_exit_handler = clean_on_exit_handler;\n \tp->next = children_to_clean;\n \tchildren_to_clean = p;\n \n@@ -422,7 +426,7 @@ int start_command(struct child_process *cmd)\n \tif (cmd->pid < 0)\n \t\terror_errno(\"cannot fork() for %s\", cmd->argv[0]);\n \telse if (cmd->clean_on_exit)\n-\t\tmark_child_for_cleanup(cmd->pid);\n+\t\tmark_child_for_cleanup(cmd->pid, cmd->clean_on_exit_handler);\n \n \t/*\n \t * Wait for child's execvp. If the execvp succeeds (or if fork()\n@@ -483,7 +487,7 @@ int start_command(struct child_process *cmd)\n \tif (cmd->pid < 0 && (!cmd->silent_exec_failure || errno != ENOENT))\n \t\terror_errno(\"cannot spawn %s\", cmd->argv[0]);\n \tif (cmd->clean_on_exit && cmd->pid >= 0)\n-\t\tmark_child_for_cleanup(cmd->pid);\n+\t\tmark_child_for_cleanup(cmd->pid, cmd->clean_on_exit_handler);\n \n \targv_array_clear(&nargv);\n \tcmd->argv = sargv;\n@@ -752,7 +756,7 @@ int start_async(struct async *async)\n \t\texit(!!async->proc(proc_in, proc_out, async->data));\n \t}\n \n-\tmark_child_for_cleanup(async->pid);\n+\tmark_child_for_cleanup(async->pid, NULL);\n \n \tif (need_in)\n \t\tclose(fdin[0]);\ndiff --git a/run-command.h b/run-command.h\nindex 5066649..59d21ea 100644\n--- a/run-command.h\n+++ b/run-command.h\n@@ -43,6 +43,7 @@ struct child_process {\n \tunsigned stdout_to_stderr:1;\n \tunsigned use_shell:1;\n \tunsigned clean_on_exit:1;\n+\tvoid (*clean_on_exit_handler)(pid_t);\n };\n \n #define CHILD_PROCESS_INIT { NULL, ARGV_ARRAY_INIT, ARGV_ARRAY_INIT }\n-- \n2.9.0\n\n"},{"id":"292915","messageId":"20160803164225.46355-4-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 03/12] pkt-line: add packet_flush_gentle()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:16Z","receivedAt":"2016-08-03T17:05:00Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\npacket_flush() would die in case of a write error even though for some callers\nan error would be acceptable. Add packet_flush_gentle() which writes a pkt-line\nflush packet and returns `0` for success and `1` for failure.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 6 ++++++\n pkt-line.h | 1 +\n 2 files changed, 7 insertions(+)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex aa158ba..c8a052a 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -91,6 +91,12 @@ void packet_flush(int fd)\n \twrite_or_die(fd, \"0000\", 4);\n }\n \n+int packet_flush_gently(int fd)\n+{\n+\tpacket_trace(\"0000\", 4, 1);\n+\treturn !write_or_whine_pipe(fd, \"0000\", 4, \"flush packet\");\n+}\n+\n void packet_buf_flush(struct strbuf *buf)\n {\n \tpacket_trace(\"0000\", 4, 1);\ndiff --git a/pkt-line.h b/pkt-line.h\nindex ed64511..2fbaee9 100644\n--- a/pkt-line.h\n+++ b/pkt-line.h\n@@ -23,6 +23,7 @@ void packet_flush(int fd);\n void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n void packet_buf_flush(struct strbuf *buf);\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n+int packet_flush_gently(int fd);\n int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n int direct_packet_write_data(int fd, const char *data, size_t size, int gentle);\n \n-- \n2.9.0\n\n"},{"id":"292916","messageId":"20160803164225.46355-3-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 02/12] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:15Z","receivedAt":"2016-08-03T17:05:02Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nSometimes pkt-line data is already available in a buffer and it would\nbe a waste of resources to write the packet using packet_write() which\nwould copy the existing buffer into a strbuf before writing it.\n\nIf the caller has control over the buffer creation then the\nPKTLINE_DATA_START macro can be used to skip the header and write\ndirectly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\nwould be the maximum). direct_packet_write() would take this buffer,\nadjust the pkt-line header and write it.\n\nIf the caller has no control over the buffer creation then\ndirect_packet_write_data() can be used. This function creates a pkt-line\nheader. Afterwards the header and the data buffer are written using two\nconsecutive write calls.\n\nBoth functions have a gentle parameter that indicates if Git should die\nin case of a write error (gentle set to 0) or return with a error (gentle\nset to 1).\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 42 ++++++++++++++++++++++++++++++++++++++++++\n pkt-line.h |  5 +++++\n 2 files changed, 47 insertions(+)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex 177dc73..aa158ba 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -136,6 +136,48 @@ void packet_write(int fd, const char *fmt, ...)\n \twrite_or_die(fd, buf.buf, buf.len);\n }\n \n+int direct_packet_write(int fd, char *buf, size_t size, int gentle)\n+{\n+\tint ret = 0;\n+\tif (size > LARGE_PACKET_MAX) {\n+\t\tif (gentle)\n+\t\t\treturn 0;\n+\t\telse\n+\t\t\tdie(\"protocol error: impossibly long line\");\n+\t}\n+\tpacket_trace(buf + 4, size - 4, 1);\n+\tset_packet_header(buf, size);\n+\tif (gentle)\n+\t\tret = !write_or_whine_pipe(fd, buf, size, \"pkt-line\");\n+\telse\n+\t\twrite_or_die(fd, buf, size);\n+\treturn ret;\n+}\n+\n+int direct_packet_write_data(int fd, const char *data, size_t size, int gentle)\n+{\n+\tint ret = 0;\n+\tchar hdr[PKTLINE_HEADER_LEN];\n+\tif (size > PKTLINE_DATA_MAXLEN) {\n+\t\tif (gentle)\n+\t\t\treturn 0;\n+\t\telse\n+\t\t\tdie(\"protocol error: impossibly long line\");\n+\t}\n+\tset_packet_header(hdr, PKTLINE_HEADER_LEN + size);\n+\tpacket_trace(data, size, 1);\n+\tif (gentle) {\n+\t\tret = (\n+\t\t\t!write_or_whine_pipe(fd, hdr, PKTLINE_HEADER_LEN, \"pkt-line header\") ||\n+\t\t\t!write_or_whine_pipe(fd, data, size, \"pkt-line data\")\n+\t\t);\n+\t} else {\n+\t\twrite_or_die(fd, hdr, PKTLINE_HEADER_LEN);\n+\t\twrite_or_die(fd, data, size);\n+\t}\n+\treturn ret;\n+}\n+\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...)\n {\n \tva_list args;\ndiff --git a/pkt-line.h b/pkt-line.h\nindex 3cb9d91..ed64511 100644\n--- a/pkt-line.h\n+++ b/pkt-line.h\n@@ -23,6 +23,8 @@ void packet_flush(int fd);\n void packet_write(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n void packet_buf_flush(struct strbuf *buf);\n void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));\n+int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n+int direct_packet_write_data(int fd, const char *data, size_t size, int gentle);\n \n /*\n  * Read a packetized line into the buffer, which must be at least size bytes\n@@ -77,6 +79,9 @@ char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);\n \n #define DEFAULT_PACKET_MAX 1000\n #define LARGE_PACKET_MAX 65520\n+#define PKTLINE_HEADER_LEN 4\n+#define PKTLINE_DATA_START(pkt) ((pkt) + PKTLINE_HEADER_LEN)\n+#define PKTLINE_DATA_MAXLEN (LARGE_PACKET_MAX - PKTLINE_HEADER_LEN)\n extern char packet_buffer[LARGE_PACKET_MAX];\n \n #endif\n-- \n2.9.0\n\n"},{"id":"292918","messageId":"20160803164225.46355-9-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 08/12] convert: quote filter names in error messages","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:21Z","receivedAt":"2016-08-03T17:05:07Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGit filter driver commands with spaces (e.g. `filter.sh foo`) are hard to\nread in error messages. Quote them to improve the readability.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n convert.c | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/convert.c b/convert.c\nindex b1614bf..522e2c5 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -397,7 +397,7 @@ static int filter_buffer_or_fd(int in, int out, void *data)\n \tchild_process.out = out;\n \n \tif (start_command(&child_process))\n-\t\treturn error(\"cannot fork to run external filter %s\", params->cmd);\n+\t\treturn error(\"cannot fork to run external filter '%s'\", params->cmd);\n \n \tsigchain_push(SIGPIPE, SIG_IGN);\n \n@@ -415,13 +415,13 @@ static int filter_buffer_or_fd(int in, int out, void *data)\n \tif (close(child_process.in))\n \t\twrite_err = 1;\n \tif (write_err)\n-\t\terror(\"cannot feed the input to external filter %s\", params->cmd);\n+\t\terror(\"cannot feed the input to external filter '%s'\", params->cmd);\n \n \tsigchain_pop(SIGPIPE);\n \n \tstatus = finish_command(&child_process);\n \tif (status)\n-\t\terror(\"external filter %s failed %d\", params->cmd, status);\n+\t\terror(\"external filter '%s' failed %d\", params->cmd, status);\n \n \tstrbuf_release(&cmd);\n \treturn (write_err || status);\n@@ -462,15 +462,15 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n \t\treturn 0;\t/* error was already reported */\n \n \tif (strbuf_read(&nbuf, async.out, len) < 0) {\n-\t\terror(\"read from external filter %s failed\", cmd);\n+\t\terror(\"read from external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \tif (close(async.out)) {\n-\t\terror(\"read from external filter %s failed\", cmd);\n+\t\terror(\"read from external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \tif (finish_async(&async)) {\n-\t\terror(\"external filter %s failed\", cmd);\n+\t\terror(\"external filter '%s' failed\", cmd);\n \t\tret = 0;\n \t}\n \n-- \n2.9.0\n\n"},{"id":"292922","messageId":"20160803164225.46355-6-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 05/12] pkt-line: add functions to read/write flush terminated packet streams","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:18Z","receivedAt":"2016-08-03T17:05:11Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\npacket_write_stream_with_flush_from_fd() and\npacket_write_stream_with_flush_from_buf() write a stream of packets. All\ncontent packets use the maximal packet size except for the last one.\nAfter the last content packet a `flush` control packet is written.\n\npacket_read_till_flush() reads arbitary sized packets until it detects\na `flush` packet.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 88 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n pkt-line.h |  7 +++++\n 2 files changed, 95 insertions(+)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex d1368e6..f115537 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -193,6 +193,44 @@ void packet_buf_write(struct strbuf *buf, const char *fmt, ...)\n \tva_end(args);\n }\n \n+int packet_write_stream_with_flush_from_fd(const int fd_in, const int fd_out)\n+{\n+\tint did_fail = 0;\n+\tssize_t bytes_to_write;\n+\twhile (!did_fail) {\n+\t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_MAXLEN);\n+\t\tif (bytes_to_write < 0)\n+\t\t\treturn COPY_READ_ERROR;\n+\t\tif (bytes_to_write == 0)\n+\t\t\tbreak;\n+\t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n+\t}\n+\tif (!did_fail)\n+\t\tdid_fail = packet_flush_gently(fd_out);\n+\treturn (did_fail ? COPY_WRITE_ERROR : 0);\n+}\n+\n+int packet_write_stream_with_flush_from_buf(const char *src_in, size_t len, int fd_out)\n+{\n+\tint did_fail = 0;\n+\tsize_t bytes_written = 0;\n+\tsize_t bytes_to_write;\n+\twhile (!did_fail) {\n+\t\tif ((len - bytes_written) > PKTLINE_DATA_MAXLEN)\n+\t\t\tbytes_to_write = PKTLINE_DATA_MAXLEN;\n+\t\telse\n+\t\t\tbytes_to_write = len - bytes_written;\n+\t\tif (bytes_to_write == 0)\n+\t\t\tbreak;\n+\t\tdid_fail |= direct_packet_write_data(fd_out, src_in + bytes_written, bytes_to_write, 1);\n+\t\tbytes_written += bytes_to_write;\n+\t}\n+\tif (!did_fail)\n+\t\tdid_fail = packet_flush_gently(fd_out);\n+\treturn did_fail;\n+}\n+\n+\n static int get_packet_data(int fd, char **src_buf, size_t *src_size,\n \t\t\t   void *dst, unsigned size, int options)\n {\n@@ -302,3 +340,53 @@ char *packet_read_line_buf(char **src, size_t *src_len, int *dst_len)\n {\n \treturn packet_read_line_generic(-1, src, src_len, dst_len);\n }\n+\n+ssize_t packet_read_till_flush(int fd_in, struct strbuf *sb_out)\n+{\n+\tint len, ret;\n+\tint options = PACKET_READ_GENTLE_ON_EOF;\n+\tchar linelen[4];\n+\n+\tsize_t oldlen = sb_out->len;\n+\tsize_t oldalloc = sb_out->alloc;\n+\n+\tfor (;;) {\n+\t\t// Read packet header\n+\t\tret = get_packet_data(fd_in, NULL, NULL, linelen, 4, options);\n+\t\tif (ret < 0)\n+\t\t\tgoto done;\n+\t\tlen = packet_length(linelen);\n+\t\tif (len < 0)\n+\t\t\tdie(\"protocol error: bad line length character: %.4s\", linelen);\n+\t\tif (!len) {\n+\t\t\t// Found a flush packet - Done!\n+\t\t\tpacket_trace(\"0000\", 4, 0);\n+\t\t\tbreak;\n+\t\t}\n+\t\tlen -= 4;\n+\n+\t\t// Read packet content\n+\t\tstrbuf_grow(sb_out, len);\n+\t\tret = get_packet_data(fd_in, NULL, NULL, sb_out->buf + sb_out->len, len, options);\n+\t\tif (ret < 0)\n+\t\t\tgoto done;\n+\n+\t\tif (ret != len) {\n+\t\t\terror(\"protocol error: incomplete read (expected %d, got %d)\", len, ret);\n+\t\t\tgoto done;\n+\t\t}\n+\n+\t\tpacket_trace(sb_out->buf + sb_out->len, len, 0);\n+\t\tsb_out->len += len;\n+\t}\n+\n+done:\n+\tif (ret < 0) {\n+\t\tif (oldalloc == 0)\n+\t\t\tstrbuf_release(sb_out);\n+\t\telse\n+\t\t\tstrbuf_setlen(sb_out, oldlen);\n+\t\treturn ret;  // unexpected EOF\n+\t}\n+\treturn sb_out->len - oldlen;\n+}\ndiff --git a/pkt-line.h b/pkt-line.h\nindex 2fbaee9..3c0821f 100644\n--- a/pkt-line.h\n+++ b/pkt-line.h\n@@ -26,6 +26,8 @@ void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((f\n int packet_flush_gently(int fd);\n int direct_packet_write(int fd, char *buf, size_t size, int gentle);\n int direct_packet_write_data(int fd, const char *data, size_t size, int gentle);\n+int packet_write_stream_with_flush_from_fd(const int fd_in, const int fd_out);\n+int packet_write_stream_with_flush_from_buf(const char *src_in, size_t len, int fd_out);\n \n /*\n  * Read a packetized line into the buffer, which must be at least size bytes\n@@ -78,6 +80,11 @@ char *packet_read_line(int fd, int *size);\n  */\n char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);\n \n+/*\n+ * Reads a stream of variable sized packets until a flush packet is detected.\n+ */\n+ssize_t packet_read_till_flush(int fd_in, struct strbuf *sb_out);\n+\n #define DEFAULT_PACKET_MAX 1000\n #define LARGE_PACKET_MAX 65520\n #define PKTLINE_HEADER_LEN 4\n-- \n2.9.0\n\n"},{"id":"292923","messageId":"20160803164225.46355-10-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 09/12] convert: modernize tests","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:22Z","receivedAt":"2016-08-03T17:05:12Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nUse `test_config` to set the config, check that files are empty with\n`test_must_be_empty`, compare files with `test_cmp`, and remove spaces\nafter \">\" and \"<\".\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh | 62 +++++++++++++++++++++++++--------------------------\n 1 file changed, 31 insertions(+), 31 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 7bac2bc..7b45136 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -13,8 +13,8 @@ EOF\n chmod +x rot13.sh\n \n test_expect_success setup '\n-\tgit config filter.rot13.smudge ./rot13.sh &&\n-\tgit config filter.rot13.clean ./rot13.sh &&\n+\ttest_config filter.rot13.smudge ./rot13.sh &&\n+\ttest_config filter.rot13.clean ./rot13.sh &&\n \n \t{\n \t    echo \"*.t filter=rot13\"\n@@ -38,8 +38,8 @@ script='s/^\\$Id: \\([0-9a-f]*\\) \\$/\\1/p'\n \n test_expect_success check '\n \n-\tcmp test.o test &&\n-\tcmp test.o test.t &&\n+\ttest_cmp test.o test &&\n+\ttest_cmp test.o test.t &&\n \n \t# ident should be stripped in the repository\n \tgit diff --raw --exit-code :test :test.i &&\n@@ -47,10 +47,10 @@ test_expect_success check '\n \tembedded=$(sed -ne \"$script\" test.i) &&\n \ttest \"z$id\" = \"z$embedded\" &&\n \n-\tgit cat-file blob :test.t > test.r &&\n+\tgit cat-file blob :test.t >test.r &&\n \n-\t./rot13.sh < test.o > test.t &&\n-\tcmp test.r test.t\n+\t./rot13.sh <test.o >test.t &&\n+\ttest_cmp test.r test.t\n '\n \n # If an expanded ident ever gets into the repository, we want to make sure that\n@@ -130,7 +130,7 @@ test_expect_success 'filter shell-escaped filenames' '\n \n \t# delete the files and check them out again, using a smudge filter\n \t# that will count the args and echo the command-line back to us\n-\tgit config filter.argc.smudge \"sh ./argc.sh %f\" &&\n+\ttest_config filter.argc.smudge \"sh ./argc.sh %f\" &&\n \trm \"$normal\" \"$special\" &&\n \tgit checkout -- \"$normal\" \"$special\" &&\n \n@@ -141,7 +141,7 @@ test_expect_success 'filter shell-escaped filenames' '\n \ttest_cmp expect \"$special\" &&\n \n \t# do the same thing, but with more args in the filter expression\n-\tgit config filter.argc.smudge \"sh ./argc.sh %f --my-extra-arg\" &&\n+\ttest_config filter.argc.smudge \"sh ./argc.sh %f --my-extra-arg\" &&\n \trm \"$normal\" \"$special\" &&\n \tgit checkout -- \"$normal\" \"$special\" &&\n \n@@ -154,9 +154,9 @@ test_expect_success 'filter shell-escaped filenames' '\n '\n \n test_expect_success 'required filter should filter data' '\n-\tgit config filter.required.smudge ./rot13.sh &&\n-\tgit config filter.required.clean ./rot13.sh &&\n-\tgit config filter.required.required true &&\n+\ttest_config filter.required.smudge ./rot13.sh &&\n+\ttest_config filter.required.clean ./rot13.sh &&\n+\ttest_config filter.required.required true &&\n \n \techo \"*.r filter=required\" >.gitattributes &&\n \n@@ -165,17 +165,17 @@ test_expect_success 'required filter should filter data' '\n \n \trm -f test.r &&\n \tgit checkout -- test.r &&\n-\tcmp test.o test.r &&\n+\ttest_cmp test.o test.r &&\n \n \t./rot13.sh <test.o >expected &&\n \tgit cat-file blob :test.r >actual &&\n-\tcmp expected actual\n+\ttest_cmp expected actual\n '\n \n test_expect_success 'required filter smudge failure' '\n-\tgit config filter.failsmudge.smudge false &&\n-\tgit config filter.failsmudge.clean cat &&\n-\tgit config filter.failsmudge.required true &&\n+\ttest_config filter.failsmudge.smudge false &&\n+\ttest_config filter.failsmudge.clean cat &&\n+\ttest_config filter.failsmudge.required true &&\n \n \techo \"*.fs filter=failsmudge\" >.gitattributes &&\n \n@@ -186,9 +186,9 @@ test_expect_success 'required filter smudge failure' '\n '\n \n test_expect_success 'required filter clean failure' '\n-\tgit config filter.failclean.smudge cat &&\n-\tgit config filter.failclean.clean false &&\n-\tgit config filter.failclean.required true &&\n+\ttest_config filter.failclean.smudge cat &&\n+\ttest_config filter.failclean.clean false &&\n+\ttest_config filter.failclean.required true &&\n \n \techo \"*.fc filter=failclean\" >.gitattributes &&\n \n@@ -197,8 +197,8 @@ test_expect_success 'required filter clean failure' '\n '\n \n test_expect_success 'filtering large input to small output should use little memory' '\n-\tgit config filter.devnull.clean \"cat >/dev/null\" &&\n-\tgit config filter.devnull.required true &&\n+\ttest_config filter.devnull.clean \"cat >/dev/null\" &&\n+\ttest_config filter.devnull.required true &&\n \tfor i in $(test_seq 1 30); do printf \"%1048576d\" 1; done >30MB &&\n \techo \"30MB filter=devnull\" >.gitattributes &&\n \tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add 30MB\n@@ -207,7 +207,7 @@ test_expect_success 'filtering large input to small output should use little mem\n test_expect_success 'filter that does not read is fine' '\n \ttest-genrandom foo $((128 * 1024 + 1)) >big &&\n \techo \"big filter=epipe\" >.gitattributes &&\n-\tgit config filter.epipe.clean \"echo xyzzy\" &&\n+\ttest_config filter.epipe.clean \"echo xyzzy\" &&\n \tgit add big &&\n \tgit cat-file blob :big >actual &&\n \techo xyzzy >expect &&\n@@ -215,20 +215,20 @@ test_expect_success 'filter that does not read is fine' '\n '\n \n test_expect_success EXPENSIVE 'filter large file' '\n-\tgit config filter.largefile.smudge cat &&\n-\tgit config filter.largefile.clean cat &&\n+\ttest_config filter.largefile.smudge cat &&\n+\ttest_config filter.largefile.clean cat &&\n \tfor i in $(test_seq 1 2048); do printf \"%1048576d\" 1; done >2GB &&\n \techo \"2GB filter=largefile\" >.gitattributes &&\n \tgit add 2GB 2>err &&\n-\t! test -s err &&\n+\ttest_must_be_empty err &&\n \trm -f 2GB &&\n \tgit checkout -- 2GB 2>err &&\n-\t! test -s err\n+\ttest_must_be_empty err\n '\n \n test_expect_success \"filter: clean empty file\" '\n-\tgit config filter.in-repo-header.clean  \"echo cleaned && cat\" &&\n-\tgit config filter.in-repo-header.smudge \"sed 1d\" &&\n+\ttest_config filter.in-repo-header.clean  \"echo cleaned && cat\" &&\n+\ttest_config filter.in-repo-header.smudge \"sed 1d\" &&\n \n \techo \"empty-in-worktree    filter=in-repo-header\" >>.gitattributes &&\n \t>empty-in-worktree &&\n@@ -240,8 +240,8 @@ test_expect_success \"filter: clean empty file\" '\n '\n \n test_expect_success \"filter: smudge empty file\" '\n-\tgit config filter.empty-in-repo.clean \"cat >/dev/null\" &&\n-\tgit config filter.empty-in-repo.smudge \"echo smudged && cat\" &&\n+\ttest_config filter.empty-in-repo.clean \"cat >/dev/null\" &&\n+\ttest_config filter.empty-in-repo.smudge \"echo smudged && cat\" &&\n \n \techo \"empty-in-repo filter=empty-in-repo\" >>.gitattributes &&\n \techo dead data walking >empty-in-repo &&\n-- \n2.9.0\n\n"},{"id":"292926","messageId":"20160803164225.46355-1-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160729233801.82844-1-larsxschneider@gmail.com","subject":"[PATCH v4 00/12] Git filter protocol","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:13Z","receivedAt":"2016-08-03T17:05:16Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nHi,\n\nthanks a lot for the very helpful reviews!\n\nPatch 1-10 are preparation. Patch 11 and 12 the real feature.\n\nDiff to v3:\n* simplify protocol, remove size information\n* run clean_on_exit_handler() only on SIGTERM (Hannes)\n* move hex() macro inside set_packet_header(), undef it after use (Jakub)\n* rename buf to data in direct_packet_write_data() (Jakub)\n* add benchmark summary (Jakub)\n* add empty file test case (Jakub)\n* rename multi_packet_read() to packet_read_till_flush()\n* expect a flush packet even after 0 content\n* move packet stream helper functions to pkt-line.c/h (Jakub)\n* add GIT_PACKET_TRACE hint to docs\n* remove SIGPIPE ignore (Jakub)\n* change to goto error handling style (Jakub)\n* cleanup test cases with helper functions (Jakub)\n* move shutdown implementation to dedicated patch\n\nThanks,\nLars\n\n\nLars Schneider (12):\n  pkt-line: extract set_packet_header()\n  pkt-line: add direct_packet_write() and direct_packet_write_data()\n  pkt-line: add packet_flush_gentle()\n  pkt-line: call packet_trace() only if a packet is actually send\n  pkt-line: add functions to read/write flush terminated packet streams\n  pack-protocol: fix maximum pkt-line size\n  run-command: add clean_on_exit_handler\n  convert: quote filter names in error messages\n  convert: modernize tests\n  convert: generate large test files only once\n  convert: add filter.<driver>.process option\n  convert: add filter.<driver>.process shutdown command option\n\n Documentation/gitattributes.txt             | 108 +++++-\n Documentation/technical/protocol-common.txt |   6 +-\n convert.c                                   | 324 ++++++++++++++++--\n pkt-line.c                                  | 156 ++++++++-\n pkt-line.h                                  |  13 +\n run-command.c                               |  12 +-\n run-command.h                               |   1 +\n t/t0021-conversion.sh                       | 503 +++++++++++++++++++++++++---\n t/t0021/rot13-filter.pl                     | 155 +++++++++\n 9 files changed, 1187 insertions(+), 91 deletions(-)\n create mode 100755 t/t0021/rot13-filter.pl\n\n--\n2.9.0\n\n"},{"id":"292933","messageId":"20160803164225.46355-13-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 12/12] convert: add filter.<driver>.process shutdown command option","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:25Z","receivedAt":"2016-08-03T17:05:24Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nAdd the \"shutdown\" capability to the `filter.<driver>.process` filter\nprotocol. If a filter supports this capability then Git will send the\n\"shutdown\" command and wait until the filter answers. This gives the\nfilter the opportunity to perform cleanup tasks. Afterwards the filter\nis expected to exit.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n Documentation/gitattributes.txt | 12 ++++++-\n convert.c                       | 35 ++++++++++++++++++++\n t/t0021-conversion.sh           | 71 +++++++++++++++++++++++++++++++++++++++++\n t/t0021/rot13-filter.pl         |  7 ++++\n 4 files changed, 124 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 49514ab..5556cc0 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -400,7 +400,8 @@ packet:          git< git-filter-protocol\\n\n packet:          git< version=2\\n\n packet:          git< capabilities=clean smudge\\n\n ------------------------\n-Supported filter capabilities are \"clean\" and \"smudge\".\n+Supported filter capabilities are \"clean\", \"smudge\", and\n+\"shutdown\".\n \n Afterwards Git sends a command (based on the supported\n capabilities), the pathname of a file relative to the\n@@ -453,6 +454,15 @@ After the filter has processed a blob it is expected to wait for\n the next command. When the Git process terminates, it will send\n a kill signal to the filter in that stage.\n \n+If the filter supports the \"shutdown\" capability then Git will\n+send the \"shutdown\" command and wait until the filter answers\n+with \"done\". This gives the filter the opportunity to perform\n+cleanup tasks. Afterwards the filter is expected to exit.\n+------------------------\n+packet:          git> command=shutdown\\n\n+packet:          git< result=success\\n\n+------------------------\n+\n A long running filter demo implementation can be found in\n `t/t0021/rot13-filter.pl` located in the Git core repository.\n If you develop your own long running filter process then the\ndiff --git a/convert.c b/convert.c\nindex 130430a..41e3229 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -478,8 +478,10 @@ static int apply_single_file_filter(const char *path, const char *src, size_t le\n \n #define FILTER_CAPABILITIES_CLEAN    (1u<<0)\n #define FILTER_CAPABILITIES_SMUDGE   (1u<<1)\n+#define FILTER_CAPABILITIES_SHUTDOWN (1u<<2)\n #define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n #define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n+#define FILTER_SUPPORTS_SHUTDOWN(type) ((type) & FILTER_CAPABILITIES_SHUTDOWN)\n \n struct cmd2process {\n \tstruct hashmap_entry ent; /* must be the first member! */\n@@ -520,6 +522,35 @@ static void kill_multi_file_filter(struct hashmap *hashmap, struct cmd2process *\n \tfree(entry);\n }\n \n+void shutdown_multi_file_filter(pid_t pid)\n+{\n+\tint did_fail;\n+\tstruct cmd2process *entry;\n+\tstruct hashmap_iter iter;\n+\tstatic const char shutdown[] = \"command=shutdown\\n\";\n+\tchar *result = NULL;\n+\n+\tif (!cmd_process_map_initialized)\n+\t\treturn;\n+\n+\thashmap_iter_init(&cmd_process_map, &iter);\n+\twhile ((entry = hashmap_iter_next(&iter))) {\n+\t\tif (entry->process.pid == pid &&\n+\t\t\tFILTER_SUPPORTS_SHUTDOWN(entry->supported_capabilities)\n+\t\t) {\n+\t\t\tdid_fail = direct_packet_write_data(\n+\t\t\t\tentry->process.in, shutdown, strlen(shutdown), 1);\n+\t\t\tif (!did_fail)\n+\t\t\t\tresult = packet_read_line(entry->process.out, NULL);\n+\t\t\tclose(entry->process.in);\n+\t\t\tclose(entry->process.out);\n+\n+\t\t\tif (did_fail || !result || strcmp(result, \"result=success\"))\n+\t\t\t\terror(\"shutdown of external filter '%s' failed\", entry->cmd);\n+\t\t}\n+\t}\n+}\n+\n static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, const char *cmd)\n {\n \tint did_fail;\n@@ -543,6 +574,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \tprocess->use_shell = 1;\n \tprocess->in = -1;\n \tprocess->out = -1;\n+\tprocess->clean_on_exit = 1;\n+\tprocess->clean_on_exit_handler = shutdown_multi_file_filter;\n \n \tif (start_command(process)) {\n \t\terror(\"cannot fork to run external filter '%s'\", cmd);\n@@ -575,6 +608,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_CLEAN;\n \t\t\t} else if (!strcmp(requested, \"smudge\")) {\n \t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SMUDGE;\n+\t\t\t} else if (!strcmp(requested, \"shutdown\")) {\n+\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SHUTDOWN;\n \t\t\t} else {\n \t\t\t\twarning(\n \t\t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex c1a22f4..613e370 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -417,6 +417,77 @@ test_expect_success PERL 'required process filter should filter data' '\n \t)\n '\n \n+test_expect_success PERL 'required process filter should filter data with shutdown' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge shutdown\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\tcat ../test2.o >test2.r &&\n+\n+\t\tcheck_filter \\\n+\t\t\tgit add . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\t1 IN: clean test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\t1 IN: shutdown -- [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_count_clean \\\n+\t\t\tgit commit . -m \"test commit\" \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tx IN: clean test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\t1 IN: shutdown -- [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\n+\t\trm -f test?.r testsubdir/test3-subdir.r &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\tIN: shutdown -- [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout empty \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: shutdown -- [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout master \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\tIN: shutdown -- [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_rot13 ../test.o test.r &&\n+\t\tcheck_rot13 ../test2.o test2.r\n+\t)\n+'\n+\n test_expect_success PERL 'required process filter should filter smudge data and one-shot filter should clean' '\n \ttest_config_global filter.protocol.clean ./../rot13.sh &&\n \ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl smudge\" &&\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nindex ca6d5e4..654741b 100755\n--- a/t/t0021/rot13-filter.pl\n+++ b/t/t0021/rot13-filter.pl\n@@ -84,6 +84,13 @@ while (1) {\n     print $debug \"IN: $command\";\n     $debug->flush();\n \n+    if ( $command eq \"shutdown\" ) {\n+        print $debug \" -- [OK]\";\n+        $debug->flush();\n+        packet_write(\"result=success\\n\");\n+        exit();\n+    }\n+\n     my ($pathname) = packet_read() =~ /^pathname=([^=]+)\\n$/;\n     print $debug \" $pathname\";\n     $debug->flush();\n-- \n2.9.0\n\n"},{"id":"292934","messageId":"20160803164225.46355-7-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 06/12] pack-protocol: fix maximum pkt-line size","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:19Z","receivedAt":"2016-08-03T17:05:25Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nAccording to LARGE_PACKET_MAX in pkt-line.h the maximal length of a\npkt-line packet is 65520 bytes. The pkt-line header takes 4 bytes and\ntherefore the pkt-line data component must not exceed 65516 bytes.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n Documentation/technical/protocol-common.txt | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/technical/protocol-common.txt b/Documentation/technical/protocol-common.txt\nindex bf30167..ecedb34 100644\n--- a/Documentation/technical/protocol-common.txt\n+++ b/Documentation/technical/protocol-common.txt\n@@ -67,9 +67,9 @@ with non-binary data the same whether or not they contain the trailing\n LF (stripping the LF if present, and not complaining when it is\n missing).\n \n-The maximum length of a pkt-line's data component is 65520 bytes.\n-Implementations MUST NOT send pkt-line whose length exceeds 65524\n-(65520 bytes of payload + 4 bytes of length data).\n+The maximum length of a pkt-line's data component is 65516 bytes.\n+Implementations MUST NOT send pkt-line whose length exceeds 65520\n+(65516 bytes of payload + 4 bytes of length data).\n \n Implementations SHOULD NOT send an empty pkt-line (\"0004\").\n \n-- \n2.9.0\n\n"},{"id":"292936","messageId":"20160803164225.46355-5-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 04/12] pkt-line: call packet_trace() only if a packet is actually send","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:17Z","receivedAt":"2016-08-03T17:05:27Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nThe packet_trace() call is not ideal in format_packet() as we would print\na trace when a packet is formatted and (potentially) when the packet is\nactually send. This was no problem up until now because format_packet()\nwas only used by one function. Fix it by moving the trace call into the\nfunction that actally sends the packet.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n pkt-line.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/pkt-line.c b/pkt-line.c\nindex c8a052a..d1368e6 100644\n--- a/pkt-line.c\n+++ b/pkt-line.c\n@@ -127,7 +127,6 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)\n \t\tdie(\"protocol error: impossibly long line\");\n \n \tset_packet_header(&out->buf[orig_len], n);\n-\tpacket_trace(out->buf + orig_len + 4, n - 4, 1);\n }\n \n void packet_write(int fd, const char *fmt, ...)\n@@ -139,6 +138,7 @@ void packet_write(int fd, const char *fmt, ...)\n \tva_start(args, fmt);\n \tformat_packet(&buf, fmt, args);\n \tva_end(args);\n+\tpacket_trace(buf.buf + 4, buf.len - 4, 1);\n \twrite_or_die(fd, buf.buf, buf.len);\n }\n \n-- \n2.9.0\n\n"},{"id":"292937","messageId":"20160803164225.46355-12-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:24Z","receivedAt":"2016-08-03T17:05:30Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGit's clean/smudge mechanism invokes an external filter process for every\nsingle blob that is affected by a filter. If Git filters a lot of blobs\nthen the startup time of the external filter processes can become a\nsignificant part of the overall Git execution time.\n\nIn a preliminary performance test this developer used a clean/smudge filter\nwritten in golang to filter 12,000 files. This process took 364s with the\nexisting filter mechanism and 5s with the new mechanism. See details here:\nhttps://github.com/github/git-lfs/pull/1382\n\nThis patch adds the `filter.<driver>.process` string option which, if used,\nkeeps the external filter process running and processes all blobs with\nthe packet format (pkt-line) based protocol over standard input and standard\noutput described below.\n\nGit starts the filter when it encounters the first file\nthat needs to be cleaned or smudged. After the filter started\nGit expects a welcome message, protocol version number, and\nfilter capabilities separated by spaces:\n------------------------\npacket:          git< git-filter-protocol\\n\npacket:          git< version=2\\n\npacket:          git< capabilities=clean smudge\\n\n------------------------\nSupported filter capabilities are \"clean\" and \"smudge\".\n\nAfterwards Git sends a command (based on the supported\ncapabilities), the pathname of a file relative to the\nrepository root, the content split in zero or more pkt-line\npackets, and a flush packet at the end:\n------------------------\npacket:          git> command=smudge\\n\npacket:          git> pathname=path/testfile.dat\\n\npacket:          git> CONTENT\npacket:          git> 0000\n------------------------\n\nThe filter is expected to respond with the result content in zero\nor more pkt-line packets and a flush packet at the end. Finally, a\n\"result=success\" packet is expected if everything went well.\n------------------------\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< result=success\\n\n------------------------\n\nIf the result content is empty then the filter is expected to respond\nonly with a flush packet and a \"result=success\" packet.\n------------------------\npacket:          git< 0000\npacket:          git< result=success\\n\n------------------------\n\nIn case the filter cannot or does not want to process the content,\nit is expected to respond with a flush packet and a \"result=reject\"\npacket. Depending on the `filter.<driver>.required` flag Git will\ninterpret that as error but it will not stop or restart the filter\nprocess.\n------------------------\npacket:          git< 0000\npacket:          git< result=reject\\n\n------------------------\n\nIf the filter experiences an error during processing, then it can\neither die or send a flush packet and a \"result=error\" packet. If\nGit receives such an error then it will stop and restart the filter\nwith the next file that needs to be processed.\n------------------------\npacket:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\npacket:          git< 0000\npacket:          git< result=error\\n\n------------------------\n\nAfter the filter has processed a blob it is expected to wait for\nthe next command. When the Git process terminates, it will send\na kill signal to the filter in that stage.\n\nIf a `filter.<driver>.clean` or `filter.<driver>.smudge` command\nis configured then these commands always take precedence over\na configured `filter.<driver>.process` command.\n\nHelped-by: Martin-Louis Bright <mlbright@gmail.com>\nReviewed-by: Jakub Narebski <jnareb@gmail.com>\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n Documentation/gitattributes.txt |  98 +++++++++++-\n convert.c                       | 277 ++++++++++++++++++++++++++++++----\n t/t0021-conversion.sh           | 322 ++++++++++++++++++++++++++++++++++++++++\n t/t0021/rot13-filter.pl         | 148 ++++++++++++++++++\n 4 files changed, 815 insertions(+), 30 deletions(-)\n create mode 100755 t/t0021/rot13-filter.pl\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 8882a3e..49514ab 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -300,7 +300,13 @@ checkout, when the `smudge` command is specified, the command is\n fed the blob object from its standard input, and its standard\n output is used to update the worktree file.  Similarly, the\n `clean` command is used to convert the contents of worktree file\n-upon checkin.\n+upon checkin. By default these commands process only a single\n+blob and terminate.  If a long running `process` filter is used\n+in place of `clean` and/or `smudge` filters, then Git can process\n+all blobs with a single filter command invocation for the entire\n+life of a single Git command, for example `git add --all`.  See\n+section below for the description of the protocol used to\n+communicate with a `process` filter.\n \n One use of the content filtering is to massage the content into a shape\n that is more convenient for the platform, filesystem, and the user to use.\n@@ -375,6 +381,96 @@ substitution.  For example:\n ------------------------\n \n \n+Long Running Filter Process\n+^^^^^^^^^^^^^^^^^^^^^^^^^^^\n+\n+If the filter command (a string value) is defined via\n+`filter.<driver>.process` then Git can process all blobs with a\n+single filter invocation for the entire life of a single Git\n+command. This is achieved by using the following packet format\n+(pkt-line, see technical/protocol-common.txt) based protocol over\n+standard input and standard output.\n+\n+Git starts the filter when it encounters the first file\n+that needs to be cleaned or smudged. After the filter started\n+Git expects a welcome message, protocol version number, and\n+filter capabilities separated by spaces:\n+------------------------\n+packet:          git< git-filter-protocol\\n\n+packet:          git< version=2\\n\n+packet:          git< capabilities=clean smudge\\n\n+------------------------\n+Supported filter capabilities are \"clean\" and \"smudge\".\n+\n+Afterwards Git sends a command (based on the supported\n+capabilities), the pathname of a file relative to the\n+repository root, the content split in zero or more pkt-line\n+packets, and a flush packet at the end:\n+------------------------\n+packet:          git> command=smudge\\n\n+packet:          git> pathname=path/testfile.dat\\n\n+packet:          git> CONTENT\n+packet:          git> 0000\n+------------------------\n+\n+The filter is expected to respond with the result content in zero\n+or more pkt-line packets and a flush packet at the end. Finally, a\n+\"result=success\" packet is expected if everything went well.\n+------------------------\n+packet:          git< SMUDGED_CONTENT\n+packet:          git< 0000\n+packet:          git< result=success\\n\n+------------------------\n+\n+If the result content is empty then the filter is expected to respond\n+only with a flush packet and a \"result=success\" packet.\n+------------------------\n+packet:          git< 0000\n+packet:          git< result=success\\n\n+------------------------\n+\n+In case the filter cannot or does not want to process the content,\n+it is expected to respond with a flush packet and a \"result=reject\"\n+packet. Depending on the `filter.<driver>.required` flag Git will\n+interpret that as error but it will not stop or restart the filter\n+process.\n+------------------------\n+packet:          git< 0000\n+packet:          git< result=reject\\n\n+------------------------\n+\n+If the filter experiences an error during processing, then it can\n+either die or send a flush packet and a \"result=error\" packet. If\n+Git receives such an error then it will stop and restart the filter\n+with the next file that needs to be processed.\n+------------------------\n+packet:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\n+packet:          git< 0000\n+packet:          git< result=error\\n\n+------------------------\n+\n+After the filter has processed a blob it is expected to wait for\n+the next command. When the Git process terminates, it will send\n+a kill signal to the filter in that stage.\n+\n+A long running filter demo implementation can be found in\n+`t/t0021/rot13-filter.pl` located in the Git core repository.\n+If you develop your own long running filter process then the\n+`GIT_TRACE_PACKET` environment variables can be very helpful\n+for debugging (see linkgit:git[1]).\n+\n+If a `filter.<driver>.clean` or `filter.<driver>.smudge` command\n+is configured then these commands always take precedence over\n+a configured `filter.<driver>.process` command.\n+\n+Please note that you cannot use an existing `filter.<driver>.clean`\n+or `filter.<driver>.smudge` command with `filter.<driver>.process`\n+because the former two use a different inter process communication\n+protocol than the latter one. As soon as Git would detect a file\n+that needs to be processed by such an invalid \"process\" filter,\n+it would wait for a proper protocol handshake and appear \"hanging\".\n+\n+\n Interaction between checkin/checkout attributes\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \ndiff --git a/convert.c b/convert.c\nindex 522e2c5..130430a 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -3,6 +3,7 @@\n #include \"run-command.h\"\n #include \"quote.h\"\n #include \"sigchain.h\"\n+#include \"pkt-line.h\"\n \n /*\n  * convert.c - convert a file when checking it out and checking it in.\n@@ -427,7 +428,7 @@ static int filter_buffer_or_fd(int in, int out, void *data)\n \treturn (write_err || status);\n }\n \n-static int apply_filter(const char *path, const char *src, size_t len, int fd,\n+static int apply_single_file_filter(const char *path, const char *src, size_t len, int fd,\n                         struct strbuf *dst, const char *cmd)\n {\n \t/*\n@@ -441,12 +442,6 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n \tstruct async async;\n \tstruct filter_params params;\n \n-\tif (!cmd || !*cmd)\n-\t\treturn 0;\n-\n-\tif (!dst)\n-\t\treturn 1;\n-\n \tmemset(&async, 0, sizeof(async));\n \tasync.proc = filter_buffer_or_fd;\n \tasync.data = &params;\n@@ -481,14 +476,245 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,\n \treturn ret;\n }\n \n+#define FILTER_CAPABILITIES_CLEAN    (1u<<0)\n+#define FILTER_CAPABILITIES_SMUDGE   (1u<<1)\n+#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n+#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n+\n+struct cmd2process {\n+\tstruct hashmap_entry ent; /* must be the first member! */\n+\tint supported_capabilities;\n+\tconst char *cmd;\n+\tstruct child_process process;\n+};\n+\n+static int cmd_process_map_initialized = 0;\n+static struct hashmap cmd_process_map;\n+\n+static int cmd2process_cmp(const struct cmd2process *e1,\n+                           const struct cmd2process *e2,\n+                           const void *unused)\n+{\n+\treturn strcmp(e1->cmd, e2->cmd);\n+}\n+\n+static struct cmd2process *find_multi_file_filter_entry(struct hashmap *hashmap, const char *cmd)\n+{\n+\tstruct cmd2process key;\n+\thashmap_entry_init(&key, strhash(cmd));\n+\tkey.cmd = cmd;\n+\treturn hashmap_get(hashmap, &key, NULL);\n+}\n+\n+static void kill_multi_file_filter(struct hashmap *hashmap, struct cmd2process *entry)\n+{\n+\tif (!entry)\n+\t\treturn;\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\tclose(entry->process.in);\n+\tclose(entry->process.out);\n+\tsigchain_pop(SIGPIPE);\n+\tfinish_command(&entry->process);\n+\tchild_process_clear(&entry->process);\n+\thashmap_remove(hashmap, entry, NULL);\n+\tfree(entry);\n+}\n+\n+static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, const char *cmd)\n+{\n+\tint did_fail;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tconst char *argv[] = { cmd, NULL };\n+\tstatic const char cap_key[] = \"capabilities=\";\n+\tint cap_key_len = strlen(cap_key);\n+\tstruct string_list cap_list = STRING_LIST_INIT_NODUP;\n+\tchar *cap_buf;\n+\tint i;\n+\n+\tentry = xmalloc(sizeof(*entry));\n+\thashmap_entry_init(entry, strhash(cmd));\n+\tentry->cmd = cmd;\n+\tentry->supported_capabilities = 0;\n+\tprocess = &entry->process;\n+\n+\tchild_process_init(process);\n+\tprocess->argv = argv;\n+\tprocess->use_shell = 1;\n+\tprocess->in = -1;\n+\tprocess->out = -1;\n+\n+\tif (start_command(process)) {\n+\t\terror(\"cannot fork to run external filter '%s'\", cmd);\n+\t\tkill_multi_file_filter(hashmap, entry);\n+\t\treturn NULL;\n+\t}\n+\n+\tdid_fail = strcmp(packet_read_line(process->out, NULL), \"git-filter-protocol\");\n+\tif (did_fail)\n+\t\tgoto done;\n+\n+\tdid_fail = strcmp(packet_read_line(process->out, NULL), \"version=2\");\n+\tif (did_fail)\n+\t\tgoto done;\n+\n+\tcap_buf = packet_read_line(process->out, NULL);\n+\tif (!cap_buf ||\n+\t\tstrlen(cap_buf) <= cap_key_len ||\n+\t\tstrncmp(cap_buf, cap_key, cap_key_len)) {\n+\t\terror(\"filter capabilities not found\");\n+\t\tdid_fail = 1;\n+\t\tgoto done;\n+\t}\n+\n+\tstring_list_split_in_place(&cap_list, &cap_buf[cap_key_len], ' ', -1);\n+\tif (cap_list.nr > 0) {\n+\t\tfor (i = 0; i < cap_list.nr; i++) {\n+\t\t\tconst char *requested = cap_list.items[i].string;\n+\t\t\tif (!strcmp(requested, \"clean\")) {\n+\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_CLEAN;\n+\t\t\t} else if (!strcmp(requested, \"smudge\")) {\n+\t\t\t\tentry->supported_capabilities |= FILTER_CAPABILITIES_SMUDGE;\n+\t\t\t} else {\n+\t\t\t\twarning(\n+\t\t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n+\t\t\t\t\tcmd, requested\n+\t\t\t\t);\n+\t\t\t}\n+\t\t}\n+\t}\n+\tstring_list_clear(&cap_list, 0);\n+\n+done:\n+\tif (did_fail) {\n+\t\terror(\"initialization for external filter '%s' failed\", cmd);\n+\t\tkill_multi_file_filter(hashmap, entry);\n+\t\treturn NULL;\n+\t}\n+\n+\thashmap_add(hashmap, entry);\n+\treturn entry;\n+}\n+\n+static int apply_multi_file_filter(const char *path, const char *src, size_t len,\n+                                   int fd, struct strbuf *dst, const char *cmd,\n+                                   const int wanted_capability)\n+{\n+\tint ret = 1;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tstruct stat file_stat;\n+\tstruct strbuf nbuf = STRBUF_INIT;\n+\tchar *filter_type;\n+\tchar *filter_result = NULL;\n+\n+\tif (!cmd_process_map_initialized) {\n+\t\tcmd_process_map_initialized = 1;\n+\t\thashmap_init(&cmd_process_map, (hashmap_cmp_fn) cmd2process_cmp, 0);\n+\t\tentry = NULL;\n+\t} else {\n+\t\tentry = find_multi_file_filter_entry(&cmd_process_map, cmd);\n+\t}\n+\n+\tfflush(NULL);\n+\n+\tif (!entry) {\n+\t\tentry = start_multi_file_filter(&cmd_process_map, cmd);\n+\t\tif (!entry)\n+\t\t\treturn 0;\n+\t}\n+\tprocess = &entry->process;\n+\n+\tif (!(wanted_capability & entry->supported_capabilities))\n+\t\treturn 1;  // it is OK if the wanted capability is not supported\n+\n+\tif (FILTER_SUPPORTS_CLEAN(wanted_capability))\n+\t\tfilter_type = \"clean\";\n+\telse if (FILTER_SUPPORTS_SMUDGE(wanted_capability))\n+\t\tfilter_type = \"smudge\";\n+\telse\n+\t\tdie(\"unexpected filter type\");\n+\n+\tif (fd >= 0 && !src) {\n+\t\tif (fstat(fd, &file_stat) == -1)\n+\t\t\treturn 0;\n+\t\tlen = file_stat.st_size;\n+\t}\n+\n+\tpacket_buf_write(&nbuf, \"command=%s\\n\", filter_type);\n+\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n+\tif (!ret)\n+\t\tgoto done;\n+\n+\tstrbuf_reset(&nbuf);\n+\tpacket_buf_write(&nbuf, \"pathname=%s\\n\", path);\n+\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n+\tif (!ret)\n+\t\tgoto done;\n+\n+\tif (fd >= 0)\n+\t\tret = !packet_write_stream_with_flush_from_fd(fd, process->in);\n+\telse\n+\t\tret = !packet_write_stream_with_flush_from_buf(src, len, process->in);\n+\tif (!ret)\n+\t\tgoto done;\n+\n+\tstrbuf_reset(&nbuf);\n+\tret = packet_read_till_flush(process->out, &nbuf) >= 0;\n+\tif (!ret)\n+\t\tgoto done;\n+\n+\tfilter_result = packet_read_line(process->out, NULL);\n+\tret = !strcmp(filter_result, \"result=success\");\n+\n+done:\n+\tif (ret) {\n+\t\tstrbuf_swap(dst, &nbuf);\n+\t} else {\n+\t\tif (!filter_result || strcmp(filter_result, \"result=reject\")) {\n+\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n+\t\t\terror(\"external filter '%s' failed\", cmd);\n+\t\t\tkill_multi_file_filter(&cmd_process_map, entry);\n+\t\t}\n+\t}\n+\tstrbuf_release(&nbuf);\n+\treturn ret;\n+}\n+\n static struct convert_driver {\n \tconst char *name;\n \tstruct convert_driver *next;\n \tconst char *smudge;\n \tconst char *clean;\n+\tconst char *process;\n \tint required;\n } *user_convert, **user_convert_tail;\n \n+static int apply_filter(const char *path, const char *src, size_t len,\n+                        int fd, struct strbuf *dst, struct convert_driver *drv,\n+                        const int wanted_capability)\n+{\n+\tconst char* cmd = NULL;\n+\n+\tif (!drv)\n+\t\treturn 0;\n+\n+\tif (!dst)\n+\t\treturn 1;\n+\n+\tif (FILTER_SUPPORTS_CLEAN(wanted_capability) && drv->clean)\n+\t\tcmd = drv->clean;\n+\telse if (FILTER_SUPPORTS_SMUDGE(wanted_capability) && drv->smudge)\n+\t\tcmd = drv->smudge;\n+\n+\tif (cmd && *cmd)\n+\t\treturn apply_single_file_filter(path, src, len, fd, dst, cmd);\n+\telse if (drv->process && *drv->process)\n+\t\treturn apply_multi_file_filter(path, src, len, fd, dst, drv->process, wanted_capability);\n+\n+\treturn 0;\n+}\n+\n static int read_convert_config(const char *var, const char *value, void *cb)\n {\n \tconst char *key, *name;\n@@ -526,6 +752,10 @@ static int read_convert_config(const char *var, const char *value, void *cb)\n \tif (!strcmp(\"clean\", key))\n \t\treturn git_config_string(&drv->clean, var, value);\n \n+\tif (!strcmp(\"process\", key)) {\n+\t\treturn git_config_string(&drv->process, var, value);\n+\t}\n+\n \tif (!strcmp(\"required\", key)) {\n \t\tdrv->required = git_config_bool(var, value);\n \t\treturn 0;\n@@ -823,7 +1053,7 @@ int would_convert_to_git_filter_fd(const char *path)\n \tif (!ca.drv->required)\n \t\treturn 0;\n \n-\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n+\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv, FILTER_CAPABILITIES_CLEAN);\n }\n \n const char *get_convert_attr_ascii(const char *path)\n@@ -856,18 +1086,12 @@ int convert_to_git(const char *path, const char *src, size_t len,\n                    struct strbuf *dst, enum safe_crlf checksafe)\n {\n \tint ret = 0;\n-\tconst char *filter = NULL;\n-\tint required = 0;\n \tstruct conv_attrs ca;\n \n \tconvert_attrs(&ca, path);\n-\tif (ca.drv) {\n-\t\tfilter = ca.drv->clean;\n-\t\trequired = ca.drv->required;\n-\t}\n \n-\tret |= apply_filter(path, src, len, -1, dst, filter);\n-\tif (!ret && required)\n+\tret |= apply_filter(path, src, len, -1, dst, ca.drv, FILTER_CAPABILITIES_CLEAN);\n+\tif (!ret && ca.drv && ca.drv->required)\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n \tif (ret && dst) {\n@@ -889,9 +1113,9 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n \tconvert_attrs(&ca, path);\n \n \tassert(ca.drv);\n-\tassert(ca.drv->clean);\n+\tassert(ca.drv->clean || ca.drv->process);\n \n-\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv->clean))\n+\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv, FILTER_CAPABILITIES_CLEAN))\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n@@ -903,15 +1127,9 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t\t\t\t    int normalizing)\n {\n \tint ret = 0, ret_filter = 0;\n-\tconst char *filter = NULL;\n-\tint required = 0;\n \tstruct conv_attrs ca;\n \n \tconvert_attrs(&ca, path);\n-\tif (ca.drv) {\n-\t\tfilter = ca.drv->smudge;\n-\t\trequired = ca.drv->required;\n-\t}\n \n \tret |= ident_to_worktree(path, src, len, dst, ca.ident);\n \tif (ret) {\n@@ -920,9 +1138,10 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t}\n \t/*\n \t * CRLF conversion can be skipped if normalizing, unless there\n-\t * is a smudge filter.  The filter might expect CRLFs.\n+\t * is a smudge or process filter (even if the process filter doesn't\n+\t * support smudge).  The filters might expect CRLFs.\n \t */\n-\tif (filter || !normalizing) {\n+\tif ((ca.drv && (ca.drv->smudge || ca.drv->process)) || !normalizing) {\n \t\tret |= crlf_to_worktree(path, src, len, dst, ca.crlf_action);\n \t\tif (ret) {\n \t\t\tsrc = dst->buf;\n@@ -930,8 +1149,8 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t}\n \t}\n \n-\tret_filter = apply_filter(path, src, len, -1, dst, filter);\n-\tif (!ret_filter && required)\n+\tret_filter = apply_filter(path, src, len, -1, dst, ca.drv, FILTER_CAPABILITIES_SMUDGE);\n+\tif (!ret_filter && ca.drv && ca.drv->required)\n \t\tdie(\"%s: smudge filter %s failed\", path, ca.drv->name);\n \n \treturn ret | ret_filter;\n@@ -1383,7 +1602,7 @@ struct stream_filter *get_stream_filter(const char *path, const unsigned char *s\n \tstruct stream_filter *filter = NULL;\n \n \tconvert_attrs(&ca, path);\n-\tif (ca.drv && (ca.drv->smudge || ca.drv->clean))\n+\tif (ca.drv && (ca.drv->process || ca.drv->smudge || ca.drv->clean))\n \t\treturn NULL;\n \n \tif (ca.crlf_action == CRLF_AUTO || ca.crlf_action == CRLF_AUTO_CRLF)\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 34c8eb9..c1a22f4 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -42,6 +42,9 @@ test_expect_success setup '\n \trm -f test test.t test.i &&\n \tgit checkout -- test test.t test.i &&\n \n+\techo \"content-test2\" >test2.o &&\n+\techo \"content-test3-subdir\" >test3-subdir.o &&\n+\n \tmkdir generated-test-data &&\n \tfor i in $(test_seq 1 $T0021_LARGE_FILE_SIZE)\n \tdo\n@@ -296,4 +299,323 @@ test_expect_success 'disable filter with empty override' '\n \ttest_must_be_empty err\n '\n \n+check_filter () {\n+\trm -f rot13-filter.log actual.log &&\n+\t\"$@\" 2> git_stderr.log &&\n+\ttest_must_be_empty git_stderr.log &&\n+\tcat >expected.log &&\n+\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" >actual.log &&\n+\ttest_cmp expected.log actual.log\n+}\n+\n+check_filter_count_clean () {\n+\trm -f rot13-filter.log actual.log &&\n+\t\"$@\" 2> git_stderr.log &&\n+\ttest_must_be_empty git_stderr.log &&\n+\tcat >expected.log &&\n+\tsort rot13-filter.log | uniq -c | sed \"s/^[ ]*//\" |\n+\t\tsed \"s/^\\([0-9]\\) IN: clean/x IN: clean/\" >actual.log &&\n+\ttest_cmp expected.log actual.log\n+}\n+\n+check_filter_ignore_clean () {\n+\trm -f rot13-filter.log actual.log &&\n+\t\"$@\" &&\n+\tcat >expected.log &&\n+\tgrep -v \"IN: clean\" rot13-filter.log >actual.log &&\n+\ttest_cmp expected.log actual.log\n+}\n+\n+check_filter_no_call () {\n+\trm -f rot13-filter.log &&\n+\t\"$@\" 2> git_stderr.log &&\n+\ttest_must_be_empty git_stderr.log &&\n+\ttest_must_be_empty rot13-filter.log\n+}\n+\n+check_rot13 () {\n+\ttest_cmp $1 $2 &&\n+\t./../rot13.sh <$1 >expected &&\n+\tgit cat-file blob :$2 >actual &&\n+\ttest_cmp expected actual\n+}\n+\n+test_expect_success PERL 'required process filter should filter data' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\tcat ../test2.o >test2.r &&\n+\t\tmkdir testsubdir &&\n+\t\tcat ../test3-subdir.o >testsubdir/test3-subdir.r &&\n+\t\t>test4-empty.r &&\n+\n+\t\tcheck_filter \\\n+\t\t\tgit add . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\t1 IN: clean test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\t1 IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]\n+\t\t\t\t\t1 IN: clean testsubdir/test3-subdir.r 21 [OK] -- OUT: 21 [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_count_clean \\\n+\t\t\tgit commit . -m \"test commit\" \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tx IN: clean test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\tx IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]\n+\t\t\t\t\tx IN: clean testsubdir/test3-subdir.r 21 [OK] -- OUT: 21 [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\n+\t\trm -f test?.r testsubdir/test3-subdir.r &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\tIN: smudge testsubdir/test3-subdir.r 21 [OK] -- OUT: 21 [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout empty \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout master \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\t\tIN: smudge test4-empty.r 0 [OK] -- OUT: 0 [OK]\n+\t\t\t\t\tIN: smudge testsubdir/test3-subdir.r 21 [OK] -- OUT: 21 [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_rot13 ../test.o test.r &&\n+\t\tcheck_rot13 ../test2.o test2.r &&\n+\t\tcheck_rot13 ../test3-subdir.o testsubdir/test3-subdir.r\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should filter smudge data and one-shot filter should clean' '\n+\ttest_config_global filter.protocol.clean ./../rot13.sh &&\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\tcat ../test2.o >test2.r &&\n+\n+\t\tcheck_filter_no_call \\\n+\t\t\tgit add . &&\n+\n+\t\tcheck_filter_no_call \\\n+\t\t\tgit commit . -m \"test commit\" &&\n+\n+\t\trm -f test?.r testsubdir/test3-subdir.r &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\tEOF\n+\n+\t\tgit checkout empty &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout master\\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_rot13 ../test.o test.r &&\n+\t\tcheck_rot13 ../test2.o test2.r\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should clean only' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\tgit branch empty &&\n+\n+\t\tcat ../test.o >test.r &&\n+\n+\t\tcheck_filter \\\n+\t\t\tgit add . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\t1 IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\n+\t\tcheck_filter_count_clean \\\n+\t\t\tgit commit . -m \"test commit\" \\\n+\t\t\t\t<<-\\EOF\n+\t\t\t\t\tx IN: clean test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\t1 start\n+\t\t\t\t\t1 wrote filter header\n+\t\t\t\tEOF\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should process files larger LARGE_PACKET_MAX' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.file filter=protocol\" >.gitattributes &&\n+\t\tcat ../generated-test-data/largish.file.rot13 >large.rot13 &&\n+\t\tcat ../generated-test-data/largish.file >large.file &&\n+\t\tcat large.file >large.original &&\n+\n+\t\tgit add large.file .gitattributes &&\n+\t\tgit commit . -m \"test commit\" &&\n+\n+\t\trm -f large.file &&\n+\t\tgit checkout -- large.file &&\n+\t\tgit cat-file blob :large.file >actual &&\n+\t\ttest_cmp large.rot13 actual\n+\t)\n+'\n+\n+test_expect_success PERL 'required process filter should with clean error should fail' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.required true &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\techo \"this is going to fail\" >clean-write-fail.r &&\n+\t\techo \"content-test3-subdir\" >test3.r &&\n+\n+\t\t# Note: There are three clean paths in convert.c we just test one here.\n+\t\ttest_must_fail git add .\n+\t)\n+'\n+\n+test_expect_success PERL 'process filter should restart after unexpected write failure' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\tcat ../test2.o >test2.r &&\n+\t\techo \"this is going to fail\" >smudge-write-fail.o &&\n+\t\tcat smudge-write-fail.o >smudge-write-fail.r &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\trm -f *.r &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge smudge-write-fail.r 22 [OK] -- OUT: 22 [WRITE FAIL]\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_rot13 ../test.o test.r &&\n+\t\tcheck_rot13 ../test2.o test2.r &&\n+\n+\t\t! test_cmp smudge-write-fail.o smudge-write-fail.r && # Smudge failed!\n+\t\t./../rot13.sh <smudge-write-fail.o >expected &&\n+\t\tgit cat-file blob :smudge-write-fail.r >actual &&\n+\t\ttest_cmp expected actual\t\t\t\t\t\t\t  # Clean worked!\n+\t)\n+'\n+\n+test_expect_success PERL 'process filter should not restart after intentionally rejected file' '\n+\ttest_config_global filter.protocol.process \"$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge\" &&\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\n+\t\techo \"*.r filter=protocol\" >.gitattributes &&\n+\n+\t\tcat ../test.o >test.r &&\n+\t\tcat ../test2.o >test2.r &&\n+\t\techo \"this is going to be rejected\" >reject.o &&\n+\t\tcat reject.o >reject.r &&\n+\t\tgit add . &&\n+\t\tgit commit . -m \"test commit\" &&\n+\t\trm -f *.r &&\n+\n+\t\tcheck_filter_ignore_clean \\\n+\t\t\tgit checkout . \\\n+\t\t\t\t<<-\\EOF &&\n+\t\t\t\t\tstart\n+\t\t\t\t\twrote filter header\n+\t\t\t\t\tIN: smudge reject.r 29 [OK] -- OUT: 0 [REJECT]\n+\t\t\t\t\tIN: smudge test.r 57 [OK] -- OUT: 57 [OK]\n+\t\t\t\t\tIN: smudge test2.r 14 [OK] -- OUT: 14 [OK]\n+\t\t\t\tEOF\n+\n+\t\tcheck_rot13 ../test.o test.r &&\n+\t\tcheck_rot13 ../test2.o test2.r\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nnew file mode 100755\nindex 0000000..ca6d5e4\n--- /dev/null\n+++ b/t/t0021/rot13-filter.pl\n@@ -0,0 +1,148 @@\n+#!/usr/bin/perl\n+#\n+# Example implementation for the Git filter protocol version 2\n+# See Documentation/gitattributes.txt, section \"Filter Protocol\"\n+#\n+# The script takes the list of supported protocol capabilities as\n+# arguments (\"clean\", \"smudge\", etc).\n+#\n+# This implementation supports three special test cases:\n+# (1) If data with the pathname \"clean-write-fail.r\" is processed with\n+#     a \"clean\" operation then the write operation will die.\n+# (2) If data with the pathname \"smudge-write-fail.r\" is processed with\n+#     a \"smudge\" operation then the write operation will die.\n+# (3) If data with the pathname \"reject.r\" is processed with any\n+#     operation then the filter signals that it does not want to process\n+#     the file.\n+#\n+\n+use strict;\n+use warnings;\n+\n+my $MAX_PACKET_CONTENT_SIZE = 65516;\n+my @capabilities            = @ARGV;\n+\n+sub rot13 {\n+    my ($str) = @_;\n+    $str =~ y/A-Za-z/N-ZA-Mn-za-m/;\n+    return $str;\n+}\n+\n+sub packet_read {\n+    my $buffer;\n+    my $bytes_read = read STDIN, $buffer, 4;\n+    if ( $bytes_read == 0 ) {\n+        return;\n+    }\n+    elsif ( $bytes_read != 4 ) {\n+        die \"invalid packet size '$bytes_read' field\";\n+    }\n+    my $pkt_size = hex($buffer);\n+    if ( $pkt_size == 0 ) {\n+        return ( 1, \"\" );\n+    }\n+    elsif ( $pkt_size > 4 ) {\n+        my $content_size = $pkt_size - 4;\n+        $bytes_read = read STDIN, $buffer, $content_size;\n+        if ( $bytes_read != $content_size ) {\n+            die \"invalid packet ($content_size expected; $bytes_read read)\";\n+        }\n+        return ( 0, $buffer );\n+    }\n+    else {\n+        die \"invalid packet size\";\n+    }\n+}\n+\n+sub packet_write {\n+    my ($packet) = @_;\n+    print STDOUT sprintf( \"%04x\", length($packet) + 4 );\n+    print STDOUT $packet;\n+    STDOUT->flush();\n+}\n+\n+sub packet_flush {\n+    print STDOUT sprintf( \"%04x\", 0 );\n+    STDOUT->flush();\n+}\n+\n+open my $debug, \">>\", \"rot13-filter.log\";\n+print $debug \"start\\n\";\n+$debug->flush();\n+\n+packet_write(\"git-filter-protocol\\n\");\n+packet_write(\"version=2\\n\");\n+packet_write( \"capabilities=\" . join( ' ', @capabilities ) . \"\\n\" );\n+print $debug \"wrote filter header\\n\";\n+$debug->flush();\n+\n+while (1) {\n+    my ($command) = packet_read() =~ /^command=([^=]+)\\n$/;\n+    unless ( defined($command) ) {\n+        exit();\n+    }\n+    print $debug \"IN: $command\";\n+    $debug->flush();\n+\n+    my ($pathname) = packet_read() =~ /^pathname=([^=]+)\\n$/;\n+    print $debug \" $pathname\";\n+    $debug->flush();\n+\n+    my $input = \"\";\n+    {\n+        binmode(STDIN);\n+        my $buffer;\n+        my $done = 0;\n+        while ( !$done ) {\n+            ( $done, $buffer ) = packet_read();\n+            $input .= $buffer;\n+        }\n+        print $debug \" \" . length($input) . \" [OK] -- \";\n+        $debug->flush();\n+    }\n+\n+    my $output;\n+    if ( $pathname eq \"reject.r\" ) {\n+        $output = \"\";\n+    }\n+    elsif ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n+        $output = rot13($input);\n+    }\n+    elsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n+        $output = rot13($input);\n+    }\n+    else {\n+        die \"bad command $command\";\n+    }\n+\n+    print $debug \"OUT: \" . length($output) . \" \";\n+    $debug->flush();\n+\n+    if ( $pathname eq \"${command}-write-fail.r\" ) {\n+        print $debug \"[WRITE FAIL]\\n\";\n+        $debug->flush();\n+        die \"write error\";\n+    }\n+    elsif ( $pathname eq \"reject.r\" ) {\n+        packet_flush();\n+        print $debug \"[REJECT]\\n\";\n+        $debug->flush();\n+        packet_write(\"result=reject\\n\");\n+    }\n+    else {\n+        while ( length($output) > 0 ) {\n+            my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n+            packet_write($packet);\n+            if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n+                $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n+            }\n+            else {\n+                $output = \"\";\n+            }\n+        }\n+        packet_flush();\n+        print $debug \"[OK]\\n\";\n+        $debug->flush();\n+        packet_write(\"result=success\\n\");\n+    }\n+}\n-- \n2.9.0\n\n"},{"id":"292941","messageId":"20160803164225.46355-11-larsxschneider@gmail.com","threadId":"42968","inReplyTo":"20160803164225.46355-1-larsxschneider@gmail.com","subject":"[PATCH v4 10/12] convert: generate large test files only once","fromName":"","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T16:42:23Z","receivedAt":"2016-08-03T17:05:34Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"From: Lars Schneider <larsxschneider@gmail.com>\n\nGenerate more interesting large test files with pseudo random characters\nin between and reuse these test files in multiple tests. Run tests formerly\nmarked as EXPENSIVE every time but with a smaller data set.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh | 48 ++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 38 insertions(+), 10 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 7b45136..34c8eb9 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -4,6 +4,15 @@ test_description='blob conversion via gitattributes'\n \n . ./test-lib.sh\n \n+if test_have_prereq EXPENSIVE\n+then\n+\tT0021_LARGE_FILE_SIZE=2048\n+\tT0021_LARGISH_FILE_SIZE=100\n+else\n+\tT0021_LARGE_FILE_SIZE=30\n+\tT0021_LARGISH_FILE_SIZE=2\n+fi\n+\n cat <<EOF >rot13.sh\n #!$SHELL_PATH\n tr \\\n@@ -31,7 +40,26 @@ test_expect_success setup '\n \tcat test >test.i &&\n \tgit add test test.t test.i &&\n \trm -f test test.t test.i &&\n-\tgit checkout -- test test.t test.i\n+\tgit checkout -- test test.t test.i &&\n+\n+\tmkdir generated-test-data &&\n+\tfor i in $(test_seq 1 $T0021_LARGE_FILE_SIZE)\n+\tdo\n+\t\tRANDOM_STRING=\"$(test-genrandom end $i | tr -dc \"A-Za-z0-9\" )\"\n+\t\tROT_RANDOM_STRING=\"$(echo $RANDOM_STRING | ./rot13.sh )\"\n+\t\t# Generate 1MB of empty data and 100 bytes of random characters\n+\t\t# printf \"$(test-genrandom start $i)\"\n+\t\tprintf \"%1048576d\" 1 >>generated-test-data/large.file &&\n+\t\tprintf \"$RANDOM_STRING\" >>generated-test-data/large.file &&\n+\t\tprintf \"%1048576d\" 1 >>generated-test-data/large.file.rot13 &&\n+\t\tprintf \"$ROT_RANDOM_STRING\" >>generated-test-data/large.file.rot13 &&\n+\n+\t\tif test $i = $T0021_LARGISH_FILE_SIZE\n+\t\tthen\n+\t\t\tcat generated-test-data/large.file >generated-test-data/largish.file &&\n+\t\t\tcat generated-test-data/large.file.rot13 >generated-test-data/largish.file.rot13\n+\t\tfi\n+\tdone\n '\n \n script='s/^\\$Id: \\([0-9a-f]*\\) \\$/\\1/p'\n@@ -199,9 +227,9 @@ test_expect_success 'required filter clean failure' '\n test_expect_success 'filtering large input to small output should use little memory' '\n \ttest_config filter.devnull.clean \"cat >/dev/null\" &&\n \ttest_config filter.devnull.required true &&\n-\tfor i in $(test_seq 1 30); do printf \"%1048576d\" 1; done >30MB &&\n-\techo \"30MB filter=devnull\" >.gitattributes &&\n-\tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add 30MB\n+\tcp generated-test-data/large.file large.file &&\n+\techo \"large.file filter=devnull\" >.gitattributes &&\n+\tGIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add large.file\n '\n \n test_expect_success 'filter that does not read is fine' '\n@@ -214,15 +242,15 @@ test_expect_success 'filter that does not read is fine' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success EXPENSIVE 'filter large file' '\n+test_expect_success 'filter large file' '\n \ttest_config filter.largefile.smudge cat &&\n \ttest_config filter.largefile.clean cat &&\n-\tfor i in $(test_seq 1 2048); do printf \"%1048576d\" 1; done >2GB &&\n-\techo \"2GB filter=largefile\" >.gitattributes &&\n-\tgit add 2GB 2>err &&\n+\techo \"large.file filter=largefile\" >.gitattributes &&\n+\tcp generated-test-data/large.file large.file &&\n+\tgit add large.file 2>err &&\n \ttest_must_be_empty err &&\n-\trm -f 2GB &&\n-\tgit checkout -- 2GB 2>err &&\n+\trm -f large.file &&\n+\tgit checkout -- large.file 2>err &&\n \ttest_must_be_empty err\n '\n \n-- \n2.9.0\n\n"},{"id":"292949","messageId":"xmqqtwf19263.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"20160803164225.46355-12-larsxschneider@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-03T17:45:40Z","receivedAt":"2016-08-03T17:46:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"larsxschneider@gmail.com writes:\n\n> packet:          git< git-filter-protocol\\n\n> packet:          git< version=2\\n\n> packet:          git< capabilities=clean smudge\\n\n\nDuring the discussion on the future of pack-protocol, it was pointed\nout that having to shove all capabilities on a single line/packet\nwas one of the things we would want to fix in the current protocol\nwhen we revamp to v2.  As this exhange between the convert machinery\nand an external process is a brand new one, I do not think you want\nto mimic the limitation in the current pack protocol like this; the\nlimitation mostly came from the constraint that we cannot break\nexisting pack protocol clients and servers before we extended the\nprotocol to add capabilities.\n\nYou may not foresee that the caps won't grow very long beyond\nclean/smudge right now, just like we did not foresee that we would\nwish to be able to convey a lot longer capability values to the\nother side when we added the capability exchange to the pack\nprotocol, so \"but but but we will never have that many\" is not a\ngood counter-argument.\n\n"},{"id":"292962","messageId":"607c07fe-5b6f-fd67-13e1-705020c267ee@gmail.com","threadId":"42968","inReplyTo":"ABE7D2DB-C45F-4F29-8CC2-8D873FD6C36A@gmail.com","subject":"Designing the filter process protocol (was: Re: [PATCH v3 10/10] convert: add filter.<driver>.process option)","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-03T18:30:18Z","receivedAt":"2016-08-03T18:31:22Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[I'm sorry for taking so long in writing this, as I see there is v4 already]\n\nGreetings,\n\n\nI'll answer to individual emails in more detail later, but I'd like to\ngo back to the drawing board, and attempt to summarize the discussion and\nthe proposal so far.\n\nThe ultimate goal is to be able to run filter drivers faster for both `clean`\nand `smudge` operations.  This is done by starting filter driver once per\ngit command invocation, instead of once per file being processed.  Git needs\nto pass actual contents of files to filter driver, and get its output.\n\nWe want the protocol between Git and filter driver process to be extensible,\nso that new features can be added without modifying protocol.\n\n\n1. CONFIGURATION\n\nAs I wrote, there are different ways of configuring new-type filter driver:\n\n * Using a separate variable to mark filter as using new protocol\n   (the original approach):\n\n   \t[filter \"protocol\"]\n\t\tprotocolVersion = v2\n\t\tclean  = rot13-clean-filter.pl\n\t\tsmudge = rot13-smudge-filter.pl\n\n   PROS: allows to have separate clean and smudge filters\n   CONS: does not allow using old-style per-file filter together with new;\n         easy to make mistake and use old-style filter, leading to hang\n\n * Creating new variables for new filter type, separate for each phase,\n   for example `cleanProcess` and `smudgeProcess` (or `processClean` and\n   `processSmudge`).\n\n   \t[filter \"protocol\"]\n   \t\tcleanProcess  = rot13-clean-filter.pl\n   \t\tsmudgeProcess = rot13-smudge-filter.pl\n\n   PROS: allows to have separate clean and smudge filters;\n         makes possible to use per-file and per-command filters together\n   CONS: proliferation of additional variables, (esp. when extending it);\n   NOTE: need to decide precedence between `clean` and `cleanProcess`, etc.\n\n # Using a single variable for new filter type, and decide on which phase\n   (which operation) is supported by filter driver during the handshake\n   *(current approach)*\n\n   \t[filter \"protocol\"]\n   \t\tprocess = rot13-filtes.pl\n\n   PROS: per-file and per-command filters possible with precedence rule;\n         extensible to other types of drivers: textconv, diff, etc.\n         only one invocation for commands which use both clean and smudge\n   CONS: need single driver to be responsible for both clean and smudge;\n         need to run driver to know that it does not support given\n           operation (workaround exists)\n\n\n2. HANDSHAKE (INITIALIZATION)\n\nNext, there is deciding on and designing the handshake between Git (between\nGit command) and the filter driver process.  With the `filter.<driver>.process`\nsolution the driver needs to tell which operations among (for now) \"clean\"\nand \"smudge\" it does support.  Plus it provides a way to extend protocol,\nadding new features, like support for streaming, cleaning from file or\nsmudging to file, providing size upfront, perhaps even progress report.\n\nCurrent handshake consist of filter driver printing a signature, version\nnumber and capabilities, in that order.  Git checks that it is well formed\nand matches expectations, and notes which of \"clean\" and \"smudge\" operations\nare supported by the filter.\n\nThere is no interaction from the Git side in the handshake, for example to\nset options and expectations common to all files being filtered.  Take\none possible extension of protocol: supporting streaming.  The filter\ndriver needs to know whether it needs to read all the input, or whether\nit can start printing output while input is incoming (e.g. to reduce\nmemory consumption)... though we may simply decide it to be next version\nof the protocol.\n\nOn the other hand if the handshake began with Git sending some initializer\ninfo to the filter driver, we probably could detect one-shot filter\nmisconfigured as process-filter.\n\nNote that we need some way of deciding where handshake ends, either by\nspecifying number of entries (currently: three lines / pkt-line packets),\nor providing some terminator (\"smart\" transport protocol uses flush packet\nfor this).\n\nCurrent handshake (in symbolic form):\n\n    git< [signature]    git-filter-protocol\n    git< [version]      version 2\n    git< [capabilites]  clean smudge\n\nIt is expected that the handshake is limited to this information, and\nthat they are in this order; so naming them doesn't buy us much\n\n    git< [capabilites]  capabilities clean smudge\n\nor\n\n    git< [capabilites]  capabilities=clean smudge\n\nor\n\n    git< [capabilites]  capabilities: clean smudge\n\nIf capabilities are to be third item, adding \"capabilities\", as if Git would\nlook at the name and select what to do based on this name, doesn't buy us\nanything.  Well, beside self-documenting of the protocol.  The \"smart\" protocol\ndo not use \"capabilities\" as prefix/name either.\n\nWe would probably do not want to move from strict-order of information, that\nis \"positional parameters\".  It would require to implement a parser, both for\nthe Git side and for the filter driver process side.\n\nOn the other hand requiring flush packet to end the handshake doesn't bring\nmuch overhead (it is 4 bytes, it is not over the network), and improves\nextendability.  Well, so does using names, be it \"<var> <value>\", \n\"<var>=<value>\", \"<var>: <value>...\", \"<var>=[<value>, <value>...]\", etc.\n\n\nLet's take a look how other parts of Git communicate with external process\n(a \"helper\").\n\nThe git-credential(1) protocol uses <variable>=<value> syntax.  But capabilities\nform a list; \"<var>=<val1> <val2>\" doesn't look that well.  Credential helper\nonly uses scalar (single) values.\n\nThe gitremote-helpers(1) protocol is command / response; for example helper\nresponds to \"capabilities\" command with the list of capabilities.  Here commands\nand parameters are space separated, e.g. \"option <name> <value>\".\n\nThe \"smart\" transport protocol (send-pack and receive-pack) had to (ab)use\na quirk of implementation to extend protocol with capabilities negotiation.\nHere the capabilities list is sent without any prefix; some capabilities\nare parametrized, and use <capability>=<value> syntax (for example\n\"symref=HEAD:refs/heads/master\").  The handshake is closed with flush\npacket, but as it consist of variable-length ref advertisement, it needs\nto have explicit terminator of the each part of the \"handshake\".\n\n\n3. SENDING CONTENTS (FILE TO BE FILTERED AND FILTER OUTPUT)\n\nNext thing to design is decision how to send contents to be filtered\nto the filter driver process, and how to get filtered output from the\nfilter driver process.\n\nOne thing I think we can agree on early, is sending data to filter\nprocess on its standard input, and receiving filtered result from its\nstandard output.\n\nBecause Git is sending (and receiving) multiple files, it needs some\nway to distinguish where one file ends and the next begins, in both\ndirections, to and from filter.  Also, the `clean` and `smudge`\nfilters support expansion of the '%f' placeholder, so at least \nsome filter drivers need name of the file being filtered.  So the\nprotocol must send it somehow to the filter driver.\n\nThere are different approaches possible; here are ones that were used,\nand ones I thought about.\n\n * Send whole data to filter at once, and receiver all data at once,\n   for example using something akin to the 'tar' archive, or \n   uncompressed 'zip' archive (both are implemented in Git for the\n   `git archive` command).  Or just list of sizes and pathnames,\n   empty entry as terminator, and then contents of all files\n   concatenated.\n\n   PROS:\n   - can use the one-shot infrastructure implemented already\n   CONS:\n   - complicates Git code and filter driver code unnecessarily\n   - difficult to implement error handling, esp. soft errors\n     on filter driver side (error for single file, perhaps during\n     output)\n   - in synchronous version (non-streaming) requires absurd amout\n     of memory / storage for the filter driver process\n\n * Send/receive data file by file, using <size> + <content>,\n   that is, send size (plus other data like the filename), then\n   file contents.\n\n   This was the protocol used in the first iteration of series.\n\n   PROS:\n   - simple to implement on Git and on filter driver side\n   NOTE:\n   - you need to loop over read / user read_in_full anyway\n   CONS:\n   - no way to signal an error encountered during output, e.g. LFS\n     network/server failure for after some contents were actually\n     sent\n   - impossible to implement streaming for filters that do not\n     know size of output without examining full input\n\n # Send/receive data file by file, using some kind of chunking,\n   with a end-of-file marker.  The solution used by Git is\n   pkt-line, with flush packet used to signal end of file.\n\n   This is protocol used by the current implementation.\n\n   PROS:\n   - no need to know size upfront, so easier streaming support\n   - you can signal error that happened during output, after\n     some data were sent, as well as error known upfront\n   - tracing support for free (GIT_TRACE_PACKET)\n   CONS:\n   - filter driver program slightly more difficult to implement\n   - some negligible amount of overhead\n\nIf we want in the end to implement streaming, then the last solution\nis the way to go.\n\n\n4. PER-FILE HANDSHAKE - SENDING FILE TO FILTER\n\nLet's assume that for simplicity we want to implement (for now) only\nthe synchronous (non-streaming) case, where we send whole contents\nof a file to filter driver process, and *then* read filter driver\noutput.  This is enough for git-LFS solutions, which were the reason\nfor this patch series.  But we want to keep the protocol flexible\nenough so that streaming and other features could be added easily.\n\nFirst, if we choose the solution where one process is responsible\nfor both \"clean\" and \"smudge\" operations (and in the future possibly\nalso \"cleanFromFile\" and \"smudgeToFile\"), Git needs to tell the\ndriver which operation to perform.\n\nTogether with operation Git can send additional information\n(sub-capabilities)... or we can use a separate line / packet to\nsend it.\n\nIf we are using pkt-line, then the convention is that text lines\nare terminated using LF (\"\\n\") character.  This needs to be stated\nexplicitly in the documentation for filter.<driver>.process writers.\n\n    git> packet:  [operation] clean size=67\\n\n\nWe could denote that it is operation name, but it is obvious from\nposition in the stream, thus not really needed.\n\nThen we need to provide the filename; some filters supposedly need\nthis ('%f' in per-file `clean` / `smudge`).  Note that filename can\ncontain internal space characters, and could contain newlines, equal\nsigns; anything that is not NUL (\"\\0\") character.\n\n    git> packet:  [pathname] subdir/sample-file.r\\n\n\nIn most cases filename would be text, so perhaps we should use \"\\n\"\nterminator (which filter driver would have to strip).  We could use\n\"filename=\" prefix, but it is not necessary.  We know where / when\nto expect the pathname (relative to project root).\n\nIf we would want to be able to add variable number of packets to\nthe handshake, then Git should send flush packet to signal the\nend of the handshake.  But IMVHO it is unnecessary complication\nof the protocol; there is enough flexibility in it.  We know\nthat handshake consists of two packets.\n\nThe Git would sent contents of the file to be filtered, using\nas many pack lines as needed (note: large file support needs\nto be tested, at least as expensive test).  Flush packet is\nused to signal the end of the file.\n\n    git> packets:  <file contents>\n    git> flush packet\n\n\n5. FILTER DRIVER PROCESS RESPONSE\n\nFirst filter should, in my opinion, reply that it received the\nrequest (or the command, in the case of streaming supported).\nAlso, in this response it can provide further information to\nGit process.\n\n    git< packet: [received]  ok size=67\\n\n\nThis response could be used to refuse to filter specific file\nupfront (for example if the file is not present in the artifactory\nfor git-LFS solutions).\n\n   git< packet: [rejected]  reject\\n\n\nWe can even provide the reasoning to Git (maybe in the future\nextension)... or filter driver can print the explanation to the\nstandard error (but then, no --quiet / --verbose support).\n\n   git< packet: [rejected]  reject with-message\\n\n   git< packet: [message]   File not found on server\\n\n   git< flush packet\n\nAnother response, which I think should be standarized, or at\nleast described in the documentation, is filter driver refusing\nto filter further (e.g. git-LFS and network is down), to be not\nrestarted by Git.\n\n   git< packet: [quit]      quit msg=Server error\\n\n\nor\n\n   git< packet: [quit]      quit Server error\\n\n\nor\n\n   git< packet: [quit]      quit with-message\\n\n   git< packet: [message]   Server error\\n\n   git< flush packet\n\nMaybe this is over-engineering, but I don't think so.\n\nNext comes the output from the filter driver (filtered contents),\nusing possibly multiple pkt-lines, ending with a flush packet:\n\n    git< packets:  <filtered contents>\n    git< flush packet\n\nNote that empty file would consist of zero pack lines of contents,\nand one flush packet.\n\nFinally, to allow handling of [resumable] errors that occurred\nduring sending file contents, especially for the future streaming\nfilters case, we want to confirm that we send whole file\nsuccessfully.\n\n    git< packet: [status]   success\\n\n\nIf there was an error during process, making data receives so far\ninvalid, filter driver should tell about it\n\n    git< packet: [status]   fail\\n\n\nor\n\n    git< packet: [status]   reject\\n\n\nThis may happen for example for UCS-2 <-> UTF-8 filter when invalid\nbyte sequence is encountered.  This may happen for git-LFS if the\nserver fails during fetch, and spare / slave server doesn't have\na file.\n\nWe may want to quit filtering at this point, and not to send another\nfile.\n\n   git< packet: [status]    quit\\n\n\nThere is place for extra information after the status, and in the\nfuture we can allow variable length information too.\n\n\nBest,\n-- \nJakub Narębski\n"},{"id":"292977","messageId":"51080fb7-2bab-a100-0971-e82063c1ac78@gmail.com","threadId":"42968","inReplyTo":"73AED3FA-C666-4C49-90B2-387E410F7D52@gmail.com","subject":"Re: [PATCH v3 01/10] pkt-line: extract set_packet_header()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-03T20:05:51Z","receivedAt":"2016-08-03T20:07:33Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[This response might have been invalidated by v4]\n\nW dniu 01.08.2016 o 13:33, Lars Schneider pisze: \n>> On 30 Jul 2016, at 12:30, Jakub Narębski <jnareb@gmail.com> wrote:\n\n>>> #define hex(a) (hexchar[(a) & 15])\n>>\n>> I guess that this is inherited from the original, but this preprocessor\n>> macro is local to the format_header() / set_packet_header() function,\n>> and would not work outside it.  Therefore I think we should #undef it\n>> after set_packet_header(), just in case somebody mistakes it for\n>> a generic hex() function.  Perhaps even put it inside set_packet_header(),\n>> together with #undef.\n>>\n>> But I might be mistaken... let's check... no, it isn't used outside it.\n> \n> Agreed. Would that be OK?\n> \n> static void set_packet_header(char *buf, const int size)\n> {\n> \tstatic char hexchar[] = \"0123456789abcdef\";\n> \t#define hex(a) (hexchar[(a) & 15])\n> \tbuf[0] = hex(size >> 12);\n> \tbuf[1] = hex(size >> 8);\n> \tbuf[2] = hex(size >> 4);\n> \tbuf[3] = hex(size);\n> \t#undef hex\n> }\n\nThat's better, though I wonder if we need to start #defines at begining\nof line.  But I think current proposal is O.K.\n\n\nEither this (which has unnecessary larger scope)\n\n  #define hex(a) (hexchar[(a) & 15])\n  static void set_packet_header(char *buf, const int size)\n  {\n  \tstatic char hexchar[] = \"0123456789abcdef\";\n\n  \tbuf[0] = hex(size >> 12);\n  \tbuf[1] = hex(size >> 8);\n  \tbuf[2] = hex(size >> 4);\n  \tbuf[3] = hex(size);\n  }\n  #undef hex\n\nor this (which looks worse)\n\n  static void set_packet_header(char *buf, const int size)\n  {\n  \tstatic char hexchar[] = \"0123456789abcdef\";\n  #define hex(a) (hexchar[(a) & 15])\n  \tbuf[0] = hex(size >> 12);\n  \tbuf[1] = hex(size >> 8);\n  \tbuf[2] = hex(size >> 4);\n  \tbuf[3] = hex(size);\n  #undef hex\n  }\n\n"},{"id":"292980","messageId":"xmqqd1lp8v2o.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"20160803164225.46355-2-larsxschneider@gmail.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-03T20:18:55Z","receivedAt":"2016-08-03T20:19:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"larsxschneider@gmail.com writes:\n\n> From: Lars Schneider <larsxschneider@gmail.com>\n>\n> set_packet_header() converts an integer to a 4 byte hex string. Make\n> this function locally available so that other pkt-line functions can\n> use it.\n\nDidn't I say that this is a bad idea already in an earlier review?\n\nThe only reason why you want it, together with direct_packet_write()\n(which I think is another bad idea), is because you use\npacket_buf_write() to create a \"<header><payload>\" in a buf in the\nusercode in step 11/12 like this:\n\n+\tpacket_buf_write(&nbuf, \"command=%s\\n\", filter_type);\n+\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n\nwhich would be totally unnecessary if you just did strbuf_addf()\ninto nbuf and used packet_write() like everybody else does.\n\nPuzzled.  Why are steps 01/12 and 02/12 an improvement?\n\n"},{"id":"292981","messageId":"e8b550ed-1765-764f-49e5-72e5a609d936@gmail.com","threadId":"42968","inReplyTo":"64783AA5-D579-4783-88E7-E0B3BDE5FDEB@gmail.com","subject":"Re: [PATCH v3 02/10] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-03T20:12:48Z","receivedAt":"2016-08-03T20:19:10Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[This response might have been invalidated by v4]\n\nW dniu 01.08.2016 o 14:00, Lars Schneider pisze:\n>> On 30 Jul 2016, at 12:49, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>>>\n>>> Sometimes pkt-line data is already available in a buffer and it would\n>>> be a waste of resources to write the packet using packet_write() which\n>>> would copy the existing buffer into a strbuf before writing it.\n>>>\n>>> If the caller has control over the buffer creation then the\n>>> PKTLINE_DATA_START macro can be used to skip the header and write\n>>> directly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\n>>> would be the maximum). direct_packet_write() would take this buffer,\n>>> adjust the pkt-line header and write it.\n>>>\n>>> If the caller has no control over the buffer creation then\n>>> direct_packet_write_data() can be used. This function creates a pkt-line\n>>> header. Afterwards the header and the data buffer are written using two\n>>> consecutive write calls.\n>>\n>> I don't quite understand what do you mean by \"caller has control\n>> over the buffer creation\".  Do you mean that caller either can write\n>> over the buffer, or cannot overwrite the buffer?  Or do you mean that\n>> caller either can allocate buffer to hold header, or is getting\n>> only the data?\n> \n> How about this:\n> \n> [...]\n> \n> If the caller creates the buffer then a proper pkt-line buffer with header\n> and data section can be created. The PKTLINE_DATA_START macro can be used \n> to skip the header section and write directly to the data section (PKTLINE_DATA_LEN \n> bytes would be the maximum). direct_packet_write() would take this buffer, \n> fill the pkt-line header section with the appropriate data length value and \n> write the entire buffer.\n> \n> If the caller does not create the buffer, and consequently cannot leave room\n> for the pkt-line header, then direct_packet_write_data() can be used. This \n> function creates an extra buffer for the pkt-line header and afterwards writes\n> the header buffer and the data buffer with two consecutive write calls.\n> \n> ---\n> Is that more clear?\n\nYes, I think it is more clear.  \n\nThe only thing that could be improved is to perhaps instead of using\n\n  \"then a proper pkt-line buffer with header and data section can be created\"\n\nit might be more clear to write\n\n  \"then a proper pkt-line buffer with data section and a place for pkt-line header\"\n \n\n>>> +{\n>>> +\tint ret = 0;\n>>> +\tchar hdr[4];\n>>> +\tset_packet_header(hdr, sizeof(hdr) + size);\n>>> +\tpacket_trace(buf, size, 1);\n>>> +\tif (gentle) {\n>>> +\t\tret = (\n>>> +\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n>>\n>> You can write '4' here, no need for sizeof(hdr)... though compiler would\n>> optimize it away.\n> \n> Right, it would be optimized. However, I don't like the 4 there either. OK to use a macro\n> instead? PKTLINE_HEADER_LEN ?\n\nDid you mean \n\n    +\tchar hdr[PKTLINE_HEADER_LEN];\n    +\tset_packet_header(hdr, sizeof(hdr) + size);\n\n \n>>> +\t\t\t!write_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>>> +\t\t);\n>>\n>> Do we want to try to write \"pkt-line data\" if \"pkt-line header\" failed?\n>> If not, perhaps De Morgan-ize it\n>>\n>>  +\t\tret = !(\n>>  +\t\t\twrite_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") &&\n>>  +\t\t\twrite_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>>  +\t\t);\n> \n> \n> Original:\n> \t\tret = (\n> \t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n> \t\t\t!write_or_whine_pipe(fd, data, size, \"pkt-line data\")\n> \t\t);\n> \n> Well, if the first write call fails (return == 0), then it is negated and evaluates to true.\n> I would think the second call is not evaluated, then?!\n\nThis is true both for || and for &&, as in C logical boolean operators\nshort-circuit.\n\n> Should I make this more explicit with a if clause?\n\nNo need.\n\n-- \nJakub Narębski\n\n"},{"id":"292985","messageId":"fb5f8496-fb85-93b7-d83d-f5d24f0bc5c2@gmail.com","threadId":"42968","inReplyTo":"52A5703D-AF9E-405D-A4F8-E30A7B923400@gmail.com","subject":"Re: [PATCH v3 04/10] pkt-line: call packet_trace() only if a packet is actually send","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-03T20:15:12Z","receivedAt":"2016-08-03T20:22:38Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[This response might have been invalidated by v4]\n\nW dniu 01.08.2016 o 14:18, Lars Schneider pisze:\n>> On 30 Jul 2016, at 14:29, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n\n>> I don't buy this explanation.  If you want to trace packets, you might\n>> do it on input (when formatting packet), or on output (when writing\n>> packet).  It's when there are more than one formatting function, but\n>> one writing function, then placing trace call in write function means\n>> less code duplication; and of course the reverse.\n>>\n>> Another issue is that something may happen between formatting packet\n>> and sending it, and we probably want to packet_trace() when packet\n>> is actually send.\n>>\n>> Neither of those is visible in commit message.\n> \n> The packet_trace() call in format_packet() is not ideal, as we would print\n> a trace when a packet is formatted and (potentially) when the same packet is\n> actually written. This was no problem up until now because packet_write(),\n> the function that uses format_packet() and writes the formatted packet,\n> did not trace the packet.\n> \n> This developer believes that trace calls should only happen when a packet\n> is actually written as the packet could be modified between formatting\n> and writing. Therefore the trace call was moved from format_packet() to \n> packet_write().\n> \n> --\n> \n> Better?\n\nYes, that's much better.\n\nP.S. Yes, this is one of those changes where commit message is much longer\n     than the change itself...\n\n-- \nJakub Narębski\n\n"},{"id":"292986","messageId":"xmqq8twd8uld.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"20160803164225.46355-12-larsxschneider@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-03T20:29:18Z","receivedAt":"2016-08-03T20:30:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"larsxschneider@gmail.com writes:\n\n> +#define FILTER_CAPABILITIES_CLEAN    (1u<<0)\n> +#define FILTER_CAPABILITIES_SMUDGE   (1u<<1)\n> +#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n> +#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n\nI would expect a lot shorter names as these are file-local;\nCAP_CLEAN and CAP_SMUDGE, perhaps, _WITHOUT_ \"supports BLAH\" macros?\n\n\tif (FILTER_SUPPORTS_CLEAN(type))\n\nis not all that more readable than\n\n\tif (CAP_CLEAN & type)\n\n\n\n> +struct cmd2process {\n> +\tstruct hashmap_entry ent; /* must be the first member! */\n> +\tint supported_capabilities;\n> +\tconst char *cmd;\n> +\tstruct child_process process;\n> +};\n> +\n> +static int cmd_process_map_initialized = 0;\n> +static struct hashmap cmd_process_map;\n\nDon't initialize statics to 0 or NULL.\n\n> +static int cmd2process_cmp(const struct cmd2process *e1,\n> +                           const struct cmd2process *e2,\n> +                           const void *unused)\n> +{\n> +\treturn strcmp(e1->cmd, e2->cmd);\n> +}\n> +\n> +static struct cmd2process *find_multi_file_filter_entry(struct hashmap *hashmap, const char *cmd)\n> +{\n> +\tstruct cmd2process key;\n> +\thashmap_entry_init(&key, strhash(cmd));\n> +\tkey.cmd = cmd;\n> +\treturn hashmap_get(hashmap, &key, NULL);\n> +}\n> +\n> +static void kill_multi_file_filter(struct hashmap *hashmap, struct cmd2process *entry)\n> +{\n> +\tif (!entry)\n> +\t\treturn;\n> +\tsigchain_push(SIGPIPE, SIG_IGN);\n> +\tclose(entry->process.in);\n> +\tclose(entry->process.out);\n> +\tsigchain_pop(SIGPIPE);\n> +\tfinish_command(&entry->process);\n\nI wonder if we want to diagnose failures from close(), which is a\nlot more interesting than usual because these are connected to\npipes.\n\n> +static int apply_multi_file_filter(const char *path, const char *src, size_t len,\n> +                                   int fd, struct strbuf *dst, const char *cmd,\n> +                                   const int wanted_capability)\n> +{\n> +\tint ret = 1;\n> + ...\n> +\tif (!(wanted_capability & entry->supported_capabilities))\n> +\t\treturn 1;  // it is OK if the wanted capability is not supported\n\nNo // comment please.\n\n> +\tfilter_result = packet_read_line(process->out, NULL);\n> +\tret = !strcmp(filter_result, \"result=success\");\n> +\n> +done:\n> +\tif (ret) {\n> +\t\tstrbuf_swap(dst, &nbuf);\n> +\t} else {\n> +\t\tif (!filter_result || strcmp(filter_result, \"result=reject\")) {\n> +\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n> +\t\t\terror(\"external filter '%s' failed\", cmd);\n> +\t\t\tkill_multi_file_filter(&cmd_process_map, entry);\n> +\t\t}\n> +\t}\n> +\tstrbuf_release(&nbuf);\n> +\treturn ret;\n> +}\n\nI think this was already pointed out in the previous review by Peff,\nbut a variable \"ret\" that says \"0 is bad\" somehow makes it hard to\nfollow the code.  Perhaps rename it to \"int error\", flip the meaning,\nand if the caller wants this function to return non-zero on success\nflip the polarity in the return statement itself, i.e. \"return !errors\",\nmay make it easier to follow?\n"},{"id":"292999","messageId":"20160803211221.t2zdhvwjum2baeqs@sigill.intra.peff.net","threadId":"42968","inReplyTo":"xmqqd1lp8v2o.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T21:12:22Z","receivedAt":"2016-08-03T21:19:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 03, 2016 at 01:18:55PM -0700, Junio C Hamano wrote:\n\n> larsxschneider@gmail.com writes:\n> \n> > From: Lars Schneider <larsxschneider@gmail.com>\n> >\n> > set_packet_header() converts an integer to a 4 byte hex string. Make\n> > this function locally available so that other pkt-line functions can\n> > use it.\n> \n> Didn't I say that this is a bad idea already in an earlier review?\n> \n> The only reason why you want it, together with direct_packet_write()\n> (which I think is another bad idea), is because you use\n> packet_buf_write() to create a \"<header><payload>\" in a buf in the\n> usercode in step 11/12 like this:\n> \n> +\tpacket_buf_write(&nbuf, \"command=%s\\n\", filter_type);\n> +\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n> \n> which would be totally unnecessary if you just did strbuf_addf()\n> into nbuf and used packet_write() like everybody else does.\n> \n> Puzzled.  Why are steps 01/12 and 02/12 an improvement?\n\nI think it is an attempt to avoid the extra memcpy() of the bytes into\nanother packet buffer.\n\nI notice that the solution does still end up a using a double-write() in\nsome cases, though.  I was curious if this made any difference, though,\nso I wrote a short test program:\n\n-- >8 --\n#include <unistd.h>\n#include <string.h>\n\nint main(int argc, char **argv)\n{\n        int type;\n\n        if (argv[1] && !strcmp(argv[1], \"prepend\"))\n                type = 0; /* size prepended to buffer */\n        else if (argv[1] && !strcmp(argv[1], \"write\"))\n                type = 1;\n        else if (argv[1] && !strcmp(argv[1], \"memcpy\"))\n                type = 2;\n        else\n                return 1;\n\n        while (1) {\n                char buf[65520];\n                int r = read(0, buf + 4, sizeof(buf));\n                if (r <= 0)\n                        break;\n                if (!type) {\n                        memcpy(buf, \"1234\", 4);\n                        write(1, buf, r + 4);\n                } else if (type == 1) {\n                        write(1, \"1234\", 4);\n                        write(1, buf + 4, r);\n                } else if (type == 2) {\n                        char packet[sizeof(buf) + 4];\n                        memcpy(packet, \"1234\", 4);\n                        memcpy(packet + 4, buf + 4, r);\n                        write(1, packet, r + 4);\n                }\n        }\n        return 0;\n}\n-- >8 --\n\nWe'd expect \"prepend\" to be the fastest, as it does a single write and\nzero-copy. And then it is a question of whether the double-write is\nworse than the extra memcpy.\n\nOn Linux, feeding 100MB of zeroes into stdin, I got (best-of-five):\n\n  - prepend: 11ms\n  - write: 11ms\n  - memcpy: 15ms\n\nSo it _does_ make a difference to avoid the memcpy, though 4ms per 100MB\ndoes not seem like it is probably worth caring about. The double-write\nalso gets worse if you use a smaller buffer size (e.g., if you drop to\n4K, that adds back in about 4ms of overhead because you're calling\nwrite() a lot more times).\n\nThe cost of write() may vary on other platforms, but the cost of memcpy\ngenerally shouldn't. So I'm inclined to say that it is not really worth\nmicro-optimizing the interface.\n\nI think the other issue is that format_packet() only lets you send\nstring data via \"%s\", so it cannot be used for arbitrary data that may\ncontain NULs. So we do need _some_ other interface to let you send a raw\ndata packet, and it's going to look similar to the direct_packet_write()\nthing.\n\nThe alternative is to hand-code it, which is what send_sideband() does\n(it uses xsnprintf(\"%04x\") to do the hex formatting, though).\n\n-Peff\n"},{"id":"293000","messageId":"20160803212433.zzdino3ivyem5a2v@sigill.intra.peff.net","threadId":"42968","inReplyTo":"20160803164225.46355-8-larsxschneider@gmail.com","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T21:24:33Z","receivedAt":"2016-08-03T21:31:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 03, 2016 at 06:42:20PM +0200, larsxschneider@gmail.com wrote:\n\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> Some commands might need to perform cleanup tasks on exit. Let's give\n> them an interface for doing this.\n> \n> Please note, that the cleanup callback is not executed if Git dies of a\n> signal. The reason is that only \"async-signal-safe\" functions would be\n> allowed to be call in that case. Since we cannot control what functions\n> the callback will use, we will not support the case. See 507d7804 for\n> more details.\n\nI'm not clear on why we want this cleanup filter. It looks like you use\nit in the final patch to send an explicit shutdown to any filters we\nstart. But I see two issues with that:\n\n  1. This shutdown may come at any time, and you have no idea what state\n     the protocol conversation with the filter is in. You could be in\n     the middle of sending another pkt-line, or in a sequence of non-command\n     pkt-lines where \"shutdown\" is not recognized.\n\n  2. If your protocol does bad things when it is cut off in the middle\n     without an explicit shutdown, then it's a bad protocol. As you\n     note, this patch doesn't cover signal death, nor could it ever\n     cover something like \"kill -9\", or a bug which prevented git from\n     saying \"shutdown\".\n\n     You're much better off to design the protocol so that a premature\n     EOF is detected as an error.  For example, if we're feeding file\n     data to the filter, and we're worried it might be writing it to\n     a data store (like LFS), we would not want it to see EOF and say\n     \"well, I guess I got all the data; time to store this!\". Instead,\n     it should know how many bytes are coming, or should have some kind\n     of framing so that the sender says \"and now you have seen all the\n     bytes\" (like a pkt-line flush).\n\n     AFAIK, your protocol _does_ do those things sensibly, so this\n     explicit shutdown isn't really accomplishing anything.\n\n-Peff\n"},{"id":"293004","messageId":"20160803213920.jg3eshy57bsldqjh@sigill.intra.peff.net","threadId":"42968","inReplyTo":"20160803164225.46355-4-larsxschneider@gmail.com","subject":"Re: [PATCH v4 03/12] pkt-line: add packet_flush_gentle()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T21:39:20Z","receivedAt":"2016-08-03T21:39:27Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 03, 2016 at 06:42:16PM +0200, larsxschneider@gmail.com wrote:\n\n> From: Lars Schneider <larsxschneider@gmail.com>\n> \n> packet_flush() would die in case of a write error even though for some callers\n> an error would be acceptable. Add packet_flush_gentle() which writes a pkt-line\n> flush packet and returns `0` for success and `1` for failure.\n\nOur normal convention would be \"0\" for success, \"-1\" for failure.\n\nI see write_or_whine_pipe(), which you use here, has a bizarre \"0 for\nfailure, 1 for success\", but that nobody actually checks it.\n\nI actually think you probably don't want to use write_or_whine_pipe()\nhere. It does two things:\n\n  1. It writes to stderr unconditionally. But if you are doing a\n     \"gently\" form, then you probably don't want unconditional errors.\n     Since the point of not dying is that you could presumably recover\n     in some way, or do some other more intelligent action.\n\n     The existing callers of write_or_whine_pipe() are all in the trace\n     code. Their use is not \"let's handle an error\", but \"we _would_ die\n     except that this is low-priority debugging code that should not\n     interrupt the normal flow\". So there it at least makes sense to\n     unconditionally complain to stderr, but not to die().\n\n     For your series, I don't think that is true (and especially for\n     most potential callers of a generic \"gently flush the packet\"\n     function).\n\n  2. It calls check_pipe(), which will turn EPIPE into death-by-SIGPIPE\n     (in case you had for some reason ignored SIGPIPE).\n\n     But I think that's the opposite of what you want. You know you're\n     writing to a pipe, and I would think EPIPE is the most common\n     reason that your writes would fail (i.e., the helper unexpectedly\n     died while you were writing to it).\n\n     So you would want to explicitly ignore SIGPIPE while talking to the\n     helper, and then handle EPIPE just as any other error.\n\nThinking about (2), I'd go so far as to say that the trace actually\nshould just be using:\n\n  if (write_in_full(...) < 0)\n\twarning(\"unable to write trace to ...: %s\", strerror(errno));\n\nand we should get rid of write_or_whine_pipe entirely.\n\n-Peff\n"},{"id":"293006","messageId":"0E3FC781-1B2C-4341-9B7B-D9D836596A35@gmail.com","threadId":"42968","inReplyTo":"xmqqtwf19263.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T21:48:00Z","receivedAt":"2016-08-03T21:48:33Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 19:45, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> larsxschneider@gmail.com writes:\n> \n>> packet:          git< git-filter-protocol\\n\n>> packet:          git< version=2\\n\n>> packet:          git< capabilities=clean smudge\\n\n> \n> During the discussion on the future of pack-protocol, it was pointed\n> out that having to shove all capabilities on a single line/packet\n> was one of the things we would want to fix in the current protocol\n> when we revamp to v2.  As this exhange between the convert machinery\n> and an external process is a brand new one, I do not think you want\n> to mimic the limitation in the current pack protocol like this; the\n> limitation mostly came from the constraint that we cannot break\n> existing pack protocol clients and servers before we extended the\n> protocol to add capabilities.\n> \n> You may not foresee that the caps won't grow very long beyond\n> clean/smudge right now, just like we did not foresee that we would\n> wish to be able to convey a lot longer capability values to the\n> other side when we added the capability exchange to the pack\n> protocol, so \"but but but we will never have that many\" is not a\n> good counter-argument.\n\nOK. Is this the v2 discussion you are referring to?\nhttp://public-inbox.org/git/1461972887-22100-1-git-send-email-sbeller%40google.com/\n\nWhat format do you suggest?\n\npacket:          git< git-filter-protocol\\n\npacket:          git< version=2\\n\npacket:          git< capability=clean\\n\npacket:          git< capability=smudge\\n\npacket:          git< 0000\n\nor\n\npacket:          git< git-filter-protocol\\n\npacket:          git< version=2\\n\npacket:          git< capability\\n\npacket:          git< clean\\n\npacket:          git< smudge\\n\npacket:          git< 0000\n\nor  ... ?\n\nI would prefer the first one, I think.\n\n- Lars"},{"id":"293007","messageId":"564CA3AD-EA8E-46D3-9564-BF468CAF32B0@gmail.com","threadId":"42968","inReplyTo":"xmqq8twd8uld.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T21:37:54Z","receivedAt":"2016-08-03T21:51:21Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 22:29, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> larsxschneider@gmail.com writes:\n> \n>> +#define FILTER_CAPABILITIES_CLEAN    (1u<<0)\n>> +#define FILTER_CAPABILITIES_SMUDGE   (1u<<1)\n>> +#define FILTER_SUPPORTS_CLEAN(type)  ((type) & FILTER_CAPABILITIES_CLEAN)\n>> +#define FILTER_SUPPORTS_SMUDGE(type) ((type) & FILTER_CAPABILITIES_SMUDGE)\n> \n> I would expect a lot shorter names as these are file-local;\n> CAP_CLEAN and CAP_SMUDGE, perhaps, _WITHOUT_ \"supports BLAH\" macros?\n> \n> \tif (FILTER_SUPPORTS_CLEAN(type))\n> \n> is not all that more readable than\n> \n> \tif (CAP_CLEAN & type)\n\nOK. I will change that.\n\n\n>> +struct cmd2process {\n>> +\tstruct hashmap_entry ent; /* must be the first member! */\n>> +\tint supported_capabilities;\n>> +\tconst char *cmd;\n>> +\tstruct child_process process;\n>> +};\n>> +\n>> +static int cmd_process_map_initialized = 0;\n>> +static struct hashmap cmd_process_map;\n> \n> Don't initialize statics to 0 or NULL.\n\nOK, statics are initialized implicitly to 0.\nI will fix it.\n\n\n>> +static int cmd2process_cmp(const struct cmd2process *e1,\n>> +                           const struct cmd2process *e2,\n>> +                           const void *unused)\n>> +{\n>> +\treturn strcmp(e1->cmd, e2->cmd);\n>> +}\n>> +\n>> +static struct cmd2process *find_multi_file_filter_entry(struct hashmap *hashmap, const char *cmd)\n>> +{\n>> +\tstruct cmd2process key;\n>> +\thashmap_entry_init(&key, strhash(cmd));\n>> +\tkey.cmd = cmd;\n>> +\treturn hashmap_get(hashmap, &key, NULL);\n>> +}\n>> +\n>> +static void kill_multi_file_filter(struct hashmap *hashmap, struct cmd2process *entry)\n>> +{\n>> +\tif (!entry)\n>> +\t\treturn;\n>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n>> +\tclose(entry->process.in);\n>> +\tclose(entry->process.out);\n>> +\tsigchain_pop(SIGPIPE);\n>> +\tfinish_command(&entry->process);\n> \n> I wonder if we want to diagnose failures from close(), which is a\n> lot more interesting than usual because these are connected to\n> pipes.\n\nIn this particular case we kill the filter. That means some error \nalready happened, therefore the result wouldn't be of interest\nanymore, I think. Wrong?\n\nThe other case is the proper shutdown (see 12/12). However, in\nthat case Git is already exiting and therefore I wonder what\nwe would do with a \"close\" error?\n\n\n>> +static int apply_multi_file_filter(const char *path, const char *src, size_t len,\n>> +                                   int fd, struct strbuf *dst, const char *cmd,\n>> +                                   const int wanted_capability)\n>> +{\n>> +\tint ret = 1;\n>> + ...\n>> +\tif (!(wanted_capability & entry->supported_capabilities))\n>> +\t\treturn 1;  // it is OK if the wanted capability is not supported\n> \n> No // comment please.\n\nOK!\n\n\n>> +\tfilter_result = packet_read_line(process->out, NULL);\n>> +\tret = !strcmp(filter_result, \"result=success\");\n>> +\n>> +done:\n>> +\tif (ret) {\n>> +\t\tstrbuf_swap(dst, &nbuf);\n>> +\t} else {\n>> +\t\tif (!filter_result || strcmp(filter_result, \"result=reject\")) {\n>> +\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n>> +\t\t\terror(\"external filter '%s' failed\", cmd);\n>> +\t\t\tkill_multi_file_filter(&cmd_process_map, entry);\n>> +\t\t}\n>> +\t}\n>> +\tstrbuf_release(&nbuf);\n>> +\treturn ret;\n>> +}\n> \n> I think this was already pointed out in the previous review by Peff,\n> but a variable \"ret\" that says \"0 is bad\" somehow makes it hard to\n> follow the code.  Perhaps rename it to \"int error\", flip the meaning,\n> and if the caller wants this function to return non-zero on success\n> flip the polarity in the return statement itself, i.e. \"return !errors\",\n> may make it easier to follow?\n\nThis follows the existing filter function. Please see Peff's later\nreply here:\n\n\"So I'm not sure if changing them is a good idea. I agree with you that\nit's probably inviting confusion to have the two sets of filter\nfunctions have opposite return codes. So I think I retract my\nsuggestion. :)\"\n\nhttp://public-inbox.org/git/20160728133523.GB21311%40sigill.intra.peff.net/\n\nThat's why I kept it the way it is. If you prefer the \"!errors\" approach\nthen I will change that.\n\n\nThanks for looking at the patch,\nLars\n\n"},{"id":"293008","messageId":"20160803212718.cdqhd2zfzce7mqfa@sigill.intra.peff.net","threadId":"42968","inReplyTo":"20160803211221.t2zdhvwjum2baeqs@sigill.intra.peff.net","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T21:27:18Z","receivedAt":"2016-08-03T21:54:19Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 03, 2016 at 05:12:21PM -0400, Jeff King wrote:\n\n> The alternative is to hand-code it, which is what send_sideband() does\n> (it uses xsnprintf(\"%04x\") to do the hex formatting, though).\n\nAfter seeing that, I wondered why we need set_packet_header() at all.\nBut we do for the case when we are filling in the size at the start of a\nbuffer, because xsnprintf() will write an extra NUL that we do not care\nabout. send_sideband() is happy to then overwrite it with data, but\ncode (like format_packet) that computes the buffer, then fills in the\nsize, must avoid overwriting the first byte of the buffer.\n\n-Peff\n"},{"id":"293010","messageId":"D116610C-F33A-43DA-A49D-0B33958822E5@gmail.com","threadId":"42968","inReplyTo":"xmqqd1lp8v2o.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T21:56:42Z","receivedAt":"2016-08-03T21:56:50Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 22:18, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> larsxschneider@gmail.com writes:\n> \n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> set_packet_header() converts an integer to a 4 byte hex string. Make\n>> this function locally available so that other pkt-line functions can\n>> use it.\n> \n> Didn't I say that this is a bad idea already in an earlier review?\n\nYes, but in that earlier version I made this function *publicly*\navailable. In this patch the function is only available and used\nwithin pkt-line.c.\n\n\n> The only reason why you want it, together with direct_packet_write()\n> (which I think is another bad idea), is because you use\n> packet_buf_write() to create a \"<header><payload>\" in a buf in the\n> usercode in step 11/12 like this:\n> \n> +\tpacket_buf_write(&nbuf, \"command=%s\\n\", filter_type);\n> +\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n> \n> which would be totally unnecessary if you just did strbuf_addf()\n> into nbuf and used packet_write() like everybody else does.\n\nThe usercode in step 11/12 could use packet_buf_write(). I am not\nworried about performance here. What I am worried about is that\npacket_buf_write() dies on error. Since direct_packet_write()\nhas a \"gentle\" parameter in can handle these cases. This is important\nbecause a filter might be configured as \"required=false\" and then\nerrors are OK.\n\nWould you prefer to see a packet_buf_write_gently() instead?\n\nThanks,\nLars\n"},{"id":"293016","messageId":"CAPc5daWH1z3am2hV_U1dE5WA7R+xrOFxgrxV4CN-vhz6uHz8Hw@mail.gmail.com","threadId":"42968","inReplyTo":"564CA3AD-EA8E-46D3-9564-BF468CAF32B0@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-03T21:43:27Z","receivedAt":"2016-08-03T22:01:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Wed, Aug 3, 2016 at 2:37 PM, Lars Schneider <larsxschneider@gmail.com> wrote:\n>>\n>> I think this was already pointed out in the previous review by Peff,\n>> but a variable \"ret\" that says \"0 is bad\" somehow makes it hard to\n>> follow the code.  Perhaps rename it to \"int error\", flip the meaning,\n>> and if the caller wants this function to return non-zero on success\n>> flip the polarity in the return statement itself, i.e. \"return !errors\",\n>> may make it easier to follow?\n>\n> This follows the existing filter function. Please see Peff's later\n> reply here:\n\nWhich I did before mentioning \"pointed out in his review\".\n\n> That's why I kept it the way it is. If you prefer the \"!errors\" approach\n> then I will change that.\n\nI am not suggesting to change the RETURN VALUE from this function.\nThat is why I mentioned \"return !errors\" to flip the polarity at the end.\nInside the function, \"ret\" variable _forces_ the readers to think \"this\nfunction unlike the others signal an error with 0\" constantly while\nreading it, and one possible approach to reduce the mental burden\nis to replace \"ret\" variable with \"errors\" variable, which is clear to\nanybody that it would be non-zero when we saw error(s).\n\nOh, I am not suggesting to _count_ the number of errors by\nmentioning a possible variable name \"errors\"; the only reason\nwhy I mentioned that name is because \"error\" is already\ntaken, and \"seen_error\" is a bit too long.\n"},{"id":"293017","messageId":"87BF726E-E91C-4A6A-82D3-1842F806BE98@gmail.com","threadId":"42968","inReplyTo":"CAPc5daWH1z3am2hV_U1dE5WA7R+xrOFxgrxV4CN-vhz6uHz8Hw@mail.gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T22:01:34Z","receivedAt":"2016-08-03T22:01:45Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 23:43, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> On Wed, Aug 3, 2016 at 2:37 PM, Lars Schneider <larsxschneider@gmail.com> wrote:\n>>> \n>>> I think this was already pointed out in the previous review by Peff,\n>>> but a variable \"ret\" that says \"0 is bad\" somehow makes it hard to\n>>> follow the code.  Perhaps rename it to \"int error\", flip the meaning,\n>>> and if the caller wants this function to return non-zero on success\n>>> flip the polarity in the return statement itself, i.e. \"return !errors\",\n>>> may make it easier to follow?\n>> \n>> This follows the existing filter function. Please see Peff's later\n>> reply here:\n> \n> Which I did before mentioning \"pointed out in his review\".\n> \n>> That's why I kept it the way it is. If you prefer the \"!errors\" approach\n>> then I will change that.\n> \n> I am not suggesting to change the RETURN VALUE from this function.\n> That is why I mentioned \"return !errors\" to flip the polarity at the end.\n> Inside the function, \"ret\" variable _forces_ the readers to think \"this\n> function unlike the others signal an error with 0\" constantly while\n> reading it, and one possible approach to reduce the mental burden\n> is to replace \"ret\" variable with \"errors\" variable, which is clear to\n> anybody that it would be non-zero when we saw error(s).\n> \n> Oh, I am not suggesting to _count_ the number of errors by\n> mentioning a possible variable name \"errors\"; the only reason\n> why I mentioned that name is because \"error\" is already\n> taken, and \"seen_error\" is a bit too long.\n\nAgreed. I got that you didn't suggest to change the return value :-)\nIn order to be consistent I would also adjust the error handling in\nthe existing apply_filter() function that I renamed to \napply_single_file_filter() in 11/12. OK?\n\nThanks,\nLars \n\n"},{"id":"293024","messageId":"826967FE-BFF8-4387-83F7-AE7036D97FEC@gmail.com","threadId":"42968","inReplyTo":"20160803212433.zzdino3ivyem5a2v@sigill.intra.peff.net","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T22:15:46Z","receivedAt":"2016-08-03T22:15:54Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 23:24, Jeff King <peff@peff.net> wrote:\n> \n> On Wed, Aug 03, 2016 at 06:42:20PM +0200, larsxschneider@gmail.com wrote:\n> \n>> From: Lars Schneider <larsxschneider@gmail.com>\n>> \n>> Some commands might need to perform cleanup tasks on exit. Let's give\n>> them an interface for doing this.\n>> \n>> Please note, that the cleanup callback is not executed if Git dies of a\n>> signal. The reason is that only \"async-signal-safe\" functions would be\n>> allowed to be call in that case. Since we cannot control what functions\n>> the callback will use, we will not support the case. See 507d7804 for\n>> more details.\n> \n> I'm not clear on why we want this cleanup filter. It looks like you use\n> it in the final patch to send an explicit shutdown to any filters we\n> start. But I see two issues with that:\n> \n>  1. This shutdown may come at any time, and you have no idea what state\n>     the protocol conversation with the filter is in. You could be in\n>     the middle of sending another pkt-line, or in a sequence of non-command\n>     pkt-lines where \"shutdown\" is not recognized.\n\nMaybe I am missing something, but I don't think that can happen because \nthe cleanup callback is *only* executed if Git exits normally without error. \nIn that case we would be in a sane protocol state, no?\n\n\n>  2. If your protocol does bad things when it is cut off in the middle\n>     without an explicit shutdown, then it's a bad protocol. As you\n>     note, this patch doesn't cover signal death, nor could it ever\n>     cover something like \"kill -9\", or a bug which prevented git from\n>     saying \"shutdown\".\n> \n>     You're much better off to design the protocol so that a premature\n>     EOF is detected as an error.  For example, if we're feeding file\n>     data to the filter, and we're worried it might be writing it to\n>     a data store (like LFS), we would not want it to see EOF and say\n>     \"well, I guess I got all the data; time to store this!\". Instead,\n>     it should know how many bytes are coming, or should have some kind\n>     of framing so that the sender says \"and now you have seen all the\n>     bytes\" (like a pkt-line flush).\n> \n>     AFAIK, your protocol _does_ do those things sensibly, so this\n>     explicit shutdown isn't really accomplishing anything.\n\nThanks. The shutdown command is not intended to be a mechanism to tell\nthe filter that everything went well. At this point - as you mentioned -\nthe filter already received all data in the right way. The shutdown\ncommand is intended to give the filter some time to perform some post\nprocessing before Git returns.\n\nSee here for some brainstorming how this feature could be useful\nin filters similar to Git LFS:\nhttps://github.com/github/git-lfs/issues/1401#issuecomment-236133991\n\n- Lars\n\n"},{"id":"293029","messageId":"20160803224619.bwtbvmslhuicx2qi@sigill.intra.peff.net","threadId":"42968","inReplyTo":"0E3FC781-1B2C-4341-9B7B-D9D836596A35@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T22:46:19Z","receivedAt":"2016-08-03T22:47:11Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 03, 2016 at 11:48:00PM +0200, Lars Schneider wrote:\n\n> OK. Is this the v2 discussion you are referring to?\n> http://public-inbox.org/git/1461972887-22100-1-git-send-email-sbeller%40google.com/\n> \n> What format do you suggest?\n> \n> packet:          git< git-filter-protocol\\n\n> packet:          git< version=2\\n\n> packet:          git< capability=clean\\n\n> packet:          git< capability=smudge\\n\n> packet:          git< 0000\n> \n> or\n> \n> packet:          git< git-filter-protocol\\n\n> packet:          git< version=2\\n\n> packet:          git< capability\\n\n> packet:          git< clean\\n\n> packet:          git< smudge\\n\n> packet:          git< 0000\n> \n> or  ... ?\n> \n> I would prefer the first one, I think.\n\nHow about:\n\n  version=2\n  clean=true\n  smudge=true\n  0000\n\n? Then we do not have to care about multiple \"capability\" keys (so\nsomething naively parsing this could just store them in a string list,\nfor example).\n\nYou could also make \"clean\" a synonym for \"clean=true\" or something, and\nhave:\n\n  version=2\n  clean\n  smudge\n  0000\n\nbut it's probably better to have the protocol err on the side of\nverbose-but-unambiguous. It's not like people are typing this routinely.\n\n-Peff\n"},{"id":"293031","messageId":"986d7918-d701-66ba-c04f-4cda56cf598e@gmail.com","threadId":"42968","inReplyTo":"ABE7D2DB-C45F-4F29-8CC2-8D873FD6C36A@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-03T22:47:35Z","receivedAt":"2016-08-03T22:48:51Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Note that some of this might have been invalidated by v4]\n\nW dniu 01.08.2016 o 15:32, Lars Schneider pisze:\n>> On 31 Jul 2016, at 00:05, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n\n>>> Git starts the filter on first usage and expects a welcome\n>>> message, protocol version number, and filter capabilities\n>>> separated by spaces:\n>>> ------------------------\n>>> packet:          git< git-filter-protocol\\n\n>>> packet:          git< version 2\\n\n>>> packet:          git< capabilities clean smudge\\n\n>>\n>> Sorry for going back and forth, but now I think that 'capabilities' are\n>> not really needed here, though they are in line with \"version\" in\n>> the second packet / line, namely \"version 2\".  If it does not make\n>> parsing more difficult...\n> \n> I don't understand what you mean with \"they are not really needed\"?\n> The field is necessary to understand the protocol, no?\n> \n> In the last roll I added the \"key=value\" format to the protocol upon\n> yours and Peff's suggestion. Would it be OK to change the startup\n> sequence accordingly?\n> \n> packet:          git< version=2\\n\n> packet:          git< capabilities=clean smudge\\n\n \nWith current implementation, Git checks second packet of the handshake\nfor version, and third packet for capabilities.  The \"capabilities\" or\n\"capabilities=\" is entirely redundant; it is the position of packet\n(it is the packet number) that matters.  At least for now.\n\nThe only thing that \"version\" in \"version 2\" and \"capabilities\"\nin \"capabilities: clean smudge\" helps is self-describing of the protocol.\n\nTo really make use of them you would have to end handshake with flush\npacket, and do a parsing: loop over every packet, and match known\npatterns.  Well, perhaps with exception of known header: it doesn't\nmakes sense to have \"version N\" anywhere else than second packet,\nand it doesn't makes sense to repeat it.\n\nWe also don't want to proliferate packets unnecessarily.  Each packet\nis a bit (a tiny bit) of a performance hit.\n \n>>> ------------------------\n>>> Supported filter capabilities are \"clean\", \"smudge\", \"stream\",\n>>> and \"shutdown\".\n>>\n>> I'd rather put \"stream\" and \"shutdown\" capabilities into separate\n>> patches, for easier review.\n> \n> I agree with \"shutdown\". I think I would like to remove the \"stream\"\n> option and make it the default for the following reasons:\n> \n> (1) As you mentioned elsewhere, \"stream\" is not really streaming at this\n> point because we don't read/write in parallel.\n\nWe could, following the example of original per-file filter drivers.\nIt is as simple as starting writer using start_async(), as if we did\nwriting from Git in a child process.\n\nThough that might be left for later (assuming that protocol is flexible\nenough), as synchronous protocol (write, then read) is a bit simpler to\nimplement.\n\n> (2) Junio and you pointed out that if we transmit size and flush packet\n> then we have redundancy in the protocol.\n\nProviding size upfront can be a hint for filter or Git.  For example\nHTTP provides Content-Length: header, though it is not strictly necessary.\n\n> (3) With the newly introduced \"success\"/\"reject\"/\"failure\" packet at the \n> end of a filter operation, a filter process has a way to signal Git that\n> something went wrong. Initially I had the idea that a filter process just\n> stops writing and Git would detect the mismatch between expected bytes\n> and received bytes. But the final status packet is a much clearer solution.\n\nThe solution with stopping writing wouldn't work, I don't think.\n\n> (4) Maintaining two slightly different protocols is a waste of resources \n> and only increases the size of this (already large) patch.\n\nRight, better to design and implement basic protocol, taking care that\nit is extensible, and only then add to it.\n\n> My only argument for the size packet was that this allows efficient buffer\n> allocation. However, in non of my benchmarks this was actually a problem.\n> Therefore this is probably a epsilon optimization and should be removed.\n> \n> OK with everyone?\n\nAll right.\n\n>>> After the filter has processed a blob it is expected to wait for\n>>> the next command. A demo implementation can be found in\n>>> `t/t0021/rot13-filter.pl` located in the Git core repository.\n>>\n>> If filter does not support \"shutdown\" capability (or if said\n>> capability is postponed for later patch), it should behave sanely\n>> when Git command reaps it (SIGTERM + wait + SIGKILL?, SIGCHLD?).\n> \n> How would you do this? Don't you think the current solution is\n> good enough for processes that don't need a proper shutdown?\n\nActually... couldn't filter driver register atexit() / signal handler\nto do a clean exit, if it is needed?\n \n \n>> I wonder if it would be worth it to explain the reasoning behind\n>> this solution and show alternate ones.\n>>\n>> * Using a separate variable to signal that filters are invoked\n>>   per-command rather than per-file, and use pkt-line interface,\n>>   like boolean-valued `useProtocol`, or `protocolVersion` set\n>>   to '2' or 'v2', or `persistence` set to 'per-command', there\n>>   is high risk of user's trying to use exiting one-shot per-file\n>>   filters... and Git hanging.\n>>\n>> * Using new variables for each capability, e.g. `processSmudge`\n>>   and `processClean` would lead to explosion of variable names;\n>>   I think.\n>>\n>> * Current solution of using `process` in addition to `clean`\n>>   and `smudge` clearly says that you need to use different\n>>   command for per-file (`clean` and `smudge`), and per-command\n>>   filter, while allowing to use them together.\n>>\n>>   The possible disadvantage is Git command starting `process`\n>>   filter, only to see that it doesn't offer required capability,\n>>   for example offering only \"clean\" but not \"smudge\".  There\n>>   is simple workaround - set `smudge` variable (same as not\n>>   present capability) to empty string.\n> \n> If you think it is necessary to have this discussion in the\n> commit message, then I will add it.\n\nI think it would be good idea (not necessary, but helpful), though\npossibly not in such exacting detail.  Just why this one, one sentence\nor more.\n \n >>> +single filter invocation for the entire life of a single Git\n>>> +command. This is achieved by using the following packet\n>>> +format (pkt-line, see protocol-common.txt) based protocol over\n>>\n>> Can we linkgit-it (to technical documentation)?\n> \n> I don't think that is possible because it was never done. See:\n> git grep \"linkgit:tech\"\n\nA pity.  Well, not your problem, anyway.\n\n>>> +Git starts the filter on first usage and expects a welcome\n>>\n>> Is \"usage\" here correct?  Perhaps it would be more readable\n>> to say that Git starts filter when encountering first file\n>> that needs cleaning or smudgeing.\n> \n> OK. How about this:\n> \n> Git starts the filter when it encounters the first file\n> that needs to be cleaned or smudged. After the filter started\n> Git expects a welcome message, protocol version number, and \n> filter capabilities separated by spaces:\n\nBetter. \n\n>>> +\n>>> +After the filter has processed a blob it is expected to wait for\n>>> +the next command. A demo implementation can be found in\n>>> +`t/t0021/rot13-filter.pl` located in the Git core repository.\n>>\n>> It is actually in Git sources.  Is it the best way to refer to\n>> such files?\n> \n> Well, I could add a github.com link but I don't think everyone\n> would like that. What would you suggest?\n\nSorry, I wasn't clear.  What I meant is if \"<file> located in the\nGit core repository\" is the best way to refer to such files, and\nif we could do better.\n\nBut I think it is all right as it is.\n\nLater we might want to provide some example filter.<driver>.process\nfilters e.g. in contrib/.  But that's for the future.\n \n>>> +\n>>> +Please note that you cannot use an existing filter.<driver>.clean\n>>> +or filter.<driver>.smudge command as filter.<driver>.process\n>>> +command. As soon as Git would detect a file that needs to be\n>>> +processed by this filter, it would stop responding.\n>>\n>> This isn't.\n> \n> Would that be better?\n> \n> \n> Please note that you cannot use an existing `filter.<driver>.clean`\n> or `filter.<driver>.smudge` command as `filter.<driver>.process`\n> command because the former two use a different inter process \n> communication protocol than the latter one. As soon as Git would detect \n> a file that needs to be processed by such an invalid \"process\" filter, \n> it would wait for a proper protocol handshake and appear \"hanging\".\n\nThis is better.\n\n-- \nJakub Narębski\n\n"},{"id":"293033","messageId":"20160803225313.pk3tfe5ovz4y3i7l@sigill.intra.peff.net","threadId":"42968","inReplyTo":"826967FE-BFF8-4387-83F7-AE7036D97FEC@gmail.com","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T22:53:13Z","receivedAt":"2016-08-03T22:53:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Aug 04, 2016 at 12:15:46AM +0200, Lars Schneider wrote:\n\n> > I'm not clear on why we want this cleanup filter. It looks like you use\n> > it in the final patch to send an explicit shutdown to any filters we\n> > start. But I see two issues with that:\n> > \n> >  1. This shutdown may come at any time, and you have no idea what state\n> >     the protocol conversation with the filter is in. You could be in\n> >     the middle of sending another pkt-line, or in a sequence of non-command\n> >     pkt-lines where \"shutdown\" is not recognized.\n> \n> Maybe I am missing something, but I don't think that can happen because \n> the cleanup callback is *only* executed if Git exits normally without error. \n> In that case we would be in a sane protocol state, no?\n\nOK, then maybe I am doubly missing the point. I thought this cleanup was\nhere to hit the case where we call die() and git exits unexpectedly.\n\nIf you only want to cover the \"we are done, no errors, goodbye\" case,\nthen why don't you just write shutdown when we're done?\n\nI realize you may have multiple filters, but I don't think it should be\nrun-command's job to iterate over them. You are presumably keeping a\nlist of active filters, and should have a function to iterate over that.\n\nOr better yet, do not require a shutdown at all. The filter sees EOF and\nknows there is nothing more to do. If we are in the middle of an\noperation, then it knows git died. If not, then presumably git had\nnothing else to say (and really, it is not the filter's business if git\nsaw an error or not).\n\nThough...\n\n> Thanks. The shutdown command is not intended to be a mechanism to tell\n> the filter that everything went well. At this point - as you mentioned -\n> the filter already received all data in the right way. The shutdown\n> command is intended to give the filter some time to perform some post\n> processing before Git returns.\n> \n> See here for some brainstorming how this feature could be useful\n> in filters similar to Git LFS:\n> https://github.com/github/git-lfs/issues/1401#issuecomment-236133991\n\nOK, so it is not really \"tell the filter to shutdown\" but \"I am done\nwith you, filter, but I will wait for you to tell me you are all done,\nso that I can tell the user\".\n\nI'm not sure if calling that \"shutdown\" makes sense, though. It's almost\nmore of a checkpoint (and I wonder if git would ever want to\n\"checkpoint\" without hanging up the connection).\n\n-Peff\n"},{"id":"293044","messageId":"74C2CEA6-EAAB-406F-8B37-969654955413@gmail.com","threadId":"42968","inReplyTo":"20160803225313.pk3tfe5ovz4y3i7l@sigill.intra.peff.net","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-03T23:09:57Z","receivedAt":"2016-08-03T23:10:12Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 00:53, Jeff King <peff@peff.net> wrote:\n> \n> On Thu, Aug 04, 2016 at 12:15:46AM +0200, Lars Schneider wrote:\n> \n>>> I'm not clear on why we want this cleanup filter. It looks like you use\n>>> it in the final patch to send an explicit shutdown to any filters we\n>>> start. But I see two issues with that:\n>>> \n>>> 1. This shutdown may come at any time, and you have no idea what state\n>>>    the protocol conversation with the filter is in. You could be in\n>>>    the middle of sending another pkt-line, or in a sequence of non-command\n>>>    pkt-lines where \"shutdown\" is not recognized.\n>> \n>> Maybe I am missing something, but I don't think that can happen because \n>> the cleanup callback is *only* executed if Git exits normally without error. \n>> In that case we would be in a sane protocol state, no?\n> \n> OK, then maybe I am doubly missing the point. I thought this cleanup was\n> here to hit the case where we call die() and git exits unexpectedly.\n> \n> If you only want to cover the \"we are done, no errors, goodbye\" case,\n> then why don't you just write shutdown when we're done?\n\nI think I tried that at some point but the filter code is called from\nmultiple places and therefore I looked into atexit() (via run-command)\nand it seemed easier. Do you have a place in mind where you would call \nthe shutdown after all blobs are processed explicitly?\n\n\n> I realize you may have multiple filters, but I don't think it should be\n> run-command's job to iterate over them. You are presumably keeping a\n> list of active filters, and should have a function to iterate over that.\n\nYes, that would be easy.\n\n\n> Or better yet, do not require a shutdown at all. The filter sees EOF and\n> knows there is nothing more to do. If we are in the middle of an\n> operation, then it knows git died. If not, then presumably git had\n> nothing else to say (and really, it is not the filter's business if git\n> saw an error or not).\n\nEOF? The filter is supposed to process multiple files. How would one EOF\nindicate that we are done?\n\n\n> Though...\n> \n>> Thanks. The shutdown command is not intended to be a mechanism to tell\n>> the filter that everything went well. At this point - as you mentioned -\n>> the filter already received all data in the right way. The shutdown\n>> command is intended to give the filter some time to perform some post\n>> processing before Git returns.\n>> \n>> See here for some brainstorming how this feature could be useful\n>> in filters similar to Git LFS:\n>> https://github.com/github/git-lfs/issues/1401#issuecomment-236133991\n> \n> OK, so it is not really \"tell the filter to shutdown\" but \"I am done\n> with you, filter, but I will wait for you to tell me you are all done,\n> so that I can tell the user\".\n\nCorrect!\n\n\n> I'm not sure if calling that \"shutdown\" makes sense, though. It's almost\n> more of a checkpoint (and I wonder if git would ever want to\n> \"checkpoint\" without hanging up the connection).\n\nOK, I agree that the naming might not be ideal. But \"checkpoint\" does not\nconvey that it is only executed once after all blobs are filtered?!\n\nI understand that Git might not want to wait for the filter...\n\n- Lars"},{"id":"293051","messageId":"20160803231506.h5mo5lah2pgwdvip@sigill.intra.peff.net","threadId":"42968","inReplyTo":"74C2CEA6-EAAB-406F-8B37-969654955413@gmail.com","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-03T23:15:06Z","receivedAt":"2016-08-03T23:42:15Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Aug 04, 2016 at 01:09:57AM +0200, Lars Schneider wrote:\n\n> > Or better yet, do not require a shutdown at all. The filter sees EOF and\n> > knows there is nothing more to do. If we are in the middle of an\n> > operation, then it knows git died. If not, then presumably git had\n> > nothing else to say (and really, it is not the filter's business if git\n> > saw an error or not).\n> \n> EOF? The filter is supposed to process multiple files. How would one EOF\n> indicate that we are done?\n\nI think we may be talking about two different EOFs.\n\nGit sends a file in pkt-line format, and the flush marks EOF for that\nfile. But the filter keeps running, waiting for more input. This can\nhappen multiple times.\n\nEventually git calls close() on the descriptor, and the filter sees the\n\"real\" EOF (i.e., read() returns 0). That is the signal that git is\ndone.\n\n> > I'm not sure if calling that \"shutdown\" makes sense, though. It's almost\n> > more of a checkpoint (and I wonder if git would ever want to\n> > \"checkpoint\" without hanging up the connection).\n> \n> OK, I agree that the naming might not be ideal. But \"checkpoint\" does not\n> convey that it is only executed once after all blobs are filtered?!\n\nDoes the filter need to care? It's told to do any deferred work, and to\nreport back when it's done. The fact that git is calling it before it\ndecides to exit is not the filter's business (and you can imagine for\nsomething like fast-import, it might want to feed files to something\nlike LFS, too; it already checkpoints occasionally to avoid lost work,\nand would presumably want to ask LFS to checkpoint, too).\n\n> I understand that Git might not want to wait for the filter...\n\nIf git _doesn't_ want to wait for the filter, I don't think you need a\ncheckpoint at all. The filter just does its deferred work when it sees\ngit hang up the connection (i.e., the \"real\" EOF from above).\n\n-Peff\n"},{"id":"293057","messageId":"744bcf80-5d7e-d149-59a3-e12dd40cbea1@gmail.com","threadId":"42968","inReplyTo":"5180D54D-92C4-4875-AEB3-801663D70A8B@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-04T00:42:09Z","receivedAt":"2016-08-04T00:55:46Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Some of those answers might have been invalidated by v4]\n\nW dniu 01.08.2016 o 19:55, Lars Schneider pisze:\n>> On 01 Aug 2016, at 00:19, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n>> [...]\n\n>>> +static int multi_packet_read(int fd_in, struct strbuf *sb, size_t expected_bytes, int is_stream)\n>>\n>> About name of this function: `multi_packet_read` is fine, though I wonder\n>> if `packet_read_in_full` with nearly the same parameters as `packet_read`,\n>> or `packet_read_till_flush`, or `read_in_full_packetized` would be better.\n> \n> I like `multi_packet_read` and will rename!\n \nErrr... what? multi_packet_read() is the current name...\n \n>> Also, the problem is that while we know that what packet_read() stores\n>> would fit in memory (in size_t), it is not true for reading whole file,\n>> which might be very large - for example huge graphical assets like raw\n>> images or raw videos, or virtual machine images.  Isn't that the goal\n>> of git-LFS solutions, which need this feature?  Shouldn't we have then\n>> both `multi_packet_read_to_fd` and `multi_packet_read_to_buf`,\n>> or whatever?\n> \n> Git LFS works well with the current clean/smudge mechanism that uses the\n> same on in memory buffers. I understand your concern but I think this\n> improvement is out of scope for this patch series.\n\nTrue.  \n\nBTW. this means that it cannot share code with fetch / push codebase,\nwhere Git spools from pkt-line to packfile on disk.\n\n\n>> Also, total_bytes_read could overflow size_t, but then we would have\n>> problems storing the result in strbuf.\n> \n> Would that check be ok?\n> \n> \t\tif (total_bytes_read > SIZE_MAX - bytes_read)\n> \t\t\treturn 1;  // `total_bytes_read` would overflow and is not representable\n\nWell, if current code doesn't have such check, then I think it would\nbe all right to not have it either.\n\nNote that we do not use C++ comments.\n \n \n\n>>> +\n>>> +\tif (is_stream)\n>>> +\t\tstrbuf_grow(sb, LARGE_PACKET_MAX);           // allocate space for at least one packet\n>>> +\telse\n>>> +\t\tstrbuf_grow(sb, st_add(expected_bytes, 1));  // add one extra byte for the packet flush\n>>> +\n>>> +\tdo {\n>>> +\t\tbytes_read = packet_read(\n>>> +\t\t\tfd_in, NULL, NULL,\n>>> +\t\t\tsb->buf + total_bytes_read, sb->len - total_bytes_read - 1,\n>>> +\t\t\tPACKET_READ_GENTLE_ON_EOF\n>>> +\t\t);\n>>> +\t\tif (bytes_read < 0)\n>>> +\t\t\treturn 1;  // unexpected EOF\n>>\n>> Don't we usually return negative numbers on error?  Ah, I see that the\n>> return is a bool, which allows to use boolean expression with 'return'.\n>> But I am still unsure if it is good API, this return value.\n> \n> According to Peff zero for success is the usual style:\n> http://public-inbox.org/git/20160728133523.GB21311%40sigill.intra.peff.net/\n\nThe usual case is 0 for success, but -1 (and not 1) for error.\nBut I agree with Peff that keeping existing API is better. \n\n>>> +\t);\n>>> +\tstrbuf_setlen(sb, total_bytes_read);\n>>> +\treturn (is_stream ? 0 : expected_bytes != total_bytes_read);\n>>> +}\n>>> +\n>>> +static int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n>>\n>> Is it equivalent of copy_fd() function, but where destination uses pkt-line\n>> and we need to pack data into pkt-lines?\n> \n> Correct!\n\nYes, and we cannot keep the naming convention.  Though maybe mentioning\nthe equivalence in the comment above function would be good idea...\n\n>>> +\treturn did_fail;\n>>\n>> Return true on fail?  Shouldn't we follow example of copy_fd()\n>> from copy.c, and return COPY_READ_ERROR, or COPY_WRITE_ERROR,\n>> or PKTLINE_WRITE_ERROR?\n> \n> OK. How about this?\n> \n> static int multi_packet_write_from_fd(const int fd_in, const int fd_out)\n> {\n> \tint did_fail = 0;\n> \tssize_t bytes_to_write;\n> \twhile (!did_fail) {\n> \t\tbytes_to_write = xread(fd_in, PKTLINE_DATA_START(packet_buffer), PKTLINE_DATA_MAXLEN);\n> \t\tif (bytes_to_write < 0)\n> \t\t\treturn COPY_READ_ERROR;\n> \t\tif (bytes_to_write == 0)\n> \t\t\tbreak;\n> \t\tdid_fail |= direct_packet_write(fd_out, packet_buffer, PKTLINE_HEADER_LEN + bytes_to_write, 1);\n> \t}\n> \tif (!did_fail)\n> \t\tdid_fail = packet_flush_gently(fd_out);\n> \treturn (did_fail ? COPY_WRITE_ERROR : 0);\n> }\n\nThat's better, I think. \n \n>>> +}\n>>> +\n>>> +static int multi_packet_write_from_buf(const char *src, size_t len, int fd_out)\n>>\n>> It is equivalent of write_in_full(), with different order of parameters,\n>> but where destination file descriptor expects pkt-line and we need to pack\n>> data into pkt-lines?\n> \n> True. Do you suggest to reorder parameters? I also would like to rename `src` to `src_in`, OK?\n\nWell, no need to reorder parameters.  Better keep it the same as for\nother function.  'src' is input ('source'), 'src_in' is tautologic.\n\n>> NOTE: function description comments?\n> \n> What do you mean here?\n\nSorry for being so cryptic.  What I meant is to think about adding comments\ndescribing new functions just above them.\n \n>>  Namely:\n>>\n>> - for git -> filter:\n>>    * read from fd,      write pkt-line to fd  (off_t)\n>>    * read from str+len, write pkt-line to fd  (size_t, ssize_t)\n>> - for filter -> git:\n>>    * read pkt-line from fd, write to fd       (off_t)\n> \n> This one does not exist.\n\nRight, because filter output goes to Git via strbuf.\n \n>>    * read pkt-line from fd, write to str+len  (size_t, ssize_t)\n[...]\n\n>>> +\tstruct child_process process;\n>>> +};\n>>> +\n>>> +static int cmd_process_map_initialized = 0;\n>>> +static struct hashmap cmd_process_map;\n>>\n>> Reading Documentation/technical/api-hashmap.txt I see that:\n>>\n>>  `tablesize` is the allocated size of the hash table. A non-0 value indicates\n>>  that the hashmap is initialized.\n>>\n>> So cmd_process_map_initialized is not really needed, is it?\n> \n> I copied that from config.c:\n> https://github.com/git/git/blob/f8f7adce9fc50a11a764d57815602dcb818d1816/config.c#L1425-L1428\n> \n> `git grep \"tablesize\"` reveals that the check for `tablesize` is only used\n> in hashmap.c ... so what approach should we use?\n\nWell, git code is not always the best example... \n\n>>> +static int apply_protocol2_filter(const char *path, const char *src, size_t len,\n>>> +\t\t\t\t\t\tint fd, struct strbuf *dst, const char *cmd,\n>>> +\t\t\t\t\t\tconst int wanted_capability)\n>>\n[...]\n\n>> This is equivalent to\n>>\n>>   static int apply_filter(const char *path, const char *src, size_t len, int fd,\n>>                           struct strbuf *dst, const char *cmd)\n>>\n>> Could we have extended that one instead?\n> \n> Initially I had one function but that got kind of long ... I prefer two for now.\n\nAll right, we could always refactor to avoid code duplication later. \n \n\n>>> +\n>>> +\tfflush(NULL);\n>>\n>> This is the same as in apply_filter(), but I wonder what it is for.\n> \n> \"If the stream argument is NULL, fflush() flushes all\n>  open output streams.\"\n> \n> http://man7.org/linux/man-pages/man3/fflush.3.html\n\nWhat I wanted to ask was not \"what it does?\",\nbut \"why we need to flush here?\".\n \n>> This is very similar to apply_filter(), but the latter uses start_async()\n>> from \"run-command.h\", with filter_buffer_or_fd() as asynchronous process,\n>> which gets passed command to run in struct filter_params.  In this\n>> function start_protocol2_filter() runs start_command(), synchronous API.\n>>\n>> Why the difference?\n> \n> The protocol V2 requires a sequential processing of the packets. See\n> discussion with Junio here:\n> http://public-inbox.org/git/xmqqbn1th5qn.fsf%40gitster.mtv.corp.google.com/\n\nI don't know what you want to refer to.  The linked email explains\nwhy we fork/start_async() Git process, and the answer was to support\nstreaming.\n\nThere isn't anything there about why protocol v2 requires sequential /\nsynchronous processing of file output, that is write file contents in\nfull, then read, instead of having child write, and Git read and ready\nto read (so filter driver can start writing immediately, and do not need\nto wait for the other ed to stop writing / finish file).\n\nBest regards,\n-- \nJakub Narębski\n\n"},{"id":"293074","messageId":"0b7d7d96-dfdc-54a4-2c24-2aead6743ae1@gmail.com","threadId":"42968","inReplyTo":"9DDA993E-2AFD-4C69-8E22-58601EEC8A40@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-08-04T10:18:33Z","receivedAt":"2016-08-04T10:19:36Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Some of this answer might have been invalidated by v4;\n I might be away from computer for a few days, so I won't be reviewing]\n\nW dniu 03.08.2016 o 15:10, Lars Schneider pisze:\n> On 01 Aug 2016, at 00:19, Jakub Narębski <jnareb@gmail.com> wrote:\n>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:\n[...]\n \n>> Could this whole \"send single file\" be put in a separate function?\n>> Or is it not worth it?\n> \n> This function would have almost the same signature as apply_protocol2_filter\n> and therefore I would say it's not worth it since the function is not\n> crazy long.\n \nAll right.  Though I would say that if it makes the function more\nreadable, then it might be worth it.\n\n[...]\n>>> +\n>>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n>>\n>> Hmmm... ignoring SIGPIPE was good for one-shot filters.  Is it still\n>> O.K. for per-command persistent ones?\n> \n> Very good question. You are right... we don't want to ignore any errors\n> during the protocol... I will remove it.\n\nI was actually just wondering.\n\nActually the default behavior if SIGPIPE is not ignored (or if the\nSIGPIPE signal is not blocked / masked out) is to *terminate* the\nwriting program, which we do not want.\n\nThe correct solution is to check for error during write, and check\nif errno is set to EPIPE.  This means that reader (filter driver\nprocess) has closed pipe, usually due to crash, and we need to handle\nthat sanely, either restarting or quitting while providing sane\ninformation about error to the user.\n\nWell, we might want to set a signal handler for SIGPIPE, not just\nsimply ignore it (especially for streaming case; stop streaming\nif filter driver crashed); though signal handlers are quite limited\nabout what might be done in them.  But that's for the future.\n\n\nRead from closed pipe returns EOF; write to closed pipe results in\nSIGPIPE and returns -1 (setting errno to EPIPE).\n \n>>\n>>> +\n>>> +\tpacket_buf_write(&nbuf, \"%s\\n\", filter_type);\n>>> +\tret &= !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>>> +\n>>> +\tif (ret) {\n>>> +\t\tstrbuf_reset(&nbuf);\n>>> +\t\tpacket_buf_write(&nbuf, \"filename=%s\\n\", path);\n>>> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>>> +\t}\n>>\n>> Perhaps a better solution would be\n>>\n>>        if (err)\n>>        \tgoto fin_error;\n>>\n>> rather than this.\n> \n> OK, I change it to goto error handling style.\n\nWell, at least try it and check if it makes code more readable.\n \n>>> +\tif (ret) {\n>>> +\t\tstrbuf_reset(&nbuf);\n>>> +\t\tpacket_buf_write(&nbuf, \"size=%\"PRIuMAX\"\\n\", (uintmax_t)len);\n>>> +\t\tret = !direct_packet_write(process->in, nbuf.buf, nbuf.len, 1);\n>>> +\t}\n>>\n>> Or maybe extract writing the header for a file into a separate function?\n>> This one gets a bit long...\n> \n> Maybe... but I think that would make it harder to understand the protocol. I\n> think I would prefer to have all the communication in one function layer.\n\nI don't understand your reasoning here (\"make it harder to understand the\nprotocol\").  If you choose good names for function writing header, then\nthe main function would be the high-level view of protocol, e.g.\n\n   git> <command>\n   git> <header>\n   git> <contents>\n   git> <flush>\n\n   git< <command accepted>\n   git< <contents>\n   git< <flush>\n   git< <sent status>\n \n[...]\n>>> +\n>>> +\tif (ret) {\n>>> +\t\tfilter_result = packet_read_line(process->out, NULL);\n>>> +\t\tret = !strcmp(filter_result, \"success\");\n>>> +\t}\n>>> +\n>>> +\tsigchain_pop(SIGPIPE);\n>>> +\n>>> +\tif (ret) {\n>>> +\t\tstrbuf_swap(dst, &nbuf);\n>>> +\t} else {\n>>> +\t\tif (!filter_result || strcmp(filter_result, \"reject\")) {\n>>> +\t\t\t// Something went wrong with the protocol filter. Force shutdown!\n\nDon't use C++ one-line comments (that's C99-ism).\n\n>>> +\t\t\terror(\"external filter '%s' failed\", cmd);\n>>> +\t\t\tkill_protocol2_filter(&cmd_process_map, entry);\n>>> +\t\t}\n>>> +\t}\n>>\n>> So if Git gets finish signal \"success\" from filter, it accepts the output.\n>> If Git gets finish signal \"reject\" from filter, it restarts filter (and\n>> reject the output - user can retry the command himself / herself).\n>> If Git gets any other finish signal, for example \"error\" (but this is not\n>> standarized), then it rejects the output, keeping the unfiltered result,\n>> but keeps filtering.\n>>\n>> I think it is not described in this detail in the documentation of the\n>> new protocol.\n> \n> Agreed, will add!\n\nThat would be nice.\n\n>>> -\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n>>> +\tif (!ca.drv->clean && ca.drv->process)\n>>> +\t\treturn apply_protocol2_filter(\n>>> +\t\t\tpath, NULL, 0, -1, NULL, ca.drv->process, FILTER_CAPABILITIES_CLEAN\n>>> +\t\t);\n>>> +\telse\n>>> +\t\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);\n>>\n>> Could we augment apply_filter() instead, so that the invocation is\n>>\n>>        return apply_filter(path, NULL, 0, -1, NULL, ca.drv, FILTER_CLEAN);\n>>\n>> Though I am not sure if moving this conditional to apply_filter would\n>> be a good idea; maybe wrapper around augmented apply_filter_do()?\n> \n> Yes, a wrapper makes it way cleaner!\n\nThat's good, because we have quite a few of those constructs. \nAnd I think the compiler would inline it, so there is no penalty.\n\n>>> diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\n[...]\n>>> +\t\tgit branch empty &&\n>>> +\n>>> +\t\tcat ../test.o >test.r &&\n>>\n>> Err, the above is just copying file, isn't it?\n>> Maybe it was copied from other tests, I have not checked.\n> \n> It was created in the \"setup\" test.\n \nWhat I meant here (among other things) is that you uselessly use\n'cat' to copy files:\n\n    +\t\tcp ../test.o test.r &&\n \n>>> +\t\techo \"test22\" >test2.r &&\n>>> +\t\tmkdir testsubdir &&\n>>> +\t\techo \"test333\" >testsubdir/test3.r &&\n>>\n>> All right, we test text file, we test binary file (I assume), we test\n>> file in a subdirectory.  What about testing empty file?  Or large file\n>> which would not fit in the stdin/stdout buffer (as EXPENSIVE test)?\n> \n> No binary file. The main reason for this test is to check multiple files.\n> I'll add a empty file. A large file is tested in the next test.\n\nI assume that this large file is binary file; what matters is that it\nincludes NUL character (\"\\0\"), i.e. zero byte, checking that there is\nno error that would terminate it at NUL.\n\nI'll end here for now.\n\n-- \nJakub Narębski\n\n"},{"id":"293098","messageId":"xmqqeg645x6b.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"20160803211221.t2zdhvwjum2baeqs@sigill.intra.peff.net","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-04T16:14:04Z","receivedAt":"2016-08-04T16:14:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> The cost of write() may vary on other platforms, but the cost of memcpy\n> generally shouldn't. So I'm inclined to say that it is not really worth\n> micro-optimizing the interface.\n>\n> I think the other issue is that format_packet() only lets you send\n> string data via \"%s\", so it cannot be used for arbitrary data that may\n> contain NULs. So we do need _some_ other interface to let you send a raw\n> data packet, and it's going to look similar to the direct_packet_write()\n> thing.\n\nOK.  That is a much better argument than \"I already stuff the length\nbytes in my buffer\" (which will invite \"How about stop doing that?\")\nto justify a new \"I have N bytes of data, send it out\", whose\nsignature would look more like write(2) and deserve to be called\npacket_write() but unfortunately the name is taken by what should\nhave called packet_fmt() or something, but that squats on a good\nname packet_write().  Sigh.\n\n\n\n\n\t\n"},{"id":"293099","messageId":"xmqqa8gs5x30.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"20160803213920.jg3eshy57bsldqjh@sigill.intra.peff.net","subject":"Re: [PATCH v4 03/12] pkt-line: add packet_flush_gentle()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-04T16:16:03Z","receivedAt":"2016-08-04T16:16:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>   2. It calls check_pipe(), which will turn EPIPE into death-by-SIGPIPE\n>      (in case you had for some reason ignored SIGPIPE).\n> ...\n>\n> Thinking about (2), I'd go so far as to say that the trace actually\n> should just be using:\n>\n>   if (write_in_full(...) < 0)\n> \twarning(\"unable to write trace to ...: %s\", strerror(errno));\n>\n> and we should get rid of write_or_whine_pipe entirely.\n\nI like the simplicity the above suggestion gives us.\n\n"},{"id":"293180","messageId":"3BA617E0-4DE6-46F6-B148-67265CABE707@gmail.com","threadId":"42968","inReplyTo":"20160802195546.GA2660@atze2.lan","subject":"Re: [PATCH v3 03/10] pkt-line: add packet_flush_gentle()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T09:59:36Z","receivedAt":"2016-08-05T09:59:44Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 02 Aug 2016, at 21:56, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> On Sun, Jul 31, 2016 at 11:45:08PM +0200, Lars Schneider wrote:\n>> \n>>> On 31 Jul 2016, at 22:36, Torstem Bögershausen <tboegi@web.de> wrote:\n>>> \n>>> \n>>> \n>>>> Am 29.07.2016 um 20:37 schrieb larsxschneider@gmail.com:\n>>>> \n>>>> From: Lars Schneider <larsxschneider@gmail.com>\n>>>> \n>>>> packet_flush() would die in case of a write error even though for some callers\n>>>> an error would be acceptable.\n>>> What happens if there is a write error ?\n>>> Basically the protocol is out of synch.\n>>> Lenght information is mixed up with payload, or the other way\n>>> around.\n>>> It may be, that the consequences of a write error are acceptable,\n>>> because a filter is allowed to fail.\n>>> What is not acceptable is a \"broken\" protocol.\n>>> The consequence schould be to close the fd and tear down all\n>>> resources. connected to it.\n>>> In our case to terminate the external filter daemon in some way,\n>>> and to never use this instance again.\n>> \n>> Correct! That is exactly what is happening in kill_protocol2_filter()\n>> here:\n> \n> Wait a second.\n> Is kill the same as shutdown ?\n> I would expect that\n\nNo, kill is used if the filter behaved strangely or signaled an error.\n\"Shutdown\" is a graceful shutdown. However, that might not be an ideal\nname. See the bottom of my discussion with Peff here:\nhttp://public-inbox.org/git/74C2CEA6-EAAB-406F-8B37-969654955413%40gmail.com/\n\n\n> The process terminates itself as soon as it detects EOF.\n> As there is nothing more read.\n> \n> Then the next question: The combination of kill & protocol in kill_protocol(),\n> what does it mean ?\n\nI renamed that function to \"kill_multi_file_filter\". Initially I called\nthe multi file filter \"protocol\" (bad decision I know) and named the\nfunctions accordingly.\n\n\n> Is it more like a graceful shutdown_protocol() ?\n\nYes.\n\nThanks,\nLars"},{"id":"293182","messageId":"FE7B04D2-C017-42AF-BD30-6160D251028C@gmail.com","threadId":"42968","inReplyTo":"607c07fe-5b6f-fd67-13e1-705020c267ee@gmail.com","subject":"Re: Designing the filter process protocol (was: Re: [PATCH v3 10/10] convert: add filter.<driver>.process option)","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T10:32:27Z","receivedAt":"2016-08-05T10:32:37Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 20:30, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> The ultimate goal is to be able to run filter drivers faster for both `clean`\n> and `smudge` operations.  This is done by starting filter driver once per\n> git command invocation, instead of once per file being processed.  Git needs\n> to pass actual contents of files to filter driver, and get its output.\n> \n> We want the protocol between Git and filter driver process to be extensible,\n> so that new features can be added without modifying protocol.\n> \n> \n> 1. CONFIGURATION\n> \n> As I wrote, there are different ways of configuring new-type filter driver:\n> \n> ...\n> \n> # Using a single variable for new filter type, and decide on which phase\n>   (which operation) is supported by filter driver during the handshake\n>   *(current approach)*\n> \n>   \t[filter \"protocol\"]\n>   \t\tprocess = rot13-filtes.pl\n> \n>   PROS: per-file and per-command filters possible with precedence rule;\n>         extensible to other types of drivers: textconv, diff, etc.\n>         only one invocation for commands which use both clean and smudge\n>   CONS: need single driver to be responsible for both clean and smudge;\n>         need to run driver to know that it does not support given\n>           operation (workaround exists)\n> \n> \n> 2. HANDSHAKE (INITIALIZATION)\n> \n> Next, there is deciding on and designing the handshake between Git (between\n> Git command) and the filter driver process.  With the `filter.<driver>.process`\n> solution the driver needs to tell which operations among (for now) \"clean\"\n> and \"smudge\" it does support.  Plus it provides a way to extend protocol,\n> adding new features, like support for streaming, cleaning from file or\n> smudging to file, providing size upfront, perhaps even progress report.\n> \n> Current handshake consist of filter driver printing a signature, version\n> number and capabilities, in that order.  Git checks that it is well formed\n> and matches expectations, and notes which of \"clean\" and \"smudge\" operations\n> are supported by the filter.\n> \n> There is no interaction from the Git side in the handshake, for example to\n> set options and expectations common to all files being filtered.  Take\n> one possible extension of protocol: supporting streaming.  The filter\n> driver needs to know whether it needs to read all the input, or whether\n> it can start printing output while input is incoming (e.g. to reduce\n> memory consumption)... though we may simply decide it to be next version\n> of the protocol.\n> \n> On the other hand if the handshake began with Git sending some initializer\n> info to the filter driver, we probably could detect one-shot filter\n> misconfigured as process-filter.\n\nOK, I'll look into this.\n\n\n> Note that we need some way of deciding where handshake ends, either by\n> specifying number of entries (currently: three lines / pkt-line packets),\n> or providing some terminator (\"smart\" transport protocol uses flush packet\n> for this).\n> \n> ...\n\nWould you be OK with Peff's suggestion?\n\n  version=2\n  clean=true\n  smudge=true\n  0000\n\nhttp://public-inbox.org/git/20160803224619.bwtbvmslhuicx2qi%40sigill.intra.peff.net/\n\n\n\n> 3. SENDING CONTENTS (FILE TO BE FILTERED AND FILTER OUTPUT)\n> \n> Next thing to design is decision how to send contents to be filtered\n> to the filter driver process, and how to get filtered output from the\n> filter driver process.\n> \n> ...\n> \n> # Send/receive data file by file, using some kind of chunking,\n>   with a end-of-file marker.  The solution used by Git is\n>   pkt-line, with flush packet used to signal end of file.\n> \n>   This is protocol used by the current implementation.\n> \n>   PROS:\n>   - no need to know size upfront, so easier streaming support\n>   - you can signal error that happened during output, after\n>     some data were sent, as well as error known upfront\n>   - tracing support for free (GIT_TRACE_PACKET)\n>   CONS:\n>   - filter driver program slightly more difficult to implement\n>   - some negligible amount of overhead\n> \n> If we want in the end to implement streaming, then the last solution\n> is the way to go.\n> \n> \n> 4. PER-FILE HANDSHAKE - SENDING FILE TO FILTER\n> \n> Let's assume that for simplicity we want to implement (for now) only\n> the synchronous (non-streaming) case, where we send whole contents\n> of a file to filter driver process, and *then* read filter driver\n> output.\n> ...\n> \n> If we are using pkt-line, then the convention is that text lines\n> are terminated using LF (\"\\n\") character.  This needs to be stated\n> explicitly in the documentation for filter.<driver>.process writers.\n> \n>    git> packet:  [operation] clean size=67\\n\n> \n> We could denote that it is operation name, but it is obvious from\n> position in the stream, thus not really needed.\n\nI would prefer not to mix command and size in one packet as it\nmakes parsing a little more difficult.\n\n\n> ...\n> \n> The Git would sent contents of the file to be filtered, using\n> as many pack lines as needed (note: large file support needs\n> to be tested, at least as expensive test).  Flush packet is\n> used to signal the end of the file.\n> \n>    git> packets:  <file contents>\n>    git> flush packet\n\nIf expensive tests are enabled the test suite will process data\nlarger then max pkt size.\n\n\n> 5. FILTER DRIVER PROCESS RESPONSE\n> \n> First filter should, in my opinion, reply that it received the\n> request (or the command, in the case of streaming supported).\n> Also, in this response it can provide further information to\n> Git process.\n> \n>    git< packet: [received]  ok size=67\\n\n\nI think this would be different for real streaming and the current\nnon-streaming... therefore it would complicate the protocol?!\nI wonder if it is truly necessary.\n\n\n> This response could be used to refuse to filter specific file\n> upfront (for example if the file is not present in the artifactory\n> for git-LFS solutions).\n> \n>   git< packet: [rejected]  reject\\n\n\nReject is already supported in v4.\n\n\n> We can even provide the reasoning to Git (maybe in the future\n> extension)... or filter driver can print the explanation to the\n> standard error (but then, no --quiet / --verbose support).\n> \n>   git< packet: [rejected]  reject with-message\\n\n>   git< packet: [message]   File not found on server\\n\n>   git< flush packet\n\nI think Git shouldn't care about these details. If the filter\nneeds to tell something then it should use stderr. \n\n\n> Another response, which I think should be standarized, or at\n> least described in the documentation, is filter driver refusing\n> to filter further (e.g. git-LFS and network is down), to be not\n> restarted by Git.\n> \n>   git< packet: [quit]      quit msg=Server error\\n\n> \n> or\n> \n>   git< packet: [quit]      quit Server error\\n\n> \n> or\n> \n>   git< packet: [quit]      quit with-message\\n\n>   git< packet: [message]   Server error\\n\n>   git< flush packet\n> \n> Maybe this is over-engineering, but I don't think so.\n\nInteresting idea! I will look into this for v5!\n\n\n> Next comes the output from the filter driver (filtered contents),\n> using possibly multiple pkt-lines, ending with a flush packet:\n> \n>    git< packets:  <filtered contents>\n>    git< flush packet\n> \n> Note that empty file would consist of zero pack lines of contents,\n> and one flush packet.\n> \n> Finally, to allow handling of [resumable] errors that occurred\n> during sending file contents, especially for the future streaming\n> filters case, we want to confirm that we send whole file\n> successfully.\n> \n>    git< packet: [status]   success\\n\n> \n> If there was an error during process, making data receives so far\n> invalid, filter driver should tell about it\n> \n>    git< packet: [status]   fail\\n\n> \n> or\n> \n>    git< packet: [status]   reject\\n\n> \n> This may happen for example for UCS-2 <-> UTF-8 filter when invalid\n> byte sequence is encountered.  This may happen for git-LFS if the\n> server fails during fetch, and spare / slave server doesn't have\n> a file.\n\nCorrect!\n\n\n> We may want to quit filtering at this point, and not to send another\n> file.\n> \n>   git< packet: [status]    quit\\n\n\nI don't get this one. Git would restart the filter as soon as it finds\nanother file that needs to be filtered, right?\n\n\nThanks a lot for this write up!\n\n- Lars"},{"id":"293183","messageId":"1DDE4AD4-C87F-42BE-AE88-BE302A0D208A@gmail.com","threadId":"42968","inReplyTo":"51080fb7-2bab-a100-0971-e82063c1ac78@gmail.com","subject":"Re: [PATCH v3 01/10] pkt-line: extract set_packet_header()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T11:52:03Z","receivedAt":"2016-08-05T11:52:16Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 22:05, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> [This response might have been invalidated by v4]\n> \n> W dniu 01.08.2016 o 13:33, Lars Schneider pisze: \n>>> On 30 Jul 2016, at 12:30, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n>>>> #define hex(a) (hexchar[(a) & 15])\n>>> \n>>> I guess that this is inherited from the original, but this preprocessor\n>>> macro is local to the format_header() / set_packet_header() function,\n>>> and would not work outside it.  Therefore I think we should #undef it\n>>> after set_packet_header(), just in case somebody mistakes it for\n>>> a generic hex() function.  Perhaps even put it inside set_packet_header(),\n>>> together with #undef.\n>>> \n>>> But I might be mistaken... let's check... no, it isn't used outside it.\n>> \n>> Agreed. Would that be OK?\n>> \n>> static void set_packet_header(char *buf, const int size)\n>> {\n>> \tstatic char hexchar[] = \"0123456789abcdef\";\n>> \t#define hex(a) (hexchar[(a) & 15])\n>> \tbuf[0] = hex(size >> 12);\n>> \tbuf[1] = hex(size >> 8);\n>> \tbuf[2] = hex(size >> 4);\n>> \tbuf[3] = hex(size);\n>> \t#undef hex\n>> }\n> \n> That's better, though I wonder if we need to start #defines at begining\n> of line.  But I think current proposal is O.K.\n> \n> \n> Either this (which has unnecessary larger scope)\n> \n>  #define hex(a) (hexchar[(a) & 15])\n>  static void set_packet_header(char *buf, const int size)\n>  {\n>  \tstatic char hexchar[] = \"0123456789abcdef\";\n> \n>  \tbuf[0] = hex(size >> 12);\n>  \tbuf[1] = hex(size >> 8);\n>  \tbuf[2] = hex(size >> 4);\n>  \tbuf[3] = hex(size);\n>  }\n>  #undef hex\n> \n> or this (which looks worse)\n> \n>  static void set_packet_header(char *buf, const int size)\n>  {\n>  \tstatic char hexchar[] = \"0123456789abcdef\";\n>  #define hex(a) (hexchar[(a) & 15])\n>  \tbuf[0] = hex(size >> 12);\n>  \tbuf[1] = hex(size >> 8);\n>  \tbuf[2] = hex(size >> 4);\n>  \tbuf[3] = hex(size);\n>  #undef hex\n>  }\n> \n\nI probably will drop this patch as Junio is not convinced that it\nis a good idea:\nhttp://public-inbox.org/git/xmqqd1lp8v2o.fsf%40gitster.mtv.corp.google.com/\n\n- Lars\n\n"},{"id":"293185","messageId":"B113DD3B-A8AD-452A-B1E6-A92C84665D65@gmail.com","threadId":"42968","inReplyTo":"e8b550ed-1765-764f-49e5-72e5a609d936@gmail.com","subject":"Re: [PATCH v3 02/10] pkt-line: add direct_packet_write() and direct_packet_write_data()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T12:02:03Z","receivedAt":"2016-08-05T12:02:28Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 22:12, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> [This response might have been invalidated by v4]\n> \n> W dniu 01.08.2016 o 14:00, Lars Schneider pisze:\n>>> On 30 Jul 2016, at 12:49, Jakub Narębski <jnareb@gmail.com> wrote:\n>>> W dniu 30.07.2016 o 01:37, larsxschneider@gmail.com pisze:\n>>>> \n>>>> Sometimes pkt-line data is already available in a buffer and it would\n>>>> be a waste of resources to write the packet using packet_write() which\n>>>> would copy the existing buffer into a strbuf before writing it.\n>>>> \n>>>> If the caller has control over the buffer creation then the\n>>>> PKTLINE_DATA_START macro can be used to skip the header and write\n>>>> directly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\n>>>> would be the maximum). direct_packet_write() would take this buffer,\n>>>> adjust the pkt-line header and write it.\n>>>> \n>>>> If the caller has no control over the buffer creation then\n>>>> direct_packet_write_data() can be used. This function creates a pkt-line\n>>>> header. Afterwards the header and the data buffer are written using two\n>>>> consecutive write calls.\n>>> \n>>> I don't quite understand what do you mean by \"caller has control\n>>> over the buffer creation\".  Do you mean that caller either can write\n>>> over the buffer, or cannot overwrite the buffer?  Or do you mean that\n>>> caller either can allocate buffer to hold header, or is getting\n>>> only the data?\n>> \n>> How about this:\n>> \n>> [...]\n>> \n>> If the caller creates the buffer then a proper pkt-line buffer with header\n>> and data section can be created. The PKTLINE_DATA_START macro can be used \n>> to skip the header section and write directly to the data section (PKTLINE_DATA_LEN \n>> bytes would be the maximum). direct_packet_write() would take this buffer, \n>> fill the pkt-line header section with the appropriate data length value and \n>> write the entire buffer.\n>> \n>> If the caller does not create the buffer, and consequently cannot leave room\n>> for the pkt-line header, then direct_packet_write_data() can be used. This \n>> function creates an extra buffer for the pkt-line header and afterwards writes\n>> the header buffer and the data buffer with two consecutive write calls.\n>> \n>> ---\n>> Is that more clear?\n> \n> Yes, I think it is more clear.  \n> \n> The only thing that could be improved is to perhaps instead of using\n> \n>  \"then a proper pkt-line buffer with header and data section can be created\"\n> \n> it might be more clear to write\n> \n>  \"then a proper pkt-line buffer with data section and a place for pkt-line header\"\n\nOK. I changed it to\n\n\"If the caller has control over the buffer creation then a proper pkt-line\nbuffer with header and data section can be allocated. The \nPKTLINE_DATA_START macro can be used to skip the header and write\ndirectly into the data section of a pkt-line (PKTLINE_DATA_LEN bytes\nwould be the maximum)...\"\n\nHowever, I am not yet sure if I can/will keep this patch:\nhttp://public-inbox.org/git/xmqqeg645x6b.fsf%40gitster.mtv.corp.google.com/\n\n\n> \n>>>> +{\n>>>> +\tint ret = 0;\n>>>> +\tchar hdr[4];\n>>>> +\tset_packet_header(hdr, sizeof(hdr) + size);\n>>>> +\tpacket_trace(buf, size, 1);\n>>>> +\tif (gentle) {\n>>>> +\t\tret = (\n>>>> +\t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n>>> \n>>> You can write '4' here, no need for sizeof(hdr)... though compiler would\n>>> optimize it away.\n>> \n>> Right, it would be optimized. However, I don't like the 4 there either. OK to use a macro\n>> instead? PKTLINE_HEADER_LEN ?\n> \n> Did you mean \n> \n>    +\tchar hdr[PKTLINE_HEADER_LEN];\n>    +\tset_packet_header(hdr, sizeof(hdr) + size);\n\nyes!\n\n\n>>>> +\t\t\t!write_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>>>> +\t\t);\n>>> \n>>> Do we want to try to write \"pkt-line data\" if \"pkt-line header\" failed?\n>>> If not, perhaps De Morgan-ize it\n>>> \n>>> +\t\tret = !(\n>>> +\t\t\twrite_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") &&\n>>> +\t\t\twrite_or_whine_pipe(fd, buf, size, \"pkt-line data\")\n>>> +\t\t);\n>> \n>> \n>> Original:\n>> \t\tret = (\n>> \t\t\t!write_or_whine_pipe(fd, hdr, sizeof(hdr), \"pkt-line header\") ||\n>> \t\t\t!write_or_whine_pipe(fd, data, size, \"pkt-line data\")\n>> \t\t);\n>> \n>> Well, if the first write call fails (return == 0), then it is negated and evaluates to true.\n>> I would think the second call is not evaluated, then?!\n> \n> This is true both for || and for &&, as in C logical boolean operators\n> short-circuit.\n\nTrue. That's why I did not get your \"de morganize\" it comment... what would de morgan change?\n\n> \n>> Should I make this more explicit with a if clause?\n> \n> No need.\n\nOK\n\n\nThanks,\nLars\n"},{"id":"293186","messageId":"C3718D5C-4E98-45AF-8105-BAF77142CDD0@gmail.com","threadId":"42968","inReplyTo":"20160803224619.bwtbvmslhuicx2qi@sigill.intra.peff.net","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T12:53:38Z","receivedAt":"2016-08-05T12:53:45Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 00:46, Jeff King <peff@peff.net> wrote:\n> \n> On Wed, Aug 03, 2016 at 11:48:00PM +0200, Lars Schneider wrote:\n> \n>> OK. Is this the v2 discussion you are referring to?\n>> http://public-inbox.org/git/1461972887-22100-1-git-send-email-sbeller%40google.com/\n>> \n>> What format do you suggest?\n>> \n>> packet:          git< git-filter-protocol\\n\n>> packet:          git< version=2\\n\n>> packet:          git< capability=clean\\n\n>> packet:          git< capability=smudge\\n\n>> packet:          git< 0000\n>> \n>> or\n>> \n>> packet:          git< git-filter-protocol\\n\n>> packet:          git< version=2\\n\n>> packet:          git< capability\\n\n>> packet:          git< clean\\n\n>> packet:          git< smudge\\n\n>> packet:          git< 0000\n>> \n>> or  ... ?\n>> \n>> I would prefer the first one, I think.\n> \n> How about:\n> \n>  version=2\n>  clean=true\n>  smudge=true\n>  0000\n> \n> ? Then we do not have to care about multiple \"capability\" keys (so\n> something naively parsing this could just store them in a string list,\n> for example).\n\nAlright. I will go with this solution.\n\nThanks,\nLars\n\n"},{"id":"293187","messageId":"6C522B0F-F8F7-4B51-8BF0-67D9EDC97B3B@gmail.com","threadId":"42968","inReplyTo":"20160803231506.h5mo5lah2pgwdvip@sigill.intra.peff.net","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T13:08:33Z","receivedAt":"2016-08-05T13:08:40Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 01:15, Jeff King <peff@peff.net> wrote:\n> \n> On Thu, Aug 04, 2016 at 01:09:57AM +0200, Lars Schneider wrote:\n> \n>>> Or better yet, do not require a shutdown at all. The filter sees EOF and\n>>> knows there is nothing more to do. If we are in the middle of an\n>>> operation, then it knows git died. If not, then presumably git had\n>>> nothing else to say (and really, it is not the filter's business if git\n>>> saw an error or not).\n>> \n>> EOF? The filter is supposed to process multiple files. How would one EOF\n>> indicate that we are done?\n> \n> I think we may be talking about two different EOFs.\n> \n> Git sends a file in pkt-line format, and the flush marks EOF for that\n> file. But the filter keeps running, waiting for more input. This can\n> happen multiple times.\n\nCorrect.\n\n> Eventually git calls close() on the descriptor, and the filter sees the\n> \"real\" EOF (i.e., read() returns 0). That is the signal that git is\n> done.\n\nRight.\n\n> \n>>> I'm not sure if calling that \"shutdown\" makes sense, though. It's almost\n>>> more of a checkpoint (and I wonder if git would ever want to\n>>> \"checkpoint\" without hanging up the connection).\n>> \n>> OK, I agree that the naming might not be ideal. But \"checkpoint\" does not\n>> convey that it is only executed once after all blobs are filtered?!\n> \n> Does the filter need to care? It's told to do any deferred work, and to\n> report back when it's done. The fact that git is calling it before it\n> decides to exit is not the filter's business (and you can imagine for\n> something like fast-import, it might want to feed files to something\n> like LFS, too; it already checkpoints occasionally to avoid lost work,\n> and would presumably want to ask LFS to checkpoint, too).\n> \n>> I understand that Git might not want to wait for the filter...\n> \n> If git _doesn't_ want to wait for the filter, I don't think you need a\n> checkpoint at all.\n\nTrue. However, I wonder if it could be useful if the filter is allowed\nto do some finishing work *before* Git returns to the user.\n\n\n> The filter just does its deferred work when it sees\n> git hang up the connection (i.e., the \"real\" EOF from above).\n\nYeah it could do that. But then the filter cannot do things like\nmodifying the index after the fact... however, that might be considered\nnasty by the Git community anyways... I am thinking about dropping\nthis patch in the next roll as it is not strictly necessary for my\ncurrent use case.\n\nThanks,\nLars\n"},{"id":"293188","messageId":"DDF7BA11-A9BC-4DEE-B305-8AD12F5BC599@gmail.com","threadId":"42968","inReplyTo":"0b7d7d96-dfdc-54a4-2c24-2aead6743ae1@gmail.com","subject":"Re: [PATCH v3 10/10] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T13:20:09Z","receivedAt":"2016-08-05T13:20:23Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 12:18, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> ...\n>>>> +\n>>>> +\tsigchain_push(SIGPIPE, SIG_IGN);\n>>> \n>>> Hmmm... ignoring SIGPIPE was good for one-shot filters.  Is it still\n>>> O.K. for per-command persistent ones?\n>> \n>> Very good question. You are right... we don't want to ignore any errors\n>> during the protocol... I will remove it.\n> \n> I was actually just wondering.\n> \n> Actually the default behavior if SIGPIPE is not ignored (or if the\n> SIGPIPE signal is not blocked / masked out) is to *terminate* the\n> writing program, which we do not want.\n> \n> The correct solution is to check for error during write, and check\n> if errno is set to EPIPE.  This means that reader (filter driver\n> process) has closed pipe, usually due to crash, and we need to handle\n> that sanely, either restarting or quitting while providing sane\n> information about error to the user.\n> \n> Well, we might want to set a signal handler for SIGPIPE, not just\n> simply ignore it (especially for streaming case; stop streaming\n> if filter driver crashed); though signal handlers are quite limited\n> about what might be done in them.  But that's for the future.\n> \n> \n> Read from closed pipe returns EOF; write to closed pipe results in\n> SIGPIPE and returns -1 (setting errno to EPIPE).\n\nOK, I think I understand. I will address that in the next round.\n\n\n>>> ...\n>>> Or maybe extract writing the header for a file into a separate function?\n>>> This one gets a bit long...\n>> \n>> Maybe... but I think that would make it harder to understand the protocol. I\n>> think I would prefer to have all the communication in one function layer.\n> \n> I don't understand your reasoning here (\"make it harder to understand the\n> protocol\").  If you choose good names for function writing header, then\n> the main function would be the high-level view of protocol, e.g.\n> \n>   git> <command>\n>   git> <header>\n>   git> <contents>\n>   git> <flush>\n> \n>   git< <command accepted>\n>   git< <contents>\n>   git< <flush>\n>   git< <sent status>\n> \n\nOK, I will move the header into a separate function.\n\n\n>>>> ...\n>>>> +\t\tcat ../test.o >test.r &&\n>>> \n>>> Err, the above is just copying file, isn't it?\n>>> Maybe it was copied from other tests, I have not checked.\n>> \n>> It was created in the \"setup\" test.\n> \n> What I meant here (among other things) is that you uselessly use\n> 'cat' to copy files:\n> \n>    +\t\tcp ../test.o test.r &&\n\nAh right. No idea why I did that. I'll use cp, of course :-)\n\n\n>>>> +\t\techo \"test22\" >test2.r &&\n>>>> +\t\tmkdir testsubdir &&\n>>>> +\t\techo \"test333\" >testsubdir/test3.r &&\n>>> \n>>> All right, we test text file, we test binary file (I assume), we test\n>>> file in a subdirectory.  What about testing empty file?  Or large file\n>>> which would not fit in the stdin/stdout buffer (as EXPENSIVE test)?\n>> \n>> No binary file. The main reason for this test is to check multiple files.\n>> I'll add a empty file. A large file is tested in the next test.\n> \n> I assume that this large file is binary file; what matters is that it\n> includes NUL character (\"\\0\"), i.e. zero byte, checking that there is\n> no error that would terminate it at NUL.\n\nGood idea! I will add a small test file with \\0 bytes in between to test binaries.\n\n\nThanks,\nLars"},{"id":"293192","messageId":"486BA59A-F53B-4893-AE37-8956FCDE7E22@gmail.com","threadId":"42968","inReplyTo":"xmqqeg645x6b.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T14:55:16Z","receivedAt":"2016-08-05T14:55:24Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 18:14, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Jeff King <peff@peff.net> writes:\n> \n>> The cost of write() may vary on other platforms, but the cost of memcpy\n>> generally shouldn't. So I'm inclined to say that it is not really worth\n>> micro-optimizing the interface.\n>> \n>> I think the other issue is that format_packet() only lets you send\n>> string data via \"%s\", so it cannot be used for arbitrary data that may\n>> contain NULs. So we do need _some_ other interface to let you send a raw\n>> data packet, and it's going to look similar to the direct_packet_write()\n>> thing.\n> \n> OK.  That is a much better argument than \"I already stuff the length\n> bytes in my buffer\" (which will invite \"How about stop doing that?\")\n> to justify a new \"I have N bytes of data, send it out\", whose\n> signature would look more like write(2) and deserve to be called\n> packet_write() but unfortunately the name is taken by what should\n> have called packet_fmt() or something, but that squats on a good\n> name packet_write().  Sigh.\n\nWell, my argument wasn't meant to be offensive. It was just an idea that\nI published the to get feedback. Now I understand that it wasn't a particular\ngood idea (thanks Peff for the performance test!).\n\nHowever, besides the bogus performance argument I introduced that function\nto allow packet writs to fail using the `gentle` parameter:\nhttp://public-inbox.org/git/D116610C-F33A-43DA-A49D-0B33958822E5%40gmail.com/\n\nWould you be OK if I introduce packet_write_gently() that returns `0` if the\nwrite was OK and `-1` if it failed?\n\nThanks,\nLars\n"},{"id":"293207","messageId":"xmqqh9azxjmk.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"486BA59A-F53B-4893-AE37-8956FCDE7E22@gmail.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-05T16:31:31Z","receivedAt":"2016-08-05T16:31:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n> However, besides the bogus performance argument I introduced that function\n> to allow packet writs to fail using the `gentle` parameter:\n> http://public-inbox.org/git/D116610C-F33A-43DA-A49D-0B33958822E5%40gmail.com/\n>\n> Would you be OK if I introduce packet_write_gently() that returns `0` if the\n> write was OK and `-1` if it failed?\n\nYes, I agree with you that it would be a good thing to have a\n_gently() variant that lets the caller deal with possible error\nconditions itself instead of dying.\n\nThanks.\n"},{"id":"293209","messageId":"02E1DAD2-8CE7-4A5C-AD28-9E08F2414BDF@gmail.com","threadId":"42968","inReplyTo":"xmqqeg645x6b.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T17:31:48Z","receivedAt":"2016-08-05T17:32:05Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 04 Aug 2016, at 18:14, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Jeff King <peff@peff.net> writes:\n> \n>> The cost of write() may vary on other platforms, but the cost of memcpy\n>> generally shouldn't. So I'm inclined to say that it is not really worth\n>> micro-optimizing the interface.\n>> \n>> I think the other issue is that format_packet() only lets you send\n>> string data via \"%s\", so it cannot be used for arbitrary data that may\n>> contain NULs. So we do need _some_ other interface to let you send a raw\n>> data packet, and it's going to look similar to the direct_packet_write()\n>> thing.\n> \n> OK.  That is a much better argument than \"I already stuff the length\n> bytes in my buffer\" (which will invite \"How about stop doing that?\")\n> to justify a new \"I have N bytes of data, send it out\", whose\n> signature would look more like write(2) and deserve to be called\n> packet_write() but unfortunately the name is taken by what should\n> have called packet_fmt() or something, but that squats on a good\n> name packet_write().  Sigh.\n\n\"Sigh\" means, a series preparation patch that renames \"packet_write()\" \nto \"paket_write_fmt()\" would not be a good idea? It is used 59 times \ncurrently...\n\n- Lars\n\n"},{"id":"293210","messageId":"CAPc5daV3Tke4qHjtpri=6QCRaOax_K3uYhpFzRcd271=GHj1+Q@mail.gmail.com","threadId":"42968","inReplyTo":"02E1DAD2-8CE7-4A5C-AD28-9E08F2414BDF@gmail.com","subject":"Re: [PATCH v4 01/12] pkt-line: extract set_packet_header()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-05T17:41:32Z","receivedAt":"2016-08-05T17:42:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Fri, Aug 5, 2016 at 10:31 AM, Lars Schneider\n<larsxschneider@gmail.com> wrote:\n>\n>> On 04 Aug 2016, at 18:14, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> signature would look more like write(2) and deserve to be called\n>> packet_write() but unfortunately the name is taken by what should\n>> have called packet_fmt() or something, but that squats on a good\n>> name packet_write().  Sigh.\n>\n> \"Sigh\" means, a series preparation patch that renames \"packet_write()\"\n> to \"paket_write_fmt()\" would not be a good idea? It is used 59 times\n> currently...\n\nIt would be a good idea in the longer term, I would think. I just wasn't\nsure if you are willing to volunteer, and in-flight topics will tolerate, such\na change right now. I have a feeling that all the current callsites are\nfairly stable and no in-flight topic touches them, so if you feel like doing\nso, please go ahead ;-)\n"},{"id":"293244","messageId":"b14c6063-cef8-56b4-eb57-7ab8577ecf0a@web.de","threadId":"42968","inReplyTo":"6C522B0F-F8F7-4B51-8BF0-67D9EDC97B3B@gmail.com","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2016-08-05T21:19:07Z","receivedAt":"2016-08-05T21:20:01Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2016-08-05 15.08, Lars Schneider wrote:\n\n[]\n> Yeah it could do that. But then the filter cannot do things like\n> modifying the index after the fact... however, that might be considered\n> nasty by the Git community anyways... I am thinking about dropping\n> this patch in the next roll as it is not strictly necessary for my\n> current use case.\n(Thanks Peff for helping me out with the EOF explanation)\n\nI would say that a filter is a filter, and should do nothing else than filtering\none file,\n(or a stream).\nWhen you want to modify the index, a hook may be your friend.\n\n\n"},{"id":"293253","messageId":"2e13c31c-5ee2-890d-1268-98fb67aba1ea@web.de","threadId":"42968","inReplyTo":"20160803164225.46355-12-larsxschneider@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2016-08-05T21:34:23Z","receivedAt":"2016-08-05T21:35:03Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n> The filter is expected to respond with the result content in zero\n> or more pkt-line packets and a flush packet at the end. Finally, a\n> \"result=success\" packet is expected if everything went well.\n> ------------------------\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< result=success\\n\n> ------------------------\nI would really send the diagnostics/return codes before the content.\n\n> If the result content is empty then the filter is expected to respond\n> only with a flush packet and a \"result=success\" packet.\n> ------------------------\n> packet:          git< 0000\n> packet:          git< result=success\\n\n> ------------------------\n\nWhich may be:\n\npacket:          git< result=success\\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\n\nor for an empty file:\n\npacket:          git< result=success\\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\n\n\nor in case of an error:\npacket:          git< result=reject\\n\n# And this will not send the \"0000\" packet\n\nDoes this makes sense ?\n\n"},{"id":"293258","messageId":"59C5366C-AB41-49D9-8FFF-F109AF242580@gmail.com","threadId":"42968","inReplyTo":"2e13c31c-5ee2-890d-1268-98fb67aba1ea@web.de","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T21:49:55Z","receivedAt":"2016-08-05T21:50:09Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 05 Aug 2016, at 23:34, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n>> The filter is expected to respond with the result content in zero\n>> or more pkt-line packets and a flush packet at the end. Finally, a\n>> \"result=success\" packet is expected if everything went well.\n>> ------------------------\n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n>> packet:          git< result=success\\n\n>> ------------------------\n> I would really send the diagnostics/return codes before the content.\n> \n>> If the result content is empty then the filter is expected to respond\n>> only with a flush packet and a \"result=success\" packet.\n>> ------------------------\n>> packet:          git< 0000\n>> packet:          git< result=success\\n\n>> ------------------------\n> \n> Which may be:\n> \n> packet:          git< result=success\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> \n> or for an empty file:\n> \n> packet:          git< result=success\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n\nI think you meant:\npacket:          git< result=success\\n\npacket:          git< 0000\n\nRight?\n\n> \n> or in case of an error:\n> packet:          git< result=reject\\n\n> # And this will not send the \"0000\" packet\n> \n> Does this makes sense ?\n\nI see your point. However, I think your suggestion would not work in the\ntrue streaming case as the filter wouldn't know upfront if the operation \nwill succeed, right?\n\nThanks for the review,\nLars\n\n"},{"id":"293259","messageId":"ECE58461-10CC-49C2-9BE7-DB80A6746F88@gmail.com","threadId":"42968","inReplyTo":"b14c6063-cef8-56b4-eb57-7ab8577ecf0a@web.de","subject":"Re: [PATCH v4 07/12] run-command: add clean_on_exit_handler","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-05T21:50:38Z","receivedAt":"2016-08-05T21:50:46Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 05 Aug 2016, at 23:19, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> On 2016-08-05 15.08, Lars Schneider wrote:\n> \n> []\n>> Yeah it could do that. But then the filter cannot do things like\n>> modifying the index after the fact... however, that might be considered\n>> nasty by the Git community anyways... I am thinking about dropping\n>> this patch in the next roll as it is not strictly necessary for my\n>> current use case.\n> (Thanks Peff for helping me out with the EOF explanation)\n> \n> I would say that a filter is a filter, and should do nothing else than filtering\n> one file,\n> (or a stream).\n> When you want to modify the index, a hook may be your friend.\n\nAgreed. I will remove that feature.\n\nThanks,\nLars\n\n"},{"id":"293271","messageId":"xmqqfuqivpjv.fsf@gitster.mtv.corp.google.com","threadId":"42968","inReplyTo":"2e13c31c-5ee2-890d-1268-98fb67aba1ea@web.de","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-08-05T22:06:28Z","receivedAt":"2016-08-05T22:06:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n>> The filter is expected to respond with the result content in zero\n>> or more pkt-line packets and a flush packet at the end. Finally, a\n>> \"result=success\" packet is expected if everything went well.\n>> ------------------------\n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n>> packet:          git< result=success\\n\n>> ------------------------\n> I would really send the diagnostics/return codes before the content.\n\nI smell the assumption \"by the time the filter starts output, it\nmust have finished everything and knows both size and the status\".\n\nI'd prefer to have a protocol that allows us to do streaming I/O on\nboth ends when possible, even if the initial version of the filters\n(and the code that sits on the Git side) hold everything in-core\nbefore starting to talk.\n\n>> If the result content is empty then the filter is expected to respond\n>> only with a flush packet and a \"result=success\" packet.\n> ...\n> Which may be:\n>\n> packet:          git< result=success\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n>\n> or for an empty file:\n>\n> packet:          git< result=success\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n\nThe above two look the same to me.\n"},{"id":"293283","messageId":"20160805222710.chefh5kiktyzketh@sigill.intra.peff.net","threadId":"42968","inReplyTo":"xmqqfuqivpjv.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-05T22:27:10Z","receivedAt":"2016-08-06T20:14:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Aug 05, 2016 at 03:06:28PM -0700, Junio C Hamano wrote:\n\n> Torsten Bögershausen <tboegi@web.de> writes:\n> \n> > On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n> >> The filter is expected to respond with the result content in zero\n> >> or more pkt-line packets and a flush packet at the end. Finally, a\n> >> \"result=success\" packet is expected if everything went well.\n> >> ------------------------\n> >> packet:          git< SMUDGED_CONTENT\n> >> packet:          git< 0000\n> >> packet:          git< result=success\\n\n> >> ------------------------\n> > I would really send the diagnostics/return codes before the content.\n> \n> I smell the assumption \"by the time the filter starts output, it\n> must have finished everything and knows both size and the status\".\n> \n> I'd prefer to have a protocol that allows us to do streaming I/O on\n> both ends when possible, even if the initial version of the filters\n> (and the code that sits on the Git side) hold everything in-core\n> before starting to talk.\n\nI think you really want to handle both cases:\n\n  - the server says \"no, I can't fulfill your request\" (e.g., HTTP 404)\n\n  - the server can abort an in-progress response to indicate that it\n    could not be fulfilled completely (in HTTP chunked encoding, this\n    requires hanging up before sending the final EOF chunk)\n\nIf we expect the second case to be rare, then hanging up before sending\nthe flush packet is probably OK. But we could also have a trailing error\ncode after the data to say \"ignore that, we saw an error, but I can\nstill handle more requests\".\n\nIt is true that you don't need the up-front status code in that case\n(you can send an empty body and say \"ignore that, we saw an error\") but\nthat feels a little weird. And I expect it makes the lives of the client\neasier to get a code up front, before it starts taking steps to handle\nwhat it _thinks_ is probably a valid response.\n\n-Peff\n\nPS I haven't followed HTTP/2 development much, but I think it solves the\n   \"hangup\" issue by putting each request/response in its own framed\n   stream. I actually wonder if that is a direction we will want to go\n   eventually, too, or the same reason that HTTP/2 did: multiple async\n   requests across a single connection.\n\n   We already have some precedent in the sideband protocol. So imagine,\n   for example, that we could ask the filter to work on several files\n   simultaneously, by sending\n\n     git> \\1[file1 content]\n     git> \\2[file2 content]\n     git> \\1[file1 content]\n\n   and so on. I don't think this is something that needs to happen in\n   the initial protocol (it's not like git can do parallel checkout\n   right now anyway). If there's a capability negotiation at the front\n   of the protocol, then an async feature can be worked out later. Just\n   food for thought at this point.\n"},{"id":"293285","messageId":"20160806121421.bs7n4lhed7phdshb@sigill.intra.peff.net","threadId":"42968","inReplyTo":"87D4BF17-67BB-4AFA-9B27-40DBB44C0456@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-06T12:14:21Z","receivedAt":"2016-08-06T20:16:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Aug 06, 2016 at 01:55:23PM +0200, Lars Schneider wrote:\n\n> > And I expect it makes the lives of the client\n> > easier to get a code up front, before it starts taking steps to handle\n> > what it _thinks_ is probably a valid response.\n> \n> I am not sure I can follow you here. Which actor are you referring to when\n> you write \"client\" -- Git, right? If the response is rejected right away\n> then Git just needs to read a single flush. If the response experiences\n> an error only later, then the filter wouldn't know about the error when\n> it starts sending. Therefore I don't see how an error code up front could\n> make it easier for Git.\n\nYes, I mean git (I see it as the \"client\" side of the connection in that\nit is making requests of the filter, which will then provide responses).\n\nWhat I mean is that the git code could look something like:\n\n  status == send_filter_request();\n  if (status == OK) {\n\tprepare_storage();\n\tread_response_into_storage();\n  } else {\n\tcomplain();\n  }\n\nBut if there's no status up front, then you probably have:\n\n  send_filter_request();\n  prepare_storage();\n  status = read_response_into_storage();\n  if (status != OK) {\n\trollback_storage();\n\tcomplain();\n  }\n\nIn the first case, we could easily avoid preparing the storage if our\nrequest wasn't going to be filled, whereas in the second we have to do\nit unconditionally. That's not a big deal if preparing the storage is\ninitializing a strbuf. It's more so if you're opening a temporary object\nfile to stream into.\n\nYou _do_ still have to deal with rollback in the first one (for the case\nthat the stream ends prematurely for whatever reason). So it's really a\nquestion of where and how often we expect the failures to come, and\nwhether it is worth git knowing up front that the request is not going\nto be fulfilled.\n\nI dunno. It's not _that_ big a deal to code around. I was just surprised\nnot to see an up-front status when responding to a request. It seems\nlike the normal thing in just about every protocol I've ever used.\n\n-Peff\n"},{"id":"293287","messageId":"90A43E20-B2F5-4377-8DC3-2298BDD7CC71@gmail.com","threadId":"42968","inReplyTo":"607c07fe-5b6f-fd67-13e1-705020c267ee@gmail.com","subject":"Re: Designing the filter process protocol (was: Re: [PATCH v3 10/10] convert: add filter.<driver>.process option)","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-06T18:24:51Z","receivedAt":"2016-08-06T20:19:24Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 03 Aug 2016, at 20:30, Jakub Narębski <jnareb@gmail.com> wrote:\n> \n> ...\n> \n> \n> 2. HANDSHAKE (INITIALIZATION)\n> \n> Next, there is deciding on and designing the handshake between Git (between\n> Git command) and the filter driver process.  With the `filter.<driver>.process`\n> solution the driver needs to tell which operations among (for now) \"clean\"\n> and \"smudge\" it does support.  Plus it provides a way to extend protocol,\n> adding new features, like support for streaming, cleaning from file or\n> smudging to file, providing size upfront, perhaps even progress report.\n> \n> Current handshake consist of filter driver printing a signature, version\n> number and capabilities, in that order.  Git checks that it is well formed\n> and matches expectations, and notes which of \"clean\" and \"smudge\" operations\n> are supported by the filter.\n> \n> There is no interaction from the Git side in the handshake, for example to\n> set options and expectations common to all files being filtered.  Take\n> one possible extension of protocol: supporting streaming.  The filter\n> driver needs to know whether it needs to read all the input, or whether\n> it can start printing output while input is incoming (e.g. to reduce\n> memory consumption)... though we may simply decide it to be next version\n> of the protocol.\n\nI would like to change the startup sequence to this:\n\nGit starts the filter when it encounters the first file\nthat needs to be cleaned or smudged. After the filter started\nGit sends a welcome message, a list of supported protocol\nversion numbers, and a flush packet. Git expects to read the\nwelcome message and one protocol version number from the\npreviously sent list. Afterwards Git sends a list of supported\ncapabilities and a flush packet. Git expects to read a list of\ndesired capabilities, which must be a subset of the supported\ncapabilities list, and a flush packet as response:\n------------------------\npacket:          git> git-filter-client\npacket:          git> version=2\npacket:          git> version=42\npacket:          git> 0000\npacket:          git< git-filter-server\npacket:          git< version=2\npacket:          git> clean=true\npacket:          git> smudge=true\npacket:          git> not-yet-invented=true\npacket:          git> 0000\npacket:          git< clean=true\npacket:          git< smudge=true\npacket:          git< 0000\n------------------------\n\nThis would allow us to detect the case if a user configures an\nexisting clean/smudge filter as `filter.<driver>.process`.\nSince Git is talking first, it would not \"hang\" in that case.\n\nWould that be ok with you?\n\nThanks,\nLars\n\n"},{"id":"293292","messageId":"A07BE78B-5A5D-41F1-A51B-5C71F3E86CCF@gmail.com","threadId":"42968","inReplyTo":"20160806121421.bs7n4lhed7phdshb@sigill.intra.peff.net","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-06T18:19:28Z","receivedAt":"2016-08-06T20:29:40Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 06 Aug 2016, at 14:14, Jeff King <peff@peff.net> wrote:\n> \n> On Sat, Aug 06, 2016 at 01:55:23PM +0200, Lars Schneider wrote:\n> \n>>> And I expect it makes the lives of the client\n>>> easier to get a code up front, before it starts taking steps to handle\n>>> what it _thinks_ is probably a valid response.\n>> \n>> I am not sure I can follow you here. Which actor are you referring to when\n>> you write \"client\" -- Git, right? If the response is rejected right away\n>> then Git just needs to read a single flush. If the response experiences\n>> an error only later, then the filter wouldn't know about the error when\n>> it starts sending. Therefore I don't see how an error code up front could\n>> make it easier for Git.\n> \n> Yes, I mean git (I see it as the \"client\" side of the connection in that\n> it is making requests of the filter, which will then provide responses).\n> \n> What I mean is that the git code could look something like:\n> \n>  status == send_filter_request();\n>  if (status == OK) {\n> \tprepare_storage();\n> \tread_response_into_storage();\n>  } else {\n> \tcomplain();\n>  }\n> \n> But if there's no status up front, then you probably have:\n> \n>  send_filter_request();\n>  prepare_storage();\n>  status = read_response_into_storage();\n>  if (status != OK) {\n> \trollback_storage();\n> \tcomplain();\n>  }\n> \n> In the first case, we could easily avoid preparing the storage if our\n> request wasn't going to be filled, whereas in the second we have to do\n> it unconditionally. That's not a big deal if preparing the storage is\n> initializing a strbuf. It's more so if you're opening a temporary object\n> file to stream into.\n> \n> You _do_ still have to deal with rollback in the first one (for the case\n> that the stream ends prematurely for whatever reason). So it's really a\n> question of where and how often we expect the failures to come, and\n> whether it is worth git knowing up front that the request is not going\n> to be fulfilled.\n> \n> I dunno. It's not _that_ big a deal to code around. I was just surprised\n> not to see an up-front status when responding to a request. It seems\n> like the normal thing in just about every protocol I've ever used.\n\nAlright. The fact that it \"surprised\" you is a bad sign. \nHow about this:\n\nHappy answer:\n------------------------\npacket:          git< status=accept\\n\npacket:          git< SMUDGED_CONTENT\npacket:          git< 0000\npacket:          git< status=success\\n\n------------------------\n\nHappy answer with no content:\n------------------------\npacket:          git< status=success\\n\n------------------------\n\nRejected content:\n------------------------\npacket:          git< status=reject\\n\n------------------------\n\nError during content response:\n------------------------\npacket:          git< status=accept\\n\npacket:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\npacket:          git< 0000\npacket:          git< status=error\\n\n------------------------\n\nCheers,\nLars\n"},{"id":"293293","messageId":"87D4BF17-67BB-4AFA-9B27-40DBB44C0456@gmail.com","threadId":"42968","inReplyTo":"20160805222710.chefh5kiktyzketh@sigill.intra.peff.net","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-06T11:55:23Z","receivedAt":"2016-08-06T20:32:01Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 06 Aug 2016, at 00:27, Jeff King <peff@peff.net> wrote:\n> \n> On Fri, Aug 05, 2016 at 03:06:28PM -0700, Junio C Hamano wrote:\n> \n>> Torsten Bögershausen <tboegi@web.de> writes:\n>> \n>>> On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n>>>> The filter is expected to respond with the result content in zero\n>>>> or more pkt-line packets and a flush packet at the end. Finally, a\n>>>> \"result=success\" packet is expected if everything went well.\n>>>> ------------------------\n>>>> packet:          git< SMUDGED_CONTENT\n>>>> packet:          git< 0000\n>>>> packet:          git< result=success\\n\n>>>> ------------------------\n>>> I would really send the diagnostics/return codes before the content.\n>> \n>> I smell the assumption \"by the time the filter starts output, it\n>> must have finished everything and knows both size and the status\".\n>> \n>> I'd prefer to have a protocol that allows us to do streaming I/O on\n>> both ends when possible, even if the initial version of the filters\n>> (and the code that sits on the Git side) hold everything in-core\n>> before starting to talk.\n> \n> I think you really want to handle both cases:\n> \n>  - the server says \"no, I can't fulfill your request\" (e.g., HTTP 404)\n\nYou can do this with the current protocol:\n\npacket:          git< 0000\npacket:          git< result=reject\\n\n\nAdmittedly the flush packet could be consider overhead but I think\nthat is neglectable.\n\n\n>  - the server can abort an in-progress response to indicate that it\n>    could not be fulfilled completely (in HTTP chunked encoding, this\n>    requires hanging up before sending the final EOF chunk)\n\nAlso already supported with the following sequence:\n\npacket:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\npacket:          git< 0000\npacket:          git< result=error\\n\n\n\n> If we expect the second case to be rare, then hanging up before sending\n> the flush packet is probably OK. But we could also have a trailing error\n> code after the data to say \"ignore that, we saw an error, but I can\n> still handle more requests\".\n> \n> It is true that you don't need the up-front status code in that case\n> (you can send an empty body and say \"ignore that, we saw an error\") but\n> that feels a little weird.\n\nI understand your argument. However, I think \"0000\" indicates \n\"I have nothing for you\" and therefore I think it would be OK in the\nreject case.\n\n\n> And I expect it makes the lives of the client\n> easier to get a code up front, before it starts taking steps to handle\n> what it _thinks_ is probably a valid response.\n\nI am not sure I can follow you here. Which actor are you referring to when\nyou write \"client\" -- Git, right? If the response is rejected right away\nthen Git just needs to read a single flush. If the response experiences\nan error only later, then the filter wouldn't know about the error when\nit starts sending. Therefore I don't see how an error code up front could\nmake it easier for Git.\n\n- Lars\n\n\n"},{"id":"293296","messageId":"526d5219-525b-457b-b533-1721a055b32c@web.de","threadId":"42968","inReplyTo":"xmqqfuqivpjv.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2016-08-06T20:40:16Z","receivedAt":"2016-08-06T20:41:05Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2016-08-06 00.06, Junio C Hamano wrote:\n> Torsten Bögershausen <tboegi@web.de> writes:\n>\n>> On 2016-08-03 18.42, larsxschneider@gmail.com wrote:\n>>> The filter is expected to respond with the result content in zero\n>>> or more pkt-line packets and a flush packet at the end. Finally, a\n>>> \"result=success\" packet is expected if everything went well.\n>>> ------------------------\n>>> packet:          git< SMUDGED_CONTENT\n>>> packet:          git< 0000\n>>> packet:          git< result=success\\n\n>>> ------------------------\n>> I would really send the diagnostics/return codes before the content.\n> I smell the assumption \"by the time the filter starts output, it\n> must have finished everything and knows both size and the status\".\n>\n> I'd prefer to have a protocol that allows us to do streaming I/O on\n> both ends when possible, even if the initial version of the filters\n> (and the code that sits on the Git side) hold everything in-core\n> before starting to talk.\n>\n>>> If the result content is empty then the filter is expected to respond\n>>> only with a flush packet and a \"result=success\" packet.\n>> ...\n>> Which may be:\n>>\n>> packet:          git< result=success\\n\n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n>>\n>> or for an empty file:\n>>\n>> packet:          git< result=success\\n\n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n> The above two look the same to me.\nCopy-paste error.\ni see that we need a status after the complete transfer,\nand after some thinking I would like to take back my comment.\n\n\n"},{"id":"293354","messageId":"20160808150255.2otm3z5fluimpiqw@sigill.intra.peff.net","threadId":"42968","inReplyTo":"A07BE78B-5A5D-41F1-A51B-5C71F3E86CCF@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-08T15:02:55Z","receivedAt":"2016-08-08T15:03:01Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Aug 06, 2016 at 08:19:28PM +0200, Lars Schneider wrote:\n\n> > I dunno. It's not _that_ big a deal to code around. I was just surprised\n> > not to see an up-front status when responding to a request. It seems\n> > like the normal thing in just about every protocol I've ever used.\n> \n> Alright. The fact that it \"surprised\" you is a bad sign. \n> How about this:\n> \n> Happy answer:\n> ------------------------\n> packet:          git< status=accept\\n\n> packet:          git< SMUDGED_CONTENT\n> packet:          git< 0000\n> packet:          git< status=success\\n\n> ------------------------\n\nI notice that the status pkt-lines are by themselves. I had assumed we'd\nbe sending other data, too (presumably before, but I guess possibly\nafter, too). Something like:\n\n  git< status=accept\n  git< 0000\n  git< SMUDGED_CONTENT\n  git< 0000\n  git< status=success\n  git< 0000\n\nI don't have any particular meta-information in mind, but I thought\nstuff like the tentative \"size\" field would be here.\n\nI had imagined it at the front, but I guess it could go in either place.\nI wonder if keys at the end could simply replace ones from the beginning\n(so if you say \"foo=bar\" at the front, that is tentative, but if you\nthen say \"foo=revised\" at the end, that takes precedence).\n\nAnd so the happy answer is really:\n\n  git< status=success\n  git< 0000\n  git< SMUDGED_CONTENT\n  git< 0000\n  git< 0000  # empty list!\n\ni.e., no second status. The original \"success\" still holds.\n\nAnd then:\n\n> Happy answer with no content:\n> ------------------------\n> packet:          git< status=success\\n\n> ------------------------\n\nThis can just be spelled:\n\n  git< status=success\n  git< 0000\n  git< 0000   # empty content!\n  git< 0000   # empty list!\n\n> Rejected content:\n> ------------------------\n> packet:          git< status=reject\\n\n> ------------------------\n\nI'd assume that an error status would end the output for that file\nimmediately, no empty lists necessary (so what you have here). I'd\nprobably just call this \"error\" (see below).\n\n> Error during content response:\n> ------------------------\n> packet:          git< status=accept\\n\n> packet:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\n> packet:          git< 0000\n> packet:          git< status=error\\n\n> ------------------------\n\nAnd then this would be:\n\n  git< status=success\n  git< 0000\n  git< HALF_OF_CONTENT\n  git< 0000\n  git< status=error\n  git< 0000\n\nAnd then you have only two status codes: success and error. Which keeps\nthings simple.\n\nThere's one other case, which is when the filter dies halfway through\nthe conversation, like:\n\n  git< status=success\n  git< 0000\n  git< CONTENT\n  git< 0000\n  ... EOF on pipe ...\n\nAny time git does not get the conversation all the way to the final\nflush after the trailers, it should be considered an error (because we\ncan never know if the filter was about to say \"whoops, status=error\").\n\n-Peff\n"},{"id":"293360","messageId":"6D2101A9-2D01-47E8-9DFF-6C85DED4269D@gmail.com","threadId":"42968","inReplyTo":"20160808150255.2otm3z5fluimpiqw@sigill.intra.peff.net","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-08-08T16:21:18Z","receivedAt":"2016-08-08T16:21:27Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 08 Aug 2016, at 17:02, Jeff King <peff@peff.net> wrote:\n> \n> On Sat, Aug 06, 2016 at 08:19:28PM +0200, Lars Schneider wrote:\n> \n>>> I dunno. It's not _that_ big a deal to code around. I was just surprised\n>>> not to see an up-front status when responding to a request. It seems\n>>> like the normal thing in just about every protocol I've ever used.\n>> \n>> Alright. The fact that it \"surprised\" you is a bad sign. \n>> How about this:\n>> \n>> Happy answer:\n>> ------------------------\n>> packet:          git< status=accept\\n\n>> packet:          git< SMUDGED_CONTENT\n>> packet:          git< 0000\n>> packet:          git< status=success\\n\n>> ------------------------\n> \n> I notice that the status pkt-lines are by themselves. I had assumed we'd\n> be sending other data, too (presumably before, but I guess possibly\n> after, too). Something like:\n> \n>  git< status=accept\n>  git< 0000\n>  git< SMUDGED_CONTENT\n>  git< 0000\n>  git< status=success\n>  git< 0000\n> \n> I don't have any particular meta-information in mind, but I thought\n> stuff like the tentative \"size\" field would be here.\n> \n> I had imagined it at the front, but I guess it could go in either place.\n> I wonder if keys at the end could simply replace ones from the beginning\n> (so if you say \"foo=bar\" at the front, that is tentative, but if you\n> then say \"foo=revised\" at the end, that takes precedence).\n> \n> And so the happy answer is really:\n> \n>  git< status=success\n>  git< 0000\n>  git< SMUDGED_CONTENT\n>  git< 0000\n>  git< 0000  # empty list!\n> \n> i.e., no second status. The original \"success\" still holds.\n\nOK, that sounds sensible to me.\n\n\n> And then:\n> \n>> Happy answer with no content:\n>> ------------------------\n>> packet:          git< status=success\\n\n>> ------------------------\n> \n> This can just be spelled:\n> \n>  git< status=success\n>  git< 0000\n>  git< 0000   # empty content!\n>  git< 0000   # empty list!\n\nIs the first flush packet one too many?\nIf there is nothing then I think we shouldn't\nsend any packets?!\n\nI agree with the remaining two flush packets.\n\n\n>> Rejected content:\n>> ------------------------\n>> packet:          git< status=reject\\n\n>> ------------------------\n> \n> I'd assume that an error status would end the output for that file\n> immediately, no empty lists necessary (so what you have here). I'd\n> probably just call this \"error\" (see below).\n\nOK!\n\n\n>> Error during content response:\n>> ------------------------\n>> packet:          git< status=accept\\n\n>> packet:          git< HALF_WRITTEN_ERRONEOUS_CONTENT\n>> packet:          git< 0000\n>> packet:          git< status=error\\n\n>> ------------------------\n> \n> And then this would be:\n> \n>  git< status=success\n>  git< 0000\n>  git< HALF_OF_CONTENT\n>  git< 0000\n>  git< status=error\n>  git< 0000\n> \n> And then you have only two status codes: success and error. Which keeps\n> things simple.\n> \n> There's one other case, which is when the filter dies halfway through\n> the conversation, like:\n> \n>  git< status=success\n>  git< 0000\n>  git< CONTENT\n>  git< 0000\n>  ... EOF on pipe ...\n> \n> Any time git does not get the conversation all the way to the final\n> flush after the trailers, it should be considered an error (because we\n> can never know if the filter was about to say \"whoops, status=error\").\n\nRight. I agree with the protocol above and I will implement it\nthat way.\n\nThere is one more thing: I introduced a return value \"status=error-all\".\nUsing this the filter can signal Git that it does not want to process\nany other file using the particular command.\n\nJakub came up with this idea here:\n\n\"Another response, which I think should be standarized, or at\nleast described in the documentation, is filter driver refusing\nto filter further (e.g. git-LFS and network is down), to be not\nrestarted by Git.\"\n\nhttp://public-inbox.org/git/607c07fe-5b6f-fd67-13e1-705020c267ee%40gmail.com/\n\nI think it is a good idea. Do you see arguments against it?\n\nThanks,\nLars\n"},{"id":"293361","messageId":"20160808162642.x4k7yjb5fxs2jp25@sigill.intra.peff.net","threadId":"42968","inReplyTo":"6D2101A9-2D01-47E8-9DFF-6C85DED4269D@gmail.com","subject":"Re: [PATCH v4 11/12] convert: add filter.<driver>.process option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-08-08T16:26:43Z","receivedAt":"2016-08-08T16:26:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Aug 08, 2016 at 06:21:18PM +0200, Lars Schneider wrote:\n\n> >> Happy answer with no content:\n> >> ------------------------\n> >> packet:          git< status=success\\n\n> >> ------------------------\n> > \n> > This can just be spelled:\n> > \n> >  git< status=success\n> >  git< 0000\n> >  git< 0000   # empty content!\n> >  git< 0000   # empty list!\n> \n> Is the first flush packet one too many?\n> If there is nothing then I think we shouldn't\n> send any packets?!\n> \n> I agree with the remaining two flush packets.\n\nThere isn't nothing, there is a \"status\" field (though I think that\nshould probably be required, so I guess you could imagine it as a\nstand-alone pkt, separate from the list terminated by the flush). But\nregardless, you need the first flush to say \"I am done telling you\nup-front keys, now I am starting the content\".\n\nOtherwise, what would:\n\n  git< status=success\n  git< foo=bar\n  git< 0000\n\nbe parsed as? Is \"foo=bar\" the first line of content, or the rest of the\npre-content header? (You could guess if you could see the total\nconversation, but you can't; you have to parse it as it comes).\n\n> There is one more thing: I introduced a return value \"status=error-all\".\n> Using this the filter can signal Git that it does not want to process\n> any other file using the particular command.\n> \n> Jakub came up with this idea here:\n> \n> \"Another response, which I think should be standarized, or at\n> least described in the documentation, is filter driver refusing\n> to filter further (e.g. git-LFS and network is down), to be not\n> restarted by Git.\"\n> \n> http://public-inbox.org/git/607c07fe-5b6f-fd67-13e1-705020c267ee%40gmail.com/\n> \n> I think it is a good idea. Do you see arguments against it?\n\nNo, that seems reasonable (I would have just implemented that by hanging\nup the connection, but explicitly communicating is more robust).\n\n-Peff\n"}]}