{"thread":{"id":"45641","subject":"[PATCH v3 1/4] t0021: keep filter log files on comparison","startedAt":"2017-04-09T19:11:24Z","lastAt":"2017-05-21T20:26:07Z","messageCount":18,"participants":["Lars Schneider","Eric Wong","Torsten Bögershausen","Taylor Blau"],"isPatch":true,"patchVersion":3,"patchTotal":4},"messages":[{"id":"316440","messageId":"20170409191107.20547-2-larsxschneider@gmail.com","threadId":"45641","inReplyTo":"20170409191107.20547-1-larsxschneider@gmail.com","subject":"[PATCH v3 1/4] t0021: keep filter log files on comparison","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-09T19:11:04Z","receivedAt":"2017-04-09T19:11:24Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"The filter log files are modified on comparison. Write the modified log files\nto temp files for comparison to fix this.\n\nThis is useful for the subsequent patch 'convert: add \"status=delayed\" to\nfilter process protocol'.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 161f560446..ff2424225b 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -42,10 +42,10 @@ test_cmp_count () {\n \tfor FILE in \"$expect\" \"$actual\"\n \tdo\n \t\tsort \"$FILE\" | uniq -c |\n-\t\tsed -e \"s/^ *[0-9][0-9]*[ \t]*IN: /x IN: /\" >\"$FILE.tmp\" &&\n-\t\tmv \"$FILE.tmp\" \"$FILE\" || return\n+\t\tsed -e \"s/^ *[0-9][0-9]*[ \t]*IN: /x IN: /\" >\"$FILE.tmp\"\n \tdone &&\n-\ttest_cmp \"$expect\" \"$actual\"\n+\ttest_cmp \"$expect.tmp\" \"$actual.tmp\" &&\n+\trm \"$expect.tmp\" \"$actual.tmp\"\n }\n \n # Compare two files but exclude all `clean` invocations because Git can\n@@ -56,10 +56,10 @@ test_cmp_exclude_clean () {\n \tactual=$2\n \tfor FILE in \"$expect\" \"$actual\"\n \tdo\n-\t\tgrep -v \"IN: clean\" \"$FILE\" >\"$FILE.tmp\" &&\n-\t\tmv \"$FILE.tmp\" \"$FILE\"\n+\t\tgrep -v \"IN: clean\" \"$FILE\" >\"$FILE.tmp\"\n \tdone &&\n-\ttest_cmp \"$expect\" \"$actual\"\n+\ttest_cmp \"$expect.tmp\" \"$actual.tmp\" &&\n+\trm \"$expect.tmp\" \"$actual.tmp\"\n }\n \n # Check that the contents of two files are equal and that their rot13 version\n-- \n2.12.2\n\n"},{"id":"316441","messageId":"20170409191107.20547-4-larsxschneider@gmail.com","threadId":"45641","inReplyTo":"20170409191107.20547-1-larsxschneider@gmail.com","subject":"[PATCH v3 3/4] t0021: write \"OUT\" only on success","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-09T19:11:06Z","receivedAt":"2017-04-09T19:11:25Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\"rot13-filter.pl\" used to write \"OUT <size>\" to the debug log even in case of\nan abort or error. Fix this by writing \"OUT <size>\" to the debug log only in\nthe successful case if output is actually written.\n\nThis is useful for the subsequent patch 'convert: add \"status=delayed\" to\nfilter process protocol'.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh   | 6 +++---\n t/t0021/rot13-filter.pl | 6 +++---\n 2 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 0139b460e7..0c04d346a1 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -588,7 +588,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge smudge-write-fail.r $SF [OK] -- OUT: $SF [WRITE FAIL]\n+\t\t\tIN: smudge smudge-write-fail.r $SF [OK] -- [WRITE FAIL]\n \t\t\tSTART\n \t\t\tinit handshake complete\n \t\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n@@ -634,7 +634,7 @@ test_expect_success PERL 'process filter should not be restarted if it signals a\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge error.r $SE [OK] -- OUT: 0 [ERROR]\n+\t\t\tIN: smudge error.r $SE [OK] -- [ERROR]\n \t\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tIN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]\n \t\t\tSTOP\n@@ -673,7 +673,7 @@ test_expect_success PERL 'process filter abort stops processing of all further f\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge abort.r $SA [OK] -- OUT: 0 [ABORT]\n+\t\t\tIN: smudge abort.r $SA [OK] -- [ABORT]\n \t\t\tSTOP\n \t\tEOF\n \t\ttest_cmp_exclude_clean expected.log debug.log &&\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nindex 0b943bb377..5e43faeec1 100644\n--- a/t/t0021/rot13-filter.pl\n+++ b/t/t0021/rot13-filter.pl\n@@ -153,9 +153,6 @@ while (1) {\n \t\tdie \"bad command '$command'\";\n \t}\n \n-\tprint $debug \"OUT: \" . length($output) . \" \";\n-\t$debug->flush();\n-\n \tif ( $pathname eq \"error.r\" ) {\n \t\tprint $debug \"[ERROR]\\n\";\n \t\t$debug->flush();\n@@ -178,6 +175,9 @@ while (1) {\n \t\t\tdie \"${command} write error\";\n \t\t}\n \n+\t\tprint $debug \"OUT: \" . length($output) . \" \";\n+\t\t$debug->flush();\n+\n \t\twhile ( length($output) > 0 ) {\n \t\t\tmy $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n \t\t\tpacket_bin_write($packet);\n-- \n2.12.2\n\n"},{"id":"316442","messageId":"20170409191107.20547-3-larsxschneider@gmail.com","threadId":"45641","inReplyTo":"20170409191107.20547-1-larsxschneider@gmail.com","subject":"[PATCH v3 2/4] t0021: make debug log file name configurable","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-09T19:11:05Z","receivedAt":"2017-04-09T19:11:27Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"The \"rot13-filter.pl\" helper wrote its debug logs always to \"rot13-filter.log\".\nMake this configurable by defining the log file as first parameter of\n\"rot13-filter.pl\".\n\nThis is useful if \"rot13-filter.pl\" is configured multiple times similar to the\nsubsequent patch 'convert: add \"status=delayed\" to filter process protocol'.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n t/t0021-conversion.sh   | 44 ++++++++++++++++++++++----------------------\n t/t0021/rot13-filter.pl |  8 +++++---\n 2 files changed, 27 insertions(+), 25 deletions(-)\n\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex ff2424225b..0139b460e7 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -28,7 +28,7 @@ file_size () {\n }\n \n filter_git () {\n-\trm -f rot13-filter.log &&\n+\trm -f *.log &&\n \tgit \"$@\"\n }\n \n@@ -342,7 +342,7 @@ test_expect_success 'diff does not reuse worktree files that need cleaning' '\n '\n \n test_expect_success PERL 'required process filter should filter data' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -375,7 +375,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: clean testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n \n \t\tgit commit -m \"test commit 2\" &&\n \t\trm -f test2.r \"testsubdir/test3 '\\''sq'\\'',\\$x=.r\" &&\n@@ -388,7 +388,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: smudge testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\tfilter_git checkout --quiet --no-progress empty-branch &&\n \t\tcat >expected.log <<-EOF &&\n@@ -397,7 +397,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\tfilter_git checkout --quiet --no-progress master &&\n \t\tcat >expected.log <<-EOF &&\n@@ -409,7 +409,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: smudge testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -419,7 +419,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \n test_expect_success PERL 'required process filter takes precedence' '\n \ttest_config_global filter.protocol.clean false &&\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -439,12 +439,12 @@ test_expect_success PERL 'required process filter takes precedence' '\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log\n+\t\ttest_cmp_count expected.log debug.log\n \t)\n '\n \n test_expect_success PERL 'required process filter should be used only for \"clean\" operation only' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -462,7 +462,7 @@ test_expect_success PERL 'required process filter should be used only for \"clean\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n \n \t\trm test.r &&\n \n@@ -474,12 +474,12 @@ test_expect_success PERL 'required process filter should be used only for \"clean\n \t\t\tinit handshake complete\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log\n+\t\ttest_cmp_exclude_clean expected.log debug.log\n \t)\n '\n \n test_expect_success PERL 'required process filter should process multiple packets' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n \n \trm -rf repo &&\n@@ -514,7 +514,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n \t\t\tIN: clean 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n \n \t\trm -f *.file &&\n \n@@ -529,7 +529,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n \t\t\tIN: smudge 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\tfor FILE in *.file\n \t\tdo\n@@ -539,7 +539,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n '\n \n test_expect_success PERL 'required process filter with clean error should fail' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -558,7 +558,7 @@ test_expect_success PERL 'required process filter with clean error should fail'\n '\n \n test_expect_success PERL 'process filter should restart after unexpected write failure' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -579,7 +579,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n \t\tgit add . &&\n \t\trm -f *.r &&\n \n-\t\trm -f rot13-filter.log &&\n+\t\trm -f debug.log &&\n \t\tgit checkout --quiet --no-progress . 2>git-stderr.log &&\n \n \t\tgrep \"smudge write error at\" git-stderr.log &&\n@@ -595,7 +595,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n \t\t\tIN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -609,7 +609,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n '\n \n test_expect_success PERL 'process filter should not be restarted if it signals an error' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -639,7 +639,7 @@ test_expect_success PERL 'process filter should not be restarted if it signals a\n \t\t\tIN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -648,7 +648,7 @@ test_expect_success PERL 'process filter should not be restarted if it signals a\n '\n \n test_expect_success PERL 'process filter abort stops processing of all further files' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -676,7 +676,7 @@ test_expect_success PERL 'process filter abort stops processing of all further f\n \t\t\tIN: smudge abort.r $SA [OK] -- OUT: 0 [ABORT]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n \n \t\ttest_cmp \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp \"$TEST_ROOT/test2.o\" test2.r &&\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nindex 617f581e56..0b943bb377 100644\n--- a/t/t0021/rot13-filter.pl\n+++ b/t/t0021/rot13-filter.pl\n@@ -2,8 +2,9 @@\n # Example implementation for the Git filter protocol version 2\n # See Documentation/gitattributes.txt, section \"Filter Protocol\"\n #\n-# The script takes the list of supported protocol capabilities as\n-# arguments (\"clean\", \"smudge\", etc).\n+# The first argument defines a debug log file that the script write to.\n+# All remaining arguments define a list of supported protocol\n+# capabilities (\"clean\", \"smudge\", etc).\n #\n # This implementation supports special test cases:\n # (1) If data with the pathname \"clean-write-fail.r\" is processed with\n@@ -24,9 +25,10 @@ use warnings;\n use IO::File;\n \n my $MAX_PACKET_CONTENT_SIZE = 65516;\n+my $log_file                = shift @ARGV;\n my @capabilities            = @ARGV;\n \n-open my $debug, \">>\", \"rot13-filter.log\" or die \"cannot open log file: $!\";\n+open my $debug, \">>\", $log_file or die \"cannot open log file: $!\";\n \n sub rot13 {\n \tmy $str = shift;\n-- \n2.12.2\n\n"},{"id":"316443","messageId":"20170409191107.20547-1-larsxschneider@gmail.com","threadId":"45641","inReplyTo":null,"subject":"[PATCH v3 0/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-09T19:11:03Z","receivedAt":"2017-04-09T19:11:30Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"Hi,\n\nin v3 \"delay filter\" became a series. Patch 1 to 3 are minor t0021 test\nadjustments and patch 4 is the actual change.\n\nMost significant change since v2:\nIf the filter delays a blob, then Git send the filter a \"delay-id\". Git uses\nthis \"delay-id\" as index in an array of delayed \"cached entries\". When Git\nrequests a previously delayed blob then it will only send the \"delay-id\"\nto identify the blob. The actual blob content will not be send to the filter,\nagain.\n\nIf you review this series then please read the \"Delay\" section in\n\"Documentation/gitattributes.txt\" first for an overview of the delay mechanism.\nThe changes in \"t/t0021/rot13-filter.pl\" are easier to review if you ignore\nwhitespace changes.\n\nThanks,\nLars\n\nRFC: http://public-inbox.org/git/D10F7C47-14E8-465B-8B7A-A09A1B28A39F@gmail.com/\nv1: http://public-inbox.org/git/20170108191736.47359-1-larsxschneider@gmail.com/\nv2: http://public-inbox.org/git/20170226184816.30010-1-larsxschneider@gmail.com/\n\n\nBase Ref: v2.12.0\nWeb-Diff: https://github.com/larsxschneider/git/commit/08a461f103\nCheckout: git fetch https://github.com/larsxschneider/git filter-process/delay-v3 && git checkout 08a461f103\n\n\nInterdiff (v2..v3):\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex f6bad8db40..329baa945f 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -425,8 +425,8 @@ packet:          git< capability=clean\n packet:          git< capability=smudge\n packet:          git< 0000\n ------------------------\n-Supported filter capabilities in version 2 are \"clean\" and\n-\"smudge\".\n+Supported filter capabilities in version 2 are \"clean\", \"smudge\",\n+and \"delay\".\n\n Afterwards Git sends a list of \"key=value\" pairs terminated with\n a flush packet. The list will contain at least the filter command\n@@ -473,15 +473,6 @@ packet:          git< 0000  # empty content!\n packet:          git< 0000  # empty list, keep \"status=success\" unchanged!\n ------------------------\n\n-If the request cannot be fulfilled within a reasonable amount of time\n-then the filter can respond with a \"delayed\" status and a flush packet.\n-Git will perform the same request at a later point in time, again. The\n-filter can delay a response multiple times for a single request.\n-------------------------\n-packet:          git< status=delayed\n-packet:          git< 0000\n-------------------------\n-\n In case the filter cannot or does not want to process the content,\n it is expected to respond with an \"error\" status.\n ------------------------\n@@ -521,12 +512,77 @@ the protocol then Git will stop the filter process and restart it\n with the next file that needs to be processed. Depending on the\n `filter.<driver>.required` flag Git will interpret that as error.\n\n-After the filter has processed a blob it is expected to wait for\n-the next \"key=value\" list containing a command. Git will close\n+After the filter has processed a command it is expected to wait for\n+a \"key=value\" list containing the next command. Git will close\n the command pipe on exit. The filter is expected to detect EOF\n and exit gracefully on its own. Git will wait until the filter\n process has stopped.\n\n+Delay\n+^^^^^\n+\n+If the filter supports the \"delay\" capability, then Git can send the\n+flag \"delay-able\" after the filter command and pathname. This flag\n+denotes that the filter can delay filtering the current blob (e.g. to\n+compensate network latencies) by responding with no content but with\n+the status \"delayed\" and a flush packet. Git will answer with a\n+\"delay-id\", a number that identifies the blob, and a flush packet. The\n+filter acknowledges this number with a \"success\" status and a flush\n+packet.\n+------------------------\n+packet:          git> command=smudge\n+packet:          git> pathname=path/testfile.dat\n+packet:          git> delay-able=1\n+packet:          git> 0000\n+packet:          git> CONTENT\n+packet:          git> 0000\n+packet:          git< status=delayed\n+packet:          git< 0000\n+packet:          git> delay-id=1\n+packet:          git> 0000\n+packet:          git< status=success\n+packet:          git< 0000\n+------------------------\n+\n+If the filter supports the \"delay\" capability then it must support the\n+\"list_available_blobs\" command. If Git sends this command, then the\n+filter is expected to return a list of \"delay_ids\" of blobs that are\n+available. The list must be terminated with a flush packet followed\n+by a \"success\" status that is also terminated with a flush packet. If\n+no blobs for the delayed paths are available, yet, then the filter is\n+expected to block the response until at least one blob becomes\n+available. The filter can tell Git that it has no more delayed blobs\n+by sending an empty list.\n+------------------------\n+packet:          git> command=list_available_blobs\n+packet:          git> 0000\n+packet:          git< 7\n+packet:          git< 13\n+packet:          git< 0000\n+packet:          git< status=success\n+packet:          git< 0000\n+------------------------\n+\n+After Git received the \"delay_ids\", it will request the corresponding\n+blobs again. These requests contain a \"delay-id\" and an empty content\n+section. The filter is expected to respond with the smudged content\n+in the usual way as explained above.\n+------------------------\n+packet:          git> command=smudge\n+packet:          git> pathname=test-delay10.a\n+packet:          git> delay-id=0\n+packet:          git> 0000\n+packet:          git> 0000  # empty content!\n+packet:          git< status=success\n+packet:          git< 0000\n+packet:          git< SMUDGED_CONTENT\n+packet:          git< 0000\n+packet:          git< 0000\n+------------------------\n+\n+Example\n+^^^^^^^\n+\n A long running filter demo implementation can be found in\n `contrib/long-running-filter/example.pl` located in the Git\n core repository. If you develop your own long running filter\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 742e8742cd..e0a0bc92d4 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -355,6 +355,8 @@ static int checkout_paths(const struct checkout_opts *opts,\n \tstate.force = 1;\n \tstate.refresh_cache = 1;\n \tstate.istate = &the_index;\n+\n+\tenable_delayed_checkout(&state);\n \tfor (pos = 0; pos < active_nr; pos++) {\n \t\tstruct cache_entry *ce = active_cache[pos];\n \t\tif (ce->ce_flags & CE_MATCHED) {\n@@ -369,7 +371,7 @@ static int checkout_paths(const struct checkout_opts *opts,\n \t\t\tpos = skip_same_name(ce, pos) - 1;\n \t\t}\n \t}\n-\terrs |= checkout_delayed_entries(&state);\n+\terrs |= finish_delayed_checkout(&state);\n\n \tif (write_locked_index(&the_index, lock_file, COMMIT_LOCK))\n \t\tdie(_(\"unable to write new index file\"));\ndiff --git a/cache.h b/cache.h\nindex 66dde99a79..46076279cf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1421,20 +1421,54 @@ const char *show_ident_date(const struct ident_split *id,\n  */\n extern int ident_cmp(const struct ident_split *, const struct ident_split *);\n\n+enum ce_delay_state {\n+\tCE_DELAY_DISABLED = 0,\n+\tCE_DELAY_AVAILABLE = 1,\n+\tCE_DELAY_APPLIED = 2,\n+\tCE_DELAY_RETRY = 3\n+};\n+\n+struct delayed_checkout {\n+\tenum ce_delay_state state;\n+\t/* The value of \"delay_id\" has different meaning depending on the\n+\t * \"state\" variable:\n+\t *   - CE_DELAY_DISABLED  => \"delay_id\" not used.\n+\t *   - CE_DELAY_AVAILABLE => \"delay_id\" is available to be presented\n+\t *                           to the filter in case the filter wants to\n+\t *                           delay the response of a blob.\n+\t *   - CE_DELAY_APPLIED   => \"delay_id\" was presented to and applied by\n+\t *                           the filter for a blob. The corresponding\n+\t *                           cache entry in stored in the \"entries\"\n+\t *                           array under the index \"delay_id\".\n+\t *   - CE_DELAY_RETRY     => Git requests a blob from the filter that\n+\t *                           was previously delayed using the \"delay_id\".\n+\t */\n+\tint delay_id;\n+\t/* List of filter drivers that have delayed blobs. */\n+\tstruct string_list filters;\n+\t/* Array of cache entries that have been delayed. */\n+\tstruct cache_entry **entries;\n+\tint entries_nr;\n+\tint entries_alloc;\n+};\n+\n struct checkout {\n \tstruct index_state *istate;\n \tconst char *base_dir;\n+\tstruct delayed_checkout *delayed_checkout;\n \tint base_dir_len;\n \tunsigned force:1,\n \t\t quiet:1,\n \t\t not_new:1,\n \t\t refresh_cache:1;\n };\n-#define CHECKOUT_INIT { NULL, \"\" }\n+#define CHECKOUT_INIT { NULL, \"\", NULL }\n+\n\n #define TEMPORARY_FILENAME_LENGTH 25\n extern int checkout_entry(struct cache_entry *ce, const struct checkout *state, char *topath);\n-extern int checkout_delayed_entries(const struct checkout *state);\n+extern void enable_delayed_checkout(struct checkout *state);\n+extern int finish_delayed_checkout(struct checkout *state);\n\n struct cache_def {\n \tstruct strbuf path;\ndiff --git a/convert.c b/convert.c\nindex 24d29f5c53..0d8fa0f833 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -4,7 +4,6 @@\n #include \"quote.h\"\n #include \"sigchain.h\"\n #include \"pkt-line.h\"\n-#include \"list.h\"\n\n /*\n  * convert.c - convert a file when checking it out and checking it in.\n@@ -39,13 +38,6 @@ struct text_stat {\n \tunsigned printable, nonprintable;\n };\n\n-static LIST_HEAD(delayed_item_queue_head);\n-\n-struct delayed_item {\n-\tvoid* item;\n-\tstruct list_head node;\n-};\n-\n static void gather_stats(const char *buf, unsigned long size, struct text_stat *stats)\n {\n \tunsigned long i;\n@@ -503,6 +495,7 @@ static int apply_single_file_filter(const char *path, const char *src, size_t le\n\n #define CAP_CLEAN    (1u<<0)\n #define CAP_SMUDGE   (1u<<1)\n+#define CAP_DELAY    (1u<<2)\n\n struct cmd2process {\n \tstruct hashmap_entry ent; /* must be the first member! */\n@@ -640,7 +633,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \tif (err)\n \t\tgoto done;\n\n-\terr = packet_write_list(process->in, \"capability=clean\", \"capability=smudge\", NULL);\n+\terr = packet_write_list(process->in,\n+\t\t\"capability=clean\", \"capability=smudge\", \"capability=delay\", NULL);\n\n \tfor (;;) {\n \t\tcap_buf = packet_read_line(process->out, NULL);\n@@ -656,6 +650,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \t\t\tentry->supported_capabilities |= CAP_CLEAN;\n \t\t} else if (!strcmp(cap_name, \"smudge\")) {\n \t\t\tentry->supported_capabilities |= CAP_SMUDGE;\n+\t\t} else if (!strcmp(cap_name, \"delay\")) {\n+\t\t\tentry->supported_capabilities |= CAP_DELAY;\n \t\t} else {\n \t\t\twarning(\n \t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n@@ -680,10 +676,12 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n }\n\n static int apply_multi_file_filter(const char *path, const char *src, size_t len,\n-\t\t\t\t   int fd, struct strbuf *dst, int *delayed, const char *cmd,\n-\t\t\t\t   const unsigned int wanted_capability)\n+\t\t\t\t   int fd, struct strbuf *dst, const char *cmd,\n+\t\t\t\t   const unsigned int wanted_capability,\n+\t\t\t\t   struct delayed_checkout *dco)\n {\n \tint err;\n+\tint is_delay_available = 0;\n \tstruct cmd2process *entry;\n \tstruct child_process *process;\n \tstruct strbuf nbuf = STRBUF_INIT;\n@@ -734,6 +732,24 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \tif (err)\n \t\tgoto done;\n\n+\tif (CAP_DELAY & entry->supported_capabilities && dco) {\n+\t\tswitch (dco->state) {\n+\t\tcase CE_DELAY_AVAILABLE:\n+\t\t\tis_delay_available = 1;\n+\t\t\terr = packet_write_fmt_gently(\n+\t\t\t\tprocess->in, \"delay-able=1\\n\");\n+\t\t\tbreak;\n+\t\tcase CE_DELAY_RETRY:\n+\t\t\terr = packet_write_fmt_gently(\n+\t\t\t\tprocess->in, \"delay-id=%i\\n\", dco->delay_id);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tbreak;\n+\t\t}\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t}\n+\n \terr = packet_flush_gently(process->in);\n \tif (err)\n \t\tgoto done;\n@@ -746,18 +762,27 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \t\tgoto done;\n\n \tread_multi_file_filter_status(process->out, &filter_status);\n-\tif (delayed && !strcmp(filter_status.buf, \"delayed\")) {\n-\t\t*delayed = 1;\n+\tif (is_delay_available && !strcmp(filter_status.buf, \"delayed\")) {\n+\t\t/* The filter wants to delay the response. Send it a delay id. */\n+\t\terr = packet_write_fmt_gently(\n+\t\t\tprocess->in, \"delay-id=%i\\n\", dco->delay_id);\n+\t\tif (err)\n \t\t\tgoto done;\n+\t\terr = packet_flush_gently(process->in);\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t\tstring_list_insert(&dco->filters, cmd);\n+\t\tdco->state = CE_DELAY_APPLIED;\n \t} else {\n+\t\t/* The filter got the blob and wants to send us a response. */\n \t\terr = strcmp(filter_status.buf, \"success\");\n \t\tif (err)\n \t\t\tgoto done;\n-\t}\n\n \t\terr = read_packetized_to_strbuf(process->out, &nbuf) < 0;\n \t\tif (err)\n \t\t\tgoto done;\n+\t}\n\n \tread_multi_file_filter_status(process->out, &filter_status);\n \terr = strcmp(filter_status.buf, \"success\");\n@@ -790,6 +815,74 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \treturn !err;\n }\n\n+\n+int async_query_available_blobs(const char *cmd, unsigned long **delay_ids,\n+\t\t\t\tint *delay_ids_nr)\n+{\n+\tint err;\n+\tchar *line;\n+\tchar *end;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tstruct strbuf filter_status = STRBUF_INIT;\n+\tunsigned long delay_id;\n+\tint delay_ids_alloc = 0;\n+\t*delay_ids_nr = 0;\n+\n+\tentry = find_multi_file_filter_entry(&cmd_process_map, cmd);\n+\tif (!entry) {\n+\t\terror(\"external filter '%s' is not available anymore although \"\n+\t\t      \"not all paths have been filtered\", cmd);\n+\t\treturn 0;\n+\t}\n+\tprocess = &entry->process;\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\n+\terr = packet_write_fmt_gently(\n+\t\tprocess->in, \"command=list_available_blobs\\n\");\n+\tif (err)\n+\t\tgoto done;\n+\n+\terr = packet_flush_gently(process->in);\n+\tif (err)\n+\t\tgoto done;\n+\n+\tfor (;;) {\n+\t\tline = packet_read_line(process->out, NULL);\n+\t\tif (!line)\n+\t\t\tbreak;\n+\t\tdelay_id = strtoul(line, &end, 10);\n+\t\terr = (line == end);\n+\t\tif (err) {\n+\t\t\terror(\"invalid delay id '%s'\", line);\n+\t\t\tgoto done;\n+\t\t}\n+\t\tALLOC_GROW(*delay_ids, *delay_ids_nr+1, delay_ids_alloc);\n+\t\t(*delay_ids)[(*delay_ids_nr)++] = delay_id;\n+\t}\n+\n+\tread_multi_file_filter_status(process->out, &filter_status);\n+\terr = strcmp(filter_status.buf, \"success\");\n+\n+done:\n+\tsigchain_pop(SIGPIPE);\n+\n+\tif (err || errno == EPIPE) {\n+\t\tif (!strcmp(filter_status.buf, \"error\")) {\n+\t\t\t/* The filter signaled a problem with the file. */\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * Something went wrong with the protocol filter.\n+\t\t\t * Force shutdown and restart if another blob requires\n+\t\t\t * filtering.\n+\t\t\t */\n+\t\t\terror(\"external filter '%s' failed\", cmd);\n+\t\t\tkill_multi_file_filter(&cmd_process_map, entry);\n+\t\t}\n+\t}\n+\treturn !err;\n+}\n+\n static struct convert_driver {\n \tconst char *name;\n \tstruct convert_driver *next;\n@@ -800,8 +893,9 @@ static struct convert_driver {\n } *user_convert, **user_convert_tail;\n\n static int apply_filter(const char *path, const char *src, size_t len,\n-\t\t\tint fd, struct strbuf *dst, int *delayed,\n-\t\t\tstruct convert_driver *drv, const unsigned int wanted_capability)\n+\t\t\tint fd, struct strbuf *dst, struct convert_driver *drv,\n+\t\t\tconst unsigned int wanted_capability,\n+\t\t\tstruct delayed_checkout *dco)\n {\n \tconst char *cmd = NULL;\n\n@@ -819,7 +913,8 @@ static int apply_filter(const char *path, const char *src, size_t len,\n \tif (cmd && *cmd)\n \t\treturn apply_single_file_filter(path, src, len, fd, dst, cmd);\n \telse if (drv->process && *drv->process)\n-\t\treturn apply_multi_file_filter(path, src, len, fd, dst, delayed, drv->process, wanted_capability);\n+\t\treturn apply_multi_file_filter(path, src, len, fd, dst,\n+\t\t\tdrv->process, wanted_capability, dco);\n\n \treturn 0;\n }\n@@ -1165,7 +1260,7 @@ int would_convert_to_git_filter_fd(const char *path)\n \tif (!ca.drv->required)\n \t\treturn 0;\n\n-\treturn apply_filter(path, NULL, 0, -1, NULL, NULL, ca.drv, CAP_CLEAN);\n+\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv, CAP_CLEAN, NULL);\n }\n\n const char *get_convert_attr_ascii(const char *path)\n@@ -1202,7 +1297,7 @@ int convert_to_git(const char *path, const char *src, size_t len,\n\n \tconvert_attrs(&ca, path);\n\n-\tret |= apply_filter(path, src, len, -1, dst, NULL, ca.drv, CAP_CLEAN);\n+\tret |= apply_filter(path, src, len, -1, dst, ca.drv, CAP_CLEAN, NULL);\n \tif (!ret && ca.drv && ca.drv->required)\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n\n@@ -1227,7 +1322,7 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n \tassert(ca.drv);\n \tassert(ca.drv->clean || ca.drv->process);\n\n-\tif (!apply_filter(path, NULL, 0, fd, dst, NULL, ca.drv, CAP_CLEAN))\n+\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv, CAP_CLEAN, NULL))\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n\n \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n@@ -1235,8 +1330,8 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n }\n\n static int convert_to_working_tree_internal(const char *path, const char *src,\n-\t\t\t\t\t    size_t len, struct strbuf *dst, int *delayed,\n-\t\t\t\t\t    int normalizing)\n+\t\t\t\t\t    size_t len, struct strbuf *dst,\n+\t\t\t\t\t    int normalizing, struct delayed_checkout *dco)\n {\n \tint ret = 0, ret_filter = 0;\n \tstruct conv_attrs ca;\n@@ -1261,7 +1356,8 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t}\n \t}\n\n-\tret_filter = apply_filter(path, src, len, -1, dst, delayed, ca.drv, CAP_SMUDGE);\n+\tret_filter = apply_filter(\n+\t\tpath, src, len, -1, dst, ca.drv, CAP_SMUDGE, dco);\n \tif (!ret_filter && ca.drv && ca.drv->required)\n \t\tdie(\"%s: smudge filter %s failed\", path, ca.drv->name);\n\n@@ -1269,42 +1365,20 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n }\n\n int async_convert_to_working_tree(const char *path, const char *src,\n-\t\t\t\t\t\t\t\t  size_t len, struct strbuf *dst, void *item)\n+\t\t\t\t  size_t len, struct strbuf *dst,\n+\t\t\t\t  void *dco)\n {\n-\tint delayed = 0;\n-\tstruct delayed_item *delayed_item;\n-\tif (convert_to_working_tree_internal(path, src, len, dst, &delayed, 0)) {\n-\t\tif (delayed) {\n-\t\t\tdelayed_item = xmalloc(sizeof(*delayed_item));\n-\t\t\tdelayed_item->item = item;\n-\t\t\tlist_add_tail(&delayed_item->node, &delayed_item_queue_head);\n-\t\t\treturn ASYNC_FILTER_DELAYED;\n-\t\t}\n-\t\treturn ASYNC_FILTER_SUCCESS;\n-\t}\n-\treturn ASYNC_FILTER_FAIL;\n-}\n-\n-void* async_filter_finish(void)\n-{\n-\tstruct delayed_item *head;\n-\tif (!list_empty(&delayed_item_queue_head)) {\n-\t\thead = list_first_entry(&delayed_item_queue_head,\n-\t\t\tstruct delayed_item, node);\n-\t\tlist_del(&head->node);\n-\t\treturn head->item;\n-\t}\n-\treturn NULL;\n+\treturn convert_to_working_tree_internal(path, src, len, dst, 0, dco);\n }\n\n int convert_to_working_tree(const char *path, const char *src, size_t len, struct strbuf *dst)\n {\n-\treturn convert_to_working_tree_internal(path, src, len, dst, NULL, 0);\n+\treturn convert_to_working_tree_internal(path, src, len, dst, 0, NULL);\n }\n\n int renormalize_buffer(const char *path, const char *src, size_t len, struct strbuf *dst)\n {\n-\tint ret = convert_to_working_tree_internal(path, src, len, dst, NULL, 1);\n+\tint ret = convert_to_working_tree_internal(path, src, len, dst, 1, NULL);\n \tif (ret) {\n \t\tsrc = dst->buf;\n \t\tlen = dst->len;\ndiff --git a/convert.h b/convert.h\nindex acc016de9f..da6c702090 100644\n--- a/convert.h\n+++ b/convert.h\n@@ -4,15 +4,6 @@\n #ifndef CONVERT_H\n #define CONVERT_H\n\n-enum async_filter {\n-\tASYNC_FILTER_SUCCESS = 0,\n-\tASYNC_FILTER_FAIL = 1,\n-\tASYNC_FILTER_DELAYED = 2\n-};\n-\n-extern enum async_filter async_filter;\n-\n-\n enum safe_crlf {\n \tSAFE_CRLF_FALSE = 0,\n \tSAFE_CRLF_FAIL = 1,\n@@ -53,8 +44,9 @@ extern int convert_to_working_tree(const char *path, const char *src,\n \t\t\t\t   size_t len, struct strbuf *dst);\n extern int async_convert_to_working_tree(const char *path, const char *src,\n \t\t\t\t\t size_t len, struct strbuf *dst,\n-\t\t\t\t\t void *item);\n-extern void* async_filter_finish(void);\n+\t\t\t\t\t void *dco);\n+extern int async_query_available_blobs(const char *cmd, unsigned long **delay_ids,\n+\t\t\t\t       int *delay_ids_nr);\n extern int renormalize_buffer(const char *path, const char *src, size_t len,\n \t\t\t      struct strbuf *dst);\n static inline int would_convert_to_git(const char *path)\ndiff --git a/entry.c b/entry.c\nindex d15e69a55e..963e94dee8 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -136,6 +136,86 @@ static int streaming_write_entry(const struct cache_entry *ce, char *path,\n \treturn result;\n }\n\n+void enable_delayed_checkout(struct checkout *state)\n+{\n+\tif (!state->delayed_checkout) {\n+\t\tstate->delayed_checkout = xmalloc(sizeof(*state->delayed_checkout));\n+\t\tstate->delayed_checkout->entries_nr = 0;\n+\t\tstate->delayed_checkout->entries_alloc = 0;\n+\t\tstate->delayed_checkout->delay_id = -1;\n+\t\tstate->delayed_checkout->state = CE_DELAY_AVAILABLE;\n+\t\tALLOC_ARRAY(state->delayed_checkout->entries, 0);\n+\t\tstring_list_init(&state->delayed_checkout->filters, 0);\n+\t}\n+}\n+\n+int finish_delayed_checkout(struct checkout *state)\n+{\n+\tint errs = 0;\n+\tstruct string_list_item *filter;\n+\tstruct delayed_checkout *dco = state->delayed_checkout;\n+\n+\tif (!state->delayed_checkout) {\n+\t\treturn errs;\n+\t}\n+\n+\twhile (dco->entries_nr > 0 && dco->filters.nr > 0) {\n+\t\tfor_each_string_list_item(filter, &dco->filters) {\n+\t\t\tint i;\n+\t\t\tint delay_ids_nr;\n+\t\t\tunsigned long *delay_ids;\n+\t\t\tALLOC_ARRAY(delay_ids, 0);\n+\t\t\tif (!async_query_available_blobs(\n+\t\t\t\tfilter->string, &delay_ids, &delay_ids_nr)) {\n+\t\t\t\t/* Filter reported an error */\n+\t\t\t\terrs = 1;\n+\t\t\t\tfilter->string = \"\";\n+\t\t\t\tfree(delay_ids);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (delay_ids_nr <= 0) {\n+\t\t\t\t/* Filter responded with no entries. That means\n+\t\t\t\t   the filter is done and we can remove the\n+\t\t\t\t   filter from the list\n+\t\t\t\t   (see \"string_list_remove_empty_items\" call\n+\t\t\t\t   below).\n+\t\t\t\t*/\n+\t\t\t\tfilter->string = \"\";\n+\t\t\t\tfree(delay_ids);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tfor (i = 0; i < delay_ids_nr; i++) {\n+\t\t\t\tstruct cache_entry* ce;\n+\t\t\t\tunsigned long delay_id = delay_ids[i];\n+\t\t\t\tassert(delay_id >= 0 && delay_id < dco->entries_nr);\n+\t\t\t\tce = dco->entries[delay_id];\n+\t\t\t\tdco->entries[delay_id] = NULL;\n+\t\t\t\tdco->delay_id = delay_id;\n+\t\t\t\tdco->state = CE_DELAY_RETRY;\n+\t\t\t\t/* Shrink entries array as much as possible */\n+\t\t\t\twhile (\n+\t\t\t\t\tdco->entries_nr > 0 &&\n+\t\t\t\t\tdelay_id == dco->entries_nr - 1 &&\n+\t\t\t\t\t!dco->entries[i]\n+\t\t\t\t) {\n+\t\t\t\t\tdelay_id--;\n+\t\t\t\t\tdco->entries_nr--;\n+\t\t\t\t}\n+\t\t\t\terrs |= (ce ? checkout_entry(ce, state, NULL) : 1);\n+\t\t\t}\n+\t\t\tfree(delay_ids);\n+\t\t}\n+\t\tstring_list_remove_empty_items(&dco->filters, 0);\n+\t}\n+\n+\tstring_list_clear(&dco->filters, 0);\n+\tfree(dco->entries);\n+\tfree(dco);\n+\tstate->delayed_checkout = NULL;\n+\n+\treturn errs;\n+}\n+\n static int write_entry(struct cache_entry *ce,\n \t\t       char *path, const struct checkout *state, int to_tempfile)\n {\n@@ -178,16 +258,44 @@ static int write_entry(struct cache_entry *ce,\n \t\t * Convert from git internal format to working tree format\n \t\t */\n \t\tif (ce_mode_s_ifmt == S_IFREG) {\n-\t\t\tret = async_convert_to_working_tree(ce->name, new, size, &buf, ce);\n-\t\t\tif (ret == ASYNC_FILTER_SUCCESS) {\n-\t\t\t\tfree(new);\n-\t\t\t\tnew = strbuf_detach(&buf, &newsize);\n-\t\t\t\tsize = newsize;\n+\t\t\tstruct delayed_checkout *dco = state->delayed_checkout;\n+\t\t\tif (dco && dco->state != CE_DELAY_DISABLED) {\n+\t\t\t\tswitch (dco->state) {\n+\t\t\t\tcase CE_DELAY_AVAILABLE:\n+\t\t\t\t\tdco->delay_id = dco->entries_nr; break;\n+\t\t\t\tcase CE_DELAY_RETRY:\n+\t\t\t\t\tnew = NULL; size = 0; break;\n+\t\t\t\tdefault: break;\n \t\t\t\t}\n-\t\t\telse if (ret == ASYNC_FILTER_DELAYED) {\n+\t\t\t\tassert(dco->delay_id >= 0);\n+\t\t\t\tret = async_convert_to_working_tree(\n+\t\t\t\t\tce->name, new, size, &buf, dco);\n+\t\t\t\tif (ret && dco->state == CE_DELAY_APPLIED) {\n+\t\t\t\t\tassert(dco->delay_id == dco->entries_nr);\n \t\t\t\t\tfree(new);\n+\t\t\t\t\tdco->entries_nr++;\n+\t\t\t\t\tALLOC_GROW(\n+\t\t\t\t\t\tdco->entries, dco->entries_nr,\n+\t\t\t\t\t\tdco->entries_alloc);\n+\t\t\t\t\tdco->entries[dco->delay_id] = ce;\n+\t\t\t\t\tdco->state = CE_DELAY_AVAILABLE;\n+\t\t\t\t\tdco->delay_id = -1;\n \t\t\t\t\tgoto finish;\n \t\t\t\t}\n+\t\t\t} else\n+\t\t\t\tret = convert_to_working_tree(\n+\t\t\t\t\tce->name, new, size, &buf);\n+\n+\t\t\tif (ret) {\n+\t\t\t\tfree(new);\n+\t\t\t\tnew = strbuf_detach(&buf, &newsize);\n+\t\t\t\tsize = newsize;\n+\t\t\t}\n+\t\t\t/*\n+\t\t\t * No \"else\" here as errors from convert are OK at this\n+\t\t\t * point. If the error would have been fatal (e.g.\n+\t\t\t * filter is required), then we would have died already.\n+\t\t\t */\n \t\t}\n\n \t\tfd = open_output_fd(path, ce, to_tempfile);\n@@ -297,16 +405,3 @@ int checkout_entry(struct cache_entry *ce,\n \tcreate_directories(path.buf, path.len, state);\n \treturn write_entry(ce, path.buf, state, 0);\n }\n-\n-int checkout_delayed_entries(const struct checkout *state)\n-{\n-\tstruct cache_entry *ce;\n-\tint errs = 0;\n-\n-\twhile ((ce = async_filter_finish())) {\n-\t\tce->ce_flags &= ~CE_UPDATE;\n-\t\terrs |= checkout_entry(ce, state, NULL);\n-\t}\n-\n-\treturn errs;\n-}\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 8ae5b1a521..21d4cd9453 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -28,7 +28,7 @@ file_size () {\n }\n\n filter_git () {\n-\trm -f rot13-filter.log &&\n+\trm -f *.log &&\n \tgit \"$@\"\n }\n\n@@ -42,10 +42,10 @@ test_cmp_count () {\n \tfor FILE in \"$expect\" \"$actual\"\n \tdo\n \t\tsort \"$FILE\" | uniq -c |\n-\t\tsed -e \"s/^ *[0-9][0-9]*[ \t]*IN: /x IN: /\" >\"$FILE.tmp\" &&\n-\t\tmv \"$FILE.tmp\" \"$FILE\" || return\n+\t\tsed -e \"s/^ *[0-9][0-9]*[ \t]*IN: /x IN: /\" >\"$FILE.tmp\"\n \tdone &&\n-\ttest_cmp \"$expect\" \"$actual\"\n+\ttest_cmp \"$expect.tmp\" \"$actual.tmp\" &&\n+\trm \"$expect.tmp\" \"$actual.tmp\"\n }\n\n # Compare two files but exclude all `clean` invocations because Git can\n@@ -56,10 +56,10 @@ test_cmp_exclude_clean () {\n \tactual=$2\n \tfor FILE in \"$expect\" \"$actual\"\n \tdo\n-\t\tgrep -v \"IN: clean\" \"$FILE\" >\"$FILE.tmp\" &&\n-\t\tmv \"$FILE.tmp\" \"$FILE\"\n+\t\tgrep -v \"IN: clean\" \"$FILE\" >\"$FILE.tmp\"\n \tdone &&\n-\ttest_cmp \"$expect\" \"$actual\"\n+\ttest_cmp \"$expect.tmp\" \"$actual.tmp\" &&\n+\trm \"$expect.tmp\" \"$actual.tmp\"\n }\n\n # Check that the contents of two files are equal and that their rot13 version\n@@ -342,7 +342,7 @@ test_expect_success 'diff does not reuse worktree files that need cleaning' '\n '\n\n test_expect_success PERL 'required process filter should filter data' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -375,7 +375,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: clean testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n\n \t\tgit commit -m \"test commit 2\" &&\n \t\trm -f test2.r \"testsubdir/test3 '\\''sq'\\'',\\$x=.r\" &&\n@@ -388,7 +388,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: smudge testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\tfilter_git checkout --quiet --no-progress empty-branch &&\n \t\tcat >expected.log <<-EOF &&\n@@ -397,7 +397,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\tfilter_git checkout --quiet --no-progress master &&\n \t\tcat >expected.log <<-EOF &&\n@@ -409,7 +409,7 @@ test_expect_success PERL 'required process filter should filter data' '\n \t\t\tIN: smudge testsubdir/test3 '\\''sq'\\'',\\$x=.r $S3 [OK] -- OUT: $S3 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -419,7 +419,7 @@ test_expect_success PERL 'required process filter should filter data' '\n\n test_expect_success PERL 'required process filter takes precedence' '\n \ttest_config_global filter.protocol.clean false &&\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -439,12 +439,12 @@ test_expect_success PERL 'required process filter takes precedence' '\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log\n+\t\ttest_cmp_count expected.log debug.log\n \t)\n '\n\n test_expect_success PERL 'required process filter should be used only for \"clean\" operation only' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -462,7 +462,7 @@ test_expect_success PERL 'required process filter should be used only for \"clean\n \t\t\tIN: clean test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n\n \t\trm test.r &&\n\n@@ -474,12 +474,12 @@ test_expect_success PERL 'required process filter should be used only for \"clean\n \t\t\tinit handshake complete\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log\n+\t\ttest_cmp_exclude_clean expected.log debug.log\n \t)\n '\n\n test_expect_success PERL 'required process filter should process multiple packets' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n\n \trm -rf repo &&\n@@ -514,7 +514,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n \t\t\tIN: clean 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log &&\n+\t\ttest_cmp_count expected.log debug.log &&\n\n \t\trm -f *.file &&\n\n@@ -529,7 +529,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n \t\t\tIN: smudge 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\tfor FILE in *.file\n \t\tdo\n@@ -539,7 +539,7 @@ test_expect_success PERL 'required process filter should process multiple packet\n '\n\n test_expect_success PERL 'required process filter with clean error should fail' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \ttest_config_global filter.protocol.required true &&\n \trm -rf repo &&\n \tmkdir repo &&\n@@ -558,7 +558,7 @@ test_expect_success PERL 'required process filter with clean error should fail'\n '\n\n test_expect_success PERL 'process filter should restart after unexpected write failure' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -579,7 +579,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n \t\tgit add . &&\n \t\trm -f *.r &&\n\n-\t\trm -f rot13-filter.log &&\n+\t\trm -f debug.log &&\n \t\tgit checkout --quiet --no-progress . 2>git-stderr.log &&\n\n \t\tgrep \"smudge write error at\" git-stderr.log &&\n@@ -588,14 +588,14 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge smudge-write-fail.r $SF [OK] -- OUT: $SF [WRITE FAIL]\n+\t\t\tIN: smudge smudge-write-fail.r $SF [OK] -- [WRITE FAIL]\n \t\t\tSTART\n \t\t\tinit handshake complete\n \t\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tIN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -609,7 +609,7 @@ test_expect_success PERL 'process filter should restart after unexpected write f\n '\n\n test_expect_success PERL 'process filter should not be restarted if it signals an error' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -634,12 +634,12 @@ test_expect_success PERL 'process filter should not be restarted if it signals a\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge error.r $SE [OK] -- OUT: 0 [ERROR]\n+\t\t\tIN: smudge error.r $SE [OK] -- [ERROR]\n \t\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n \t\t\tIN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -648,7 +648,7 @@ test_expect_success PERL 'process filter should not be restarted if it signals a\n '\n\n test_expect_success PERL 'process filter abort stops processing of all further files' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n+\ttest_config_global filter.protocol.process \"rot13-filter.pl debug.log clean smudge\" &&\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n@@ -673,10 +673,10 @@ test_expect_success PERL 'process filter abort stops processing of all further f\n \t\tcat >expected.log <<-EOF &&\n \t\t\tSTART\n \t\t\tinit handshake complete\n-\t\t\tIN: smudge abort.r $SA [OK] -- OUT: 0 [ABORT]\n+\t\t\tIN: smudge abort.r $SA [OK] -- [ABORT]\n \t\t\tSTOP\n \t\tEOF\n-\t\ttest_cmp_exclude_clean expected.log rot13-filter.log &&\n+\t\ttest_cmp_exclude_clean expected.log debug.log &&\n\n \t\ttest_cmp \"$TEST_ROOT/test.o\" test.r &&\n \t\ttest_cmp \"$TEST_ROOT/test2.o\" test2.r &&\n@@ -702,55 +702,75 @@ test_expect_success PERL 'invalid process filter must fail (and not hang!)' '\n '\n\n test_expect_success PERL 'delayed checkout in process filter' '\n-\ttest_config_global filter.protocol.process \"rot13-filter.pl clean smudge\" &&\n-\ttest_config_global filter.protocol.required true &&\n+\ttest_config_global filter.a.process \"rot13-filter.pl a.log clean smudge delay\" &&\n+\ttest_config_global filter.a.required true &&\n+\ttest_config_global filter.b.process \"rot13-filter.pl b.log clean smudge delay\" &&\n+\ttest_config_global filter.b.required true &&\n+\n \trm -rf repo &&\n \tmkdir repo &&\n \t(\n \t\tcd repo &&\n \t\tgit init &&\n-\t\techo \"*.r filter=protocol\" >.gitattributes &&\n-\t\tcp \"$TEST_ROOT/test.o\" test.r &&\n-\t\tcp \"$TEST_ROOT/test.o\" test-delay1.r &&\n-\t\tcp \"$TEST_ROOT/test.o\" test-delay3.r &&\n+\t\techo \"*.a filter=a\" >.gitattributes &&\n+\t\techo \"*.b filter=b\" >>.gitattributes &&\n+\t\tcp \"$TEST_ROOT/test.o\" test.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay10.b &&\n \t\tgit add . &&\n \t\tgit commit -m \"test commit 1\"\n \t) &&\n\n-\tS=$(file_size repo/test.r) &&\n-\trm -rf repo-cloned &&\n-\tfilter_git clone repo repo-cloned &&\n-\tcat >expected.log <<-EOF &&\n+\tS=$(file_size \"$TEST_ROOT/test.o\") &&\n+\tcat >a.exp <<-EOF &&\n \t\tSTART\n \t\tinit handshake complete\n-\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n-\t\tIN: smudge test-delay1.r $S [OK] -- OUT: $S [DELAYED]\n-\t\tIN: smudge test-delay1.r $S [OK] -- OUT: $S . [OK]\n-\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test.a $S [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test-delay10.a $S [OK] -- [DELAYED]\n+\t\tIN: smudge test-delay11.a $S [OK] -- [DELAYED]\n+\t\tIN: smudge test-delay20.a $S [OK] -- [DELAYED]\n+\t\tIN: list_available_blobs test-delay10.a test-delay11.a [OK]\n+\t\tIN: smudge test-delay10.a 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test-delay11.a 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: list_available_blobs test-delay20.a [OK]\n+\t\tIN: smudge test-delay20.a 0 [OK] -- OUT: $S . [OK]\n \t\tSTOP\n \tEOF\n-\ttest_cmp_count expected.log repo-cloned/rot13-filter.log &&\n-\n-\t(\n-\t\tcd repo-cloned &&\n-\t\trm *.r rot13-filter.log &&\n-\t\tfilter_git checkout . &&\n-\t\tcat >expected.log <<-EOF &&\n+\tcat >b.exp <<-EOF &&\n \t\tSTART\n \t\tinit handshake complete\n-\t\t\tIN: smudge test.r $S [OK] -- OUT: $S . [OK]\n-\t\t\tIN: smudge test-delay1.r $S [OK] -- OUT: $S [DELAYED]\n-\t\t\tIN: smudge test-delay1.r $S [OK] -- OUT: $S . [OK]\n-\t\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S [DELAYED]\n-\t\t\tIN: smudge test-delay3.r $S [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test-delay10.b $S [OK] -- [DELAYED]\n+\t\tIN: list_available_blobs test-delay10.b [OK]\n+\t\tIN: smudge test-delay10.b 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: list_available_blobs [OK]\n \t\tSTOP\n \tEOF\n-\t\ttest_cmp_count expected.log rot13-filter.log\n+\n+\trm -rf repo-cloned &&\n+\tfilter_git clone repo repo-cloned &&\n+\ttest_cmp_count a.exp repo-cloned/a.log &&\n+\ttest_cmp_count b.exp repo-cloned/b.log &&\n+\n+\t(\n+\t\tcd repo-cloned &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.b &&\n+\n+\t\trm *.a *.b &&\n+\t\tfilter_git checkout . &&\n+\t\ttest_cmp_count ../a.exp a.log &&\n+\t\ttest_cmp_count ../b.exp b.log &&\n+\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.b\n \t)\n '\n\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nindex ece0d314b4..05024c6e1b 100644\n--- a/t/t0021/rot13-filter.pl\n+++ b/t/t0021/rot13-filter.pl\n@@ -2,8 +2,9 @@\n # Example implementation for the Git filter protocol version 2\n # See Documentation/gitattributes.txt, section \"Filter Protocol\"\n #\n-# The script takes the list of supported protocol capabilities as\n-# arguments (\"clean\", \"smudge\", etc).\n+# The first argument defines a debug log file that the script write to.\n+# All remaining arguments define a list of supported protocol\n+# capabilities (\"clean\", \"smudge\", etc).\n #\n # This implementation supports special test cases:\n # (1) If data with the pathname \"clean-write-fail.r\" is processed with\n@@ -18,9 +19,10 @@\n #     to process the file and any file after that is processed with the\n #     same command.\n # (5) If data with a pathname that is a key in the DELAY hash is\n-#     processed (e.g. 'test-delay1.r') then the filter signals n times\n-#     to Git that the processing is delayed (n being the value of the\n-#     DELAY hash key).\n+#     requested (e.g. 'test-delay10.a') then the filter responds with\n+#     a \"delay\" status and sets the \"requested\" field in the DELAY hash.\n+#     The filter will signal the availability of this object after\n+#     \"count\" (field in DELAY hash) \"list_available_blobs\" commands.\n #\n\n use strict;\n@@ -28,15 +30,17 @@ use warnings;\n use IO::File;\n\n my $MAX_PACKET_CONTENT_SIZE = 65516;\n+my $log_file                = shift @ARGV;\n my @capabilities            = @ARGV;\n-my $DELAY3 = 3;\n-my $DELAY1 = 1;\n\n-my %DELAY;\n-$DELAY{'test-delay1.r'} = 1;\n-$DELAY{'test-delay3.r'} = 3;\n+open my $debug, \">>\", $log_file or die \"cannot open log file: $!\";\n\n-open my $debug, \">>\", \"rot13-filter.log\" or die \"cannot open log file: $!\";\n+my %DELAY = (\n+\t'test-delay10.a' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+\t'test-delay11.a' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+\t'test-delay20.a' => { \"requested\" => 0, \"count\" => 2, \"delay_id\" => -1 },\n+\t'test-delay10.b' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+);\n\n sub rot13 {\n \tmy $str = shift;\n@@ -74,7 +78,7 @@ sub packet_bin_read {\n\n sub packet_txt_read {\n \tmy ( $res, $buf ) = packet_bin_read();\n-\tunless ( $buf =~ s/\\n$// ) {\n+\tunless ( $buf eq '' or $buf =~ s/\\n$// ) {\n \t\tdie \"A non-binary line MUST be terminated by an LF.\";\n \t}\n \treturn ( $res, $buf );\n@@ -109,6 +113,7 @@ packet_flush();\n\n ( packet_txt_read() eq ( 0, \"capability=clean\" ) )  || die \"bad capability\";\n ( packet_txt_read() eq ( 0, \"capability=smudge\" ) ) || die \"bad capability\";\n+( packet_txt_read() eq ( 0, \"capability=delay\" ) )  || die \"bad capability\";\n ( packet_bin_read() eq ( 1, \"\" ) )                  || die \"bad capability end\";\n\n foreach (@capabilities) {\n@@ -123,6 +128,31 @@ while (1) {\n \tprint $debug \"IN: $command\";\n \t$debug->flush();\n\n+\tif ( $command eq \"list_available_blobs\" ) {\n+\t\t# Flush\n+\t\tpacket_bin_read();\n+\n+\t\tforeach my $pathname (sort keys %DELAY) {\n+\t\t\tif ( $DELAY{$pathname}{\"requested\"} == 1 ) {\n+\n+\t\t\t\t# die $pathname;\n+\t\t\t\t$DELAY{$pathname}{\"count\"} = $DELAY{$pathname}{\"count\"} - 1;\n+\t\t\t\tif ($DELAY{$pathname}{\"count\"} == 0 ) {\n+\t\t\t\t\tprint $debug \" $pathname\";\n+\t\t\t\t\t# packet_txt_write($pathname);\n+\t\t\t\t\tpacket_txt_write($DELAY{$pathname}{\"delay_id\"});\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n+\n+\t\tpacket_flush();\n+\n+\t\tprint $debug \" [OK]\\n\";\n+\t\t$debug->flush();\n+\t\tpacket_txt_write(\"status=success\");\n+\t\tpacket_flush();\n+\t}\n+\telse {\n \t\tmy ($pathname) = packet_txt_read() =~ /^pathname=(.+)$/;\n \t\tprint $debug \" $pathname\";\n \t\t$debug->flush();\n@@ -131,8 +161,25 @@ while (1) {\n \t\t\tdie \"bad pathname '$pathname'\";\n \t\t}\n\n-\t# Flush\n-\tpacket_bin_read();\n+\t\t# Read until flush\n+\t\tmy ( $done, $buffer ) = packet_txt_read();\n+\t\twhile ( $buffer ne '' ) {\n+\t\t\tif ( $buffer eq \"delay-able=1\" ) {\n+\t\t\t\tif ( exists $DELAY{$pathname} and $DELAY{$pathname}{\"requested\"} == 0 ) {\n+\t\t\t\t\t$DELAY{$pathname}{\"requested\"} = 1;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t\telsif ( $buffer =~ m/^delay-id=/ ) {\n+\t\t\t\tmy ($delay_id) = $buffer =~ /^delay-id=(.+)$/;\n+\t\t\t\tif ( $DELAY{$pathname}{\"delay_id\"} != $delay_id ) {\n+\t\t\t\t\tdie \"unexpected delay-id for '$pathname'\";\n+\t\t\t\t}\n+\t\t\t} else {\n+\t\t\t\tdie \"Unknown message '$buffer'\";\n+\t\t\t}\n+\n+\t\t\t( $done, $buffer ) = packet_txt_read();\n+\t\t}\n\n \t\tmy $input = \"\";\n \t\t{\n@@ -148,7 +195,10 @@ while (1) {\n \t\t}\n\n \t\tmy $output;\n-\tif ( $pathname eq \"error.r\" or $pathname eq \"abort.r\" ) {\n+\t\tif ( exists $DELAY{$pathname} and exists $DELAY{$pathname}{\"output\"} ) {\n+\t\t\t$output = $DELAY{$pathname}{\"output\"}\n+\t\t}\n+\t\telsif ( $pathname eq \"error.r\" or $pathname eq \"abort.r\" ) {\n \t\t\t$output = \"\";\n \t\t}\n \t\telsif ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n@@ -161,9 +211,6 @@ while (1) {\n \t\t\tdie \"bad command '$command'\";\n \t\t}\n\n-\tprint $debug \"OUT: \" . length($output) . \" \";\n-\t$debug->flush();\n-\n \t\tif ( $pathname eq \"error.r\" ) {\n \t\t\tprint $debug \"[ERROR]\\n\";\n \t\t\t$debug->flush();\n@@ -178,12 +225,28 @@ while (1) {\n \t\t}\n \t\telsif ( $command eq \"smudge\" and\n \t\t\texists $DELAY{$pathname} and\n-\t\t    $DELAY{$pathname} > 0 ) {\n-\t\t$DELAY{$pathname} = $DELAY{$pathname} - 1;\n+\t\t\t$DELAY{$pathname}{\"requested\"} == 1 and\n+\t\t\t$DELAY{$pathname}{\"delay_id\"} < 0\n+\t\t) {\n \t\t\tprint $debug \"[DELAYED]\\n\";\n \t\t\t$debug->flush();\n \t\t\tpacket_txt_write(\"status=delayed\");\n \t\t\tpacket_flush();\n+\t\t\tmy ( $done, $buffer ) = packet_txt_read();\n+\t\t\tif ( $buffer =~ m/^delay-id=/ ) {\n+\t\t\t\tmy ($delay_id) = $buffer =~ /^delay-id=(.+)$/;\n+\t\t\t\tif ( $delay_id eq \"\" ) {\n+\t\t\t\t\tdie \"bad delay_id '$delay_id'\";\n+\t\t\t\t}\n+\t\t\t\t$DELAY{$pathname}{\"delay_id\"} = $delay_id;\n+\t\t\t\t$DELAY{$pathname}{\"output\"} = $output;\n+\t\t\t}\n+\n+\t\t\t# Flush\n+\t\t\tpacket_bin_read();\n+\n+\t\t\tpacket_txt_write(\"status=success\");\n+\t\t\tpacket_flush();\n \t\t}\n \t\telse {\n \t\t\tpacket_txt_write(\"status=success\");\n@@ -195,6 +258,9 @@ while (1) {\n \t\t\t\tdie \"${command} write error\";\n \t\t\t}\n\n+\t\t\tprint $debug \"OUT: \" . length($output) . \" \";\n+\t\t\t$debug->flush();\n+\n \t\t\twhile ( length($output) > 0 ) {\n \t\t\t\tmy $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n \t\t\t\tpacket_bin_write($packet);\n@@ -213,3 +279,4 @@ while (1) {\n \t\t\tpacket_flush();\n \t\t}\n \t}\n+}\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 6b3246db03..9f50a417ec 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -301,6 +301,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tremove_marked_cache_entries(index);\n \tremove_scheduled_dirs();\n\n+\tenable_delayed_checkout(&state);\n \tfor (i = 0; i < index->cache_nr; i++) {\n \t\tstruct cache_entry *ce = index->cache[i];\n\n@@ -315,7 +316,7 @@ static int check_updates(struct unpack_trees_options *o)\n \t\t\t}\n \t\t}\n \t}\n-\terrs |= checkout_delayed_entries(&state);\n+\terrs |= finish_delayed_checkout(&state);\n \tstop_progress(&progress);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n\n\nLars Schneider (4):\n  t0021: keep filter log files on comparison\n  t0021: make debug log file name configurable\n  t0021: write \"OUT\" only on success\n  convert: add \"status=delayed\" to filter process protocol\n\n Documentation/gitattributes.txt |  73 ++++++++++++-\n builtin/checkout.c              |   3 +\n cache.h                         |  37 ++++++-\n convert.c                       | 150 +++++++++++++++++++++++----\n convert.h                       |   5 +\n entry.c                         | 124 +++++++++++++++++++++-\n t/t0021-conversion.sh           | 135 ++++++++++++++++++------\n t/t0021/rot13-filter.pl         | 220 ++++++++++++++++++++++++++++------------\n unpack-trees.c                  |   2 +\n 9 files changed, 624 insertions(+), 125 deletions(-)\n\n\nbase-commit: e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n--\n2.12.2\n\n"},{"id":"316444","messageId":"20170409191107.20547-5-larsxschneider@gmail.com","threadId":"45641","inReplyTo":"20170409191107.20547-1-larsxschneider@gmail.com","subject":"[PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-09T19:11:07Z","receivedAt":"2017-04-09T19:11:36Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"Some `clean` / `smudge` filters might require a significant amount of\ntime to process a single blob. During this process the Git checkout\noperation is blocked and Git needs to wait until the filter is done to\ncontinue with the checkout.\n\nTeach the filter process protocol (introduced in edcc858) to accept the\nstatus \"delayed\" as response to a filter request. Upon this response Git\ncontinues with the checkout operation. After the checkout operation Git\ncalls \"finish_delayed_checkout\" which queries the filter for remaining\nblobs. If the filter is still working on the completion, then the filter\nis expected to block. If the filter has completed all remaining blobs\nthen an empty response is expected.\n\nGit has a multiple code paths that checkout a blob. Support delayed\ncheckouts only in `clone` (in unpack-trees.c) and `checkout` operations.\n\nSigned-off-by: Lars Schneider <larsxschneider@gmail.com>\n---\n Documentation/gitattributes.txt |  73 +++++++++++++-\n builtin/checkout.c              |   3 +\n cache.h                         |  37 ++++++-\n convert.c                       | 150 ++++++++++++++++++++++++----\n convert.h                       |   5 +\n entry.c                         | 124 +++++++++++++++++++++++-\n t/t0021-conversion.sh           |  73 ++++++++++++++\n t/t0021/rot13-filter.pl         | 210 ++++++++++++++++++++++++++++------------\n unpack-trees.c                  |   2 +\n 9 files changed, 587 insertions(+), 90 deletions(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex e0b66c1220..329baa945f 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -425,8 +425,8 @@ packet:          git< capability=clean\n packet:          git< capability=smudge\n packet:          git< 0000\n ------------------------\n-Supported filter capabilities in version 2 are \"clean\" and\n-\"smudge\".\n+Supported filter capabilities in version 2 are \"clean\", \"smudge\",\n+and \"delay\".\n \n Afterwards Git sends a list of \"key=value\" pairs terminated with\n a flush packet. The list will contain at least the filter command\n@@ -512,12 +512,77 @@ the protocol then Git will stop the filter process and restart it\n with the next file that needs to be processed. Depending on the\n `filter.<driver>.required` flag Git will interpret that as error.\n \n-After the filter has processed a blob it is expected to wait for\n-the next \"key=value\" list containing a command. Git will close\n+After the filter has processed a command it is expected to wait for\n+a \"key=value\" list containing the next command. Git will close\n the command pipe on exit. The filter is expected to detect EOF\n and exit gracefully on its own. Git will wait until the filter\n process has stopped.\n \n+Delay\n+^^^^^\n+\n+If the filter supports the \"delay\" capability, then Git can send the\n+flag \"delay-able\" after the filter command and pathname. This flag\n+denotes that the filter can delay filtering the current blob (e.g. to\n+compensate network latencies) by responding with no content but with\n+the status \"delayed\" and a flush packet. Git will answer with a\n+\"delay-id\", a number that identifies the blob, and a flush packet. The\n+filter acknowledges this number with a \"success\" status and a flush\n+packet.\n+------------------------\n+packet:          git> command=smudge\n+packet:          git> pathname=path/testfile.dat\n+packet:          git> delay-able=1\n+packet:          git> 0000\n+packet:          git> CONTENT\n+packet:          git> 0000\n+packet:          git< status=delayed\n+packet:          git< 0000\n+packet:          git> delay-id=1\n+packet:          git> 0000\n+packet:          git< status=success\n+packet:          git< 0000\n+------------------------\n+\n+If the filter supports the \"delay\" capability then it must support the\n+\"list_available_blobs\" command. If Git sends this command, then the\n+filter is expected to return a list of \"delay_ids\" of blobs that are\n+available. The list must be terminated with a flush packet followed\n+by a \"success\" status that is also terminated with a flush packet. If\n+no blobs for the delayed paths are available, yet, then the filter is\n+expected to block the response until at least one blob becomes\n+available. The filter can tell Git that it has no more delayed blobs\n+by sending an empty list.\n+------------------------\n+packet:          git> command=list_available_blobs\n+packet:          git> 0000\n+packet:          git< 7\n+packet:          git< 13\n+packet:          git< 0000\n+packet:          git< status=success\n+packet:          git< 0000\n+------------------------\n+\n+After Git received the \"delay_ids\", it will request the corresponding\n+blobs again. These requests contain a \"delay-id\" and an empty content\n+section. The filter is expected to respond with the smudged content\n+in the usual way as explained above.\n+------------------------\n+packet:          git> command=smudge\n+packet:          git> pathname=test-delay10.a\n+packet:          git> delay-id=0\n+packet:          git> 0000\n+packet:          git> 0000  # empty content!\n+packet:          git< status=success\n+packet:          git< 0000\n+packet:          git< SMUDGED_CONTENT\n+packet:          git< 0000\n+packet:          git< 0000\n+------------------------\n+\n+Example\n+^^^^^^^\n+\n A long running filter demo implementation can be found in\n `contrib/long-running-filter/example.pl` located in the Git\n core repository. If you develop your own long running filter\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex f174f50303..e0a0bc92d4 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -355,6 +355,8 @@ static int checkout_paths(const struct checkout_opts *opts,\n \tstate.force = 1;\n \tstate.refresh_cache = 1;\n \tstate.istate = &the_index;\n+\n+\tenable_delayed_checkout(&state);\n \tfor (pos = 0; pos < active_nr; pos++) {\n \t\tstruct cache_entry *ce = active_cache[pos];\n \t\tif (ce->ce_flags & CE_MATCHED) {\n@@ -369,6 +371,7 @@ static int checkout_paths(const struct checkout_opts *opts,\n \t\t\tpos = skip_same_name(ce, pos) - 1;\n \t\t}\n \t}\n+\terrs |= finish_delayed_checkout(&state);\n \n \tif (write_locked_index(&the_index, lock_file, COMMIT_LOCK))\n \t\tdie(_(\"unable to write new index file\"));\ndiff --git a/cache.h b/cache.h\nindex 61fc86e6d7..46076279cf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1421,19 +1421,54 @@ const char *show_ident_date(const struct ident_split *id,\n  */\n extern int ident_cmp(const struct ident_split *, const struct ident_split *);\n \n+enum ce_delay_state {\n+\tCE_DELAY_DISABLED = 0,\n+\tCE_DELAY_AVAILABLE = 1,\n+\tCE_DELAY_APPLIED = 2,\n+\tCE_DELAY_RETRY = 3\n+};\n+\n+struct delayed_checkout {\n+\tenum ce_delay_state state;\n+\t/* The value of \"delay_id\" has different meaning depending on the\n+\t * \"state\" variable:\n+\t *   - CE_DELAY_DISABLED  => \"delay_id\" not used.\n+\t *   - CE_DELAY_AVAILABLE => \"delay_id\" is available to be presented\n+\t *                           to the filter in case the filter wants to\n+\t *                           delay the response of a blob.\n+\t *   - CE_DELAY_APPLIED   => \"delay_id\" was presented to and applied by\n+\t *                           the filter for a blob. The corresponding\n+\t *                           cache entry in stored in the \"entries\"\n+\t *                           array under the index \"delay_id\".\n+\t *   - CE_DELAY_RETRY     => Git requests a blob from the filter that\n+\t *                           was previously delayed using the \"delay_id\".\n+\t */\n+\tint delay_id;\n+\t/* List of filter drivers that have delayed blobs. */\n+\tstruct string_list filters;\n+\t/* Array of cache entries that have been delayed. */\n+\tstruct cache_entry **entries;\n+\tint entries_nr;\n+\tint entries_alloc;\n+};\n+\n struct checkout {\n \tstruct index_state *istate;\n \tconst char *base_dir;\n+\tstruct delayed_checkout *delayed_checkout;\n \tint base_dir_len;\n \tunsigned force:1,\n \t\t quiet:1,\n \t\t not_new:1,\n \t\t refresh_cache:1;\n };\n-#define CHECKOUT_INIT { NULL, \"\" }\n+#define CHECKOUT_INIT { NULL, \"\", NULL }\n+\n \n #define TEMPORARY_FILENAME_LENGTH 25\n extern int checkout_entry(struct cache_entry *ce, const struct checkout *state, char *topath);\n+extern void enable_delayed_checkout(struct checkout *state);\n+extern int finish_delayed_checkout(struct checkout *state);\n \n struct cache_def {\n \tstruct strbuf path;\ndiff --git a/convert.c b/convert.c\nindex 4e17e45ed2..0d8fa0f833 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -495,6 +495,7 @@ static int apply_single_file_filter(const char *path, const char *src, size_t le\n \n #define CAP_CLEAN    (1u<<0)\n #define CAP_SMUDGE   (1u<<1)\n+#define CAP_DELAY    (1u<<2)\n \n struct cmd2process {\n \tstruct hashmap_entry ent; /* must be the first member! */\n@@ -632,7 +633,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \tif (err)\n \t\tgoto done;\n \n-\terr = packet_write_list(process->in, \"capability=clean\", \"capability=smudge\", NULL);\n+\terr = packet_write_list(process->in,\n+\t\t\"capability=clean\", \"capability=smudge\", \"capability=delay\", NULL);\n \n \tfor (;;) {\n \t\tcap_buf = packet_read_line(process->out, NULL);\n@@ -648,6 +650,8 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \t\t\tentry->supported_capabilities |= CAP_CLEAN;\n \t\t} else if (!strcmp(cap_name, \"smudge\")) {\n \t\t\tentry->supported_capabilities |= CAP_SMUDGE;\n+\t\t} else if (!strcmp(cap_name, \"delay\")) {\n+\t\t\tentry->supported_capabilities |= CAP_DELAY;\n \t\t} else {\n \t\t\twarning(\n \t\t\t\t\"external filter '%s' requested unsupported filter capability '%s'\",\n@@ -673,9 +677,11 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons\n \n static int apply_multi_file_filter(const char *path, const char *src, size_t len,\n \t\t\t\t   int fd, struct strbuf *dst, const char *cmd,\n-\t\t\t\t   const unsigned int wanted_capability)\n+\t\t\t\t   const unsigned int wanted_capability,\n+\t\t\t\t   struct delayed_checkout *dco)\n {\n \tint err;\n+\tint is_delay_available = 0;\n \tstruct cmd2process *entry;\n \tstruct child_process *process;\n \tstruct strbuf nbuf = STRBUF_INIT;\n@@ -726,6 +732,24 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \tif (err)\n \t\tgoto done;\n \n+\tif (CAP_DELAY & entry->supported_capabilities && dco) {\n+\t\tswitch (dco->state) {\n+\t\tcase CE_DELAY_AVAILABLE:\n+\t\t\tis_delay_available = 1;\n+\t\t\terr = packet_write_fmt_gently(\n+\t\t\t\tprocess->in, \"delay-able=1\\n\");\n+\t\t\tbreak;\n+\t\tcase CE_DELAY_RETRY:\n+\t\t\terr = packet_write_fmt_gently(\n+\t\t\t\tprocess->in, \"delay-id=%i\\n\", dco->delay_id);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tbreak;\n+\t\t}\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t}\n+\n \terr = packet_flush_gently(process->in);\n \tif (err)\n \t\tgoto done;\n@@ -738,13 +762,27 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \t\tgoto done;\n \n \tread_multi_file_filter_status(process->out, &filter_status);\n-\terr = strcmp(filter_status.buf, \"success\");\n-\tif (err)\n-\t\tgoto done;\n+\tif (is_delay_available && !strcmp(filter_status.buf, \"delayed\")) {\n+\t\t/* The filter wants to delay the response. Send it a delay id. */\n+\t\terr = packet_write_fmt_gently(\n+\t\t\tprocess->in, \"delay-id=%i\\n\", dco->delay_id);\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t\terr = packet_flush_gently(process->in);\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t\tstring_list_insert(&dco->filters, cmd);\n+\t\tdco->state = CE_DELAY_APPLIED;\n+\t} else {\n+\t\t/* The filter got the blob and wants to send us a response. */\n+\t\terr = strcmp(filter_status.buf, \"success\");\n+\t\tif (err)\n+\t\t\tgoto done;\n \n-\terr = read_packetized_to_strbuf(process->out, &nbuf) < 0;\n-\tif (err)\n-\t\tgoto done;\n+\t\terr = read_packetized_to_strbuf(process->out, &nbuf) < 0;\n+\t\tif (err)\n+\t\t\tgoto done;\n+\t}\n \n \tread_multi_file_filter_status(process->out, &filter_status);\n \terr = strcmp(filter_status.buf, \"success\");\n@@ -777,6 +815,74 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len\n \treturn !err;\n }\n \n+\n+int async_query_available_blobs(const char *cmd, unsigned long **delay_ids,\n+\t\t\t\tint *delay_ids_nr)\n+{\n+\tint err;\n+\tchar *line;\n+\tchar *end;\n+\tstruct cmd2process *entry;\n+\tstruct child_process *process;\n+\tstruct strbuf filter_status = STRBUF_INIT;\n+\tunsigned long delay_id;\n+\tint delay_ids_alloc = 0;\n+\t*delay_ids_nr = 0;\n+\n+\tentry = find_multi_file_filter_entry(&cmd_process_map, cmd);\n+\tif (!entry) {\n+\t\terror(\"external filter '%s' is not available anymore although \"\n+\t\t      \"not all paths have been filtered\", cmd);\n+\t\treturn 0;\n+\t}\n+\tprocess = &entry->process;\n+\tsigchain_push(SIGPIPE, SIG_IGN);\n+\n+\terr = packet_write_fmt_gently(\n+\t\tprocess->in, \"command=list_available_blobs\\n\");\n+\tif (err)\n+\t\tgoto done;\n+\n+\terr = packet_flush_gently(process->in);\n+\tif (err)\n+\t\tgoto done;\n+\n+\tfor (;;) {\n+\t\tline = packet_read_line(process->out, NULL);\n+\t\tif (!line)\n+\t\t\tbreak;\n+\t\tdelay_id = strtoul(line, &end, 10);\n+\t\terr = (line == end);\n+\t\tif (err) {\n+\t\t\terror(\"invalid delay id '%s'\", line);\n+\t\t\tgoto done;\n+\t\t}\n+\t\tALLOC_GROW(*delay_ids, *delay_ids_nr+1, delay_ids_alloc);\n+\t\t(*delay_ids)[(*delay_ids_nr)++] = delay_id;\n+\t}\n+\n+\tread_multi_file_filter_status(process->out, &filter_status);\n+\terr = strcmp(filter_status.buf, \"success\");\n+\n+done:\n+\tsigchain_pop(SIGPIPE);\n+\n+\tif (err || errno == EPIPE) {\n+\t\tif (!strcmp(filter_status.buf, \"error\")) {\n+\t\t\t/* The filter signaled a problem with the file. */\n+\t\t} else {\n+\t\t\t/*\n+\t\t\t * Something went wrong with the protocol filter.\n+\t\t\t * Force shutdown and restart if another blob requires\n+\t\t\t * filtering.\n+\t\t\t */\n+\t\t\terror(\"external filter '%s' failed\", cmd);\n+\t\t\tkill_multi_file_filter(&cmd_process_map, entry);\n+\t\t}\n+\t}\n+\treturn !err;\n+}\n+\n static struct convert_driver {\n \tconst char *name;\n \tstruct convert_driver *next;\n@@ -788,7 +894,8 @@ static struct convert_driver {\n \n static int apply_filter(const char *path, const char *src, size_t len,\n \t\t\tint fd, struct strbuf *dst, struct convert_driver *drv,\n-\t\t\tconst unsigned int wanted_capability)\n+\t\t\tconst unsigned int wanted_capability,\n+\t\t\tstruct delayed_checkout *dco)\n {\n \tconst char *cmd = NULL;\n \n@@ -806,7 +913,8 @@ static int apply_filter(const char *path, const char *src, size_t len,\n \tif (cmd && *cmd)\n \t\treturn apply_single_file_filter(path, src, len, fd, dst, cmd);\n \telse if (drv->process && *drv->process)\n-\t\treturn apply_multi_file_filter(path, src, len, fd, dst, drv->process, wanted_capability);\n+\t\treturn apply_multi_file_filter(path, src, len, fd, dst,\n+\t\t\tdrv->process, wanted_capability, dco);\n \n \treturn 0;\n }\n@@ -1152,7 +1260,7 @@ int would_convert_to_git_filter_fd(const char *path)\n \tif (!ca.drv->required)\n \t\treturn 0;\n \n-\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv, CAP_CLEAN);\n+\treturn apply_filter(path, NULL, 0, -1, NULL, ca.drv, CAP_CLEAN, NULL);\n }\n \n const char *get_convert_attr_ascii(const char *path)\n@@ -1189,7 +1297,7 @@ int convert_to_git(const char *path, const char *src, size_t len,\n \n \tconvert_attrs(&ca, path);\n \n-\tret |= apply_filter(path, src, len, -1, dst, ca.drv, CAP_CLEAN);\n+\tret |= apply_filter(path, src, len, -1, dst, ca.drv, CAP_CLEAN, NULL);\n \tif (!ret && ca.drv && ca.drv->required)\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n@@ -1214,7 +1322,7 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n \tassert(ca.drv);\n \tassert(ca.drv->clean || ca.drv->process);\n \n-\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv, CAP_CLEAN))\n+\tif (!apply_filter(path, NULL, 0, fd, dst, ca.drv, CAP_CLEAN, NULL))\n \t\tdie(\"%s: clean filter '%s' failed\", path, ca.drv->name);\n \n \tcrlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);\n@@ -1223,7 +1331,7 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,\n \n static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t\t\t\t    size_t len, struct strbuf *dst,\n-\t\t\t\t\t    int normalizing)\n+\t\t\t\t\t    int normalizing, struct delayed_checkout *dco)\n {\n \tint ret = 0, ret_filter = 0;\n \tstruct conv_attrs ca;\n@@ -1248,21 +1356,29 @@ static int convert_to_working_tree_internal(const char *path, const char *src,\n \t\t}\n \t}\n \n-\tret_filter = apply_filter(path, src, len, -1, dst, ca.drv, CAP_SMUDGE);\n+\tret_filter = apply_filter(\n+\t\tpath, src, len, -1, dst, ca.drv, CAP_SMUDGE, dco);\n \tif (!ret_filter && ca.drv && ca.drv->required)\n \t\tdie(\"%s: smudge filter %s failed\", path, ca.drv->name);\n \n \treturn ret | ret_filter;\n }\n \n+int async_convert_to_working_tree(const char *path, const char *src,\n+\t\t\t\t  size_t len, struct strbuf *dst,\n+\t\t\t\t  void *dco)\n+{\n+\treturn convert_to_working_tree_internal(path, src, len, dst, 0, dco);\n+}\n+\n int convert_to_working_tree(const char *path, const char *src, size_t len, struct strbuf *dst)\n {\n-\treturn convert_to_working_tree_internal(path, src, len, dst, 0);\n+\treturn convert_to_working_tree_internal(path, src, len, dst, 0, NULL);\n }\n \n int renormalize_buffer(const char *path, const char *src, size_t len, struct strbuf *dst)\n {\n-\tint ret = convert_to_working_tree_internal(path, src, len, dst, 1);\n+\tint ret = convert_to_working_tree_internal(path, src, len, dst, 1, NULL);\n \tif (ret) {\n \t\tsrc = dst->buf;\n \t\tlen = dst->len;\ndiff --git a/convert.h b/convert.h\nindex 82871a11d5..da6c702090 100644\n--- a/convert.h\n+++ b/convert.h\n@@ -42,6 +42,11 @@ extern int convert_to_git(const char *path, const char *src, size_t len,\n \t\t\t  struct strbuf *dst, enum safe_crlf checksafe);\n extern int convert_to_working_tree(const char *path, const char *src,\n \t\t\t\t   size_t len, struct strbuf *dst);\n+extern int async_convert_to_working_tree(const char *path, const char *src,\n+\t\t\t\t\t size_t len, struct strbuf *dst,\n+\t\t\t\t\t void *dco);\n+extern int async_query_available_blobs(const char *cmd, unsigned long **delay_ids,\n+\t\t\t\t       int *delay_ids_nr);\n extern int renormalize_buffer(const char *path, const char *src, size_t len,\n \t\t\t      struct strbuf *dst);\n static inline int would_convert_to_git(const char *path)\ndiff --git a/entry.c b/entry.c\nindex c6eea240b6..963e94dee8 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -136,6 +136,86 @@ static int streaming_write_entry(const struct cache_entry *ce, char *path,\n \treturn result;\n }\n \n+void enable_delayed_checkout(struct checkout *state)\n+{\n+\tif (!state->delayed_checkout) {\n+\t\tstate->delayed_checkout = xmalloc(sizeof(*state->delayed_checkout));\n+\t\tstate->delayed_checkout->entries_nr = 0;\n+\t\tstate->delayed_checkout->entries_alloc = 0;\n+\t\tstate->delayed_checkout->delay_id = -1;\n+\t\tstate->delayed_checkout->state = CE_DELAY_AVAILABLE;\n+\t\tALLOC_ARRAY(state->delayed_checkout->entries, 0);\n+\t\tstring_list_init(&state->delayed_checkout->filters, 0);\n+\t}\n+}\n+\n+int finish_delayed_checkout(struct checkout *state)\n+{\n+\tint errs = 0;\n+\tstruct string_list_item *filter;\n+\tstruct delayed_checkout *dco = state->delayed_checkout;\n+\n+\tif (!state->delayed_checkout) {\n+\t\treturn errs;\n+\t}\n+\n+\twhile (dco->entries_nr > 0 && dco->filters.nr > 0) {\n+\t\tfor_each_string_list_item(filter, &dco->filters) {\n+\t\t\tint i;\n+\t\t\tint delay_ids_nr;\n+\t\t\tunsigned long *delay_ids;\n+\t\t\tALLOC_ARRAY(delay_ids, 0);\n+\t\t\tif (!async_query_available_blobs(\n+\t\t\t\tfilter->string, &delay_ids, &delay_ids_nr)) {\n+\t\t\t\t/* Filter reported an error */\n+\t\t\t\terrs = 1;\n+\t\t\t\tfilter->string = \"\";\n+\t\t\t\tfree(delay_ids);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (delay_ids_nr <= 0) {\n+\t\t\t\t/* Filter responded with no entries. That means\n+\t\t\t\t   the filter is done and we can remove the\n+\t\t\t\t   filter from the list\n+\t\t\t\t   (see \"string_list_remove_empty_items\" call\n+\t\t\t\t   below).\n+\t\t\t\t*/\n+\t\t\t\tfilter->string = \"\";\n+\t\t\t\tfree(delay_ids);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tfor (i = 0; i < delay_ids_nr; i++) {\n+\t\t\t\tstruct cache_entry* ce;\n+\t\t\t\tunsigned long delay_id = delay_ids[i];\n+\t\t\t\tassert(delay_id >= 0 && delay_id < dco->entries_nr);\n+\t\t\t\tce = dco->entries[delay_id];\n+\t\t\t\tdco->entries[delay_id] = NULL;\n+\t\t\t\tdco->delay_id = delay_id;\n+\t\t\t\tdco->state = CE_DELAY_RETRY;\n+\t\t\t\t/* Shrink entries array as much as possible */\n+\t\t\t\twhile (\n+\t\t\t\t\tdco->entries_nr > 0 &&\n+\t\t\t\t\tdelay_id == dco->entries_nr - 1 &&\n+\t\t\t\t\t!dco->entries[i]\n+\t\t\t\t) {\n+\t\t\t\t\tdelay_id--;\n+\t\t\t\t\tdco->entries_nr--;\n+\t\t\t\t}\n+\t\t\t\terrs |= (ce ? checkout_entry(ce, state, NULL) : 1);\n+\t\t\t}\n+\t\t\tfree(delay_ids);\n+\t\t}\n+\t\tstring_list_remove_empty_items(&dco->filters, 0);\n+\t}\n+\n+\tstring_list_clear(&dco->filters, 0);\n+\tfree(dco->entries);\n+\tfree(dco);\n+\tstate->delayed_checkout = NULL;\n+\n+\treturn errs;\n+}\n+\n static int write_entry(struct cache_entry *ce,\n \t\t       char *path, const struct checkout *state, int to_tempfile)\n {\n@@ -177,11 +257,45 @@ static int write_entry(struct cache_entry *ce,\n \t\t/*\n \t\t * Convert from git internal format to working tree format\n \t\t */\n-\t\tif (ce_mode_s_ifmt == S_IFREG &&\n-\t\t    convert_to_working_tree(ce->name, new, size, &buf)) {\n-\t\t\tfree(new);\n-\t\t\tnew = strbuf_detach(&buf, &newsize);\n-\t\t\tsize = newsize;\n+\t\tif (ce_mode_s_ifmt == S_IFREG) {\n+\t\t\tstruct delayed_checkout *dco = state->delayed_checkout;\n+\t\t\tif (dco && dco->state != CE_DELAY_DISABLED) {\n+\t\t\t\tswitch (dco->state) {\n+\t\t\t\tcase CE_DELAY_AVAILABLE:\n+\t\t\t\t\tdco->delay_id = dco->entries_nr; break;\n+\t\t\t\tcase CE_DELAY_RETRY:\n+\t\t\t\t\tnew = NULL; size = 0; break;\n+\t\t\t\tdefault: break;\n+\t\t\t\t}\n+\t\t\t\tassert(dco->delay_id >= 0);\n+\t\t\t\tret = async_convert_to_working_tree(\n+\t\t\t\t\tce->name, new, size, &buf, dco);\n+\t\t\t\tif (ret && dco->state == CE_DELAY_APPLIED) {\n+\t\t\t\t\tassert(dco->delay_id == dco->entries_nr);\n+\t\t\t\t\tfree(new);\n+\t\t\t\t\tdco->entries_nr++;\n+\t\t\t\t\tALLOC_GROW(\n+\t\t\t\t\t\tdco->entries, dco->entries_nr,\n+\t\t\t\t\t\tdco->entries_alloc);\n+\t\t\t\t\tdco->entries[dco->delay_id] = ce;\n+\t\t\t\t\tdco->state = CE_DELAY_AVAILABLE;\n+\t\t\t\t\tdco->delay_id = -1;\n+\t\t\t\t\tgoto finish;\n+\t\t\t\t}\n+\t\t\t} else\n+\t\t\t\tret = convert_to_working_tree(\n+\t\t\t\t\tce->name, new, size, &buf);\n+\n+\t\t\tif (ret) {\n+\t\t\t\tfree(new);\n+\t\t\t\tnew = strbuf_detach(&buf, &newsize);\n+\t\t\t\tsize = newsize;\n+\t\t\t}\n+\t\t\t/*\n+\t\t\t * No \"else\" here as errors from convert are OK at this\n+\t\t\t * point. If the error would have been fatal (e.g.\n+\t\t\t * filter is required), then we would have died already.\n+\t\t\t */\n \t\t}\n \n \t\tfd = open_output_fd(path, ce, to_tempfile);\ndiff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh\nindex 0c04d346a1..21d4cd9453 100755\n--- a/t/t0021-conversion.sh\n+++ b/t/t0021-conversion.sh\n@@ -701,4 +701,77 @@ test_expect_success PERL 'invalid process filter must fail (and not hang!)' '\n \t)\n '\n \n+test_expect_success PERL 'delayed checkout in process filter' '\n+\ttest_config_global filter.a.process \"rot13-filter.pl a.log clean smudge delay\" &&\n+\ttest_config_global filter.a.required true &&\n+\ttest_config_global filter.b.process \"rot13-filter.pl b.log clean smudge delay\" &&\n+\ttest_config_global filter.b.required true &&\n+\n+\trm -rf repo &&\n+\tmkdir repo &&\n+\t(\n+\t\tcd repo &&\n+\t\tgit init &&\n+\t\techo \"*.a filter=a\" >.gitattributes &&\n+\t\techo \"*.b filter=b\" >>.gitattributes &&\n+\t\tcp \"$TEST_ROOT/test.o\" test.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\tcp \"$TEST_ROOT/test.o\" test-delay10.b &&\n+\t\tgit add . &&\n+\t\tgit commit -m \"test commit 1\"\n+\t) &&\n+\n+\tS=$(file_size \"$TEST_ROOT/test.o\") &&\n+\tcat >a.exp <<-EOF &&\n+\t\tSTART\n+\t\tinit handshake complete\n+\t\tIN: smudge test.a $S [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test-delay10.a $S [OK] -- [DELAYED]\n+\t\tIN: smudge test-delay11.a $S [OK] -- [DELAYED]\n+\t\tIN: smudge test-delay20.a $S [OK] -- [DELAYED]\n+\t\tIN: list_available_blobs test-delay10.a test-delay11.a [OK]\n+\t\tIN: smudge test-delay10.a 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: smudge test-delay11.a 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: list_available_blobs test-delay20.a [OK]\n+\t\tIN: smudge test-delay20.a 0 [OK] -- OUT: $S . [OK]\n+\t\tSTOP\n+\tEOF\n+\tcat >b.exp <<-EOF &&\n+\t\tSTART\n+\t\tinit handshake complete\n+\t\tIN: smudge test-delay10.b $S [OK] -- [DELAYED]\n+\t\tIN: list_available_blobs test-delay10.b [OK]\n+\t\tIN: smudge test-delay10.b 0 [OK] -- OUT: $S . [OK]\n+\t\tIN: list_available_blobs [OK]\n+\t\tSTOP\n+\tEOF\n+\n+\trm -rf repo-cloned &&\n+\tfilter_git clone repo repo-cloned &&\n+\ttest_cmp_count a.exp repo-cloned/a.log &&\n+\ttest_cmp_count b.exp repo-cloned/b.log &&\n+\n+\t(\n+\t\tcd repo-cloned &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.b &&\n+\n+\t\trm *.a *.b &&\n+\t\tfilter_git checkout . &&\n+\t\ttest_cmp_count ../a.exp a.log &&\n+\t\ttest_cmp_count ../b.exp b.log &&\n+\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay11.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay20.a &&\n+\t\ttest_cmp_committed_rot13 \"$TEST_ROOT/test.o\" test-delay10.b\n+\t)\n+'\n+\n test_done\ndiff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl\nindex 5e43faeec1..05024c6e1b 100644\n--- a/t/t0021/rot13-filter.pl\n+++ b/t/t0021/rot13-filter.pl\n@@ -18,6 +18,11 @@\n #     operation then the filter signals that it cannot or does not want\n #     to process the file and any file after that is processed with the\n #     same command.\n+# (5) If data with a pathname that is a key in the DELAY hash is\n+#     requested (e.g. 'test-delay10.a') then the filter responds with\n+#     a \"delay\" status and sets the \"requested\" field in the DELAY hash.\n+#     The filter will signal the availability of this object after\n+#     \"count\" (field in DELAY hash) \"list_available_blobs\" commands.\n #\n \n use strict;\n@@ -30,6 +35,13 @@ my @capabilities            = @ARGV;\n \n open my $debug, \">>\", $log_file or die \"cannot open log file: $!\";\n \n+my %DELAY = (\n+\t'test-delay10.a' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+\t'test-delay11.a' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+\t'test-delay20.a' => { \"requested\" => 0, \"count\" => 2, \"delay_id\" => -1 },\n+\t'test-delay10.b' => { \"requested\" => 0, \"count\" => 1, \"delay_id\" => -1 },\n+);\n+\n sub rot13 {\n \tmy $str = shift;\n \t$str =~ y/A-Za-z/N-ZA-Mn-za-m/;\n@@ -66,7 +78,7 @@ sub packet_bin_read {\n \n sub packet_txt_read {\n \tmy ( $res, $buf ) = packet_bin_read();\n-\tunless ( $buf =~ s/\\n$// ) {\n+\tunless ( $buf eq '' or $buf =~ s/\\n$// ) {\n \t\tdie \"A non-binary line MUST be terminated by an LF.\";\n \t}\n \treturn ( $res, $buf );\n@@ -101,6 +113,7 @@ packet_flush();\n \n ( packet_txt_read() eq ( 0, \"capability=clean\" ) )  || die \"bad capability\";\n ( packet_txt_read() eq ( 0, \"capability=smudge\" ) ) || die \"bad capability\";\n+( packet_txt_read() eq ( 0, \"capability=delay\" ) )  || die \"bad capability\";\n ( packet_bin_read() eq ( 1, \"\" ) )                  || die \"bad capability end\";\n \n foreach (@capabilities) {\n@@ -115,84 +128,155 @@ while (1) {\n \tprint $debug \"IN: $command\";\n \t$debug->flush();\n \n-\tmy ($pathname) = packet_txt_read() =~ /^pathname=(.+)$/;\n-\tprint $debug \" $pathname\";\n-\t$debug->flush();\n+\tif ( $command eq \"list_available_blobs\" ) {\n+\t\t# Flush\n+\t\tpacket_bin_read();\n \n-\tif ( $pathname eq \"\" ) {\n-\t\tdie \"bad pathname '$pathname'\";\n-\t}\n+\t\tforeach my $pathname (sort keys %DELAY) {\n+\t\t\tif ( $DELAY{$pathname}{\"requested\"} == 1 ) {\n \n-\t# Flush\n-\tpacket_bin_read();\n-\n-\tmy $input = \"\";\n-\t{\n-\t\tbinmode(STDIN);\n-\t\tmy $buffer;\n-\t\tmy $done = 0;\n-\t\twhile ( !$done ) {\n-\t\t\t( $done, $buffer ) = packet_bin_read();\n-\t\t\t$input .= $buffer;\n+\t\t\t\t# die $pathname;\n+\t\t\t\t$DELAY{$pathname}{\"count\"} = $DELAY{$pathname}{\"count\"} - 1;\n+\t\t\t\tif ($DELAY{$pathname}{\"count\"} == 0 ) {\n+\t\t\t\t\tprint $debug \" $pathname\";\n+\t\t\t\t\t# packet_txt_write($pathname);\n+\t\t\t\t\tpacket_txt_write($DELAY{$pathname}{\"delay_id\"});\n+\t\t\t\t}\n+\t\t\t}\n \t\t}\n-\t\tprint $debug \" \" . length($input) . \" [OK] -- \";\n-\t\t$debug->flush();\n-\t}\n \n-\tmy $output;\n-\tif ( $pathname eq \"error.r\" or $pathname eq \"abort.r\" ) {\n-\t\t$output = \"\";\n-\t}\n-\telsif ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n-\t\t$output = rot13($input);\n-\t}\n-\telsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n-\t\t$output = rot13($input);\n-\t}\n-\telse {\n-\t\tdie \"bad command '$command'\";\n-\t}\n-\n-\tif ( $pathname eq \"error.r\" ) {\n-\t\tprint $debug \"[ERROR]\\n\";\n-\t\t$debug->flush();\n-\t\tpacket_txt_write(\"status=error\");\n \t\tpacket_flush();\n-\t}\n-\telsif ( $pathname eq \"abort.r\" ) {\n-\t\tprint $debug \"[ABORT]\\n\";\n+\n+\t\tprint $debug \" [OK]\\n\";\n \t\t$debug->flush();\n-\t\tpacket_txt_write(\"status=abort\");\n+\t\tpacket_txt_write(\"status=success\");\n \t\tpacket_flush();\n \t}\n \telse {\n-\t\tpacket_txt_write(\"status=success\");\n-\t\tpacket_flush();\n+\t\tmy ($pathname) = packet_txt_read() =~ /^pathname=(.+)$/;\n+\t\tprint $debug \" $pathname\";\n+\t\t$debug->flush();\n+\n+\t\tif ( $pathname eq \"\" ) {\n+\t\t\tdie \"bad pathname '$pathname'\";\n+\t\t}\n+\n+\t\t# Read until flush\n+\t\tmy ( $done, $buffer ) = packet_txt_read();\n+\t\twhile ( $buffer ne '' ) {\n+\t\t\tif ( $buffer eq \"delay-able=1\" ) {\n+\t\t\t\tif ( exists $DELAY{$pathname} and $DELAY{$pathname}{\"requested\"} == 0 ) {\n+\t\t\t\t\t$DELAY{$pathname}{\"requested\"} = 1;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t\telsif ( $buffer =~ m/^delay-id=/ ) {\n+\t\t\t\tmy ($delay_id) = $buffer =~ /^delay-id=(.+)$/;\n+\t\t\t\tif ( $DELAY{$pathname}{\"delay_id\"} != $delay_id ) {\n+\t\t\t\t\tdie \"unexpected delay-id for '$pathname'\";\n+\t\t\t\t}\n+\t\t\t} else {\n+\t\t\t\tdie \"Unknown message '$buffer'\";\n+\t\t\t}\n+\n+\t\t\t( $done, $buffer ) = packet_txt_read();\n+\t\t}\n \n-\t\tif ( $pathname eq \"${command}-write-fail.r\" ) {\n-\t\t\tprint $debug \"[WRITE FAIL]\\n\";\n+\t\tmy $input = \"\";\n+\t\t{\n+\t\t\tbinmode(STDIN);\n+\t\t\tmy $buffer;\n+\t\t\tmy $done = 0;\n+\t\t\twhile ( !$done ) {\n+\t\t\t\t( $done, $buffer ) = packet_bin_read();\n+\t\t\t\t$input .= $buffer;\n+\t\t\t}\n+\t\t\tprint $debug \" \" . length($input) . \" [OK] -- \";\n \t\t\t$debug->flush();\n-\t\t\tdie \"${command} write error\";\n \t\t}\n \n-\t\tprint $debug \"OUT: \" . length($output) . \" \";\n-\t\t$debug->flush();\n+\t\tmy $output;\n+\t\tif ( exists $DELAY{$pathname} and exists $DELAY{$pathname}{\"output\"} ) {\n+\t\t\t$output = $DELAY{$pathname}{\"output\"}\n+\t\t}\n+\t\telsif ( $pathname eq \"error.r\" or $pathname eq \"abort.r\" ) {\n+\t\t\t$output = \"\";\n+\t\t}\n+\t\telsif ( $command eq \"clean\" and grep( /^clean$/, @capabilities ) ) {\n+\t\t\t$output = rot13($input);\n+\t\t}\n+\t\telsif ( $command eq \"smudge\" and grep( /^smudge$/, @capabilities ) ) {\n+\t\t\t$output = rot13($input);\n+\t\t}\n+\t\telse {\n+\t\t\tdie \"bad command '$command'\";\n+\t\t}\n+\n+\t\tif ( $pathname eq \"error.r\" ) {\n+\t\t\tprint $debug \"[ERROR]\\n\";\n+\t\t\t$debug->flush();\n+\t\t\tpacket_txt_write(\"status=error\");\n+\t\t\tpacket_flush();\n+\t\t}\n+\t\telsif ( $pathname eq \"abort.r\" ) {\n+\t\t\tprint $debug \"[ABORT]\\n\";\n+\t\t\t$debug->flush();\n+\t\t\tpacket_txt_write(\"status=abort\");\n+\t\t\tpacket_flush();\n+\t\t}\n+\t\telsif ( $command eq \"smudge\" and\n+\t\t\texists $DELAY{$pathname} and\n+\t\t\t$DELAY{$pathname}{\"requested\"} == 1 and\n+\t\t\t$DELAY{$pathname}{\"delay_id\"} < 0\n+\t\t) {\n+\t\t\tprint $debug \"[DELAYED]\\n\";\n+\t\t\t$debug->flush();\n+\t\t\tpacket_txt_write(\"status=delayed\");\n+\t\t\tpacket_flush();\n+\t\t\tmy ( $done, $buffer ) = packet_txt_read();\n+\t\t\tif ( $buffer =~ m/^delay-id=/ ) {\n+\t\t\t\tmy ($delay_id) = $buffer =~ /^delay-id=(.+)$/;\n+\t\t\t\tif ( $delay_id eq \"\" ) {\n+\t\t\t\t\tdie \"bad delay_id '$delay_id'\";\n+\t\t\t\t}\n+\t\t\t\t$DELAY{$pathname}{\"delay_id\"} = $delay_id;\n+\t\t\t\t$DELAY{$pathname}{\"output\"} = $output;\n+\t\t\t}\n \n-\t\twhile ( length($output) > 0 ) {\n-\t\t\tmy $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n-\t\t\tpacket_bin_write($packet);\n-\t\t\t# dots represent the number of packets\n-\t\t\tprint $debug \".\";\n-\t\t\tif ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n-\t\t\t\t$output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n+\t\t\t# Flush\n+\t\t\tpacket_bin_read();\n+\n+\t\t\tpacket_txt_write(\"status=success\");\n+\t\t\tpacket_flush();\n+\t\t}\n+\t\telse {\n+\t\t\tpacket_txt_write(\"status=success\");\n+\t\t\tpacket_flush();\n+\n+\t\t\tif ( $pathname eq \"${command}-write-fail.r\" ) {\n+\t\t\t\tprint $debug \"[WRITE FAIL]\\n\";\n+\t\t\t\t$debug->flush();\n+\t\t\t\tdie \"${command} write error\";\n \t\t\t}\n-\t\t\telse {\n-\t\t\t\t$output = \"\";\n+\n+\t\t\tprint $debug \"OUT: \" . length($output) . \" \";\n+\t\t\t$debug->flush();\n+\n+\t\t\twhile ( length($output) > 0 ) {\n+\t\t\t\tmy $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );\n+\t\t\t\tpacket_bin_write($packet);\n+\t\t\t\t# dots represent the number of packets\n+\t\t\t\tprint $debug \".\";\n+\t\t\t\tif ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {\n+\t\t\t\t\t$output = substr( $output, $MAX_PACKET_CONTENT_SIZE );\n+\t\t\t\t}\n+\t\t\t\telse {\n+\t\t\t\t\t$output = \"\";\n+\t\t\t\t}\n \t\t\t}\n+\t\t\tpacket_flush();\n+\t\t\tprint $debug \" [OK]\\n\";\n+\t\t\t$debug->flush();\n+\t\t\tpacket_flush();\n \t\t}\n-\t\tpacket_flush();\n-\t\tprint $debug \" [OK]\\n\";\n-\t\t$debug->flush();\n-\t\tpacket_flush();\n \t}\n }\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 3a8ee19fe8..9f50a417ec 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -301,6 +301,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tremove_marked_cache_entries(index);\n \tremove_scheduled_dirs();\n \n+\tenable_delayed_checkout(&state);\n \tfor (i = 0; i < index->cache_nr; i++) {\n \t\tstruct cache_entry *ce = index->cache[i];\n \n@@ -315,6 +316,7 @@ static int check_updates(struct unpack_trees_options *o)\n \t\t\t}\n \t\t}\n \t}\n+\terrs |= finish_delayed_checkout(&state);\n \tstop_progress(&progress);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n-- \n2.12.2\n\n"},{"id":"316464","messageId":"16331164-8E8C-4CDA-B319-AB8092BD7188@gmail.com","threadId":"45641","inReplyTo":"20170409191107.20547-5-larsxschneider@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-10T10:00:01Z","receivedAt":"2017-04-10T10:00:10Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 09 Apr 2017, at 21:11, Lars Schneider <larsxschneider@gmail.com> wrote:\n> \n> Some `clean` / `smudge` filters might require a significant amount of\n> time to process a single blob. During this process the Git checkout\n> operation is blocked and Git needs to wait until the filter is done to\n> continue with the checkout.\n> \n> Teach the filter process protocol (introduced in edcc858) to accept the\n> status \"delayed\" as response to a filter request. Upon this response Git\n> continues with the checkout operation. After the checkout operation Git\n> calls \"finish_delayed_checkout\" which queries the filter for remaining\n> blobs. If the filter is still working on the completion, then the filter\n> is expected to block. If the filter has completed all remaining blobs\n> then an empty response is expected.\n> \n> Git has a multiple code paths that checkout a blob. Support delayed\n> checkouts only in `clone` (in unpack-trees.c) and `checkout` operations.\n> \n> Signed-off-by: Lars Schneider <larsxschneider@gmail.com>\n> ---\n> ...\n> diff --git a/convert.h b/convert.h\n> index 82871a11d5..da6c702090 100644\n> --- a/convert.h\n> +++ b/convert.h\n> @@ -42,6 +42,11 @@ extern int convert_to_git(const char *path, const char *src, size_t len,\n> \t\t\t  struct strbuf *dst, enum safe_crlf checksafe);\n> extern int convert_to_working_tree(const char *path, const char *src,\n> \t\t\t\t   size_t len, struct strbuf *dst);\n> +extern int async_convert_to_working_tree(const char *path, const char *src,\n> +\t\t\t\t\t size_t len, struct strbuf *dst,\n> +\t\t\t\t\t void *dco);\n> \n\nI don't like the void pointer here. However, \"cache.h\" includes \"convert.h\" and\ntherefore \"convert.h\" cannot include \"cache.h\". That's why \"convert.h\" doesn't\nknow about \"struct delayed_checkout\". \n\nI just realized that I could move \"struct delayed_checkout\" and \"enum ce_delay_state\"\ndefinition from \"cache.h\" to \"convert.h\" to solve the problem nicely.\n\nAny objection to this approach?\n\nThanks,\nLars"},{"id":"316492","messageId":"20170410142850.GA23068@starla","threadId":"45641","inReplyTo":"16331164-8E8C-4CDA-B319-AB8092BD7188@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2017-04-10T14:28:50Z","receivedAt":"2017-04-10T14:28:57Z","isPatch":true,"sender":{"key":"e@80x24.org","avatar":null},"body":"Lars Schneider <larsxschneider@gmail.com> wrote:\n> > diff --git a/convert.h b/convert.h\n> > index 82871a11d5..da6c702090 100644\n> > --- a/convert.h\n> > +++ b/convert.h\n> > @@ -42,6 +42,11 @@ extern int convert_to_git(const char *path, const char *src, size_t len,\n> > \t\t\t  struct strbuf *dst, enum safe_crlf checksafe);\n> > extern int convert_to_working_tree(const char *path, const char *src,\n> > \t\t\t\t   size_t len, struct strbuf *dst);\n> > +extern int async_convert_to_working_tree(const char *path, const char *src,\n> > +\t\t\t\t\t size_t len, struct strbuf *dst,\n> > +\t\t\t\t\t void *dco);\n> > \n> \n> I don't like the void pointer here. However, \"cache.h\" includes \"convert.h\" and\n> therefore \"convert.h\" cannot include \"cache.h\". That's why \"convert.h\" doesn't\n> know about \"struct delayed_checkout\". \n\nYou can forward declare the struct without fields in convert.h:\n\ndiff --git a/convert.h b/convert.h\nindex da6c702090..3fb6b420b2 100644\n--- a/convert.h\n+++ b/convert.h\n@@ -32,6 +32,8 @@ enum eol {\n #endif\n };\n \n+struct delayed_checkout;\n+\n extern enum eol core_eol;\n extern const char *get_cached_convert_stats_ascii(const char *path);\n extern const char *get_wt_convert_stats_ascii(const char *path);\n@@ -44,7 +46,7 @@ extern int convert_to_working_tree(const char *path, const char *src,\n \t\t\t\t   size_t len, struct strbuf *dst);\n extern int async_convert_to_working_tree(const char *path, const char *src,\n \t\t\t\t\t size_t len, struct strbuf *dst,\n-\t\t\t\t\t void *dco);\n+\t\t\t\t\t struct delayed_checkout *dco);\n extern int async_query_available_blobs(const char *cmd, unsigned long **delay_ids,\n \t\t\t\t       int *delay_ids_nr);\n extern int renormalize_buffer(const char *path, const char *src, size_t len,\n\n> \n> I just realized that I could move \"struct delayed_checkout\" and \"enum ce_delay_state\"\n> definition from \"cache.h\" to \"convert.h\" to solve the problem nicely.\n> \n\nBut yeah, maybe you can reduce cache.h size, too :)\n\n> Any objection to this approach?\n> \n"},{"id":"316493","messageId":"F4C4D5CD-5A96-404C-88D0-00BCC78A1C65@gmail.com","threadId":"45641","inReplyTo":"20170410142850.GA23068@starla","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-10T14:52:12Z","receivedAt":"2017-04-10T14:52:19Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 10 Apr 2017, at 16:28, Eric Wong <e@80x24.org> wrote:\n> \n> Lars Schneider <larsxschneider@gmail.com> wrote:\n>>> diff --git a/convert.h b/convert.h\n>>> index 82871a11d5..da6c702090 100644\n>>> --- a/convert.h\n>>> +++ b/convert.h\n>>> @@ -42,6 +42,11 @@ extern int convert_to_git(const char *path, const char *src, size_t len,\n>>> \t\t\t  struct strbuf *dst, enum safe_crlf checksafe);\n>>> extern int convert_to_working_tree(const char *path, const char *src,\n>>> \t\t\t\t   size_t len, struct strbuf *dst);\n>>> +extern int async_convert_to_working_tree(const char *path, const char *src,\n>>> +\t\t\t\t\t size_t len, struct strbuf *dst,\n>>> +\t\t\t\t\t void *dco);\n>>> \n>> \n>> I don't like the void pointer here. However, \"cache.h\" includes \"convert.h\" and\n>> therefore \"convert.h\" cannot include \"cache.h\". That's why \"convert.h\" doesn't\n>> know about \"struct delayed_checkout\". \n> \n> You can forward declare the struct without fields in convert.h:\n\nOMG. Of course. Now I feel stupid.\n\n>> I just realized that I could move \"struct delayed_checkout\" and \"enum ce_delay_state\"\n>> definition from \"cache.h\" to \"convert.h\" to solve the problem nicely.\n>> \n> \n> But yeah, maybe you can reduce cache.h size, too :)\n\nYeah, then I will do this in the next round.\n\nThanks,\nLars"},{"id":"316516","messageId":"a7fd3bef-49b2-0b0a-8ca4-89e41a402661@web.de","threadId":"45641","inReplyTo":"20170409191107.20547-5-larsxschneider@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-10T20:54:19Z","receivedAt":"2017-04-10T20:54:40Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2017-04-09 21:11, Lars Schneider wrote:\n[]\n> +------------------------\n> +packet:          git> command=smudge\n> +packet:          git> pathname=path/testfile.dat\n> +packet:          git> delay-able=1\n> +packet:          git> 0000\n> +packet:          git> CONTENT\n> +packet:          git> 0000\n> +packet:          git< status=delayed\n> +packet:          git< 0000\n> +packet:          git> delay-id=1\n> +packet:          git> 0000\n> +packet:          git< status=success\n> +packet:          git< 0000\n\n(not sure if this was mentioned before)\nIf a filter uses the delayed feature, I would read it as\na response from the filter in the style:\n\"Hallo Git, I need some time to process it, but as I have\nCPU capacity available, please send another blob,\nso I can chew them in parallel.\"\n\nCan we save one round trip ?\n\npacket:          git> command=smudge\npacket:          git> pathname=path/testfile.dat\npacket:          git> delay-id=1\npacket:          git> 0000\npacket:          git> CONTENT\npacket:          git> 0000\npacket:          git< status=delayed # this means: Git, please feed more\npacket:          git> 0000\n\n# Git feeds the next blob.\n# This may be repeated some rounds.\n# (We may want to restrict the number of rounds for Git, see below)\n# After these some rounds, the filter needs to signal:\n# no more fresh blobs please, collect some data and I can free memory\n# and after that I am able to get a fresh blob.\npacket:          git> command=smudge\npacket:          git> pathname=path/testfile.dat\npacket:          git> delay-id=2\npacket:          git> 0000\npacket:          git> CONTENT\npacket:          git> 0000\npacket:          git< status=pleaseWait\npacket:          git> 0000\n\n# Now Git needs to ask for ready blobs.\n\n> +------------------------\n> +\n> +If the filter supports the \"delay\" capability then it must support the\n> +\"list_available_blobs\" command. If Git sends this command, then the\n> +filter is expected to return a list of \"delay_ids\" of blobs that are\n> +available. The list must be terminated with a flush packet followed\n> +by a \"success\" status that is also terminated with a flush packet. If\n> +no blobs for the delayed paths are available, yet, then the filter is\n> +expected to block the response until at least one blob becomes\n> +available. The filter can tell Git that it has no more delayed blobs\n> +by sending an empty list.\n> +------------------------\n> +packet:          git> command=list_available_blobs\n> +packet:          git> 0000\n> +packet:          git< 7\n\nIs the \"7\" the same as the \"delay-id=1\" from above?\nIt may be easier to understand, even if it costs some bytes, to answer instead\npacket:          git< delay-id=1\n(And at this point, may I suggest to change \"delay-id\" into \"request-id=1\" ?\n\n> +packet:          git< 13\n\nSame question here: is this the delay-id ?\n\n> +packet:          git< 0000\n> +packet:          git< status=success\n> +packet:          git< 0000\n> +------------------------\n> +\n> +After Git received the \"delay_ids\", it will request the corresponding\n> +blobs again. These requests contain a \"delay-id\" and an empty content\n> +section. The filter is expected to respond with the smudged content\n> +in the usual way as explained above.\n> +------------------------\n> +packet:          git> command=smudge\n> +packet:          git> pathname=test-delay10.a\n> +packet:          git> delay-id=0\n\nMinor question: Where does the \"id=0\" come from ?\n\n> +packet:          git> 0000\n> +packet:          git> 0000  # empty content!\n> +packet:          git< status=success\n> +packet:          git< 0000\n> +packet:          git< SMUDGED_CONTENT\n> +packet:          git< 0000\n> +packet:          git< 0000\n\nOK, good.\n\nThe quest is: what happens next ?\n\n2 things, kind of in parallel, but we need to prioritize and serialize:\n- Send the next blob\n- Fetch ready blobs\n- And of course: ask for more ready blobs.\n(it looks as if Peff and Jakub had useful comments already,\n  so I can stop here?)\n\n\nIn general, Git should not have a unlimited number of blobs outstanding,\nas memory constraints may apply.\nThere may be a config variable for the number of outstanding blobs,\n(similar to the window size in other protocols) and a variable\nfor the number of \"send bytes in outstanding blobs\"\n(similar to window size (again!) in e.g TCP)\n\nThe number of outstanding blobs is may be less important, and it is more\nimportant to monitor the number of bytes we keep in memory in some way.\n\nSomething like \"we set a limit to 500K of out standng data\", once we are\nabove the limit, don't send any new blobs.\n\n\n\n(No, I didn't look at the code at all, this protocol discussion\nis much more interesting)\n\n\n"},{"id":"316631","messageId":"388C3F2A-AC77-499F-9C74-216F5DC00FD8@gmail.com","threadId":"45641","inReplyTo":"a7fd3bef-49b2-0b0a-8ca4-89e41a402661@web.de","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-11T19:50:14Z","receivedAt":"2017-04-11T19:50:22Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 10 Apr 2017, at 22:54, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> On 2017-04-09 21:11, Lars Schneider wrote:\n> []\n>> +------------------------\n>> +packet:          git> command=smudge\n>> +packet:          git> pathname=path/testfile.dat\n>> +packet:          git> delay-able=1\n>> +packet:          git> 0000\n>> +packet:          git> CONTENT\n>> +packet:          git> 0000\n>> +packet:          git< status=delayed\n>> +packet:          git< 0000\n>> +packet:          git> delay-id=1\n>> +packet:          git> 0000\n>> +packet:          git< status=success\n>> +packet:          git< 0000\n> \n> (not sure if this was mentioned before)\n> If a filter uses the delayed feature, I would read it as\n> a response from the filter in the style:\n> \"Hallo Git, I need some time to process it, but as I have\n> CPU capacity available, please send another blob,\n> so I can chew them in parallel.\"\n> \n> Can we save one round trip ?\n> \n> packet:          git> command=smudge\n> packet:          git> pathname=path/testfile.dat\n> packet:          git> delay-id=1\n> packet:          git> 0000\n> packet:          git> CONTENT\n> packet:          git> 0000\n> packet:          git< status=delayed # this means: Git, please feed more\n> packet:          git> 0000\n\nActually, this is how I implemented it first.\n\nHowever, I didn't like that because we associate a\npathname with a delay-id. If the filter does not\ndelay the item then we associate a different\npathname with the same delay-id in the next request. \nTherefore I think it is better to present the delay-id \n*only* to the filter if the item is actually delayed.\n\nI would be surprised if the extra round trip does impact\nthe performance in any meaningful way.\n\n\n> # Git feeds the next blob.\n> # This may be repeated some rounds.\n> # (We may want to restrict the number of rounds for Git, see below)\n> # After these some rounds, the filter needs to signal:\n> # no more fresh blobs please, collect some data and I can free memory\n> # and after that I am able to get a fresh blob.\n> packet:          git> command=smudge\n> packet:          git> pathname=path/testfile.dat\n> packet:          git> delay-id=2\n> packet:          git> 0000\n> packet:          git> CONTENT\n> packet:          git> 0000\n> packet:          git< status=pleaseWait\n> packet:          git> 0000\n> \n> # Now Git needs to ask for ready blobs.\n\nWe could do this but I think this would only complicate\nthe protocol. I expect the filter to spool results to the\ndisk or something.\n\n\n>> +------------------------\n>> +\n>> +If the filter supports the \"delay\" capability then it must support the\n>> +\"list_available_blobs\" command. If Git sends this command, then the\n>> +filter is expected to return a list of \"delay_ids\" of blobs that are\n>> +available. The list must be terminated with a flush packet followed\n>> +by a \"success\" status that is also terminated with a flush packet. If\n>> +no blobs for the delayed paths are available, yet, then the filter is\n>> +expected to block the response until at least one blob becomes\n>> +available. The filter can tell Git that it has no more delayed blobs\n>> +by sending an empty list.\n>> +------------------------\n>> +packet:          git> command=list_available_blobs\n>> +packet:          git> 0000\n>> +packet:          git< 7\n> \n> Is the \"7\" the same as the \"delay-id=1\" from above?\n\nYes! Sorry, I will make this more clear in the next round.\n\n\n> It may be easier to understand, even if it costs some bytes, to answer instead\n> packet:          git< delay-id=1\n\nAgreed!\n\n\n> (And at this point, may I suggest to change \"delay-id\" into \"request-id=1\" ?\n\nIf there is no objection by another reviewer then I am happy to change it.\n\n\n>> +packet:          git< 13\n> \n> Same question here: is this the delay-id ?\n\nYes.\n\n\n>> +packet:          git< 0000\n>> +packet:          git< status=success\n>> +packet:          git< 0000\n>> +------------------------\n>> +\n>> +After Git received the \"delay_ids\", it will request the corresponding\n>> +blobs again. These requests contain a \"delay-id\" and an empty content\n>> +section. The filter is expected to respond with the smudged content\n>> +in the usual way as explained above.\n>> +------------------------\n>> +packet:          git> command=smudge\n>> +packet:          git> pathname=test-delay10.a\n>> +packet:          git> delay-id=0\n> \n> Minor question: Where does the \"id=0\" come from ?\n\nThat's the delay-id (aka request-id) that Git gave to the filter\non the first request (which was delayed).\n\n\n>> +packet:          git> 0000\n>> +packet:          git> 0000  # empty content!\n>> +packet:          git< status=success\n>> +packet:          git< 0000\n>> +packet:          git< SMUDGED_CONTENT\n>> +packet:          git< 0000\n>> +packet:          git< 0000\n> \n> OK, good.\n> \n> The quest is: what happens next ?\n> \n> 2 things, kind of in parallel, but we need to prioritize and serialize:\n> - Send the next blob\n> - Fetch ready blobs\n> - And of course: ask for more ready blobs.\n> (it looks as if Peff and Jakub had useful comments already,\n>  so I can stop here?)\n\nI would like to keep the mechanism as follows:\n\n1. sends all blobs to the filter\n2. fetch blobs until we are done\n\n@Taylor: Do you think that would be OK for LFS?\n\n\n> In general, Git should not have a unlimited number of blobs outstanding,\n> as memory constraints may apply.\n> There may be a config variable for the number of outstanding blobs,\n> (similar to the window size in other protocols) and a variable\n> for the number of \"send bytes in outstanding blobs\"\n> (similar to window size (again!) in e.g TCP)\n> \n> The number of outstanding blobs is may be less important, and it is more\n> important to monitor the number of bytes we keep in memory in some way.\n> \n> Something like \"we set a limit to 500K of out standng data\", once we are\n> above the limit, don't send any new blobs.\n\nI don't expect the filter to keep everything in memory. If there is no memory\nanymore then I expect the filter to spool to disk. This keeps the protocol simple. \nIf this turns out to be not sufficient then we could improve that later, too.\n\n\nThanks,\nLars"},{"id":"316667","messageId":"106c2be9-c558-edcc-2d97-5091c15010d1@web.de","threadId":"45641","inReplyTo":"388C3F2A-AC77-499F-9C74-216F5DC00FD8@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-12T04:37:11Z","receivedAt":"2017-04-12T04:37:44Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2017-04-11 21:50, Lars Schneider wrote:\n\n[]\n>> packet:          git> command=smudge\n>> packet:          git> pathname=path/testfile.dat\n>> packet:          git> delay-id=1\n>> packet:          git> 0000\n>> packet:          git> CONTENT\n>> packet:          git> 0000\n>> packet:          git< status=delayed # this means: Git, please feed more\n>> packet:          git> 0000\n> Actually, this is how I implemented it first.\n> \n> However, I didn't like that because we associate a\n> pathname with a delay-id. If the filter does not\n> delay the item then we associate a different\n> pathname with the same delay-id in the next request. \n> Therefore I think it is better to present the delay-id \n> *only* to the filter if the item is actually delayed.\n> \n> I would be surprised if the extra round trip does impact\n> the performance in any meaningful way.\n> \n\n2 spontanous remarks:\n\n- Git can simply use a counter which is incremented by each blob\n  that is send to the filter.\n  Regardless what the filter answers (delayed or not), simply increment a\n  counter. (or is this too simple and I miss something?)\n\n- I was thinking that the filter code is written as either \"never delay\" or\n  \"always delay\".\n  \"Never delay\" is the existing code.\n  What is your idea, when should a filter respond with delayed ?\n  My thinking was \"always\", silently assuming the more than one core can be\n  used, so that blobs can be filtered in parallel.\n\n>We could do this but I think this would only complicate\n>the protocol. I expect the filter to spool results to the\n>disk or something.\n  Spooling things to disk was not part of my picture, to be honest.\n  This means additional execution time when a SSD is used, the chips\n  are more worn out...\n  There may be situations, where this is OK for some users (but not for others)\n  How can we prevent Git from (over-) flooding the filter?\n  The protocol response from the filter would be just \"delayed\", and the filter\n  would block Git, right ?\n  But, in any case, it would still be nice if Git would collect converted blobs\n  from the filter, to free resource here.\n  This is more like the \"streaming model\", but on a higher level:\n  Send 4 blobs to the filter, collect the ready one, send the 5th blob to\n  the filter, collect the ready one, send the 6th blob to the filter, collect\n  ready one....\n\n\n(Back to the roots)\nWhich criteria do you have in mind: When should a filter process the blob\nand return it immediately, and when would it respond \"delayed\" ?\n\n\n\n\n"},{"id":"316693","messageId":"20170412173404.GA49694@Ida","threadId":"45641","inReplyTo":"388C3F2A-AC77-499F-9C74-216F5DC00FD8@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Taylor Blau","fromEmail":"ttaylorr@github.com","sentAt":"2017-04-12T17:34:04Z","receivedAt":"2017-04-12T17:34:17Z","isPatch":true,"sender":{"key":"ttaylorr@github.com","avatar":"https://gravatar.com/avatar/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?d=mp&s=160"},"body":"> > (And at this point, may I suggest to change \"delay-id\" into \"request-id=1\" ?\n>\n> If there is no objection by another reviewer then I am happy to change it.\n\nI think \"delay-id\" may be more illustrative of what's occurring in this request.\nThat being said, my preference would be that we remove the\n\"delay-id\"/\"request-id\" entirely from the protocol, and make Git responsible for\nhandling the path lookup by a hashmap.\n\nIs the concern that a hashmap covering all entries in a large checkout would be\ntoo large to keep in memory? If so, keeping an opaque ID as a part of the\nprotocol is something I would not object to.\n\n> >> +packet:          git> 0000\n> >> +packet:          git> 0000  # empty content!\n> >> +packet:          git< status=success\n> >> +packet:          git< 0000\n> >> +packet:          git< SMUDGED_CONTENT\n> >> +packet:          git< 0000\n> >> +packet:          git< 0000\n> >\n> > OK, good.\n> >\n> > The quest is: what happens next ?\n> >\n> > 2 things, kind of in parallel, but we need to prioritize and serialize:\n> > - Send the next blob\n> > - Fetch ready blobs\n> > - And of course: ask for more ready blobs.\n> > (it looks as if Peff and Jakub had useful comments already,\n> >  so I can stop here?)\n>\n> I would like to keep the mechanism as follows:\n>\n> 1. sends all blobs to the filter\n> 2. fetch blobs until we are done\n>\n> @Taylor: Do you think that would be OK for LFS?\n\nI think that this would be fine for LFS and filters of this kind in general. For\nLFS in particular, my initial inclination would be to have the protocol open\nsupport writing blob data back to Git at anytime during the checkout process,\nnot just after all blobs have been sent to the filter.\n\nThat being said, I don't think this holds up in practice. The blobs are too big\nto fit in memory anyway, and will just end up getting written to LFS's object\ncache in .git/lfs/objects.\n\nSince they're already in there, all we would have to do is keep the list of\n`readyIds map[int]*os.File` in memory (or even map int -> LFS OID, and open the\nfile later), and then `io.Copy()` from the open file back to Git.\n\nThis makes me think of adding another capability to the protocol, which would\njust be exchanging paths on disk in `/tmp` or any other directory so that we\nwouldn't have to stream content over the pipe. Instead of responding with\n\n    packet:          git< status=success\n    packet:          git< 0000\n    packet:          git< SMUDGED_CONTENT\n    packet:          git< 0000\n    packet:          git< 0000\n\nWe could respond with:\n\n    packet:          git< status=success\n    packet:          git< 0000\n    packet:          git< /path/to/contents.dat # <-\n    packet:          git< 0000\n    packet:          git< 0000\n\nGit would then be responsible for opening that file on disk (the filter would\nguarantee that to be possible), and then copying its contents into the working\ntree.\n\nI think that's a topic for later discussion, though :-).\n\n> > In general, Git should not have a unlimited number of blobs outstanding,\n> > as memory constraints may apply.\n> > There may be a config variable for the number of outstanding blobs,\n> > (similar to the window size in other protocols) and a variable\n> > for the number of \"send bytes in outstanding blobs\"\n> > (similar to window size (again!) in e.g TCP)\n> >\n> > The number of outstanding blobs is may be less important, and it is more\n> > important to monitor the number of bytes we keep in memory in some way.\n> >\n> > Something like \"we set a limit to 500K of out standng data\", once we are\n> > above the limit, don't send any new blobs.\n>\n> I don't expect the filter to keep everything in memory. If there is no memory\n> anymore then I expect the filter to spool to disk. This keeps the protocol simple.\n> If this turns out to be not sufficient then we could improve that later, too.\n\nAgree.\n\n\n--\nThanks,\nTaylor Blau\n"},{"id":"316694","messageId":"20170412174610.GB49694@Ida","threadId":"45641","inReplyTo":"20170409191107.20547-5-larsxschneider@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Taylor Blau","fromEmail":"ttaylorr@github.com","sentAt":"2017-04-12T17:46:10Z","receivedAt":"2017-04-12T17:47:27Z","isPatch":true,"sender":{"key":"ttaylorr@github.com","avatar":"https://gravatar.com/avatar/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?d=mp&s=160"},"body":"I think this is a great approach and one that I'd be happy to implement in LFS.\nThe additional capability isn't too complex, so I think other similar filters to\nLFS shouldn't have a hard time implementing it either.\n\nI left a few comments, mostly expressing approval to the documentation changes.\nI'll leave the C review to someone more expert than me.\n\n+1 from me on the protocol changes.\n\n> +Delay\n> +^^^^^\n> +\n> +If the filter supports the \"delay\" capability, then Git can send the\n> +flag \"delay-able\" after the filter command and pathname.\n\nNit: I think either way is fine, but `can_delay` will save us 1 byte per each\nnew checkout entry.\n\n> +\"delay-id\", a number that identifies the blob, and a flush packet. The\n> +filter acknowledges this number with a \"success\" status and a flush\n> +packet.\n\nI mentioned this in another thread, but I'd prefer, if possible, that we use the\npathname as a unique identifier for referring back to a particular checkout\nentry. I think adding an additional identifier adds unnecessary complication to\nthe protocol and introduces a forced mapping on the filter side from id to\npath.\n\nBoth Git and the filter are going to have to keep these paths in memory\nsomewhere, be that in-process, or on disk. That being said, I can see potential\ntroubles with a large number of long paths that exceed the memory available to\nGit or the filter when stored in a hashmap/set.\n\nOn Git's side, I think trading that for some CPU time might make sense. If Git\nwere to SHA1 each path and store that in a hashmap, it would consume more CPU\ntime, but less memory to store each path. Git and the filter could then exchange\npath names, and Git would simply SHA1 the pathname each time it needed to refer\nback to memory associated with that entry in a hashmap.\n\n> +by a \"success\" status that is also terminated with a flush packet. If\n> +no blobs for the delayed paths are available, yet, then the filter is\n> +expected to block the response until at least one blob becomes\n> +available. The filter can tell Git that it has no more delayed blobs\n> +by sending an empty list.\n\nI think the blocking nature of the `list_available_blobs` command is a great\nsynchronization mechanism for telling the filter when it can and can't send new\nblob information to Git. I was tenatively thinking of suggesting to remove this\ncommand and instead allow the filter to send readied blobs in sequence after all\nunique checkout entries had been sent from Git to the filter. But I think this\nallows approach allows us more flexibility, and isn't that much extra\ncomplication or bytes across the pipe.\n\n\n--\nThanks,\nTaylor Blau\n"},{"id":"317088","messageId":"638E6914-6B66-4C66-996F-F04A285A2129@gmail.com","threadId":"45641","inReplyTo":"106c2be9-c558-edcc-2d97-5091c15010d1@web.de","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-18T08:53:45Z","receivedAt":"2017-04-18T08:54:09Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 12. Apr 2017, at 06:37, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> On 2017-04-11 21:50, Lars Schneider wrote:\n> \n> []\n>>> packet:          git> command=smudge\n>>> packet:          git> pathname=path/testfile.dat\n>>> packet:          git> delay-id=1\n>>> packet:          git> 0000\n>>> packet:          git> CONTENT\n>>> packet:          git> 0000\n>>> packet:          git< status=delayed # this means: Git, please feed more\n>>> packet:          git> 0000\n>> Actually, this is how I implemented it first.\n>> \n>> However, I didn't like that because we associate a\n>> pathname with a delay-id. If the filter does not\n>> delay the item then we associate a different\n>> pathname with the same delay-id in the next request. \n>> Therefore I think it is better to present the delay-id \n>> *only* to the filter if the item is actually delayed.\n>> \n>> I would be surprised if the extra round trip does impact\n>> the performance in any meaningful way.\n>> \n> \n> 2 spontanous remarks:\n> \n> - Git can simply use a counter which is incremented by each blob\n>  that is send to the filter.\n>  Regardless what the filter answers (delayed or not), simply increment a\n>  counter. (or is this too simple and I miss something?)\n\nI am not sure I understand what you mean. Do you want to say that the filter just starts to delay items with an increasing counter by itself (first item = 0, second item = 1, ...). That could work but I would prefer a more explicitly defined solution. The self counting implicit way could easily get confused/mixed up I think.\n\n\n> - I was thinking that the filter code is written as either \"never delay\" or\n>  \"always delay\".\n>  \"Never delay\" is the existing code.\n>  What is your idea, when should a filter respond with delayed ?\n>  My thinking was \"always\", silently assuming the more than one core can be\n>  used, so that blobs can be filtered in parallel.\n\nIn case of Git LFS I expect the filter to answer with \"delay\" if a network call is required to fulfill the (smudge-) filter request.\n\n\n>> We could do this but I think this would only complicate\n>> the protocol. I expect the filter to spool results to the\n>> disk or something.\n>  Spooling things to disk was not part of my picture, to be honest.\n>  This means additional execution time when a SSD is used, the chips\n>  are more worn out...\n>  There may be situations, where this is OK for some users (but not for others)\n>  How can we prevent Git from (over-) flooding the filter?\n\nI don't think this is Git's responsibility. If the filter can't handle the data then I expect the filter to *not* ask Git for a delay.\n\n\n>  The protocol response from the filter would be just \"delayed\", and the filter\n>  would block Git, right ?\n>  But, in any case, it would still be nice if Git would collect converted blobs\n>  from the filter, to free resource here.\n>  This is more like the \"streaming model\", but on a higher level:\n>  Send 4 blobs to the filter, collect the ready one, send the 5th blob to\n>  the filter, collect the ready one, send the 6th blob to the filter, collect\n>  ready one....\n\nThat sounds nice, but I don't think that it is necessary in this (first) iteration. I think it is acceptable for the filter to spool data to disk.\n\n\n> (Back to the roots)\n> Which criteria do you have in mind: When should a filter process the blob\n> and return it immediately, and when would it respond \"delayed\" ?\n\nSee above: it's up to the filter. In case of Git LFS: delay if a network call is required.\n\nThanks,\nLars"},{"id":"317112","messageId":"1D510C6F-A830-48BE-880B-62F4212F4A7F@gmail.com","threadId":"45641","inReplyTo":"20170412174610.GB49694@Ida","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-04-18T16:14:36Z","receivedAt":"2017-04-18T16:14:49Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 12. Apr 2017, at 19:46, Taylor Blau <ttaylorr@github.com> wrote:\n> \n> I think this is a great approach and one that I'd be happy to implement in LFS.\n> The additional capability isn't too complex, so I think other similar filters to\n> LFS shouldn't have a hard time implementing it either.\n> \n> I left a few comments, mostly expressing approval to the documentation changes.\n> I'll leave the C review to someone more expert than me.\n> \n> +1 from me on the protocol changes.\n\nThanks!\n\n\n>> +Delay\n>> +^^^^^\n>> +\n>> +If the filter supports the \"delay\" capability, then Git can send the\n>> +flag \"delay-able\" after the filter command and pathname.\n> \n> Nit: I think either way is fine, but `can_delay` will save us 1 byte per each\n> new checkout entry.\n\n1 byte is no convincing argument to me but since you are a native speaker I trust your \"can-delay\" suggestion. I prefer dashes over underscores, though, for consistency with the rest of the protocol.\n\n\n>> +\"delay-id\", a number that identifies the blob, and a flush packet. The\n>> +filter acknowledges this number with a \"success\" status and a flush\n>> +packet.\n> \n> I mentioned this in another thread, but I'd prefer, if possible, that we use the\n> pathname as a unique identifier for referring back to a particular checkout\n> entry. I think adding an additional identifier adds unnecessary complication to\n> the protocol and introduces a forced mapping on the filter side from id to\n> path.\n\nI agree! I answered in the other thread. Let's keep the discussion there.\n\n\n> Both Git and the filter are going to have to keep these paths in memory\n> somewhere, be that in-process, or on disk. That being said, I can see potential\n> troubles with a large number of long paths that exceed the memory available to\n> Git or the filter when stored in a hashmap/set.\n> \n> On Git's side, I think trading that for some CPU time might make sense. If Git\n> were to SHA1 each path and store that in a hashmap, it would consume more CPU\n> time, but less memory to store each path. Git and the filter could then exchange\n> path names, and Git would simply SHA1 the pathname each time it needed to refer\n> back to memory associated with that entry in a hashmap.\n\nI would be surprised if this would be necessary. If we filter delay 50,000 files (= a lot!) with a path length of 1000 characters (= very long!) then we would use 50MB plus some hashmap data structures. Modern machines should have enough RAM I would think...\n\nThanks,\nLars"},{"id":"317123","messageId":"20170418174209.GA92973@Ida","threadId":"45641","inReplyTo":"1D510C6F-A830-48BE-880B-62F4212F4A7F@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Taylor Blau","fromEmail":"ttaylorr@github.com","sentAt":"2017-04-18T17:42:09Z","receivedAt":"2017-04-18T17:42:16Z","isPatch":true,"sender":{"key":"ttaylorr@github.com","avatar":"https://gravatar.com/avatar/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?d=mp&s=160"},"body":"On Tue, Apr 18, 2017 at 06:14:36PM +0200, Lars Schneider wrote:\n> > Both Git and the filter are going to have to keep these paths in memory\n> > somewhere, be that in-process, or on disk. That being said, I can see potential\n> > troubles with a large number of long paths that exceed the memory available to\n> > Git or the filter when stored in a hashmap/set.\n> >\n> > On Git's side, I think trading that for some CPU time might make sense. If Git\n> > were to SHA1 each path and store that in a hashmap, it would consume more CPU\n> > time, but less memory to store each path. Git and the filter could then exchange\n> > path names, and Git would simply SHA1 the pathname each time it needed to refer\n> > back to memory associated with that entry in a hashmap.\n>\n> I would be surprised if this would be necessary. If we filter delay 50,000\n> files (= a lot!) with a path length of 1000 characters (= very long!) then we\n> would use 50MB plus some hashmap data structures. Modern machines should have\n> enough RAM I would think...\n\nI agree, and thanks for correcting my thinking here. I ran a simple command to\nget the longest path names in a large repository, as:\n\n  $ find . -type f | awk '{ print length($1) }' | sort -r -n | uniq -c\n\nAnd found a few files close to the 200 character mark as the longest pathnames\nin the repository. I think 50k files at 1k bytes per pathname is quite enough\nhead-room :-).\n\n\n--\nThanks,\nTaylor Blau\n"},{"id":"317262","messageId":"48d9e7f0-15bf-ce0e-aae8-4c3ebd5c07cf@web.de","threadId":"45641","inReplyTo":"638E6914-6B66-4C66-996F-F04A285A2129@gmail.com","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-19T18:55:12Z","receivedAt":"2017-04-19T18:55:44Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"\n>> (Back to the roots)\n>> Which criteria do you have in mind: When should a filter process the blob\n>> and return it immediately, and when would it respond \"delayed\" ?\n> \n> See above: it's up to the filter. In case of Git LFS: delay if a network call is required.\n> \nThat make sense.\nI try to understand the big picture, and from here try to review\nthe details.\nDoes it make sense to mention \"git lfs\" in the commit message,\nand/or add some test code ?\n\n\n"},{"id":"320380","messageId":"7D282F93-6131-4A1E-BCFD-62EA9DDD7F9F@gmail.com","threadId":"45641","inReplyTo":"48d9e7f0-15bf-ce0e-aae8-4c3ebd5c07cf@web.de","subject":"Re: [PATCH v3 4/4] convert: add \"status=delayed\" to filter process protocol","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-05-21T20:25:58Z","receivedAt":"2017-05-21T20:26:07Z","isPatch":true,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 19 Apr 2017, at 20:55, Torsten Bögershausen <tboegi@web.de> wrote:\n> \n> \n>>> (Back to the roots)\n>>> Which criteria do you have in mind: When should a filter process the blob\n>>> and return it immediately, and when would it respond \"delayed\" ?\n>> \n>> See above: it's up to the filter. In case of Git LFS: delay if a network call is required.\n>> \n> That make sense.\n> I try to understand the big picture, and from here try to review\n> the details.\n> Does it make sense to mention \"git lfs\" in the commit message,\n> and/or add some test code ?\n\nI'll mention Git LFS in the commit message. The test code (t/t0021/rot13-filter.pl)\nshould mimic the behavior of Git LFS already.\n\nThanks,\nLars"}]}