git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v3 10/10] convert: add filter.<driver>.process option

From
Lars Schneider <larsxschneider@gmail.com>
Date
Jul 31, 2016, 19:49 UTC
Message-ID
<6765D972-876A-4F94-A170-468002498296@gmail.com>
In-Reply-To
<69988611-06ec-048d-12e7-7b87882ddc6a@gmail.com>
> On 31 Jul 2016, at 11:42, Jakub Narębski <jnareb@gmail.com> wrote:
> 
> [Excuse me replying to myself, but there are a few things I forgot,
> or realized only later]
No worries :)
Show 25 quoted lines
> 
> W dniu 31.07.2016 o 00:05, Jakub Narębski pisze:
>> W dniu 30.07.2016 o 01:38, larsxschneider@gmail.com pisze:
>>> From: Lars Schneider <larsxschneider@gmail.com>
>>> 
>>> Git's clean/smudge mechanism invokes an external filter process for every
>>> single blob that is affected by a filter. If Git filters a lot of blobs
>>> then the startup time of the external filter processes can become a
>>> significant part of the overall Git execution time.
>>> 
>>> This patch adds the filter.<driver>.process string option which, if used,
>>> keeps the external filter process running and processes all blobs with
>>> the following packet format (pkt-line) based protocol over standard input
>>> and standard output.
>> 
>> I think it would be nice to have here at least summary of the benchmarks
>> you did in https://github.com/github/git-lfs/pull/1382
> 
> Note that this feature is especially useful if startup time is long,
> that is if you are using an operating system with costly fork / new process
> startup time like MS Windows (which you have mentioned), or writing
> filter in a programming language with large startup time like Java
> or Python (the latter may have changed since).
> 
>  https://gnustavo.wordpress.com/2012/06/28/programming-languages-start-up-times/

OK, I will add this. Is it OK to add the link to the commit message? (since I don't know how long the link will be available).

Show 45 quoted lines
> [...]
>> I was thinking about having possible responses to receiving file
>> contents (or starting receiving in the streaming case) to be:
>> 
>>  packet:          git< ok size=7\n    (or "ok 7\n", if size is known)
>> 
>> or
>> 
>>  packet:          git< ok\n           (if filter does not know size upfront)
>> 
>> or
>> 
>>  packet:          git< fail <msg>\n   (or just "fail" + packet with msg)
>> 
>> The last would be when filter knows upfront that it cannot perform
>> the operation.  Though sending an empty file with non-"success" final
>> would work as well.
> 
> [...]
> 
>>> In case the filter cannot process the content, it is expected
>>> to respond with the result content size 0 (only if "stream" is
>>> not defined) and a "reject" packet.
>>> ------------------------
>>> packet:          git< size=0\n    (omitted with capability "stream")
>>> packet:          git< reject\n
>>> ------------------------
>> 
>> This is *wrong* idea!  Empty file, with size=0, can be a perfectly
>> legitimate response.  
> 
> Actually, I think I have misunderstood your intent.  If you want to have
> simpler protocol, with only one place to signal errors, that is after
> sending a response, then proper way of signaling the error condition
> would be to send empty file and then "reject" instead of "success":
> 
>   packet:          git< size=0\n    (omitted with capability "stream")
>   packet:          git< 0000        (we need this flush packet)
>   packet:          git< reject\n
> 
> Otherwise in the case without size upfront (capability "stream")
> file with contents "reject" would be mistaken for the "reject" packet.
> 
> See below for proposal with two places to signal errors: before sending
> first byte, and after.
Right now the protocol is implemented covering the following cases:
## CASE 1 - no stream success

packet: git< size=57\n packet: git< SMUDGED_CONTENT packet: git< 0000 packet: git< success\n

## CASE 2 - no stream success but 0 byte response

packet: git< size=0\n packet: git< success\n

## CASE 3 - no stream filter; filter doesn't want to process the file

packet: git< size=0\n packet: git< reject\n

## CASE 4 - no stream filter; filter error

packet: git< size=57\n packet: git< SMUDGED_CONTENT packet: git< 0000 packet: git< error\n

CASE 4 is not explicitly checked. If a final message is neither "success" nor "reject" then it is interpreted as error. If that happens then Git will shutdown and restart the filter process if there is another file to filter.

Alternatively a filter process can shutdown itself, too, to signal an error.

The corresponding stream filter look like this:
## CASE 1 - stream success

packet: git< SMUDGED_CONTENT packet: git< 0000 packet: git< success\n

## CASE 2 - stream success but 0 byte response

packet: git< 0000 packet: git< success\n

## CASE 3 - stream filter; filter doesn't want to process the file

packet: git< 0000 packet: git< reject\n

## CASE 4 - stream filter; filter error

packet: git< SMUDGED_CONTENT packet: git< 0000 packet: git< error\n

--

I just realized that the size 0 case is a bit inconsistent in the no stream case as it has no flush packet. Maybe I should indeed remove the flush packet in the no stream case completely?!

Do the cases above make sense to you?

Regarding error handling. I would prefer it if the filter prints all errors to STDERR by itself. I think that is the safest option to communicate errors to the users because if the communication got into a bad state then Git might not be able to read the errors properly.

See Peff's response on the topic, too: http://public-inbox.org/git/20160729165018.GA6553%40sigill.intra.peff.net/

> NOTE: there is a bit of mixed and possibly confusing notation, that
> is 0000 is flush packet, not packet with 0000 as content.  Perhaps
> write pkt-line in full?

I am not sure I understand what you mean (maybe it's too late for me...). Can you try to rephrase or give an example?

Thank you, Lars

Show 46 quoted lines
> 
> 
> [...]
>>> ---
>>> Documentation/gitattributes.txt |  84 ++++++++-
>>> convert.c                       | 400 +++++++++++++++++++++++++++++++++++++--
>>> t/t0021-conversion.sh           | 405 ++++++++++++++++++++++++++++++++++++++++
>>> t/t0021/rot13-filter.pl         | 177 ++++++++++++++++++
>>> 4 files changed, 1053 insertions(+), 13 deletions(-)
>>> create mode 100755 t/t0021/rot13-filter.pl
> 
> Wouldn't it be better for easier review to split it into separate patches?
> Perhaps at least the new test...
> 
> [...]
>> I would assume that we have two error conditions.  
>> 
>> First situation is when the filter knows upfront (after receiving name
>> and size of file, and after receiving contents for not-streaming filters)
>> that it cannot process the file (like e.g. LFS filter with artifactory
>> replica/shard being a bit behind master, and not including contents of
>> the file being filtered).
>> 
>> My proposal is to reply with "fail" _in place of_ size of reply:
>> 
>>   packet:         git< fail\n       (any case: size known or not, stream or not)
>> 
>> It could be "reject", or "error" instead of "fail".
>> 
>> 
>> Another situation is if filter encounters error during output,
>> either with streaming filter (or non-stream, but not storing whole
>> input upfront) realizing in the middle of output that there is something
>> wrong with input (e.g. converting between encoding, and encountering
>> character that cannot be represented in output encoding), or e.g. filter
>> process being killed, or network connection dropping with LFS filter, etc.
>> The filter has send some packets with output already.  In this case
>> filter should flush, and send "reject" or "error" packet.
>> 
>>   <error condition>
>>   packet:         git< "0000"       (flush packet)
>>   packet:         git< reject\n
>> 
>> Should there be a place for an error message, or would standard error
>> (stderr) be used for this?
> 
Previous: Jakub NarębskiNext: Jakub Narębski
Message 37 of 100 in “Git filter protocol”
  1. 00/10 Git filter protocollarsxschneider@gmail.com, Jul 29, 2016
  2. 01/10 pkt-line: extract set_packet_header()larsxschneider@gmail.com, Jul 29, 2016
  3. Jakub NarębskiJul 30, 2016
  4. Lars SchneiderAug 1, 2016
  5. Jakub NarębskiAug 3, 2016
  6. Lars SchneiderAug 5, 2016
  7. 02/10 pkt-line: add direct_packet_write() and direct_packet_write_data()larsxschneider@gmail.com, Jul 29, 2016
  8. Jakub NarębskiJul 30, 2016
  9. Lars SchneiderAug 1, 2016
  10. Jakub NarębskiAug 3, 2016
  11. Lars SchneiderAug 5, 2016
  12. 03/10 pkt-line: add packet_flush_gentle()larsxschneider@gmail.com, Jul 29, 2016
  13. Jakub NarębskiJul 30, 2016
  14. Lars SchneiderAug 1, 2016
  15. Torstem BögershausenJul 31, 2016
  16. Lars SchneiderJul 31, 2016
  17. Torsten BögershausenAug 2, 2016
  18. Lars SchneiderAug 5, 2016
  19. 04/10 pkt-line: call packet_trace() only if a packet is actually sendlarsxschneider@gmail.com, Jul 29, 2016
  20. Jakub NarębskiJul 30, 2016
  21. Lars SchneiderAug 1, 2016
  22. Jakub NarębskiAug 3, 2016
  23. 05/10 pack-protocol: fix maximum pkt-line sizelarsxschneider@gmail.com, Jul 29, 2016
  24. Jakub NarębskiJul 30, 2016
  25. Lars SchneiderAug 1, 2016
  26. 07/10 convert: quote filter names in error messageslarsxschneider@gmail.com, Jul 29, 2016
  27. 06/10 run-command: add clean_on_exit_handlerlarsxschneider@gmail.com, Jul 29, 2016
  28. Johannes SixtJul 30, 2016
  29. Lars SchneiderAug 1, 2016
  30. Johannes SixtAug 2, 2016
  31. Lars SchneiderAug 2, 2016
  32. 08/10 convert: modernize testslarsxschneider@gmail.com, Jul 29, 2016
  33. 09/10 convert: generate large test files only oncelarsxschneider@gmail.com, Jul 29, 2016
  34. 10/10 convert: add filter.<driver>.process optionlarsxschneider@gmail.com, Jul 29, 2016
  35. Jakub NarębskiJul 30, 2016
  36. Jakub NarębskiJul 31, 2016
  37. Lars SchneiderJul 31, 2016
  38. Jakub NarębskiJul 31, 2016
  39. Lars SchneiderAug 1, 2016
  40. Designing the filter process protocol (was: Re: [PATCH v3 10/10] convert: add filter.<driver>.process option)Jakub Narębski, Aug 3, 2016
  41. Lars SchneiderAug 5, 2016
  42. Lars SchneiderAug 6, 2016
  43. Jakub NarębskiAug 3, 2016
  44. Jakub NarębskiJul 31, 2016
  45. Lars SchneiderAug 1, 2016
  46. Jakub NarębskiAug 4, 2016
  47. Lars SchneiderAug 3, 2016
  48. Jakub NarębskiAug 4, 2016
  49. Lars SchneiderAug 5, 2016
  50. 00/12 Git filter protocollarsxschneider@gmail.com, Aug 3, 2016
  51. 01/12 pkt-line: extract set_packet_header()larsxschneider@gmail.com, Aug 3, 2016
  52. Junio C HamanoAug 3, 2016
  53. Jeff KingAug 3, 2016
  54. Jeff KingAug 3, 2016
  55. Junio C HamanoAug 4, 2016
  56. Lars SchneiderAug 5, 2016
  57. Junio C HamanoAug 5, 2016
  58. Lars SchneiderAug 5, 2016
  59. Junio C HamanoAug 5, 2016
  60. Lars SchneiderAug 3, 2016
  61. 07/12 run-command: add clean_on_exit_handlerlarsxschneider@gmail.com, Aug 3, 2016
  62. Jeff KingAug 3, 2016
  63. Lars SchneiderAug 3, 2016
  64. Jeff KingAug 3, 2016
  65. Lars SchneiderAug 3, 2016
  66. Jeff KingAug 3, 2016
  67. Lars SchneiderAug 5, 2016
  68. Torsten BögershausenAug 5, 2016
  69. Lars SchneiderAug 5, 2016
  70. 03/12 pkt-line: add packet_flush_gentle()larsxschneider@gmail.com, Aug 3, 2016
  71. Jeff KingAug 3, 2016
  72. Junio C HamanoAug 4, 2016
  73. 02/12 pkt-line: add direct_packet_write() and direct_packet_write_data()larsxschneider@gmail.com, Aug 3, 2016
  74. 08/12 convert: quote filter names in error messageslarsxschneider@gmail.com, Aug 3, 2016
  75. 05/12 pkt-line: add functions to read/write flush terminated packet streamslarsxschneider@gmail.com, Aug 3, 2016
  76. 09/12 convert: modernize testslarsxschneider@gmail.com, Aug 3, 2016
  77. 12/12 convert: add filter.<driver>.process shutdown command optionlarsxschneider@gmail.com, Aug 3, 2016
  78. 06/12 pack-protocol: fix maximum pkt-line sizelarsxschneider@gmail.com, Aug 3, 2016
  79. 04/12 pkt-line: call packet_trace() only if a packet is actually sendlarsxschneider@gmail.com, Aug 3, 2016
  80. 11/12 convert: add filter.<driver>.process optionlarsxschneider@gmail.com, Aug 3, 2016
  81. Junio C HamanoAug 3, 2016
  82. Lars SchneiderAug 3, 2016
  83. Jeff KingAug 3, 2016
  84. Lars SchneiderAug 5, 2016
  85. Junio C HamanoAug 3, 2016
  86. Lars SchneiderAug 3, 2016
  87. Junio C HamanoAug 3, 2016
  88. Lars SchneiderAug 3, 2016
  89. Torsten BögershausenAug 5, 2016
  90. Lars SchneiderAug 5, 2016
  91. Junio C HamanoAug 5, 2016
  92. Jeff KingAug 5, 2016
  93. Lars SchneiderAug 6, 2016
  94. Jeff KingAug 6, 2016
  95. Lars SchneiderAug 6, 2016
  96. Jeff KingAug 8, 2016
  97. Lars SchneiderAug 8, 2016
  98. Jeff KingAug 8, 2016
  99. Torsten BögershausenAug 6, 2016
  100. 10/12 convert: generate large test files only oncelarsxschneider@gmail.com, Aug 3, 2016

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.