Show 73 quoted lines
> Le 4 nov. 2025 à 19:17, Justin Tobler <jltobler@gmail.com> a écrit :
>
> On 25/11/03 08:44PM, Junio C Hamano wrote:
>> Junio C Hamano <gitster@pobox.com> writes:
>>
>>> Justin Tobler <jltobler@gmail.com> writes:
>>>
>>>> I have a usecase where I would like to know exactly which files in a
>>>> diff pair are considered binary by Git when computing diffs. When
>>>> computing patch diff output, Git already omits filepair diffs where at
>>>> least one side is considered binary and prints a "binary files differ"
>>>> message instead. From this message we cannot discern exactly which files
>>>> were considered binary by Git though.
>>>
>>> I have a usecase where I would like to know exactly which side of a
>>> diff filepair ends in an incomplete line in a concise format.
>>>
>>> Should we add yet another column to the raw output to indicate who
>>> is complete and who is incomplete?
>>>
>>> Where does it lead us and when will it stop?
>>>
>>> IOW, yuck ;-).
>>
>> My point being that it will be a huge mistake to do this only by
>> singling a trait that is not so special as if it is very special,
>> only because you have been thinking about it too long (the "ends in
>> an incomplete line" trait is what has been on my mind for the past
>> few days, "this side is binary" may be what you've been thinking
>> about). There are many other things people would want to learn
>> concisely in machine readable format, like "where did the file stop
>> using CRLF line endings and swithced to LF line endings", that are
>> equally plausible as the question you are asking, or the question I
>> would be asking "which commit lost the final newline?"
>
> Completely fair. Having a bunch specific options for special info we
> want to add to the raw diff format would get messy quickly and is not
> very extensible.
>
>> Perhaps an extensible command line option syntax like
>>
>> $ git log --raw-extended=binary,incomplete,crlf,...
>
> I quite like this and agree it would be better to have a single
> extensible option.
>
>> is in order, and the presense of these options would add "tt,ic,cl"
>> somewhere in the output to signal that both sides are text, preimage
>> ends in an incomplete line but not postimage, and preimage uses crlf
>> but postimage uses lf, or something?
>
> Maybe the output should be something like:
>
> binary=tt,incomplete=ic,crlf=cl
>
> or something along those lines. That way we could freely extend in the
> future without having to worry about a specific order. If we think all
> of the raw diff extension modes would only report with yes/no for each
> file we could just do:
>
> binary=yn,incomplete=yy,crlf=nn
>
> but maybe we should be more flexible and leave it up to the mode to
> decide what its values can be?
>
> Also, maybe this info could be on a newline following each raw diff
> entry? Something like:
>
> :100644 100644 a1961526 e231acb1 M foo
> binary=yy
> :100644 100644 31eedd5c 402a70d7 M bar
> binary=nn
>
Whether combined or separate, self-documenting output is nice. Separate might be easier for line-oriented tools? Having to split on commas and loop looking for keywords seems like more work than just processing a line at a time. Idk.