{"thread":{"id":"56032","subject":"Structured (ie: json) output for query commands?","startedAt":"2021-06-30T17:00:41Z","lastAt":"2021-07-02T13:21:10Z","messageCount":11,"participants":["Martin Langhoff","Jeff King","brian m. carlson","Han-Wen Nienhuys","Ævar Arnfjörð Bjarmason"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"428833","messageId":"CACPiFCLzsiUjx-vm-dcd=0E8HezMWkErPyS==OQ7OhaXqR6CUA@mail.gmail.com","threadId":"56032","inReplyTo":"CACPiFC++fG-WL8uvTkiydf3wD8TY6dStVpuLcKA9cX_EnwoHGA@mail.gmail.com","subject":"Structured (ie: json) output for query commands?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2021-06-30T17:00:24Z","receivedAt":"2021-06-30T17:00:41Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"Hi Git Gang,\n\nI'm used to automating git by parsing stuff in Unix shell style. But\nthe golden days that gave us Perl are well behind us.\n\nBefore I write (unreliable, leaky, likely buggy) text parsing bits one\nmore time, has git grown formatted output? I'm aware of fmt and\nfriends, I'm thinking of something that handles escaping, nested data\nstructures, etc.\n\nSomething like \"git shortlog --json\".\n\nHave there been discussions? Patches? (I don't spot anything in what's cooking).\n\n\nm\n-- \n martin.langhoff@gmail.com\n - ask interesting questions  ~  http://linkedin.com/in/martinlanghoff\n - don't be distracted        ~  http://github.com/martin-langhoff\n   by shiny stuff\n"},{"id":"428844","messageId":"YNyxD4qAHmbluNRe@coredump.intra.peff.net","threadId":"56032","inReplyTo":"CACPiFCLzsiUjx-vm-dcd=0E8HezMWkErPyS==OQ7OhaXqR6CUA@mail.gmail.com","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-06-30T17:59:43Z","receivedAt":"2021-06-30T17:59:46Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jun 30, 2021 at 01:00:24PM -0400, Martin Langhoff wrote:\n\n> I'm used to automating git by parsing stuff in Unix shell style. But\n> the golden days that gave us Perl are well behind us.\n> \n> Before I write (unreliable, leaky, likely buggy) text parsing bits one\n> more time, has git grown formatted output? I'm aware of fmt and\n> friends, I'm thinking of something that handles escaping, nested data\n> structures, etc.\n> \n> Something like \"git shortlog --json\".\n> \n> Have there been discussions? Patches? (I don't spot anything in what's cooking).\n\nIt's been discussed off-and-on over the years, but I don't think\nanybody's actively working on it.\n\nThe trace2 facility has some json output. That's probably not helpful\nfor what you're doing, but it does mean we have basic json output\nroutines available.\n\nOne complication we faced is that a lot of Git's data is bag-of-bytes,\nnot utf8. And json technically requires utf8. I don't remember if we\nsimply fudged that and output possibly non-utf8 sequences, or if we\nactually encode them.\n\n-Peff\n"},{"id":"428848","messageId":"CACPiFC+F9P1DY_Dgt4+Z-U4o5WRbRduq60+Df0+FHBn6=XL2hw@mail.gmail.com","threadId":"56032","inReplyTo":"YNyxD4qAHmbluNRe@coredump.intra.peff.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2021-06-30T18:20:09Z","receivedAt":"2021-06-30T18:20:22Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Jun 30, 2021 at 1:59 PM Jeff King <peff@peff.net> wrote:\n> It's been discussed off-and-on over the years, but I don't think\n> anybody's actively working on it.\n\nthanks!\n\n> The trace2 facility has some json output. That's probably not helpful\n> for what you're doing, but it does mean we have basic json output\n> routines available.\n\nCool. I see json-writer.c; right.\n\n> One complication we faced is that a lot of Git's data is bag-of-bytes,\n\nGreat point -- hadn't thought of that. Don't see anything in\njson-writer.c but we do use iconv already.\n\ncheers,\n\n\nm\n-- \n martin.langhoff@gmail.com\n - ask interesting questions  ~  http://linkedin.com/in/martinlanghoff\n - don't be distracted        ~  http://github.com/martin-langhoff\n   by shiny stuff\n"},{"id":"428851","messageId":"YNzR5ZZDTfcN2Q+s@camp.crustytoothpaste.net","threadId":"56032","inReplyTo":"YNyxD4qAHmbluNRe@coredump.intra.peff.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-06-30T20:19:49Z","receivedAt":"2021-06-30T20:20:39Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-06-30 at 17:59:43, Jeff King wrote:\n> One complication we faced is that a lot of Git's data is bag-of-bytes,\n> not utf8. And json technically requires utf8. I don't remember if we\n> simply fudged that and output possibly non-utf8 sequences, or if we\n> actually encode them.\n\nI think we just emit invalid UTF-8 in that case, which is a problem.\nThat's why Git is not well suited to JSON output and why it isn't a good\nchoice for structured data here.  I'd like us not to do more JSON in our\ncodebase, since it's practically impossible for users to depend on our\noutput if we do that due to encoding issues[0].\n\nWe could emit data in a different format, such as YAML, which does have\nencoding for arbitrary byte sequences.  However, in YAML, binary data is\nalways base64 encoded, which is less readable, although still\ninterchangeable.  CBOR is also a possibility, although it's not human\nreadable at all.\n\nI'm personally fine with the ad-hoc approach we use now, which is\nactually very convenient to script and, in my view, not to terrible to\nparse in other tools and languages.  Your mileage may vary, though.\n\n[0] I worked on a codebase for many years that exploited its JSON parser\nnot requiring UTF-8 and it was a colossal mess that I'd like us not to\nrepeat.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"428861","messageId":"CACPiFCLAm+e_JRc=eAt_rqwCWASqr19nXdk0zWqLbTEgELL81w@mail.gmail.com","threadId":"56032","inReplyTo":"YNzR5ZZDTfcN2Q+s@camp.crustytoothpaste.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2021-06-30T23:27:40Z","receivedAt":"2021-06-30T23:27:57Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Jun 30, 2021 at 4:20 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n> I'm personally fine with the ad-hoc approach we use now, which is\n> actually very convenient to script and, in my view, not to terrible to\n> parse in other tools and languages.  Your mileage may vary, though.\n\nIt's of one piece with the Unix tradition that gave us Perl. That's\nboth a blessing and a curse.\n\nOutputting fields, escapes, etc and the ability to use xpath to\nextract, manipulate and query specific data is... very useful. There's\na hard tradeoff with utf-8, and any attempt at papering over that -\niconv conversions, etc – will inevitably munge data in some cases.\n\nhmmm,\n\n\nm\n-- \n martin.langhoff@gmail.com\n - ask interesting questions  ~  http://linkedin.com/in/martinlanghoff\n - don't be distracted        ~  http://github.com/martin-langhoff\n   by shiny stuff\n"},{"id":"428877","messageId":"CAFQ2z_MGJ7Qo5fq7bnwXi3m32iVeaC2WfdsKjRW0JO4Tq+M37Q@mail.gmail.com","threadId":"56032","inReplyTo":"CACPiFCLzsiUjx-vm-dcd=0E8HezMWkErPyS==OQ7OhaXqR6CUA@mail.gmail.com","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-07-01T08:18:48Z","receivedAt":"2021-07-01T08:19:02Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Wed, Jun 30, 2021 at 7:00 PM Martin Langhoff\n<martin.langhoff@gmail.com> wrote:\n> I'm used to automating git by parsing stuff in Unix shell style. But\n> the golden days that gave us Perl are well behind us.\n>\n> Before I write (unreliable, leaky, likely buggy) text parsing bits one\n> more time, has git grown formatted output? I'm aware of fmt and\n> friends, I'm thinking of something that handles escaping, nested data\n> structures, etc.\n>\n> Something like \"git shortlog --json\".\n>\n> Have there been discussions? Patches? (I don't spot anything in what's cooking).\n\n<unpopular opinion>\n\ntext parsing is fine if you're doing a one-off, but if it's anything\nthat needs to be just moderately complex, robust or long-lived,\nwriting a proper program is a better idea. I'd recommend either go-git\n(single file programs can easily be run with \"go run\"), or Dulwhich\n(no experience, though) in python.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"428943","messageId":"YN3jhlXyTEmoBOon@coredump.intra.peff.net","threadId":"56032","inReplyTo":"CACPiFC+F9P1DY_Dgt4+Z-U4o5WRbRduq60+Df0+FHBn6=XL2hw@mail.gmail.com","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-07-01T15:47:18Z","receivedAt":"2021-07-01T15:47:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jun 30, 2021 at 02:20:09PM -0400, Martin Langhoff wrote:\n\n> > One complication we faced is that a lot of Git's data is bag-of-bytes,\n> \n> Great point -- hadn't thought of that. Don't see anything in\n> json-writer.c but we do use iconv already.\n\nWe do, but the problem is deeper than that. We don't always know the\nintended encoding of bytes in the repository. For commits, there's an\n\"encoding\" header and we default to utf8 if it's not specified. But\nfilenames in trees do not have an encoding (nor are two entries in a\nsingle tree even required to be in the same encoding). They really are\njust sequences of NUL-terminated binary bytes from Git's perspective.\n\nMost of the time that just works, of course. People tend to use utf8\nthese days anyway. And even if they aren't utf8, as long as the user's\nterminal is configured to match, then everything will look OK to them\n(you do have to turn off core.quotepath to see any high-bit characters\nin filenames).\n\nSo in practice I suspect it is fine to just output them as-is in json.\nThings will Just Work for people using utf8 consistently. People using\nother encodings will have things look OK in their terminal, but probably\nJSON parsers would choke. We could provide an option to say \"when you\ngenerate json, assume paths are in encoding XYZ (say, latin1) and\nconvert to utf8\". That wouldn't help people who have mix-and-match\nencodings in their trees, but that seems even more rare.\n\n-Peff\n"},{"id":"428946","messageId":"YN3mk0LnyJyuQ+9T@coredump.intra.peff.net","threadId":"56032","inReplyTo":"YNzR5ZZDTfcN2Q+s@camp.crustytoothpaste.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-07-01T16:00:19Z","receivedAt":"2021-07-01T16:00:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jun 30, 2021 at 08:19:49PM +0000, brian m. carlson wrote:\n\n> On 2021-06-30 at 17:59:43, Jeff King wrote:\n> > One complication we faced is that a lot of Git's data is bag-of-bytes,\n> > not utf8. And json technically requires utf8. I don't remember if we\n> > simply fudged that and output possibly non-utf8 sequences, or if we\n> > actually encode them.\n> \n> I think we just emit invalid UTF-8 in that case, which is a problem.\n> That's why Git is not well suited to JSON output and why it isn't a good\n> choice for structured data here.  I'd like us not to do more JSON in our\n> codebase, since it's practically impossible for users to depend on our\n> output if we do that due to encoding issues[0].\n> \n> We could emit data in a different format, such as YAML, which does have\n> encoding for arbitrary byte sequences.  However, in YAML, binary data is\n> always base64 encoded, which is less readable, although still\n> interchangeable.  CBOR is also a possibility, although it's not human\n> readable at all.\n\nI don't love the invalid-utf8-in-json thing in general. But I think it\nmay be the least-bad solution. I seem to recall that YAML has its own\ncomplexities, and losing human-readability (even to base64) is a pretty\nbig downside. And the tooling for working with json seems more common\nand mature (certainly over something like CBOR, but I think even YAML\ndoesn't have anything nearly as nice as jq).\n\nOur sloppy json encoding does work correctly if you use utf8 paths, and\nI think we could provide options to cover other common cases (e.g., a\nsingle option for \"assume my paths are latin1\"). I think life is hardest\non somebody writing a script/service which is meant to process arbitrary\nrepositories (and isn't in control of the strictness of whatever is\nparsing the json).\n\nI'm sensitive to the issue of implementing something that works most of\nthe time, but then fails spectacularly when somebody does something\nunusual. But it also sucks for many users not to have that \"something\nthat works most of the time\" if it would make their lives easier.\n\n> I'm personally fine with the ad-hoc approach we use now, which is\n> actually very convenient to script and, in my view, not to terrible to\n> parse in other tools and languages.  Your mileage may vary, though.\n\nThere are a lot of gotchas, there, too. When the data gets complex, \"-z\"\nsplitting becomes ambiguous (e.g., \"git log -z --raw\" uses a NUL both to\nseparate commits from their diffs, diffs from each other, and diffs from\nsubsequent commits, so you have to pattern-match each type). It's also\ncontext-dependent (e.g., you can't parse a \"--raw -z\" entry without\ninterpreting its type character, since \"R\" and \"C\" will have multiple\npath fields; there are almost certainly a lot of \"works most of the\ntime\" parsers out there).\n\n-Peff\n"},{"id":"428983","messageId":"YN4xCRDi3JwMc+S0@camp.crustytoothpaste.net","threadId":"56032","inReplyTo":"YN3mk0LnyJyuQ+9T@coredump.intra.peff.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-07-01T21:18:01Z","receivedAt":"2021-07-01T21:18:08Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-07-01 at 16:00:19, Jeff King wrote:\n> On Wed, Jun 30, 2021 at 08:19:49PM +0000, brian m. carlson wrote:\n> \n> > On 2021-06-30 at 17:59:43, Jeff King wrote:\n> > > One complication we faced is that a lot of Git's data is bag-of-bytes,\n> > > not utf8. And json technically requires utf8. I don't remember if we\n> > > simply fudged that and output possibly non-utf8 sequences, or if we\n> > > actually encode them.\n> > \n> > I think we just emit invalid UTF-8 in that case, which is a problem.\n> > That's why Git is not well suited to JSON output and why it isn't a good\n> > choice for structured data here.  I'd like us not to do more JSON in our\n> > codebase, since it's practically impossible for users to depend on our\n> > output if we do that due to encoding issues[0].\n> > \n> > We could emit data in a different format, such as YAML, which does have\n> > encoding for arbitrary byte sequences.  However, in YAML, binary data is\n> > always base64 encoded, which is less readable, although still\n> > interchangeable.  CBOR is also a possibility, although it's not human\n> > readable at all.\n> \n> I don't love the invalid-utf8-in-json thing in general. But I think it\n> may be the least-bad solution. I seem to recall that YAML has its own\n> complexities, and losing human-readability (even to base64) is a pretty\n> big downside. And the tooling for working with json seems more common\n> and mature (certainly over something like CBOR, but I think even YAML\n> doesn't have anything nearly as nice as jq).\n\nI'm not opposed to JSON as long as we don't write landmines.  We could\nURI-encode anything that contains a bag-of-bytes, which lets people have\nthe niceties of JSON without the breakage when people don't write valid\nUTF-8.  Most things will still be human-readable.\n\nWe could even have --json be an alias for --json=encoded (URI-encoding)\nand also have --json=strict for the situation where you assert\neverything is valid UTF-8 and explicitly said you wanted us to die() if\nwe saw non-UTF-8.  I don't want us to say that something is JSON and\nthen emit junk, since that's a bad user experience.\n\nIdeally, we'd have some generic serializer support for this case, so if\npeople _do_ want to add YAML or CBOR output, it can be stuffed in.\n\n> Our sloppy json encoding does work correctly if you use utf8 paths, and\n> I think we could provide options to cover other common cases (e.g., a\n> single option for \"assume my paths are latin1\"). I think life is hardest\n> on somebody writing a script/service which is meant to process arbitrary\n> repositories (and isn't in control of the strictness of whatever is\n> parsing the json).\n\nI think I'd rather provide a general encoding functionality than try to\nhandle random encodings.  I _do_ want people to be able to do things\nlike store arbitrary bytes in paths, because many people do use that\nfunctionality for shipping test files that verify their code works\ncorrectly on Unix systems.  I also want us to handle arbitrary bytes\nwhere we've stated that's a thing we support (e.g., in refs).  I _don't_\nwant to encourage people to use non-UTF-8 text encodings, because I\nfirmly believe those are obsolete.\n\nSo, correct binary data support, yes; non-UTF-8 text, no.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"428989","messageId":"YN44NB3I/nDiTVg8@coredump.intra.peff.net","threadId":"56032","inReplyTo":"YN4xCRDi3JwMc+S0@camp.crustytoothpaste.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-07-01T21:48:36Z","receivedAt":"2021-07-01T21:48:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 01, 2021 at 09:18:01PM +0000, brian m. carlson wrote:\n\n> > I don't love the invalid-utf8-in-json thing in general. But I think it\n> > may be the least-bad solution. I seem to recall that YAML has its own\n> > complexities, and losing human-readability (even to base64) is a pretty\n> > big downside. And the tooling for working with json seems more common\n> > and mature (certainly over something like CBOR, but I think even YAML\n> > doesn't have anything nearly as nice as jq).\n> \n> I'm not opposed to JSON as long as we don't write landmines.  We could\n> URI-encode anything that contains a bag-of-bytes, which lets people have\n> the niceties of JSON without the breakage when people don't write valid\n> UTF-8.  Most things will still be human-readable.\n> \n> We could even have --json be an alias for --json=encoded (URI-encoding)\n> and also have --json=strict for the situation where you assert\n> everything is valid UTF-8 and explicitly said you wanted us to die() if\n> we saw non-UTF-8.  I don't want us to say that something is JSON and\n> then emit junk, since that's a bad user experience.\n> \n> Ideally, we'd have some generic serializer support for this case, so if\n> people _do_ want to add YAML or CBOR output, it can be stuffed in.\n\nYep, I'd agree with all of that. I think we're on more-or-less the same\npage.\n\nOne annoying thing about JSON is that (to my knowledge) it doesn't have\na binary data type. So you have to encode things and shove them into\n\"string\". I guess that is not too bad if you are using backslash or\npercent-encoding, as only a minority of characters get encoded. But it\nsure would be nice for readers if the values, once extracted from the\njson, could be used without further munging. That's most of the benefit\nof using json in the first place. But it may be the best we can do.\n\n-Peff\n"},{"id":"429040","messageId":"87lf6oam2m.fsf@evledraar.gmail.com","threadId":"56032","inReplyTo":"YN4xCRDi3JwMc+S0@camp.crustytoothpaste.net","subject":"Re: Structured (ie: json) output for query commands?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-07-02T13:13:25Z","receivedAt":"2021-07-02T13:21:10Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Jul 01 2021, brian m. carlson wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On 2021-07-01 at 16:00:19, Jeff King wrote:\n>> On Wed, Jun 30, 2021 at 08:19:49PM +0000, brian m. carlson wrote:\n>> \n>> > On 2021-06-30 at 17:59:43, Jeff King wrote:\n>> > > One complication we faced is that a lot of Git's data is bag-of-bytes,\n>> > > not utf8. And json technically requires utf8. I don't remember if we\n>> > > simply fudged that and output possibly non-utf8 sequences, or if we\n>> > > actually encode them.\n>> > \n>> > I think we just emit invalid UTF-8 in that case, which is a problem.\n>> > That's why Git is not well suited to JSON output and why it isn't a good\n>> > choice for structured data here.  I'd like us not to do more JSON in our\n>> > codebase, since it's practically impossible for users to depend on our\n>> > output if we do that due to encoding issues[0].\n>> > \n>> > We could emit data in a different format, such as YAML, which does have\n>> > encoding for arbitrary byte sequences.  However, in YAML, binary data is\n>> > always base64 encoded, which is less readable, although still\n>> > interchangeable.  CBOR is also a possibility, although it's not human\n>> > readable at all.\n>> \n>> I don't love the invalid-utf8-in-json thing in general. But I think it\n>> may be the least-bad solution. I seem to recall that YAML has its own\n>> complexities, and losing human-readability (even to base64) is a pretty\n>> big downside. And the tooling for working with json seems more common\n>> and mature (certainly over something like CBOR, but I think even YAML\n>> doesn't have anything nearly as nice as jq).\n>\n> I'm not opposed to JSON as long as we don't write landmines.  We could\n> URI-encode anything that contains a bag-of-bytes, which lets people have\n> the niceties of JSON without the breakage when people don't write valid\n> UTF-8.  Most things will still be human-readable.\n>\n> We could even have --json be an alias for --json=encoded (URI-encoding)\n> and also have --json=strict for the situation where you assert\n> everything is valid UTF-8 and explicitly said you wanted us to die() if\n> we saw non-UTF-8.  I don't want us to say that something is JSON and\n> then emit junk, since that's a bad user experience.\n>\n> Ideally, we'd have some generic serializer support for this case, so if\n> people _do_ want to add YAML or CBOR output, it can be stuffed in.\n\nI'd think the ideal end-state is for us to have some standardized way to\npass structs of structured data around at the C-level.\n\nThen everything that now supports a format such as git-log,\nfor-each-ref, cat-file --batch etc. could share the same formatting\nlogic. Our human-readable output would just be a special-case of\nproviding a default format, as it is in the case of some of these\ncommands.\n\nIf we had a bit of an extension of the %(if) etc. syntax that\nfor-each-ref uses to handle such nested structures we could emit\narbitrary structured data, e.g. the formatting language would be\nsufficient to start a nested structure, only emit commas between\nelements etc.\n\nYou could then emit JSON, XML or whatever you'd like with a \"simple\"\n(well, it would be quite verbose) format specification. We could then\nship some default formats.\n\nA related (but not quite the same) benefit would be to make the logic\ndriving the built-ins reeantrant, so a formatting feature like this\ncould be combined with a \"--batch\" mode supported by every (or most)\ncommands. So you could also e.g. run \"log --batch\" not just \"cat-file\n--batch\" if you needed some of the formatting it provides.\n\nThis would be immensely useful to editor implementations and things\ninvoking git on the server, where we often need to pay the startup cost\nfor invoking N number of commands that are built into the \"git\" binary\nanyway. So if they could be sent as a --batch ...\n\n>> Our sloppy json encoding does work correctly if you use utf8 paths, and\n>> I think we could provide options to cover other common cases (e.g., a\n>> single option for \"assume my paths are latin1\"). I think life is hardest\n>> on somebody writing a script/service which is meant to process arbitrary\n>> repositories (and isn't in control of the strictness of whatever is\n>> parsing the json).\n>\n> I think I'd rather provide a general encoding functionality than try to\n> handle random encodings.  I _do_ want people to be able to do things\n> like store arbitrary bytes in paths, because many people do use that\n> functionality for shipping test files that verify their code works\n> correctly on Unix systems.  I also want us to handle arbitrary bytes\n> where we've stated that's a thing we support (e.g., in refs).  I _don't_\n> want to encourage people to use non-UTF-8 text encodings, because I\n> firmly believe those are obsolete.\n>\n> So, correct binary data support, yes; non-UTF-8 text, no.\n\nI don't know how widely they're used with gits, but there's several\nnon-Unicode encodings in wide use, and e.g. non-UTF-8 but Unicode\nencodings like UTF-16 in some contexts/platforms:\n\nhttps://stackoverflow.com/questions/1200063/why-does-anyone-use-an-encoding-other-than-utf-8/2470079\n"}]}