{"thread":{"id":"62188","subject":"Pretty output in JSON format","startedAt":"2024-09-24T21:52:46Z","lastAt":"2025-08-04T21:19:11Z","messageCount":8,"participants":["Ron Ziroby Romero","brian m. carlson","Sean Allred","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"503414","messageId":"CAGW8g7=21pPAgCixjpayEvmw_ns-hcB4e59NP476TKtCRXHPXQ@mail.gmail.com","threadId":"62188","inReplyTo":null,"subject":"Pretty output in JSON format","fromName":"Ron Ziroby Romero","fromEmail":"ziroby@gmail.com","sentAt":"2024-09-24T21:52:35Z","receivedAt":"2024-09-24T21:52:46Z","isPatch":false,"sender":{"key":"ziroby@gmail.com","avatar":"https://gravatar.com/avatar/5b42738b00e3d753c7373bde65683bea6fc16cf5a2fc922fc34039e88c0ae4bf?d=mp&s=160"},"body":"Howdy git folk,\n\nI want to revive the discussion on JSON output. I see a discussion in\n2021 about it, but it didnt come to a resolution. That discussion was\ntalking about adding a --json flag. I have a slightly different\napproach.\n\nI see online that many people have tried to make various hacks to\nconvert git output into JSON, but they all lack completeness,\nespecially with log messages with arbitrary text. I believe the best\nand most correct way to get JSON output from git is to add it as a new\nformat to the pretty option. Then, it would be easy to pipe the output\ninto something like jq to parse the JSON.  Trying to convert git's\noutput into JSON is a losing proposition. You've already lost some of\nthe context of the output by getting it out of the git program itself.\nA pretty option would provide a standard way to get correct JSON\noutput, with git's code handling the weird corner cases.\n\nWhat do y'all think?\n\nCheers,\nRon Ziroby Romero\n"},{"id":"503435","messageId":"ZvM39VNFptcfwMGk@tapette.crustytoothpaste.net","threadId":"62188","inReplyTo":"CAGW8g7=21pPAgCixjpayEvmw_ns-hcB4e59NP476TKtCRXHPXQ@mail.gmail.com","subject":"Re: Pretty output in JSON format","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2024-09-24T22:06:45Z","receivedAt":"2024-09-24T22:06:47Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n> Howdy git folk,\n> \n> I want to revive the discussion on JSON output. I see a discussion in\n> 2021 about it, but it didnt come to a resolution. That discussion was\n> talking about adding a --json flag. I have a slightly different\n> approach.\n> \n> I see online that many people have tried to make various hacks to\n> convert git output into JSON, but they all lack completeness,\n> especially with log messages with arbitrary text. I believe the best\n> and most correct way to get JSON output from git is to add it as a new\n> format to the pretty option. Then, it would be easy to pipe the output\n> into something like jq to parse the JSON.  Trying to convert git's\n> output into JSON is a losing proposition. You've already lost some of\n> the context of the output by getting it out of the git program itself.\n> A pretty option would provide a standard way to get correct JSON\n> output, with git's code handling the weird corner cases.\n> \n> What do y'all think?\n\nI think this is ultimately a bad idea.  JSON requires that the output be\nUTF-8, but Git processes a large amount of data, including file names,\nref names, commit messages, author and committer identities, diff\noutput, and other file contents, that are not restricted to UTF-8.  In\nfact, despite my recommendation, the trace2 JSON output simply outputs\ninvalid UTF-8, which just doesn't work in nearly any tool, if it\nencounters such data.  We shouldn't add more broken-by-default\nfunctionality.\n\nHowever, if you were interested in CBOR output, which isn't\nhuman-readable but is capable of handling byte strings, then I don't see\na problem.  CBOR is used in FIDO2 and a variety of other protocols and\nis interoperable, so it should be a fine choice here.\n-- \nbrian m. carlson (they/them or he/him)\nToronto, Ontario, CA\n"},{"id":"503491","messageId":"m0r097mv19.fsf@epic96565.epic.com","threadId":"62188","inReplyTo":"ZvM39VNFptcfwMGk@tapette.crustytoothpaste.net","subject":"Re: Pretty output in JSON format","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2024-09-25T18:45:54Z","receivedAt":"2024-09-25T18:45:57Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n>> What do y'all think?\n>\n> I think this is ultimately a bad idea.  JSON requires that the output be\n> UTF-8, but Git processes a large amount of data, including file names,\n> ref names, commit messages, author and committer identities, diff\n> output, and other file contents, that are not restricted to UTF-8.\n\nThis strikes me with a little bit of 'perfect as the enemy of good'\nhere. I'm sure there are ways to signal an encoding failure. I would,\nhowever, caution against trying to provide diff output in JSON. That\njust seems... odd. Maybe base64 it first? (I don't know -- I just\nstruggle to see the use-case here.)\n\n> However, if you were interested in CBOR output, which isn't\n> human-readable but is capable of handling byte strings, then I don't\n> see a problem. CBOR is used in FIDO2 and a variety of other protocols\n> and is interoperable, so it should be a fine choice here.\n\nCBOR would certainly solve the byte stream problem, but I think it would\nprimarily be only useful for 'serious' toolsmiths that need to handle\nwildly unpredictable data. For most uses, JSON would get the job done.\n\n>> What do y'all think?\nAs with all things, I'd suggest you draw up a more formal proposal of\nexactly how this would work, and then that proposal can be discussed.\nHow would you use this option? What would its behavior be? What's in\nscope? What's _not_ in scope? :-)\n\n-- \nSean Allred\n"},{"id":"503587","messageId":"ZvXMSKaUWWA-MG9J@tapette.crustytoothpaste.net","threadId":"62188","inReplyTo":"m0r097mv19.fsf@epic96565.epic.com","subject":"Re: Pretty output in JSON format","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2024-09-26T21:04:08Z","receivedAt":"2024-09-26T21:04:10Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2024-09-25 at 18:45:54, Sean Allred wrote:\n> \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> \n> > On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n> >> What do y'all think?\n> >\n> > I think this is ultimately a bad idea.  JSON requires that the output be\n> > UTF-8, but Git processes a large amount of data, including file names,\n> > ref names, commit messages, author and committer identities, diff\n> > output, and other file contents, that are not restricted to UTF-8.\n> \n> This strikes me with a little bit of 'perfect as the enemy of good'\n> here. I'm sure there are ways to signal an encoding failure. I would,\n> however, caution against trying to provide diff output in JSON. That\n> just seems... odd. Maybe base64 it first? (I don't know -- I just\n> struggle to see the use-case here.)\n\nI understand JSON output would be useful, but it's also not useful to\nrandomly fail to do git for-each-ref (for example) because someone has a\nnon-UTF-8 ref, or to fail to do a git log because of encoding problems\n(which absolutely is a problem in the Linux kernel tree).  \"It works\nmost of the time, but seemingly randomly fails\" is not a good user\nexperience, and I'm opposed to adding serialization formats that do\nthat.  (For that reason, just-send-bytes that produces invalid JSON on\noccasion is also unacceptable.)\n\nIf we always base64-encoded or percent-encoded the things that aren't\nguaranteed to be UTF-8, then we could well create JSON.  However, that\nmakes working with the data structure in most scripting languages a pain\nsince there's no automatic decoding of this data.  In strongly typed\nlanguages like Rust, it's possible to do this decoding with no problem,\nbut I expect that's not most users who'd want this feature.\n-- \nbrian m. carlson (they/them or he/him)\nToronto, Ontario, CA\n"},{"id":"503604","messageId":"CAGW8g7mmbGRWfieVK5HL=LNF7tSAU7WOUqe83HqQk-wStqQ+Bw@mail.gmail.com","threadId":"62188","inReplyTo":"ZvXMSKaUWWA-MG9J@tapette.crustytoothpaste.net","subject":"Re: Pretty output in JSON format","fromName":"Ron Ziroby Romero","fromEmail":"ziroby@gmail.com","sentAt":"2024-09-27T06:49:51Z","receivedAt":"2024-09-27T06:50:03Z","isPatch":false,"sender":{"key":"ziroby@gmail.com","avatar":"https://gravatar.com/avatar/5b42738b00e3d753c7373bde65683bea6fc16cf5a2fc922fc34039e88c0ae4bf?d=mp&s=160"},"body":"On Thu, 26 Sept 2024 at 22:04, brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2024-09-25 at 18:45:54, Sean Allred wrote:\n> > \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n> >\n> > > On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n> > >> What do y'all think?\n> > >\n> > > I think this is ultimately a bad idea.  JSON requires that the output be\n> > > UTF-8, but Git processes a large amount of data, including file names,\n> > > ref names, commit messages, author and committer identities, diff\n> > > output, and other file contents, that are not restricted to UTF-8.\n> >\n> > This strikes me with a little bit of 'perfect as the enemy of good'\n> > here. I'm sure there are ways to signal an encoding failure. I would,\n> > however, caution against trying to provide diff output in JSON. That\n> > just seems... odd. Maybe base64 it first? (I don't know -- I just\n> > struggle to see the use-case here.)\n>\n> I understand JSON output would be useful, but it's also not useful to\n> randomly fail to do git for-each-ref (for example) because someone has a\n> non-UTF-8 ref, or to fail to do a git log because of encoding problems\n> (which absolutely is a problem in the Linux kernel tree).  \"It works\n> most of the time, but seemingly randomly fails\" is not a good user\n> experience, and I'm opposed to adding serialization formats that do\n> that.  (For that reason, just-send-bytes that produces invalid JSON on\n> occasion is also unacceptable.)\n>\n> If we always base64-encoded or percent-encoded the things that aren't\n> guaranteed to be UTF-8, then we could well create JSON.  However, that\n> makes working with the data structure in most scripting languages a pain\n> since there's no automatic decoding of this data.  In strongly typed\n> languages like Rust, it's possible to do this decoding with no problem,\n> but I expect that's not most users who'd want this feature.\n\nI do plan on percent-encoding all non-UTF-8 data.  It sounds like a\ngood way to check this feature would be to call \"git log\n--pretty:json\" on the Linux kernel and ensure we get a valid, though\nmassive, UTF-8 JSON file. (Not as an automated test, but as a way to\ncheck that we've covered everything. Any stumbling blocks should be\nput into an automated test.) The use case I'm thinking of is piping\ndata to jq to process it.\n\nCBOR output seems useful, but I see it as a follow-up project. JSON\noutput would be more beneficial to more people, so I feel we should\ntackle it first.\n\n> >> What do y'all think?\n> As with all things, I'd suggest you draw up a more formal proposal of\n> exactly how this would work, and then that proposal can be discussed.\n> How would you use this option? What would its behavior be? What's in\n> scope? What's _not_ in scope? :-)\n\nOK, I'll start working on a more formal proposal.\n\n--\nRon Ziroby Romero\n"},{"id":"523124","messageId":"CAGW8g7mjJ8+aRkg5nf1c6CCAWTnxery86uLuN-9n3nt_VDaZvA@mail.gmail.com","threadId":"62188","inReplyTo":"CAGW8g7=xK0S-i_Ekfwwo_NjMbngO_5m4LERtWRhSCgA0vf+ZAg@mail.gmail.com","subject":"Re: Pretty output in JSON format","fromName":"Ron Ziroby Romero","fromEmail":"ziroby@gmail.com","sentAt":"2025-07-31T20:19:05Z","receivedAt":"2025-07-31T20:19:18Z","isPatch":false,"sender":{"key":"ziroby@gmail.com","avatar":"https://gravatar.com/avatar/5b42738b00e3d753c7373bde65683bea6fc16cf5a2fc922fc34039e88c0ae4bf?d=mp&s=160"},"body":"> On Fri, 27 Sept 2024, 10:30 demerphq, <demerphq@gmail.com> wrote:\n>>\n>> On Thu, 26 Sept 2024 at 23:04, brian m. carlson <sandals@crustytoothpaste.net> wrote:\n>>>\n>>> On 2024-09-25 at 18:45:54, Sean Allred wrote:\n>>> > \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n>>> >\n>>> > > On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n>>> > >> What do y'all think?\n>>> > >\n>>> > > I think this is ultimately a bad idea.  JSON requires that the output be\n>>> > > UTF-8, but Git processes a large amount of data, including file names,\n>>> > > ref names, commit messages, author and committer identities, diff\n>>> > > output, and other file contents, that are not restricted to UTF-8.\n>>> >\n>>> > This strikes me with a little bit of 'perfect as the enemy of good'\n>>> > here. I'm sure there are ways to signal an encoding failure. I would,\n>>> > however, caution against trying to provide diff output in JSON. That\n>>> > just seems... odd. Maybe base64 it first? (I don't know -- I just\n>>> > struggle to see the use-case here.)\n>>>\n>>> I understand JSON output would be useful, but it's also not useful to\n>>> randomly fail to do git for-each-ref (for example) because someone has a\n>>> non-UTF-8 ref, or to fail to do a git log because of encoding problems\n>>\n>>\n>> I dont really follow your argument, and I find it weird how you are talking about a specific encoding of unicode instead of Unicode itself.\n>>\n>> It is possible to represent every binary string as Unicode encoded as UTF-8 (or any of the UTF encodings). It may not be bytewise equivalent with the original, but why should that matter? There are a set of clear rules for doing the required transformations, and there is a huge body of tooling to do so. As long as you know the target encoding, you should be able to round trip data properly.\n>>\n>> IMO CBOR would just complicate what should be a relatively simple problem to solve.\n>>\n>> cheers,\n>> Yves\n>>\n>>\n>>\n>> --\n>> perl -Mre=debug -e \"/just|another|perl|hacker/\"\n\n\n\nHi. I've been working with the code and trying to figure out how to do\nthis. I've also started work on a formal proposal. Two things have\ncome up that I wanted to discuss:\n\nFirst, I'm questioning my approach of hacking pretty.c with a series\nof 'if json' blocks. Would it be better to make a new file, json-log.,\nand divorce myself from the pretty flow entirely? This would also go\nhand in hand with changing from \"--pretty=json\" to simply \"--json\"\n\nSecond, I see that someone is adding a --json flag to git status[1]. I\nfigure that argues for git log to use the --json flag. I don't think\nthat affects me other than making the case for this JSON output.\n\n ## References\n\n [1] Patrick Steinhardt, “Re: [PATCH] diff: add --json output format,”\nmessage to git@vger.kernel.org, July 29, 2025.\nhttps://public-inbox.org/git/pull.1937.git.1753856826464.gitgitgadget@gmail.com/\n\n Thanks,\n Ziroby Ron Romero\n"},{"id":"523480","messageId":"CAGW8g7kMVqsi6+JkdjDS-czKJQ=01ULUz36sZrGom+QPVtRF3A@mail.gmail.com","threadId":"62188","inReplyTo":"CANgJU+Xs-sQgAOCPL-5skaZGq7eHmhg0MaFGDr8N57=CK67iog@mail.gmail.com","subject":"Re: Pretty output in JSON format","fromName":"Ron Ziroby Romero","fromEmail":"ziroby@gmail.com","sentAt":"2025-08-04T20:39:02Z","receivedAt":"2025-08-04T20:39:16Z","isPatch":false,"sender":{"key":"ziroby@gmail.com","avatar":"https://gravatar.com/avatar/5b42738b00e3d753c7373bde65683bea6fc16cf5a2fc922fc34039e88c0ae4bf?d=mp&s=160"},"body":"On Fri, 27 Sept 2024 at 10:30, demerphq <demerphq@gmail.com> wrote:\n>\n> On Thu, 26 Sept 2024 at 23:04, brian m. carlson <sandals@crustytoothpaste.net> wrote:\n>>\n>> On 2024-09-25 at 18:45:54, Sean Allred wrote:\n>> > \"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n>> >\n>> > > On 2024-09-24 at 21:52:35, Ron Ziroby Romero wrote:\n>> > >> What do y'all think?\n>> > >\n>> > > I think this is ultimately a bad idea.  JSON requires that the output be\n>> > > UTF-8, but Git processes a large amount of data, including file names,\n>> > > ref names, commit messages, author and committer identities, diff\n>> > > output, and other file contents, that are not restricted to UTF-8.\n>> >\n>> > This strikes me with a little bit of 'perfect as the enemy of good'\n>> > here. I'm sure there are ways to signal an encoding failure. I would,\n>> > however, caution against trying to provide diff output in JSON. That\n>> > just seems... odd. Maybe base64 it first? (I don't know -- I just\n>> > struggle to see the use-case here.)\n>>\n>> I understand JSON output would be useful, but it's also not useful to\n>> randomly fail to do git for-each-ref (for example) because someone has a\n>> non-UTF-8 ref, or to fail to do a git log because of encoding problems\n>\n>\n> I dont really follow your argument, and I find it weird how you are talking about a specific encoding of unicode instead of Unicode itself.\n>\n> It is possible to represent every binary string as Unicode encoded as UTF-8 (or any of the UTF encodings). It may not be bytewise equivalent with the original, but why should that matter? There are a set of clear rules for doing the required transformations, and there is a huge body of tooling to do so. As long as you know the target encoding, you should be able to round trip data properly.\n>\n> IMO CBOR would just complicate what should be a relatively simple problem to solve.\n\nHi. I've been working with the code and trying to figure out how to do\nthis. I've also started work on a formal proposal. Two things have\ncome up that I wanted to discuss:\n\nFirst, I'm questioning my approach of hacking pretty.c with a series\nof 'if json' blocks. Would it be better to make a new file,\njson-log.c, and divorce myself from the pretty flow entirely? This\nwould also go hand in hand with changing from \"--pretty=json\" to\nsimply \"--json\"\n\nSecond, I see that someone is adding a --json flag to git status[1]. I\nfigure that argues for git log to use the --json flag. I don't think\nthat affects me other than making the case for this JSON output.\n\n## References\n\n[1] Patrick Steinhardt, “Re: [PATCH] diff: add --json output format,”\nmessage to git@vger.kernel.org, July 29, 2025.\nhttps://public-inbox.org/git/pull.1937.git.1753856826464.gitgitgadget@gmail.com/\n\n>\n> cheers,\n> Yves\n>\n>\n>\n> --\n> perl -Mre=debug -e \"/just|another|perl|hacker/\"\n\nCheers,\nZiroby Ron Romero\n"},{"id":"523481","messageId":"xmqqldny4wfn.fsf@gitster.g","threadId":"62188","inReplyTo":"CAGW8g7kMVqsi6+JkdjDS-czKJQ=01ULUz36sZrGom+QPVtRF3A@mail.gmail.com","subject":"Re: Pretty output in JSON format","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-08-04T21:19:08Z","receivedAt":"2025-08-04T21:19:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ron Ziroby Romero <ziroby@gmail.com> writes:\n\n> First, I'm questioning my approach of hacking pretty.c with a series\n> of 'if json' blocks. Would it be better to make a new file,\n> json-log.c, and divorce myself from the pretty flow entirely?\n\nThe same question to the other thread applies: why json?\n\nIf the objective is to give a parseable output for machines to\nrobustly read, then I do not think you want to use any of the\ninfrastructure laid by and for the pretty_print_commit() function,\nwhose purpose is quite the opposite, like squeezing inter paragraph\nspaces, trimming trailing whitespaces, indenting even an empty line\nby 4 spaces, etc., etc.\n\n> Second, I see that someone is adding a --json flag to git status[1]. I\n> figure that argues for git log to use the --json flag. I don't think\n> that affects me other than making the case for this JSON output.\n\nPlease don't.\n\nThat other thread is getting discouraged from introducing a new\noption just for a single new format.  Unfortunately \"status\" does\nnot have the --format={short,long,...} so we need to add one new\noption to allow new formats to be added in a more generic way, but\nonce that is done, the next new format would not have to add a new\noption.  Compared to it, \"log\" already has --pretty={...}, so we do\nnot have to add --json just for this single format, which makes us\nluckier than the other thread.\n"}]}