{"thread":{"id":"23395","subject":"git status --porcelain is a mess that needs fixing","startedAt":"2010-04-09T18:46:08Z","lastAt":"2010-04-10T23:33:15Z","messageCount":19,"participants":["Eric Raymond","Junio C Hamano","Jeff King","Jonathan Nieder","Julian Phillips","Jon Seymour"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"139059","messageId":"20100409184608.C7C61475FEF@snark.thyrsus.com","threadId":"23395","inReplyTo":null,"subject":"git status --porcelain is a mess that needs fixing","fromName":"Eric Raymond","fromEmail":"esr@snark.thyrsus.com","sentAt":"2010-04-09T18:46:08Z","receivedAt":"2010-04-09T18:46:08Z","isPatch":false,"sender":{"key":"esr@snark.thyrsus.com","avatar":null},"body":"I'm going to gripe a lot in this mail, possibly verging on flaming.\nTherefore I want to start by making clear that I am not here to\ncomplain without pitching in to help fix the problems.  If I can get\nresponsive answers to my questions, I will take responsibility for\nediting them into the relevant git documentation,\n\nShort version: \"git status --porcelain\" is horribly badly documented\nand appears to be seriously maldesigned.  Both these problems need to\nbe fixed before git causes a lot of unnecessary grief for people\ntrying to use it.\n\nHere is the entire documentation on this feature in HEAD:\n\n=============================================================================\n\nIn short-format, the status of each path is shown as\n\n\tXY PATH1 -> PATH2\n\nwhere `PATH1` is the path in the `HEAD`, and ` -> PATH2` part is\nshown only when `PATH1` corresponds to a different path in the\nindex/worktree (i.e. renamed).\n\nFor unmerged entries, `X` shows the status of stage #2 (i.e. ours) and `Y`\nshows the status of stage #3 (i.e. theirs).\n\nFor entries that do not have conflicts, `X` shows the status of the index,\nand `Y` shows the status of the work tree.  For untracked paths, `XY` are\n`??`.\n\n    X          Y     Meaning\n    -------------------------------------------------\n              [MD]   not updated\n    M        [ MD]   updated in index\n    A        [ MD]   added to index\n    D        [ MD]   deleted from index\n    R        [ MD]   renamed in index\n    C        [ MD]   copied in index\n    [MARC]           index and work tree matches\n    [ MARC]     M    work tree changed since index\n    [ MARC]     D    deleted in work tree\n    -------------------------------------------------\n    D           D    unmerged, both deleted\n    A           U    unmerged, added by us\n    U           D    unmerged, deleted by them\n    U           A    unmerged, added by them\n    D           U    unmerged, deleted by us\n    A           A    unmerged, both added\n    U           U    unmerged, both modified\n    -------------------------------------------------\n    ?           ?    untracked\n    -------------------------------------------------\n\n=============================================================================\n\nThis was clearly written as an aide-memoire by someone intimately\nfamiliar with the system, but I have to tell you it is so confusing\nto me as to be nearly worse than useless.  \n\nIn addition, some of the design choices it appears to imply are quite\nbad - so I hope I am wrong about those implications.  If I am not, you\nhave specified a misdesigned format that will frustrate and annoy your\ncustomers (script and front-end writers). And that would be a problem.\n\nAs I criticize, bear in mind that (a) none of my issues are VC\nspecific, and (b) I am the author of several version-control front\nends - *I have done this before.* My objections are *not*\ntheoretical!\n\nFirst, the documentation issues, in roughly increasing order of severity:\n\n1. What separates the XY column from the first path?\n\nI'd assume a tab, but it's not documented. It needs to be documented.\n\n2. What separates the '->' on either side from the path columns?\n\nNot documented.  Needs to be documented.  \n\n3. What do the status codes M A D R C mean?\n\nI can guess, but I should not have to guess.  They should be documented.\n\n4. Some columns in the table have sets of codes enclosed by [].  Is\nthis indicating alternation?\n\nMy guess is yes, but I should not have to guess.  This should be documented.\n\n5. What is 'us' versus 'them'? What are \"stage #2\" and \"stage #3\"?\n\nIt makes my brain hurt just trying to list all the things \"us\"\nand \"them\" could mean.\n\nRemember that because you're advertising a format for script use, your\naudience for this page is not git hackers.  It's not git power\nusers. It's not even ordinary git users. It's people whose main\nexpertise is is *other tools*.  They want to get in, write their\nscript and get out, having learned as little about git as they can get\naway with.\n\nIf 'us'/'them'/'stage #2'/'stage #3' are git terms of art that are well\ndefined elsewhere, you must reference that elsewhere.  If they are\nnot, you need to define them here.  And because of the special\naudience for this page, it needs to be more self-contained and make\nfewer assumptions about the reader's knowledge than usual.\n\nNote: I, personally, read very fast and don't mind the mental effort\nof skimming 50-100 pages of other documentation.  But you must *not*\nassume I am anything but an exception.  This *particular* section on\nthis *particular* page needs (more than others) to be written so it\nwould be comprehensible to a lazy idiot who vaguely knows about\notther version-control systems and can't be bothered to read \nabout this one, either.  \n\n\nNow to the functional problems, again in roughly increasing order of\nseverity:\n\nA. The '->' separator considered harmful\n\nThe '->' was superfluous and thus a poor design choice; the\ndistinction between two columns and three columns is easy enough to\nmake in any scripting language.  As it is, it's meaningless and\nscripts will actually have to go to some extra effort to throw it\naway.\n\nI think the underlying problem here is that whoever designed this\nnever got past the idea that it needed to have cues for human\neyeballs in it.  That was a mistake.  If you're serious about it\nbeing easily parseable, design it that way.\n\nB. Does \"untracked\" include \"ignored\"?  \n\nIf so, that is a problem -- front ends care about the difference, for\nexample when C-x v v is trying to compute the logical next action.\nFor an unregistered file, it's to register it.  For an ignored file,\nit's to throw a user-visible error.\n\nC. If \"untracked\" does not include \"ignored\", how is an ignored file tagged?\n\nIf ignored files are not listed, that's another problem. Even more\nserious, actually.\n\nD. How do I tell the conflict/no-conflict cases apart?\n\nYou have three divisions in the table.  The first two are supposed\nto pertain to \"entries that do not have conflicts\" and \"unmerged\nentries\".  \n\nThey share code letters.  *How do I tell them apart?* \n\nIllustrative case: I see the status code \"DD\". How do I distinguish\nbetween case 4 (\"deleted from index\") and case 10 (\"unmerged, both\ndeleted\")?\n\nIf the distinction is meaningless, then why are they listed\nseparately?\n\nE. Are you *really* using a space as a status character?\n\nIt certainly appears so from the first and seventh rows of the table.\nIf so, this was a major blunder. It complicates parsing code\nunnecessarily, because the easiest way to separate columns is with the\nequivalent of a Python or Perl split() operation that will eat that\nspace.  Then we have to special-case depending on the field width.\n\nThe correct way to design a format like this for script parseability\nis to (a) never make the difference between space and tab significant,\nand (b) never use whitespace as anything but a field separator.  If\nyou want the equivalent of \"blank\" you use '-', as in Unix ls -l\noutput.\n\nThis may sound like a nitpick, but it's actually a crash landing, or\nclose to it.  Front-end writers look at things like this and think\n\"Idiots.  Can't trust them an inch...\".  And git already has a bad\nreputation for interface spikiness to live down.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\nA right is not what someone gives you; it's what no one can take from you. \n\t-- Ramsey Clark\n"},{"id":"139072","messageId":"7vmxxc1i8g.fsf@alter.siamese.dyndns.org","threadId":"23395","inReplyTo":"20100409184608.C7C61475FEF@snark.thyrsus.com","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-04-09T20:30:39Z","receivedAt":"2010-04-09T20:30:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Eric Raymond <esr@snark.thyrsus.com> writes:\n\n> First, the documentation issues, in roughly increasing order of severity:\n> ...\n\nThese all fall within \"Patches welcome\" category (meaning: I agree the\ndocumentation can be improved and I don't object to changing them).\n\n> D. How do I tell the conflict/no-conflict cases apart?\n> ...\n> Illustrative case: I see the status code \"DD\". How do I distinguish\n> between case 4 (\"deleted from index\") and case 10 (\"unmerged, both\n> deleted\")?\n\nIs that DD really \"illustrative\", or did you mean to say \"only/sole\"?\n\nYou should never get \"DD\" in non-conflicting case.  I think I was fairly\ncareful not to make them ambiguous when I did that code, but apparently I\nwasn't so careful about the documentation.\n\nThanks for going through this area with fine comb.\n\ndiff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\nindex 1cab91b..313dd04 100644\n--- a/Documentation/git-status.txt\n+++ b/Documentation/git-status.txt\n@@ -86,7 +86,7 @@ and `Y` shows the status of the work tree.  For untracked paths, `XY` are\n               [MD]   not updated\n     M        [ MD]   updated in index\n     A        [ MD]   added to index\n-    D        [ MD]   deleted from index\n+    D         [ M]    deleted from index\n     R        [ MD]   renamed in index\n     C        [ MD]   copied in index\n     [MARC]           index and work tree matches\n"},{"id":"139102","messageId":"20100410040959.GA11977@coredump.intra.peff.net","threadId":"23395","inReplyTo":"20100409184608.C7C61475FEF@snark.thyrsus.com","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-04-10T04:09:59Z","receivedAt":"2010-04-10T04:09:59Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 09, 2010 at 02:46:08PM -0400, Eric Raymond wrote:\n\n> First, the documentation issues, in roughly increasing order of severity:\n\nNote that \"status --porcelain\" is brand new in v1.7.0, so you may be\namong the first to be seriously reading the documentation. As Junio\nsaid, I think patches in this area are very welcome.\n\nMy answers below are meant to help you understand. I omitted the \"...and\nyes, this should be documented better\" from the end of each, but you can\nsay it in your head if you want.\n\n> 1. What separates the XY column from the first path?\n> \n> I'd assume a tab, but it's not documented. It needs to be documented.\n\nIt's a space.\n\n> 2. What separates the '->' on either side from the path columns?\n> \n> Not documented.  Needs to be documented.\n\nIt's a space. But more importantly, the path columns are actually\nC-quoted. E.g.:\n\n  $ perl -e 'open foo, \">\", \"foo\\n\"'\\\n  $ git add .\n  $ git status --porcelain\n  A  \"foo\\n\"\n\nIf your parser supports it, it will almost certainly be easier to use\n\"-z\":\n\n  $ git status --porcelain -z | cat -A\n  A  foo$\n  ^@\n\nDo note that for the 'R'ename status, you will get _two_ NUL-terminated\nentries, and they will be in the order of \"to\\0from\\0\", whereas the\nnon-NUL form is \"from -> to\" (and no, I doubt this is adequately\ndocumented, either).\n\n> 3. What do the status codes M A D R C mean?\n> \n> I can guess, but I should not have to guess.  They should be documented.\n\nThey are the same as in \"git diff --name-status\", which in turn has kind\nof crappy documentation. Patches welcome for both issues.\n\n> 5. What is 'us' versus 'them'? What are \"stage #2\" and \"stage #3\"?\n> \n> It makes my brain hurt just trying to list all the things \"us\"\n> and \"them\" could mean.\n\nThe terms \"us / ours\" and \"them / theirs\" are frequently used in the git\ndocumentation.  I'm not sure if they are ever defined rigorously. They\nare only meaningful in a merging context, and basically refer to the two\nsides of a merge. If I am on branch \"master\" and do \"git merge foo\",\nthen \"us\" refers to the master branch and the the contents of index\nstage 2 (bear with me a moment, I'll define that in a second). \"Them\"\nrefers to branch \"foo\" and index stage 3.\n\nGit's \"index\" is where it keeps uncommitted state about files it tracks\n(sort of like CVS/Entries, if that helps, except that git exposes the\nconcept much more). Most of the time, you use it for building a commit\nincrementally. You \"git add\" files to the index, and then \"git commit\"\ncreates a new commit from the contents of your index.\n\nBut the index actually has several different slots for each file entry,\nwhich are called stages, and each has a number. \"Stage #0\" is the\n\"normal\" stage, which you use as described in the last paragraph. During\na merge, entries with conflicts use the other stages. The stage #1 entry\ncontains the common ancestor. Stage #2 contains our original version\nfrom before the merge. Stage #3 contains the other side's original\nversion from before the merge.\n\nThe details of how they are used is discussed in \"git help read-tree\",\nunder \"3-Way Merge\". Those details are way too gory for somebody\ninterested in \"git status\" output, but you might find them interesting.\nFor the \"git status\" documentation, it probably makes sense to keep\nthings simple and just indicate that XY shows what each side of a 2-way\nmerge did to the file.\n\n> A. The '->' separator considered harmful\n> \n> The '->' was superfluous and thus a poor design choice; the\n> distinction between two columns and three columns is easy enough to\n> make in any scripting language.  As it is, it's meaningless and\n> scripts will actually have to go to some extra effort to throw it\n> away.\n> \n> I think the underlying problem here is that whoever designed this\n> never got past the idea that it needed to have cues for human\n> eyeballs in it.  That was a mistake.  If you're serious about it\n> being easily parseable, design it that way.\n\nShort answer: use -z.\n\nLong answer:\n\nThis is my fault, to some degree. The \"short-status\" form _is_ meant for\nhuman eyeballs, and was designed by Junio. Some people wanted a\nscriptable status output, too, so I slapped a \"--porcelain\" on the same\nformat that turns off configurable features like relative pathnames and\ncolorizing, and makes an implicit promise that we won't make further\nchanges to the format. The idea was to prevent people from scripting\naround --short, because it was never intended to be stable.\n\nSo yeah, while --porcelain by itself _is_ stable and scriptable, it is\nperhaps not the most friendly to parsers. The \"-z --porcelain\" format is\nmuch more so, and I would recommend it to anyone scripting around\n\"git status\". I think a note in the documentation to that effect would\nbe helpful.\n\n> B. Does \"untracked\" include \"ignored\"?\n> \n> If so, that is a problem -- front ends care about the difference, for\n> example when C-x v v is trying to compute the logical next action.\n> For an unregistered file, it's to register it.  For an ignored file,\n> it's to throw a user-visible error.\n\nNo. Ignored files are not listed at all.\n\nIf you really want a list of ignored files, I think you are stuck\ncomparing the output of \"git ls-files -o\" and \"git ls-files -o\n--exclude-standard\".\n\n> C. If \"untracked\" does not include \"ignored\", how is an ignored file tagged?\n>\n> If ignored files are not listed, that's another problem. Even more\n> serious, actually.\n\nSee above.\n\nIt wouldn't be too hard to add them in, and would look something like\nthe patch below. But it would still need:\n\n  1. For me to investigate that \"ugh\" comment below.\n\n  2. It should be conditional on a command-line option. Many users won't\n     want to see it.\n\n  3. For full-length status (i.e., \"git status\") in my patch the\n     information is just ignored. If the user specified it on the\n     command-line, I guess we should show it.\n\n> D. How do I tell the conflict/no-conflict cases apart?\n\nJunio already answered this one, and I agree with his analysis that it\nis a documentation bug.\n\n> E. Are you *really* using a space as a status character?\n\nYes. I agree that a \"-\" probably would have been nicer for parsing, but:\n\n> It certainly appears so from the first and seventh rows of the table.\n> If so, this was a major blunder. It complicates parsing code\n> unnecessarily, because the easiest way to separate columns is with the\n> equivalent of a Python or Perl split() operation that will eat that\n> space.  Then we have to special-case depending on the field width.\n\nYour parser is already broken if you are calling split, as the filenames\nmay contain spaces (and will be quoted in that case, and you need to\nunmangle). You should use \"-z\".\n\nYou will probably then realize that the \"-z\" format looks like:\n\n  XY file1\\0file2\\0\n\nwhich still sucks. It would be more friendly as:\n\n  XY\\0file1\\0file2\\0\n\nSo you could split on \"\\0\". But even with that, you can't just blindly\nsplit, as the column and record separators are the same, and you might\nhave one or two filenames.\n\nSo you really are stuck parsing it as one char of X, one char of Y, a\njunk space, and then depending on X/Y, either one or two filenames.\n\n> This may sound like a nitpick, but it's actually a crash landing, or\n> close to it.  Front-end writers look at things like this and think\n> \"Idiots.  Can't trust them an inch...\".  And git already has a bad\n> reputation for interface spikiness to live down.\n\nI agree with most of your criticisms. The question is what we want to do\nabout it.\n\nDocumentation fixes and an optional --show-ignored are easy. Output\nformat problems are harder. We can:\n\n  (1) Ignore it. The problems make parsing harder, but it isn't\n      completely broken (which I would consider it to be if, for\n      example, there was impossibly ambiguous output).\n\n  (2) Quietly change it.  The --porcelain format has been released in\n      one major version. I don't know if anybody is actually using it\n      yet. It is tempting to just fix it and say \"we botched v1.7.0,\n      don't use it\". It is such a new feature that script writers\n      already have to check the version to see if we even support\n      --porcelain at all (or accept breakage for older versions).\n\n      But usually we have more restraint than that about backwards\n      incompatible changes. And given that it _isn't_ totally broken,\n      I don't think it's justified.\n\n  (3) Introduce --porcelain=v2 with an alternate format.\n\nPersonally, I think I am in favor of (1). Option (3) is going to\nintroduce maintenance headaches, but more importantly, I wonder if it is\njust going to confuse people more with \"Which porcelain version should I\nuse?  Which versions of git support which porcelain versions?\"\nquestions.\n\n-Peff\n"},{"id":"139116","messageId":"20100410054645.GA17711@progeny.tock","threadId":"23395","inReplyTo":"20100410040959.GA11977@coredump.intra.peff.net","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-04-10T05:46:45Z","receivedAt":"2010-04-10T05:46:45Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n\n> If you really want a list of ignored files, I think you are stuck\n> comparing the output of \"git ls-files -o\" and \"git ls-files -o\n> --exclude-standard\".\n\n\"git clean -n -d\" may help.\n\nJust my 2¢,\nJonathan\n"},{"id":"139117","messageId":"20100410055124.GA17778@progeny.tock","threadId":"23395","inReplyTo":"20100410054645.GA17711@progeny.tock","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-04-10T05:51:24Z","receivedAt":"2010-04-10T05:51:24Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jonathan Nieder wrote:\n> Jeff King wrote:\n\n>> If you really want a list of ignored files, I think you are stuck\n>> comparing the output of \"git ls-files -o\" and \"git ls-files -o\n>> --exclude-standard\".\n>\n> \"git clean -n -d\" may help.\n\nerr, \"git clean -n -d -X\".\n\nI am also not sure how stable the \"Would remove \" output format is,\nor how stable we want it to be.  Probably not stable at all, so\nsorry about that.\n\nJonathan\n"},{"id":"139118","messageId":"20100410055940.GA396@thyrsus.com","threadId":"23395","inReplyTo":"20100410040959.GA11977@coredump.intra.peff.net","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Eric Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2010-04-10T05:59:40Z","receivedAt":"2010-04-10T05:59:40Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jeff King <peff@peff.net>:\n> My answers below are meant to help you understand.\n\nThey do that quite well.  Thank you.\n\nI've got a couple of other things on my plate, including prepping for \na GPSD point release early next week, so I can't respond immediately.\nExpect a response and some patches Tueday, Wednesday, or Thursday\nof next week.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"139119","messageId":"20100410060353.GA4585@coredump.intra.peff.net","threadId":"23395","inReplyTo":"20100410055124.GA17778@progeny.tock","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-04-10T06:03:53Z","receivedAt":"2010-04-10T06:03:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 10, 2010 at 12:51:24AM -0500, Jonathan Nieder wrote:\n\n> >> If you really want a list of ignored files, I think you are stuck\n> >> comparing the output of \"git ls-files -o\" and \"git ls-files -o\n> >> --exclude-standard\".\n> >\n> > \"git clean -n -d\" may help.\n> \n> err, \"git clean -n -d -X\".\n> \n> I am also not sure how stable the \"Would remove \" output format is,\n> or how stable we want it to be.  Probably not stable at all, so\n> sorry about that.\n\nThat's the same information, isn't it? You do \"git clean -ndX\" to see\n_everything_ that is untracked, and \"git clean -nd\" to see things that\nare untracked but not ignored. So I think it is just as painful to use\nas ls-files, but as you noted, it is not really plumbing.\n\n-Peff\n"},{"id":"139121","messageId":"20100410061224.GA18715@progeny.tock","threadId":"23395","inReplyTo":"20100410060353.GA4585@coredump.intra.peff.net","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-04-10T06:12:24Z","receivedAt":"2010-04-10T06:12:24Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n\n> You do \"git clean -ndX\" to see\n> _everything_ that is untracked, and \"git clean -nd\" to see things that\n> are untracked but not ignored.\n\nNo, the capital X tells clean to only list excluded files.  The\nstandard use is as a poor man’s “make maintainer-clean”, leaving\nunrelated files alone.\n\nI only learned about it just now.  I’m glad I did (I often use the\nlowercase version for this because I just didn’t know about -X), but\nas you mentioned, it is not so applicable here because not plumbing.\n\nJonathan\n"},{"id":"139122","messageId":"20100410063215.GA6260@coredump.intra.peff.net","threadId":"23395","inReplyTo":"20100410061224.GA18715@progeny.tock","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-04-10T06:32:15Z","receivedAt":"2010-04-10T06:32:15Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 10, 2010 at 01:12:24AM -0500, Jonathan Nieder wrote:\n\n> Jeff King wrote:\n> \n> > You do \"git clean -ndX\" to see\n> > _everything_ that is untracked, and \"git clean -nd\" to see things that\n> > are untracked but not ignored.\n> \n> No, the capital X tells clean to only list excluded files.  The\n> standard use is as a poor man’s “make maintainer-clean”, leaving\n> unrelated files alone.\n\nAh, I read it as \"-x\" (probably because I had never heard of \"-X\"\neither...).\n\nSo yes, it would do the right thing. I still think a --show-ignored\noption to git-status would probably be better (in addition to being\nsanctioned plumbing, it means we only have to traverse the tree once\nfor Eric's case, instead of twice).\n\n> I only learned about it just now.  I’m glad I did (I often use the\n> lowercase version for this because I just didn’t know about -X), but\n> as you mentioned, it is not so applicable here because not plumbing.\n\nThe \"-X\" mode seems much safer to me, as you are less likely to blow\naway things you actually wanted to keep while cleaning the tree of\ncrufty build products. It seems like it should have been the\neasier-to-type \"-x\", but it is far too late for such bikeshedding at\nthis point.\n\nThanks for the pointer.\n\n-Peff\n"},{"id":"139150","messageId":"9c7e1f33b7ec0dab68a92aa8f067989e@212.159.54.234","threadId":"23395","inReplyTo":"20100410040959.GA11977@coredump.intra.peff.net","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2010-04-10T13:35:25Z","receivedAt":"2010-04-10T13:35:25Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Sat, 10 Apr 2010 00:09:59 -0400, Jeff King <peff@peff.net> wrote:\n> Your parser is already broken if you are calling split, as the filenames\n> may contain spaces (and will be quoted in that case, and you need to\n> unmangle). You should use \"-z\".\n> \n> You will probably then realize that the \"-z\" format looks like:\n> \n>   XY file1\\0file2\\0\n> \n> which still sucks. It would be more friendly as:\n> \n>   XY\\0file1\\0file2\\0\n> \n> So you could split on \"\\0\". But even with that, you can't just blindly\n> split, as the column and record separators are the same, and you might\n> have one or two filenames.\n\nNot true.  If the second form was used, then you _can_ split on \\0.  It\nwill tokenise the data for you, and then you consume ether two or three\ntokens depending on the status flags.  So it would make the parsing\nsimpler.  But to make it even easier, how about adding a -Z that makes the\noutput format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per record,\nwith the third token being empty if there is no second filename)?  Though\nif future expandability was wanted you could end each record with \\0\\0 and\nthen parsing would be a two stages of split on \\0\\0 for records and then\nsplit on \\0 for entries?  The is already precedence for the -z option to\nchange the output format, so a second similar switch should be ok?  Then\nthe updated documentation could recommend --porcelain -Z for new users\nwithout affecting old ones.\n\n-- \nJulian\n"},{"id":"139155","messageId":"20100410144334.GB23959@thyrsus.com","threadId":"23395","inReplyTo":"9c7e1f33b7ec0dab68a92aa8f067989e@212.159.54.234","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Eric Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2010-04-10T14:43:34Z","receivedAt":"2010-04-10T14:43:34Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk>:\n> On Sat, 10 Apr 2010 00:09:59 -0400, Jeff King <peff@peff.net> wrote:\n> > Your parser is already broken if you are calling split, as the filenames\n> > may contain spaces (and will be quoted in that case, and you need to\n> > unmangle). You should use \"-z\".\n> > \n> > You will probably then realize that the \"-z\" format looks like:\n> > \n> >   XY file1\\0file2\\0\n> > \n> > which still sucks. It would be more friendly as:\n> > \n> >   XY\\0file1\\0file2\\0\n> > \n> > So you could split on \"\\0\". But even with that, you can't just blindly\n> > split, as the column and record separators are the same, and you might\n> > have one or two filenames.\n> \n> Not true.  If the second form was used, then you _can_ split on \\0.  It\n> will tokenise the data for you, and then you consume ether two or three\n> tokens depending on the status flags.  So it would make the parsing\n> simpler.  But to make it even easier, how about adding a -Z that makes the\n> output format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per record,\n> with the third token being empty if there is no second filename)?  Though\n> if future expandability was wanted you could end each record with \\0\\0 and\n> then parsing would be a two stages of split on \\0\\0 for records and then\n> split on \\0 for entries?  The is already precedence for the -z option to\n> change the output format, so a second similar switch should be ok?  Then\n> the updated documentation could recommend --porcelain -Z for new users\n> without affecting old ones.\n\n+1\n\n-Z could fix some of the other issues, as well, like use of space\nas a flag character.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"139161","messageId":"h2j2cfc40321004100756g15ad7f12jf37e500f924e7b96@mail.gmail.com","threadId":"23395","inReplyTo":"9c7e1f33b7ec0dab68a92aa8f067989e@212.159.54.234","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jon Seymour","fromEmail":"jon.seymour@gmail.com","sentAt":"2010-04-10T14:56:47Z","receivedAt":"2010-04-10T14:56:47Z","isPatch":false,"sender":{"key":"jon.seymour@gmail.com","avatar":"https://avatars.githubusercontent.com/u/207131?v=4"},"body":"On Sat, Apr 10, 2010 at 11:35 PM, Julian Phillips\n<julian@quantumfyre.co.uk> wrote:\n> On Sat, 10 Apr 2010 00:09:59 -0400, Jeff King <peff@peff.net> wrote:\n>> Your parser is already broken if you are calling split, as the filenames\n>> may contain spaces (and will be quoted in that case, and you need to\n>> unmangle). You should use \"-z\".\n>>\n>> You will probably then realize that the \"-z\" format looks like:\n>>\n>>   XY file1\\0file2\\0\n>>\n>> which still sucks. It would be more friendly as:\n>>\n>>   XY\\0file1\\0file2\\0\n>>\n>> So you could split on \"\\0\". But even with that, you can't just blindly\n>> split, as the column and record separators are the same, and you might\n>> have one or two filenames.\n>\n> Not true.  If the second form was used, then you _can_ split on \\0.  It\n> will tokenise the data for you, and then you consume ether two or three\n> tokens depending on the status flags.  So it would make the parsing\n> simpler.  But to make it even easier, how about adding a -Z that makes the\n> output format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per record,\n> with the third token being empty if there is no second filename)?  Though\n> if future expandability was wanted you could end each record with \\0\\0 and\n> then parsing would be a two stages of split on \\0\\0 for records and then\n> split on \\0 for entries?\n\nSurely that won't work - if file2 can be empty, \\0[file2]\\0 reduces to\n\\0\\0 which would be confused with the \\0\\0 proposed as a record\nseparator.\n\njon.\n"},{"id":"139162","messageId":"8871c8959d3ea4cd71452400e4c60dd0@212.159.54.234","threadId":"23395","inReplyTo":"h2j2cfc40321004100756g15ad7f12jf37e500f924e7b96@mail.gmail.com","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2010-04-10T15:50:59Z","receivedAt":"2010-04-10T15:50:59Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Sun, 11 Apr 2010 00:56:47 +1000, Jon Seymour <jon.seymour@gmail.com>\nwrote:\n> On Sat, Apr 10, 2010 at 11:35 PM, Julian Phillips\n> <julian@quantumfyre.co.uk> wrote:\n>> On Sat, 10 Apr 2010 00:09:59 -0400, Jeff King <peff@peff.net> wrote:\n>>> Your parser is already broken if you are calling split, as the\nfilenames\n>>> may contain spaces (and will be quoted in that case, and you need to\n>>> unmangle). You should use \"-z\".\n>>>\n>>> You will probably then realize that the \"-z\" format looks like:\n>>>\n>>>   XY file1\\0file2\\0\n>>>\n>>> which still sucks. It would be more friendly as:\n>>>\n>>>   XY\\0file1\\0file2\\0\n>>>\n>>> So you could split on \"\\0\". But even with that, you can't just blindly\n>>> split, as the column and record separators are the same, and you might\n>>> have one or two filenames.\n>>\n>> Not true.  If the second form was used, then you _can_ split on \\0.  It\n>> will tokenise the data for you, and then you consume ether two or three\n>> tokens depending on the status flags.  So it would make the parsing\n>> simpler.  But to make it even easier, how about adding a -Z that makes\n>> the\n>> output format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per\n>> record,\n>> with the third token being empty if there is no second filename)?\n>>  Though\n>> if future expandability was wanted you could end each record with \\0\\0\n>> and\n>> then parsing would be a two stages of split on \\0\\0 for records and\nthen\n>> split on \\0 for entries?\n> \n> Surely that won't work - if file2 can be empty, \\0[file2]\\0 reduces to\n> \\0\\0 which would be confused with the \\0\\0 proposed as a record\n> separator.\n\nYes.  But they were alternative suggestions, so if using \\0\\0 as the\nrecord marker you would omit the second filename when empty as is currently\ndone.\n\n-- \nJulian\n"},{"id":"139186","messageId":"20100410192529.23731.79803.julian@quantumfyre.co.uk","threadId":"23395","inReplyTo":"9c7e1f33b7ec0dab68a92aa8f067989e@212.159.54.234","subject":"[RFC/PATCH] status: Add a new NUL separated output format","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2010-04-10T19:25:28Z","receivedAt":"2010-04-10T19:25:28Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"Add a new output format option to git-status that is a more extreme\nform of the -z output that places a NUL between all parts of the\nrecord, and always has three entries per record, even when only two\nare relevant.  This make the parsing of --porcelain output much\nsimpler for the consumer.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nOn Sat, 10 Apr 2010, Julian Phillips wrote:\n\n> Not true.  If the second form was used, then you _can_ split on \\0.  It\n> will tokenise the data for you, and then you consume ether two or three\n> tokens depending on the status flags.  So it would make the parsing\n> simpler.  But to make it even easier, how about adding a -Z that makes the\n> output format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per record,\n> with the third token being empty if there is no second filename)?  Though\n> if future expandability was wanted you could end each record with \\0\\0 and\n> then parsing would be a two stages of split on \\0\\0 for records and then\n> split on \\0 for entries?  The is already precedence for the -z option to\n> change the output format, so a second similar switch should be ok?  Then\n> the updated documentation could recommend --porcelain -Z for new users\n> without affecting old ones.\n\nSomething like this for the first variant (fixed three entries per record)\nperhaps ... (though a proper patch would probably want some tests too)\n\n builtin/commit.c |    6 ++++--\n wt-status.c      |   19 ++++++++++++++-----\n 2 files changed, 18 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex c5ab683..acbcefc 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -1025,8 +1025,10 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n \t\tOPT_SET_INT(0, \"porcelain\", &status_format,\n \t\t\t    \"show porcelain output format\",\n \t\t\t    STATUS_FORMAT_PORCELAIN),\n-\t\tOPT_BOOLEAN('z', \"null\", &null_termination,\n-\t\t\t    \"terminate entries with NUL\"),\n+\t\tOPT_SET_INT('z', \"null\", &null_termination,\n+\t\t\t    \"terminate entries with NUL\", 1),\n+\t\tOPT_SET_INT('Z', \"intense-null\", &null_termination,\n+\t\t\t    \"use NUL for all seperators, including absent values\", 2),\n \t\t{ OPTION_STRING, 'u', \"untracked-files\", &untracked_files_arg,\n \t\t  \"mode\",\n \t\t  \"show untracked files, optional modes: all, normal, no. (Default: all)\",\ndiff --git a/wt-status.c b/wt-status.c\nindex 8ca59a2..9f23ec6 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -663,7 +663,9 @@ static void wt_shortstatus_unmerged(int null_termination, struct string_list_ite\n \tcase 7: how = \"UU\"; break; /* both modified */\n \t}\n \tcolor_fprintf(s->fp, color(WT_STATUS_UNMERGED, s), \"%s\", how);\n-\tif (null_termination) {\n+\tif (null_termination == 2) {\n+\t\tfprintf(stdout, \"%c%s%c%c\", 0, it->string, 0, 0);\n+\t} else if (null_termination) {\n \t\tfprintf(stdout, \" %s%c\", it->string, 0);\n \t} else {\n \t\tstruct strbuf onebuf = STRBUF_INIT;\n@@ -687,14 +689,19 @@ static void wt_shortstatus_status(int null_termination, struct string_list_item\n \t\tcolor_fprintf(s->fp, color(WT_STATUS_CHANGED, s), \"%c\", d->worktree_status);\n \telse\n \t\tputchar(' ');\n-\tputchar(' ');\n-\tif (null_termination) {\n-\t\tfprintf(stdout, \"%s%c\", it->string, 0);\n+\tif (null_termination == 2) {\n+\t\tchar *file2 = \"\";\n+\t\tif (d->head_path)\n+\t\t\tfile2 = d->head_path;\n+\t\tfprintf(stdout, \"%c%s%c%s%c\", 0, it->string, 0, file2, 0);\n+\t} else if (null_termination) {\n+\t\tfprintf(stdout, \" %s%c\", it->string, 0);\n \t\tif (d->head_path)\n \t\t\tfprintf(stdout, \"%s%c\", d->head_path, 0);\n \t} else {\n \t\tstruct strbuf onebuf = STRBUF_INIT;\n \t\tconst char *one;\n+\t\tputchar(' ');\n \t\tif (d->head_path) {\n \t\t\tone = quote_path(d->head_path, -1, &onebuf, s->prefix);\n \t\t\tprintf(\"%s -> \", one);\n@@ -709,7 +716,9 @@ static void wt_shortstatus_status(int null_termination, struct string_list_item\n static void wt_shortstatus_untracked(int null_termination, struct string_list_item *it,\n \t\t\t    struct wt_status *s)\n {\n-\tif (null_termination) {\n+\tif (null_termination == 2) {\n+\t\tfprintf(stdout, \"??%c%s%c%c\", 0, it->string, 0, 0);\n+\t} else if (null_termination) {\n \t\tfprintf(stdout, \"?? %s%c\", it->string, 0);\n \t} else {\n \t\tstruct strbuf onebuf = STRBUF_INIT;\n-- \n1.7.0.4\n"},{"id":"139189","messageId":"20100410195003.GA28871@thyrsus.com","threadId":"23395","inReplyTo":"20100410192529.23731.79803.julian@quantumfyre.co.uk","subject":"Re: [RFC/PATCH] status: Add a new NUL separated output format","fromName":"Eric Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2010-04-10T19:50:03Z","receivedAt":"2010-04-10T19:50:03Z","isPatch":true,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk>:\n> Add a new output format option to git-status that is a more extreme\n> form of the -z output that places a NUL between all parts of the\n> record, and always has three entries per record, even when only two\n> are relevant.  This make the parsing of --porcelain output much\n> simpler for the consumer.\n\nIf you're open to changing this to lose the exiguous \"-> \" and use \"-\"\ninstead of \" \" as a status character, that would make me happy \nand fix the rest of the design problems with the format.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"139195","messageId":"1702f7c7b0e0689149702335c9efad3f@212.159.54.234","threadId":"23395","inReplyTo":"20100410195003.GA28871@thyrsus.com","subject":"Re: [RFC/PATCH] status: Add a new NUL separated output format","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2010-04-10T20:34:56Z","receivedAt":"2010-04-10T20:34:56Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Sat, 10 Apr 2010 15:50:03 -0400, Eric Raymond <esr@thyrsus.com> wrote:\n> Julian Phillips <julian@quantumfyre.co.uk>:\n>> Add a new output format option to git-status that is a more extreme\n>> form of the -z output that places a NUL between all parts of the\n>> record, and always has three entries per record, even when only two\n>> are relevant.  This make the parsing of --porcelain output much\n>> simpler for the consumer.\n> \n> If you're open to changing this to lose the exiguous \"-> \" and use \"-\"\n> instead of \" \" as a status character, that would make me happy \n> and fix the rest of the design problems with the format.\n\nIf you use \"--porcelain -Z\" then you don't get the \"->\", the format is\nalways XY<NUL><file1><NUL><file2><NUL>, with <file2> being an empty string\nif only file1 is relevant.\n\nI didn't use \"-\" instead of \" \" as that seemed out of scope for a output\nformatting option.  Though I don't personally have an objection to it, I\nalso don't see a particularly strong need for it as with the -Z format\nthere is no ambiguity.\n\nIf you're talking about the output without -Z, then changing the format\nraises compatibility issues, and were talking about something more like\n--porcelain2 or --porcelain=new and I don't know if that would be\nconsidered acceptable.\n\n-- \nJulian\n"},{"id":"139198","messageId":"20100410211223.GA29067@thyrsus.com","threadId":"23395","inReplyTo":"1702f7c7b0e0689149702335c9efad3f@212.159.54.234","subject":"Re: [RFC/PATCH] status: Add a new NUL separated output format","fromName":"Eric Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2010-04-10T21:12:23Z","receivedAt":"2010-04-10T21:12:23Z","isPatch":true,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk>:\n> I didn't use \"-\" instead of \" \" as that seemed out of scope for a output\n> formatting option.  Though I don't personally have an objection to it, I\n> also don't see a particularly strong need for it as with the -Z format\n> there is no ambiguity.\n\nGood point.  OK, the combinaation of -Z and a switch to list ignored\nfiles should solve Emacs VC's problem.  \n\nHaving some sort of JSON dump might still not be a bad idea.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"139216","messageId":"20100410230349.43948.68755.julian@quantumfyre.co.uk","threadId":"23395","inReplyTo":"20100410211223.GA29067@thyrsus.com","subject":"[RFC/PATCH] status: Add json output format","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2010-04-10T23:03:48Z","receivedAt":"2010-04-10T23:03:48Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"This adds a --json switch to status, which enables a json output\nformat.  This provides a standard output format that should be easily\nparsed by scripts using any of the large number of readily available\njson libraries.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\nOn Sat, 10 Apr 2010, Eric Raymond wrote:\n\n> Julian Phillips <julian@quantumfyre.co.uk>:\n>> I didn't use \"-\" instead of \" \" as that seemed out of scope for a output\n>> formatting option.  Though I don't personally have an objection to it, I\n>> also don't see a particularly strong need for it as with the -Z format\n>> there is no ambiguity.\n>\n> Good point.  OK, the combinaation of -Z and a switch to list ignored\n> files should solve Emacs VC's problem.\n>\n> Having some sort of JSON dump might still not be a bad idea.\n\nStarter for 10 ...\n\n builtin/commit.c |   10 ++++\n wt-status.c      |  132 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n wt-status.h      |    1 +\n 3 files changed, 143 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex c5ab683..f2b5cfa 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -91,6 +91,7 @@ static enum {\n \tSTATUS_FORMAT_LONG,\n \tSTATUS_FORMAT_SHORT,\n \tSTATUS_FORMAT_PORCELAIN,\n+\tSTATUS_FORMAT_JSON,\n } status_format = STATUS_FORMAT_LONG;\n \n static int opt_parse_m(const struct option *opt, const char *arg, int unset)\n@@ -422,6 +423,9 @@ static int run_status(FILE *fp, const char *index_file, const char *prefix, int\n \tcase STATUS_FORMAT_PORCELAIN:\n \t\twt_porcelain_print(s, null_termination);\n \t\tbreak;\n+\tcase STATUS_FORMAT_JSON:\n+\t\twt_json_print(s);\n+\t\tbreak;\n \tcase STATUS_FORMAT_LONG:\n \t\twt_status_print(s);\n \t\tbreak;\n@@ -1025,6 +1029,9 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n \t\tOPT_SET_INT(0, \"porcelain\", &status_format,\n \t\t\t    \"show porcelain output format\",\n \t\t\t    STATUS_FORMAT_PORCELAIN),\n+\t\tOPT_SET_INT(0, \"json\", &status_format,\n+\t\t\t    \"show json output format\",\n+\t\t\t    STATUS_FORMAT_JSON),\n \t\tOPT_BOOLEAN('z', \"null\", &null_termination,\n \t\t\t    \"terminate entries with NUL\"),\n \t\t{ OPTION_STRING, 'u', \"untracked-files\", &untracked_files_arg,\n@@ -1068,6 +1075,9 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n \tcase STATUS_FORMAT_PORCELAIN:\n \t\twt_porcelain_print(&s, null_termination);\n \t\tbreak;\n+\tcase STATUS_FORMAT_JSON:\n+\t\twt_json_print(&s);\n+\t\tbreak;\n \tcase STATUS_FORMAT_LONG:\n \t\ts.verbose = verbose;\n \t\twt_status_print(&s);\ndiff --git a/wt-status.c b/wt-status.c\nindex 8ca59a2..ab934e7 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -750,3 +750,135 @@ void wt_porcelain_print(struct wt_status *s, int null_termination)\n \ts->prefix = NULL;\n \twt_shortstatus_print(s, null_termination);\n }\n+\n+static char *json_quote(char *s)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\twhile (*s) {\n+\t\tswitch (*s) {\n+\t\tcase '\"':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\\\\"\");\n+\t\t\tbreak;\n+\t\tcase '\\\\':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\\\\\\");\n+\t\t\tbreak;\n+\t\tcase '\\b':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\b\");\n+\t\t\tbreak;\n+\t\tcase '\\f':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\f\");\n+\t\t\tbreak;\n+\t\tcase '\\n':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\n\");\n+\t\t\tbreak;\n+\t\tcase '\\r':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\r\");\n+\t\t\tbreak;\n+\t\tcase '\\t':\n+\t\t\tstrbuf_addstr(&buf, \"\\\\t\");\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\t/* All control characters must be encode, even if they\n+\t\t\t * don't have a specific escape character of their own */\n+\t\t\tif (*s < 0x20)\n+\t\t\t\tstrbuf_addf(&buf, \"\\\\u%04x\", *s);\n+\t\t\telse\n+\t\t\t\tstrbuf_addch(&buf, *s);\n+\t\t\tbreak;\n+\t\t}\n+\t\ts++;\n+\t}\n+\n+\treturn strbuf_detach(&buf, NULL);\n+}\n+\n+static void wt_json_unmerged(struct string_list_item *it,\n+\t\t\t   struct wt_status *s)\n+{\n+\tstruct wt_status_change_data *d = it->util;\n+\tchar ours = '?', theirs = '?';\n+\tchar *name = json_quote(it->string);\n+\n+\tswitch (d->stagemask) {\n+\tcase 1: ours = 'D'; theirs = 'D'; break; /* both deleted */\n+\tcase 2: ours = 'A'; theirs = 'U'; break; /* added by us */\n+\tcase 3: ours = 'U'; theirs = 'D'; break; /* deleted by them */\n+\tcase 4: ours = 'U'; theirs = 'A'; break; /* added by them */\n+\tcase 5: ours = 'D'; theirs = 'U'; break; /* deleted by us */\n+\tcase 6: ours = 'A'; theirs = 'A'; break; /* both added */\n+\tcase 7: ours = 'U'; theirs = 'U'; break; /* both modified */\n+\t}\n+\n+\tfprintf(stdout, \"{\");\n+\tfprintf(stdout, \"\\\"ours\\\" : \\\"%c\\\", \", ours);\n+\tfprintf(stdout, \"\\\"theirs\\\" : \\\"%c\\\", \", theirs);\n+\tfprintf(stdout, \"\\\"name\\\" : \\\"%s\\\"\", name);\n+\tfprintf(stdout, \"}\");\n+\n+\tfree(name);\n+}\n+\n+static void wt_json_status(struct string_list_item *it,\n+\t\t\t struct wt_status *s)\n+{\n+\tstruct wt_status_change_data *d = it->util;\n+\tchar index = '-', worktree = '-';\n+\tchar *name = json_quote(it->string);\n+\n+\tif (d->index_status)\n+\t\tindex = d->index_status;\n+\tif (d->worktree_status)\n+\t\tworktree = d->worktree_status;\n+\n+\tfprintf(stdout, \"{\");\n+\tfprintf(stdout, \"\\\"index\\\" : \\\"%c\\\", \", index);\n+\tfprintf(stdout, \"\\\"worktree\\\" : \\\"%c\\\", \", worktree);\n+\tfprintf(stdout, \"\\\"name\\\" : \\\"%s\\\"\", name);\n+\n+\tif (d->head_path) {\n+\t\tfree(name);\n+\t\tname = json_quote(d->head_path);\n+\t\tfprintf(stdout, \", \\\"orig_name\\\" : \\\"%s\\\"\", name);\n+\t}\n+\n+\tfprintf(stdout, \"}\");\n+\n+\tfree(name);\n+}\n+\n+void wt_json_print(struct wt_status *s)\n+{\n+\tint i;\n+\tfprintf(stdout, \"[\");\n+\tfor (i = 0; i < s->change.nr; i++) {\n+\t\tstruct wt_status_change_data *d;\n+\t\tstruct string_list_item *it;\n+\n+\t\tif (i > 0)\n+\t\t\tfprintf(stdout, \",\\n\");\n+\t\tit = &(s->change.items[i]);\n+\t\td = it->util;\n+\t\tif (d->stagemask)\n+\t\t\twt_json_unmerged(it, s);\n+\t\telse\n+\t\t\twt_json_status(it, s);\n+\t}\n+\tif (s->change.nr > 0 && s->untracked.nr > 0)\n+\t\tfprintf(stdout, \",\\n\");\n+\tfor (i = 0; i < s->untracked.nr; i++) {\n+\t\tchar *name = json_quote(s->untracked.items[i].string);\n+\n+\t\tif (i > 0)\n+\t\t\tfprintf(stdout, \",\\n\");\n+\n+\t\tfprintf(stdout, \"{\");\n+\t\tfprintf(stdout, \"\\\"index\\\" : \\\"?\\\", \");\n+\t\tfprintf(stdout, \"\\\"worktree\\\" : \\\"?\\\", \");\n+\t\tfprintf(stdout, \"\\\"name\\\" : \\\"%s\\\"\", name);\n+\t\tfprintf(stdout, \"}\");\n+\n+\t\tfree(name);\n+\t}\n+\tfprintf(stdout, \"]\\n\");\n+}\ndiff --git a/wt-status.h b/wt-status.h\nindex 9120673..effd3fd 100644\n--- a/wt-status.h\n+++ b/wt-status.h\n@@ -61,5 +61,6 @@ void wt_status_collect(struct wt_status *s);\n \n void wt_shortstatus_print(struct wt_status *s, int null_termination);\n void wt_porcelain_print(struct wt_status *s, int null_termination);\n+void wt_json_print(struct wt_status *s);\n \n #endif /* STATUS_H */\n-- \n1.7.0.4\n"},{"id":"139218","messageId":"s2l2cfc40321004101633y2857f592q2a62c5c90ea7a9de@mail.gmail.com","threadId":"23395","inReplyTo":"8871c8959d3ea4cd71452400e4c60dd0@212.159.54.234","subject":"Re: git status --porcelain is a mess that needs fixing","fromName":"Jon Seymour","fromEmail":"jon.seymour@gmail.com","sentAt":"2010-04-10T23:33:15Z","receivedAt":"2010-04-10T23:33:15Z","isPatch":false,"sender":{"key":"jon.seymour@gmail.com","avatar":"https://avatars.githubusercontent.com/u/207131?v=4"},"body":"On Sun, Apr 11, 2010 at 1:50 AM, Julian Phillips\n<julian@quantumfyre.co.uk> wrote:\n> On Sun, 11 Apr 2010 00:56:47 +1000, Jon Seymour <jon.seymour@gmail.com>\n> wrote:\n>> On Sat, Apr 10, 2010 at 11:35 PM, Julian Phillips\n>> <julian@quantumfyre.co.uk> wrote:\n>>> ...\n>>> tokens depending on the status flags.  So it would make the parsing\n>>> simpler.  But to make it even easier, how about adding a -Z that makes\n>>> the\n>>> output format \"XY\\0file1\\0[file2]\\0\" (i.e. always three tokens per\n>>> record,\n>>> with the third token being empty if there is no second filename)?\n>>>  Though\n>>> if future expandability was wanted you could end each record with \\0\\0\n>>> and\n>>> then parsing would be a two stages of split on \\0\\0 for records and\n> then\n>>> split on \\0 for entries?\n>>\n>> Surely that won't work - if file2 can be empty, \\0[file2]\\0 reduces to\n>> \\0\\0 which would be confused with the \\0\\0 proposed as a record\n>> separator.\n>\n> Yes.  But they were alternative suggestions, so if using \\0\\0 as the\n> record marker you would omit the second filename when empty as is currently\n> done.\n\nAh, apologies. I appear to have failed to parse a necessary disjunctive :-)\n\njon.\n"}]}