{"thread":{"id":"55739","subject":"RFC: error codes on exit","startedAt":"2021-05-19T23:34:31Z","lastAt":"2021-05-26T09:10:25Z","messageCount":26,"participants":["Jonathan Nieder","Felipe Contreras","Junio C Hamano","Jeff King","Jeff Hostetler","brian m. carlson","Alex Henrie","H. Peter Anvin","Bagas Sanjaya","Philip Oakley","Ævar Arnfjörð Bjarmason"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"425003","messageId":"YKWggLGDhTOY+lcy@google.com","threadId":"55739","inReplyTo":null,"subject":"RFC: error codes on exit","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2021-05-19T23:34:24Z","receivedAt":"2021-05-19T23:34:31Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\n(Danger, jrn is wading into error handling again...)\n\nAt $DAYJOB we are setting up some alerting for some bot fleets and\ndeveloper workstations, using trace2 as the data source.  Having\ntrace2 has been great --- combined with gradual weekly rollouts of\n\"next\", it helps us to understand quickly when a change is creating a\nregression for users, which hopefully improves the quality of Git for\neveryone.\n\nOne kind of signal we haven't been able to make good use of is error\nrates.  The problem is that a die() call can be an indication of\n\n a. the user asked to do something that isn't sensible, and we kindly\n    rebuked the user\n\n b. we contacted a server, and the server was not happy with our\n    request\n\n c. the local Git repository is corrupt\n\n d. we ran out of resources (e.g., disk space)\n\n e. we encountered an internal error in handling the user's\n    legitimate request\n\nand these different cases do not all motivate the same response.\n(E.g., if (c) affects just a single bot but produces a high error rate\nfrom that bot, we shouldn't be alarmed; if (d) is happening on a bot,\nthen we should look into giving it more disk; if (e) is increasing\nsignificantly during a rollout then we should roll back quickly.)\n\nIn order to do this, I would like to annotate \"exit\" events with a\nclassification of the error.  I'm not too opinionated about what that\nclassification looks like (bikeshedding welcome!) --- e.g., something\nlike the enumeration at\nhttps://github.com/googleapis/googleapis/blob/master/google/rpc/code.proto\nis likely to work fine.\n\n(I'm particularly fond of how that maps to HTTP statuses.  See also\nhttps://github.com/abseil/abseil-cpp/blob/HEAD/absl/status/status.h\nfor an example of using that kind of enumeration within a single\nprocess.)\n\nThe API could look something like\n\n\t--- a/cache.h\n\t+++ b/cache.h\n\t@@ -590,6 +590,15 @@ int is_git_directory(const char *path);\n\t  */\n\t int is_nonbare_repository_dir(struct strbuf *path);\n\n\t+enum git_error_code {\n\t+\t/*\n\t+\t * Not an error (= HTTP 200)\n\t+\t */\n\t+\tOK = 0,\n\t+};\n\t+NORETURN void fatal(enum git_error_code code, const char *err, ...)\n\t+\t__attribute__((format (printf, 2, 3)));\n\t+\n\t #define READ_GITFILE_ERR_STAT_FAILED 1\n\t #define READ_GITFILE_ERR_NOT_A_FILE 2\n\t #define READ_GITFILE_ERR_OPEN_FAILED 3\n\n(with new error codes added when they first get used) and a typical\ncaller could look like\n\n\tSubject: xsize_t: tag \"cannot handle files this big\" as a failed precondition\n\n\tUnlike retriable errors, failed preconditions indicate that some\n\taspect of the state needs to be changed in order to recover.  Mark\n\tthis error as such to make signals from monitoring in controlled\n\tenvironments (e.g., bot fleets or corporate installations of Git)\n\teasier to understand.\n\n\tSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n[...]\n\t+       /*\n\t+        * The system is not in a state required for the operation to succeed.\n\t+        * For example, a file on disk is larger than we can handle.\n\t+        * (= HTTP 400)\n\t+        */\n\t+       FAILED_PRECONDITION = 9,\n[...]\n\t static inline size_t xsize_t(off_t len)\n\t {\n\t\tif (len < 0 || len > SIZE_MAX)\n\t-               die(\"Cannot handle files this big\");\n\t+               fatal(FAILED_PRECONDITION, \"Cannot handle files this big\");\n\nFurther down the line I can imagine making use of git_error_code\nelsewhere for e.g. some limited retries of the corresponding\ntransaction when we fail to lock a file.\n\nThoughts?  Good idea?  Bad idea?\n\nThanks,\nJonathan\n"},{"id":"425009","messageId":"60a5afeeb13b4_1d8f2208a5@natae.notmuch","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"RE: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-20T00:40:14Z","receivedAt":"2021-05-20T00:40:20Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Jonathan Nieder wrote:\n> The API could look something like\n> \n> \t--- a/cache.h\n> \t+++ b/cache.h\n> \t@@ -590,6 +590,15 @@ int is_git_directory(const char *path);\n> \t  */\n> \t int is_nonbare_repository_dir(struct strbuf *path);\n> \n> \t+enum git_error_code {\n> \t+\t/*\n> \t+\t * Not an error (= HTTP 200)\n> \t+\t */\n> \t+\tOK = 0,\n\nIt's good to not include many initial codes, but I would start with at\nleast three:\n\n  OK = 0,\n  UNKNOWN = 1,\n  NORMAL = 2,\n\ndie() could be mapped to UNKNOWN.\n\n> \t+};\n> \t+NORETURN void fatal(enum git_error_code code, const char *err, ...)\n> \t+\t__attribute__((format (printf, 2, 3)));\n> \t+\n\nfatal() for me sounds 1) very dramatic, 2), not a verb, 3) and not a\ncomplete thing (fatal what?)\n\nI would prefer \"fail\", or \"fault\", or anything that is a verb.\n\n> Thoughts?  Good idea?  Bad idea?\n\nGreat idea.\n\n-- \nFelipe Contreras\n"},{"id":"425011","messageId":"xmqqeee2w7ov.fsf@gitster.g","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"Re: RFC: error codes on exit","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-20T00:49:20Z","receivedAt":"2021-05-20T00:49:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> One kind of signal we haven't been able to make good use of is error\n> rates.  The problem is that a die() call can be an indication of\n>\n>  a. the user asked to do something that isn't sensible, and we kindly\n>     rebuked the user\n> ...\n>  e. we encountered an internal error in handling the user's\n>     legitimate request\n>\n> and these different cases do not all motivate the same response.\n> ...\n> In order to do this, I would like to annotate \"exit\" events with a\n> classification of the error.\n\nWe already have BUG() for e. and die() for everything else, and\n\"everything else\" may be overly broad for your purpose.\n\nI am sympathetic to the cause and I agree that introducing a\nfiner-grained classification might be a solution.  I however am not\nsure how we can enforce developers to apply such a manually assigned\n\"error code\" cosistently.\n\nJust to throw in a totally different alternative to see if it works\nbetter, I wonder if you can teach die() to report to the trace2\nstream where in the code it was called from and which vintage of Git\nit is running.\n\nThe stat collection side that cares about certain class of failures\ncan have function that maps \"die() at <filename>:<lineno>@<version>\"\nto \"what kind of die() it is\".  \n\nE.g.  blame.c:50@v2.32.0-rc0-184-gbbde7e6616\" may be BUG(), while\nblame.c:2740@v2.32.0-rc0-184-gbbde7e6616 may be an user-error.\n\nThat way, our developers do not have to do anything special and\ncannot do anything to screw up the classification.\n"},{"id":"425012","messageId":"60a5b9165fd5b_1e27520848@natae.notmuch","threadId":"55739","inReplyTo":"xmqqeee2w7ov.fsf@gitster.g","subject":"Re: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-20T01:19:18Z","receivedAt":"2021-05-20T01:19:22Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> > In order to do this, I would like to annotate \"exit\" events with a\n> > classification of the error.\n> \n> We already have BUG() for e. and die() for everything else, and\n> \"everything else\" may be overly broad for your purpose.\n\nBUG() and die() can call fatal().\n\n> I am sympathetic to the cause and I agree that introducing a\n> finer-grained classification might be a solution.  I however am not\n> sure how we can enforce developers to apply such a manually assigned\n> \"error code\" cosistently.\n\nYou don't enforce developers to do this--just like you don't enforce\ndevelopers to use advice() instead of fprintf(stderr, )... You nudge\nthem in that direction, and eventually it becomes a habit.\n\nDevelopers in other languages and stacks have no problem with this\ngranularity. They do this in languages like JavaScript, C++, Ruby and\nPython regularlly. And developers dealing with HTTP have no trouble with\nerror codes (like 200, 400, and 404).\n\n-- \nFelipe Contreras\n"},{"id":"425013","messageId":"YKXBhDbWMyB6A7z4@google.com","threadId":"55739","inReplyTo":"xmqqeee2w7ov.fsf@gitster.g","subject":"Re: RFC: error codes on exit","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2021-05-20T01:55:16Z","receivedAt":"2021-05-20T01:56:57Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJunio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> In order to do this, I would like to annotate \"exit\" events with a\n>> classification of the error.\n>\n> We already have BUG() for e. and die() for everything else, and\n> \"everything else\" may be overly broad for your purpose.\n>\n> I am sympathetic to the cause and I agree that introducing a\n> finer-grained classification might be a solution.  I however am not\n> sure how we can enforce developers to apply such a manually assigned\n> \"error code\" cosistently.\n\nI think two things you're hinting at are \"what about maintainability?\"\nand \"what is the migration path?\"\n\nI suspect that a number of error paths will remain unclassified for a\nlong time, possibly indefinitely.  The way I've seen this treated in\nother tools is that it's okay for something to show up as an INTERNAL\nerror if it doesn't happen frequently: sure, that can cause us a bit\nof unnecessary worry when it starts occuring more often, but at that\npoint we're in a good place to replace it with something more\nappropriate.\n\nThat means that we would still want to keep die() or some equivalent.\nThat in turn might suggest that the API I suggested is overly verbose;\nit might make sense to have a different die()-style helper for each\ntype of error, matching what we do with die() and BUG().\n\nSide note: you might wonder why keeping die() would even be a\nquestion.  For example, there are all the outstanding patches that\nstill use die(); changing such a fundamental API would seem to be a\nnonstarter.  Fortunately, though, the tools in\ncontrib/coccinelle/README allow changing an API in three steps:\n\n 1. Introduce the new API.  Keep the old API around for backward\n    compatibility.\n\n 2. Add a \"pending\" coccinelle semantic patch to automatically\n    update callers to the new API.  Update existing callers using\n    'make coccicheck-pending'.\n\n 3. Remove the old API and mark the semantic patch as no longer\n    pending.  Patches using the old API can be fixed using 'make\n    coccicheck'\n\nSo we can make this decision based on whether the resulting API is one\nwe like more; in this example, I suspect that keeping die() is\npreferable _even though_ it would be possible to remove by staying in\nstep 2 for a while without too much fuss.\n\n> Just to throw in a totally different alternative to see if it works\n> better, I wonder if you can teach die() to report to the trace2\n> stream where in the code it was called from and which vintage of Git\n> it is running.\n>\n> The stat collection side that cares about certain class of failures\n> can have function that maps \"die() at <filename>:<lineno>@<version>\"\n> to \"what kind of die() it is\".\n>\n> E.g.  blame.c:50@v2.32.0-rc0-184-gbbde7e6616\" may be BUG(), while\n> blame.c:2740@v2.32.0-rc0-184-gbbde7e6616 may be an user-error.\n\nFor ad hoc queries, this is a rather nice tool.  Traces already record\nfilename, version, and line number, though I believe in the die() case\nit currently just points to the implementation of die(). ;-)\n\nHowever, for analysis in aggregate (for example, to define an SLO[1])\nthat would require us to maintain a database that maps\n<filename>:<lineno> to error code.  That database would be essentially\nthe same as patches to record the error codes, so what it would really\namount to is having these deployments using a permanent fork of Git.\nIt would also get rid of the chance to discuss and improve common\nerror paths on-list.\n\nIf we expect the error codes to not be useful to anyone else, then\nthat is the right choice to make (or rather, we'd have to use other\nheuristics, such as having the traces record a collection of offsets\nin the binary and a build-id so we can key off of stack trace\nsignatures).  Part of the reason I started this thread is to get a\nsense of whether these can be useful to others.\n\nThanks,\nJonathan\n\n[1] https://sre.google/sre-book/service-level-objectives/\n"},{"id":"425021","messageId":"xmqqy2cauojm.fsf@gitster.g","threadId":"55739","inReplyTo":"YKXBhDbWMyB6A7z4@google.com","subject":"Re: RFC: error codes on exit","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-20T02:28:13Z","receivedAt":"2021-05-20T02:33:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> I am sympathetic to the cause and I agree that introducing a\n>> finer-grained classification might be a solution.  I however am not\n>> sure how we can enforce developers to apply such a manually assigned\n>> \"error code\" cosistently.\n>\n> I think two things you're hinting at are \"what about maintainability?\"\n> and \"what is the migration path?\"\n\nNot really.  There is not much to migrate from, and mislabeling by\ndevelopers, or more realistically disagreement between developers\nwhat label is appropriate for which condition that leads to die(),\nwould be a usability problem and not maintainability that I would be\nworried about.\n\n> However, for analysis in aggregate (for example, to define an SLO[1])\n> that would require us to maintain a database that maps\n> <filename>:<lineno> to error code.  That database would be essentially\n> the same as patches to record the error codes, so what it would really\n> amount to is having these deployments using a permanent fork of Git.\n\nYes, the suggestion was to start from there to gain experience,\nbecause ...\n\n> If we expect the error codes to not be useful to anyone else, then\n> that is the right choice to make (or rather, we'd have to use other\n> heuristics, such as having the traces record a collection of offsets\n> in the binary and a build-id so we can key off of stack trace\n> signatures).  Part of the reason I started this thread is to get a\n> sense of whether these can be useful to others.\n\n... others will find it useful if the classification matches _their_\nneeds, but I suspect the \"bin\" a single die() location wants to be\nclassfied into would end up to be different depending on what the\nlog collecting entity is after.\n"},{"id":"425102","messageId":"YKZj/s/9dp4Oo7aB@coredump.intra.peff.net","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"Re: RFC: error codes on exit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-05-20T13:28:30Z","receivedAt":"2021-05-20T13:28:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 19, 2021 at 04:34:24PM -0700, Jonathan Nieder wrote:\n\n> One kind of signal we haven't been able to make good use of is error\n> rates.  The problem is that a die() call can be an indication of\n> \n>  a. the user asked to do something that isn't sensible, and we kindly\n>     rebuked the user\n> \n>  b. we contacted a server, and the server was not happy with our\n>     request\n> \n>  c. the local Git repository is corrupt\n> \n>  d. we ran out of resources (e.g., disk space)\n> \n>  e. we encountered an internal error in handling the user's\n>     legitimate request\n\nI've run into this problem, too. If you run a website that runs Git\ncommands on behalf of users and try to get metrics on failing exit\ncodes, it's hard to tell the difference between \"the repo is broken\",\n\"Git has a bug\", \"the user (or other caller) asked for something\nstupid\", and \"some transient error occurred\".\n\nBut I'm not sure that even Git can always tell the difference between\nthose things. Some real-world examples I've run into:\n\n  - \"rev-list $oid\" can't find object $oid. Is the repo corrupt? Or is\n    the caller unreasonable to ask for that object? Or was there a race\n    or other transient error which made the object invisible?\n\n  - upload-pack is writing out a packfile, but gets EPIPE. Did the\n    network drop out? Or is a Git bug causing one side to break\n    protocol?\n\nSome rough categorization may help, but a lot of those need to propagate\nthe specific errors back to the caller. For instance, the rev-list\nexample could be FAILED_PRECONDITION in your terminology. But really, we\nwant to tell the caller \"the object you asked for doesn't exist\". And\nthen it can decide if that was user error (somebody hitting a URL for an\nobject that we have no reason to think exists), or a sign of problems\nelsewhere in the system (if we just got $oid from Git, we expect it to\nbe there).\n\nSo it seems like the most useful thing is specific error codes for\nspecific cases. And that gets very daunting to think about annotating\nand communicating about each such case (we don't even pass that level of\ndetailed information inside the program in a machine-readable way;\nscraping stderr is the best way to figure this stuff out now).\n\nI dunno. Maybe a rougher categorization would help your case, but not\nmine. But I'm a bit skeptical that we'll have enough coverage of various\nconditions to be useful, and that it won't turn into a headache trying\nto categorize everything.\n\n-Peff\n"},{"id":"425110","messageId":"795fd316-2bb5-e382-b104-85d1aaa09a1c@jeffhostetler.com","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"Re: RFC: error codes on exit","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2021-05-20T15:09:44Z","receivedAt":"2021-05-20T15:09:48Z","isPatch":false,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 5/19/21 7:34 PM, Jonathan Nieder wrote:\n> Hi,\n> \n> (Danger, jrn is wading into error handling again...)\n> \n> At $DAYJOB we are setting up some alerting for some bot fleets and\n> developer workstations, using trace2 as the data source.  Having\n> trace2 has been great --- combined with gradual weekly rollouts of\n> \"next\", it helps us to understand quickly when a change is creating a\n> regression for users, which hopefully improves the quality of Git for\n> everyone.\n> \n> One kind of signal we haven't been able to make good use of is error\n> rates.  The problem is that a die() call can be an indication of\n> \n>   a. the user asked to do something that isn't sensible, and we kindly\n>      rebuked the user\n> \n>   b. we contacted a server, and the server was not happy with our\n>      request\n> \n>   c. the local Git repository is corrupt\n> \n>   d. we ran out of resources (e.g., disk space)\n> \n>   e. we encountered an internal error in handling the user's\n>      legitimate request\n...\n\nFor the error event that `error()` and `die()` and friends generate,\nI emit both the fully formatted error message and the format string.\n\nThe latter, if used as a dictionary key, would let you group like\nevents from different processes without worrying about the filename\nor blob id or remote name or etc. in any one particular instance.\n\nWould that be sufficient as an error classification and something\nthat you can key off of in your post-processing ?\n\nGranted the same format message might be used in multiple places in\nthe source, but I also provide the source filename and line number.\n\nIf it turns out that all of the error events come from \"usage.c\"\n(i.e. error_builtin() or die_builtin()), then maybe we need to look\nat another way of wrapping those calls to pass the F/L of actual\ncaller.  I hesitated to do that because of the existing indirection\ntricks in usage.c WRT the `set_error_routine()` and friends.\n(And that assumes that the format string is a viable solution for\nyou problem.)\n\nJeff\n\n"},{"id":"425136","messageId":"YKaguiSjewjpvOj5@google.com","threadId":"55739","inReplyTo":"YKZj/s/9dp4Oo7aB@coredump.intra.peff.net","subject":"Re: RFC: error codes on exit","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2021-05-20T17:47:38Z","receivedAt":"2021-05-20T17:47:44Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJeff King wrote:\n> On Wed, May 19, 2021 at 04:34:24PM -0700, Jonathan Nieder wrote:\n\n>> One kind of signal we haven't been able to make good use of is error\n>> rates.  The problem is that a die() call can be an indication of\n[... list of some categories snipped ...]\n> I've run into this problem, too. If you run a website that runs Git\n> commands on behalf of users and try to get metrics on failing exit\n> codes, it's hard to tell the difference between \"the repo is broken\",\n> \"Git has a bug\", \"the user (or other caller) asked for something\n> stupid\", and \"some transient error occurred\".\n>\n> But I'm not sure that even Git can always tell the difference between\n> those things. Some real-world examples I've run into:\n>\n>   - \"rev-list $oid\" can't find object $oid. Is the repo corrupt? Or is\n>     the caller unreasonable to ask for that object? Or was there a race\n>     or other transient error which made the object invisible?\n>\n>   - upload-pack is writing out a packfile, but gets EPIPE. Did the\n>     network drop out? Or is a Git bug causing one side to break\n>     protocol?\n>\n> Some rough categorization may help, but a lot of those need to propagate\n> the specific errors back to the caller. For instance, the rev-list\n> example could be FAILED_PRECONDITION in your terminology. But really, we\n> want to tell the caller \"the object you asked for doesn't exist\". And\n> then it can decide if that was user error (somebody hitting a URL for an\n> object that we have no reason to think exists), or a sign of problems\n> elsewhere in the system (if we just got $oid from Git, we expect it to\n> be there).\n>\n> So it seems like the most useful thing is specific error codes for\n> specific cases.\n\nAs a bit of precedent: in the server we have both of these: we have\napplication-specific error codes like AUTHENTICATOR_EXPIRED, and we\nhave generic codes that those map to like UNAUTHENTICATED.  The\napplication-specific codes tend to be useful for ad hoc queries as\npart of incident response, versus the generic codes that have been\nmore useful for defining an SLO (because they are about \"how do I want\nto respond to this error\" instead of about the cause).\n\nMore specifically for the \"missing object\" case: it's common enough\nfor a user to ask for an object that doesn't exist that we indeed call\nit FAILED_PRECONDITION, which has worked well.  We have other\nmonitoring in place for checking that all repositories pass fsck and\nthat fetching an object after pushing it succeeds.  (In general, this\nkind of case is very common in monitoring any service that has state.)\n\n>                 And that gets very daunting to think about annotating\n> and communicating about each such case (we don't even pass that level of\n> detailed information inside the program in a machine-readable way;\n> scraping stderr is the best way to figure this stuff out now).\n\nThis feels like good news to me: it sounds like if we add\napplication-specific codes like MISSING_OBJECT to Git, then it would\nbe useful to both of us.\n\nThe mapping to HTTP-status-style generic codes could then wait for\nlater, to be submitted if and when others have interest.  (I.e., that\npart is easy to keep maintained internally.)\n\nSo I'm feeling encouraged. :)\n\n> I dunno. Maybe a rougher categorization would help your case, but not\n> mine. But I'm a bit skeptical that we'll have enough coverage of various\n> conditions to be useful, and that it won't turn into a headache trying\n> to categorize everything.\n\nTwo more points I want to emphasize:\n\n 1. We don't have to be exhaustive: as Felipe suggested, it's fine for\n    some errors (even most error paths!) to use a code such as\n    UNKNOWN.  I care more about coverage of commonly occuring errors\n    than categorizing everything, especially because this sets up a\n    feedback loop that can lead to improved coverage over time.\n\n 2. By focusing on the practical and ignoring everything else, I think\n    we can avoid this becoming an unbounded taxonomy exercise.  That's\n    part of the appeal of code.proto /\n    https://github.com/abseil/abseil-cpp/blob/HEAD/absl/status/status.h\n    for me: by using a preexisting list of codes based on \"here is\n    what a user would be expected to do in response to this error\",\n    they make the error classification decision relatively simple.  I\n    think we can maintain that kind of simplicity with a Git-specific\n    enum, too, so I think this is doable (and I'd make sure to be\n    available over time to help answer questions about the\n    classification as the project gets used to it).\n\nThanks,\nJonathan\n"},{"id":"425174","messageId":"YKcK2elWqBQ+IVDt@camp.crustytoothpaste.net","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"Re: RFC: error codes on exit","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-05-21T01:20:25Z","receivedAt":"2021-05-21T01:20:32Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-05-19 at 23:34:24, Jonathan Nieder wrote:\n> Hi,\n> \n> (Danger, jrn is wading into error handling again...)\n> \n> At $DAYJOB we are setting up some alerting for some bot fleets and\n> developer workstations, using trace2 as the data source.  Having\n> trace2 has been great --- combined with gradual weekly rollouts of\n> \"next\", it helps us to understand quickly when a change is creating a\n> regression for users, which hopefully improves the quality of Git for\n> everyone.\n> \n> One kind of signal we haven't been able to make good use of is error\n> rates.  The problem is that a die() call can be an indication of\n> \n>  a. the user asked to do something that isn't sensible, and we kindly\n>     rebuked the user\n> \n>  b. we contacted a server, and the server was not happy with our\n>     request\n> \n>  c. the local Git repository is corrupt\n> \n>  d. we ran out of resources (e.g., disk space)\n> \n>  e. we encountered an internal error in handling the user's\n>     legitimate request\n> \n> and these different cases do not all motivate the same response.\n> (E.g., if (c) affects just a single bot but produces a high error rate\n> from that bot, we shouldn't be alarmed; if (d) is happening on a bot,\n> then we should look into giving it more disk; if (e) is increasing\n> significantly during a rollout then we should roll back quickly.)\n\nIn general, I'm in favor of adding some sort of error code here.  Even\nthough I don't normally use trace2, I think there's a lot of benefit to\nhaving a standardized set of error codes, and this seems like as good a\nplace as any to introduce them.\n\nA future iteration of this might look like us returning a negative error\ncode from a function instead of -1 for us to signal to the caller that a\nparticular error case occurred.  We need not implement that now, of\ncourse, but I bring it up in case we want to accommodate that in our\ndesign now for future us.\n\nI do agree with Peff that this may not necessarily provide all of the\ninsight you want, since it can be hard to distinguish why the error\noccurred.  For example, in Git LFS, we sometimes will pass objects we\ndon't have to git rev-list (with --missing) and that's completely\nexpected, whereas a missing object with git fsck would generally be\ncause for alarm.  Provided you're comfortable with some ambiguity, I\nthink this would be a nice improvement.\n-- \nbrian m. carlson (he/him or they/them)\nHouston, Texas, US\n"},{"id":"425175","messageId":"YKcOA7iVstB7LtZH@camp.crustytoothpaste.net","threadId":"55739","inReplyTo":"795fd316-2bb5-e382-b104-85d1aaa09a1c@jeffhostetler.com","subject":"Re: RFC: error codes on exit","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-05-21T01:33:55Z","receivedAt":"2021-05-21T01:34:33Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-05-20 at 15:09:44, Jeff Hostetler wrote:\n> For the error event that `error()` and `die()` and friends generate,\n> I emit both the fully formatted error message and the format string.\n> \n> The latter, if used as a dictionary key, would let you group like\n> events from different processes without worrying about the filename\n> or blob id or remote name or etc. in any one particular instance.\n> \n> Would that be sufficient as an error classification and something\n> that you can key off of in your post-processing ?\n\nI don't think that's going to be sufficient.  Many calls to die()\ncontain a translated string.  I am a native Anglophone and work in a\ncompany where English is the sole language of communication, but I have\nmy computer configured in French, and I have had it configured in\nSpanish as well.\n\nIt's totally possible that one of my colleagues who has a non-English\nnative language might be using a different language as well, so it would\nbe difficult to reliably map the format string into a fixed error case\nin a typical corporate setting, since many languages might be in use.\n\n> Granted the same format message might be used in multiple places in\n> the source, but I also provide the source filename and line number.\n\nThe source filename and line number would be more helpful, but\ninconveniently, people frequently change the code of Git, so the line\nnumbers aren't always stable over time.  So I think an error code would\nbe helpful nevertheless.\n-- \nbrian m. carlson (he/him or they/them)\nHouston, Texas, US\n"},{"id":"425193","messageId":"YKeAtj7wwW0Qskm+@coredump.intra.peff.net","threadId":"55739","inReplyTo":"YKaguiSjewjpvOj5@google.com","subject":"Re: RFC: error codes on exit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-05-21T09:43:18Z","receivedAt":"2021-05-21T09:43:37Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 20, 2021 at 10:47:38AM -0700, Jonathan Nieder wrote:\n\n> >                 And that gets very daunting to think about annotating\n> > and communicating about each such case (we don't even pass that level of\n> > detailed information inside the program in a machine-readable way;\n> > scraping stderr is the best way to figure this stuff out now).\n> \n> This feels like good news to me: it sounds like if we add\n> application-specific codes like MISSING_OBJECT to Git, then it would\n> be useful to both of us.\n\nPerhaps. I think the context matters between \"missing an object from the\ncommand line\" and \"missing an object I expected to find while\ntraversing\". And I'm not sure all spots which look up an object will\nknow that context.\n\nIn some sense that's \"just\" a programming problem; surfacing the errors\nto the right spot that can decide how to exit. But I worry a bit that\nit's fighting uphill against the current code structure. There's\nprobably going to be a period where MISSING_OBJECT versus UNKNOWN is\nwildly inaccurate, and a long tail of cases to fix.\n\nErring to say \"UNKNOWN\" is probably OK for most callers (they are happy\nto learn of a specific error and act on it appropriately, but if Git\ncan't tell it to them, they have a generic path). But erring in the\nother direction might be bad (you fail to realize a repo is corrupt, and\ninstead attribute it to caller error).\n\nSo again, I return \"I dunno\". Something of this magnitude probably has\nto be done incrementally and over time. But I'd be loathe to trust it\nand convert existing callers use it for a while. And that creates a\nchicken-and-egg problem for finding the places which need improvement.\n\n-Peff\n"},{"id":"425221","messageId":"CAMMLpeScunGg5WM4N90vG+yN3tOATqhsL2iRLsJ43ksNyTx_wQ@mail.gmail.com","threadId":"55739","inReplyTo":"60a5afeeb13b4_1d8f2208a5@natae.notmuch","subject":"Re: RFC: error codes on exit","fromName":"Alex Henrie","fromEmail":"alexhenrie24@gmail.com","sentAt":"2021-05-21T16:53:58Z","receivedAt":"2021-05-21T16:54:17Z","isPatch":false,"sender":{"key":"alexhenrie24@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5951993?v=4"},"body":"On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n<felipe.contreras@gmail.com> wrote:\n>\n> It's good to not include many initial codes, but I would start with at\n> least three:\n>\n>   OK = 0,\n>   UNKNOWN = 1,\n>   NORMAL = 2,\n\nIf you go that route, could you please pick a word other than \"normal\"\nto describe errors that are not entirely unexpected? I'm worried that\nsomeone will see \"normal\" and use it instead of \"OK\" to indicate\nsuccess.\n\n-Alex\n"},{"id":"425280","messageId":"dc14c50d-c626-19f8-e615-52ca3c9051dc@zytor.com","threadId":"55739","inReplyTo":"CAMMLpeScunGg5WM4N90vG+yN3tOATqhsL2iRLsJ43ksNyTx_wQ@mail.gmail.com","subject":"Re: RFC: error codes on exit","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2021-05-21T23:20:01Z","receivedAt":"2021-05-21T23:21:37Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"\n\nOn 5/21/21 9:53 AM, Alex Henrie wrote:\n> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n> <felipe.contreras@gmail.com> wrote:\n>>\n>> It's good to not include many initial codes, but I would start with at\n>> least three:\n>>\n>>    OK = 0,\n>>    UNKNOWN = 1,\n>>    NORMAL = 2,\n> \n> If you go that route, could you please pick a word other than \"normal\"\n> to describe errors that are not entirely unexpected? I'm worried that\n> someone will see \"normal\" and use it instead of \"OK\" to indicate\n> success.\n> \n\n<sysexits.h>\n\n\t-hpa\n\n"},{"id":"425286","messageId":"657d3d24-2f08-f076-5c84-9ae434149530@gmail.com","threadId":"55739","inReplyTo":"dc14c50d-c626-19f8-e615-52ca3c9051dc@zytor.com","subject":"Re: RFC: error codes on exit","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2021-05-22T04:06:11Z","receivedAt":"2021-05-22T04:06:20Z","isPatch":false,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 22/05/21 06.20, H. Peter Anvin wrote:\n> \n> <sysexits.h>\n> \n>      -hpa\n> \nLooking that header file you mention, I saw:\n\n> #define EX_OK\t\t0\t/* successful termination */\n> \n> #define EX__BASE\t64\t/* base value for error messages */\n> \n> #define EX_USAGE\t64\t/* command line usage error */\n> #define EX_DATAERR\t65\t/* data format error */\n> #define EX_NOINPUT\t66\t/* cannot open input */\n> #define EX_NOUSER\t67\t/* addressee unknown */\n> #define EX_NOHOST\t68\t/* host name unknown */\n> #define EX_UNAVAILABLE\t69\t/* service unavailable */\n> #define EX_SOFTWARE\t70\t/* internal software error */\n> #define EX_OSERR\t71\t/* system error (e.g., can't fork) */\n> #define EX_OSFILE\t72\t/* critical OS file missing */\n> #define EX_CANTCREAT\t73\t/* can't create (user) output file */\n> #define EX_IOERR\t74\t/* input/output error */\n> #define EX_TEMPFAIL\t75\t/* temp failure; user is invited to retry */\n> #define EX_PROTOCOL\t76\t/* remote error in protocol */\n> #define EX_NOPERM\t77\t/* permission denied */\n> #define EX_CONFIG\t78\t/* configuration error */\n\nFor EX_USAGE case, we may sometimes display correct usage syntax so that\nusers can fix their typing.\n\nWe may use EX_CONFIG when we encounter any errors when parsing .gitconfig.\n\nEX_OSFILE isn't necessary for Git, though.\n\n-- \nAn old man doll... just what I always wanted! - Clara\n"},{"id":"425291","messageId":"xmqqfsyfqhkq.fsf@gitster.g","threadId":"55739","inReplyTo":"dc14c50d-c626-19f8-e615-52ca3c9051dc@zytor.com","subject":"Re: RFC: error codes on exit","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-22T08:49:09Z","receivedAt":"2021-05-22T08:49:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> On 5/21/21 9:53 AM, Alex Henrie wrote:\n>> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n>> <felipe.contreras@gmail.com> wrote:\n>>>\n>>> It's good to not include many initial codes, but I would start with at\n>>> least three:\n>>>\n>>>    OK = 0,\n>>>    UNKNOWN = 1,\n>>>    NORMAL = 2,\n>> If you go that route, could you please pick a word other than\n>> \"normal\"\n>> to describe errors that are not entirely unexpected? I'm worried that\n>> someone will see \"normal\" and use it instead of \"OK\" to indicate\n>> success.\n>> \n>\n> <sysexits.h>\n\nIs the value assignment standardized across systems?\n\nWe want human-readable names in the source to help developers while\nwe want platform neutral output in the log so that log collectors\ncan do some \"intelligent\" things about the output.  If EX_USAGE is\nalways 64 everywhere, that is great---we can emit \"64\" in the log\nand log collectors can take it as if they saw \"EX_USAGE\".  But if\nthe value assignment is platform-dependent, it does not help all\nthat much.\n\n    Side note.  We had a similar discussion on <errno.h> and\n    strerror(); the numbers do not help without knowing which\n    platform the error came from, and strerror() output is localized\n    and not suitable for machine consumption.\n\nIn a sense, it is worse than we keep a central mapping between names\nprogrammers use to give to the new fatal() helper function and the\nstring the tracing machinery will emit for these names.\n\nThanks.\n"},{"id":"425293","messageId":"357C5DA0-6A1B-4A69-8BBD-5D12327C136A@zytor.com","threadId":"55739","inReplyTo":"xmqqfsyfqhkq.fsf@gitster.g","subject":"Re: RFC: error codes on exit","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2021-05-22T09:08:10Z","receivedAt":"2021-05-22T09:09:48Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"sysexits.h numbers are the same across all platforms which have it, I believe. I think it originated with Sendmail wanting to know a bit more about why any subprocesses exited.\n\nOn May 22, 2021 1:49:09 AM PDT, Junio C Hamano <gitster@pobox.com> wrote:\n>\"H. Peter Anvin\" <hpa@zytor.com> writes:\n>\n>> On 5/21/21 9:53 AM, Alex Henrie wrote:\n>>> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n>>> <felipe.contreras@gmail.com> wrote:\n>>>>\n>>>> It's good to not include many initial codes, but I would start with\n>at\n>>>> least three:\n>>>>\n>>>>    OK = 0,\n>>>>    UNKNOWN = 1,\n>>>>    NORMAL = 2,\n>>> If you go that route, could you please pick a word other than\n>>> \"normal\"\n>>> to describe errors that are not entirely unexpected? I'm worried\n>that\n>>> someone will see \"normal\" and use it instead of \"OK\" to indicate\n>>> success.\n>>> \n>>\n>> <sysexits.h>\n>\n>Is the value assignment standardized across systems?\n>\n>We want human-readable names in the source to help developers while\n>we want platform neutral output in the log so that log collectors\n>can do some \"intelligent\" things about the output.  If EX_USAGE is\n>always 64 everywhere, that is great---we can emit \"64\" in the log\n>and log collectors can take it as if they saw \"EX_USAGE\".  But if\n>the value assignment is platform-dependent, it does not help all\n>that much.\n>\n>    Side note.  We had a similar discussion on <errno.h> and\n>    strerror(); the numbers do not help without knowing which\n>    platform the error came from, and strerror() output is localized\n>    and not suitable for machine consumption.\n>\n>In a sense, it is worse than we keep a central mapping between names\n>programmers use to give to the new fatal() helper function and the\n>string the tracing machinery will emit for these names.\n>\n>Thanks.\n\n-- \nSent from my Android device with K-9 Mail. Please excuse my brevity.\n"},{"id":"425294","messageId":"7f0c9ab8-c1ca-171b-8247-6d921702f3bc@iee.email","threadId":"55739","inReplyTo":"CAMMLpeScunGg5WM4N90vG+yN3tOATqhsL2iRLsJ43ksNyTx_wQ@mail.gmail.com","subject":"Re: RFC: error codes on exit","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2021-05-22T09:12:03Z","receivedAt":"2021-05-22T09:12:07Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 21/05/2021 17:53, Alex Henrie wrote:\n> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n> <felipe.contreras@gmail.com> wrote:\n>> It's good to not include many initial codes, but I would start with at\n>> least three:\n>>\n>>   OK = 0,\n>>   UNKNOWN = 1,\n>>   NORMAL = 2,\n> If you go that route, could you please pick a word other than \"normal\"\n> to describe errors that are not entirely unexpected? I'm worried that\n> someone will see \"normal\" and use it instead of \"OK\" to indicate\n> success.\n>\n> -Alex\nTypical <== Normal\n\nThough abnormal and atypical often have different implications ;-)\nP.\n"},{"id":"425347","messageId":"60a97550287b3_857e9208b8@natae.notmuch","threadId":"55739","inReplyTo":"7f0c9ab8-c1ca-171b-8247-6d921702f3bc@iee.email","subject":"Re: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-22T21:19:12Z","receivedAt":"2021-05-22T21:19:23Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Philip Oakley wrote:\n> On 21/05/2021 17:53, Alex Henrie wrote:\n> > On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n> > <felipe.contreras@gmail.com> wrote:\n> >> It's good to not include many initial codes, but I would start with at\n> >> least three:\n> >>\n> >>   OK = 0,\n> >>   UNKNOWN = 1,\n> >>   NORMAL = 2,\n> > If you go that route, could you please pick a word other than \"normal\"\n> > to describe errors that are not entirely unexpected? I'm worried that\n> > someone will see \"normal\" and use it instead of \"OK\" to indicate\n> > success.\n> >\n> > -Alex\n> Typical <== Normal\n> \n> Though abnormal and atypical often have different implications ;-)\n> P.\n\nOr USUAL.\n\n-- \nFelipe Contreras\n"},{"id":"425348","messageId":"60a976221c390_857e920812@natae.notmuch","threadId":"55739","inReplyTo":"xmqqfsyfqhkq.fsf@gitster.g","subject":"Re: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-22T21:22:42Z","receivedAt":"2021-05-22T21:22:46Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Junio C Hamano wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n> \n> > On 5/21/21 9:53 AM, Alex Henrie wrote:\n> >> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n> >> <felipe.contreras@gmail.com> wrote:\n> >>>\n> >>> It's good to not include many initial codes, but I would start with at\n> >>> least three:\n> >>>\n> >>>    OK = 0,\n> >>>    UNKNOWN = 1,\n> >>>    NORMAL = 2,\n> >> If you go that route, could you please pick a word other than\n> >> \"normal\"\n> >> to describe errors that are not entirely unexpected? I'm worried that\n> >> someone will see \"normal\" and use it instead of \"OK\" to indicate\n> >> success.\n> >> \n> >\n> > <sysexits.h>\n> \n> Is the value assignment standardized across systems?\n\nI think his intention was to suggest to use that list as inspiration...\nAs in have USAGE, NOINPUT, UNAVAILABE, etc.\n\nI would prefer to start with something easy... UNKNOWN = 1, USUAL = 2.\n\n-- \nFelipe Contreras\n"},{"id":"425350","messageId":"3C6468D1-3E14-4600-BC8E-86CCCB84E74C@zytor.com","threadId":"55739","inReplyTo":"60a976221c390_857e920812@natae.notmuch","subject":"Re: RFC: error codes on exit","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2021-05-22T21:29:39Z","receivedAt":"2021-05-22T21:31:30Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"No, please use the standardized numbers when they apply.\n\nOn May 22, 2021 2:22:42 PM PDT, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n>Junio C Hamano wrote:\n>> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n>> \n>> > On 5/21/21 9:53 AM, Alex Henrie wrote:\n>> >> On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n>> >> <felipe.contreras@gmail.com> wrote:\n>> >>>\n>> >>> It's good to not include many initial codes, but I would start\n>with at\n>> >>> least three:\n>> >>>\n>> >>>    OK = 0,\n>> >>>    UNKNOWN = 1,\n>> >>>    NORMAL = 2,\n>> >> If you go that route, could you please pick a word other than\n>> >> \"normal\"\n>> >> to describe errors that are not entirely unexpected? I'm worried\n>that\n>> >> someone will see \"normal\" and use it instead of \"OK\" to indicate\n>> >> success.\n>> >> \n>> >\n>> > <sysexits.h>\n>> \n>> Is the value assignment standardized across systems?\n>\n>I think his intention was to suggest to use that list as inspiration...\n>As in have USAGE, NOINPUT, UNAVAILABE, etc.\n>\n>I would prefer to start with something easy... UNKNOWN = 1, USUAL = 2.\n\n-- \nSent from my Android device with K-9 Mail. Please excuse my brevity.\n"},{"id":"425352","messageId":"60a97d51b9a7_8572320883@natae.notmuch","threadId":"55739","inReplyTo":"3C6468D1-3E14-4600-BC8E-86CCCB84E74C@zytor.com","subject":"Re: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-22T21:53:21Z","receivedAt":"2021-05-22T21:53:25Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"H. Peter Anvin wrote:\n> No, please use the standardized numbers when they apply.\n\nI wasn't talking about the numbers, but the names.\n\nDo you see something wrong with USAGE = EX__USAGE?\n\n-- \nFelipe Contreras\n"},{"id":"425353","messageId":"25358AF3-E6DF-49FE-9F41-2D81EE794227@zytor.com","threadId":"55739","inReplyTo":"60a97d51b9a7_8572320883@natae.notmuch","subject":"Re: RFC: error codes on exit","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2021-05-22T23:02:31Z","receivedAt":"2021-05-22T23:16:04Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Why avoid the standard symbols?\n\nOn May 22, 2021 2:53:21 PM PDT, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n>H. Peter Anvin wrote:\n>> No, please use the standardized numbers when they apply.\n>\n>I wasn't talking about the numbers, but the names.\n>\n>Do you see something wrong with USAGE = EX__USAGE?\n\n-- \nSent from my Android device with K-9 Mail. Please excuse my brevity.\n"},{"id":"425531","messageId":"CAMMLpeR5S3Ps4C2V4QuTxrCRB_iRsUKyCNOJ4G7Fy7jGe98ZbA@mail.gmail.com","threadId":"55739","inReplyTo":"60a97550287b3_857e9208b8@natae.notmuch","subject":"Re: RFC: error codes on exit","fromName":"Alex Henrie","fromEmail":"alexhenrie24@gmail.com","sentAt":"2021-05-25T17:24:35Z","receivedAt":"2021-05-25T17:24:55Z","isPatch":false,"sender":{"key":"alexhenrie24@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5951993?v=4"},"body":"On Sat, May 22, 2021 at 3:19 PM Felipe Contreras\n<felipe.contreras@gmail.com> wrote:\n>\n> Philip Oakley wrote:\n> > On 21/05/2021 17:53, Alex Henrie wrote:\n> > > On Wed, May 19, 2021 at 6:40 PM Felipe Contreras\n> > > <felipe.contreras@gmail.com> wrote:\n> > >> It's good to not include many initial codes, but I would start with at\n> > >> least three:\n> > >>\n> > >>   OK = 0,\n> > >>   UNKNOWN = 1,\n> > >>   NORMAL = 2,\n> > > If you go that route, could you please pick a word other than \"normal\"\n> > > to describe errors that are not entirely unexpected? I'm worried that\n> > > someone will see \"normal\" and use it instead of \"OK\" to indicate\n> > > success.\n> > >\n> > > -Alex\n> > Typical <== Normal\n> >\n> > Though abnormal and atypical often have different implications ;-)\n> > P.\n>\n> Or USUAL.\n\nThe words \"typical\" and \"usual\" have the same problem of making it\nsound like there was no error. I would suggest terms like \"user\nerror\", \"network error\", etc. instead.\n\n-Alex\n"},{"id":"425532","messageId":"60ad453d2153e_2901820821@natae.notmuch","threadId":"55739","inReplyTo":"CAMMLpeR5S3Ps4C2V4QuTxrCRB_iRsUKyCNOJ4G7Fy7jGe98ZbA@mail.gmail.com","subject":"Re: RFC: error codes on exit","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-25T18:43:09Z","receivedAt":"2021-05-25T18:43:14Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Alex Henrie wrote:\n> On Sat, May 22, 2021 at 3:19 PM Felipe Contreras\n> <felipe.contreras@gmail.com> wrote:\n\n> > > Though abnormal and atypical often have different implications ;-)\n> >\n> > Or USUAL.\n> \n> The words \"typical\" and \"usual\" have the same problem of making it\n> sound like there was no error.\n\nNot to me. \"Usual\" is typically an adjective, which makes me think: usual\nwhat? (something is missing). Although both can be used as nouns (\"the\nusual\", \"the new normal\") in this context \"usual\" is much less confusing\n(I don't recall hearing \"the command returned usual status\").\n\n-- \nFelipe Contreras\n"},{"id":"425553","messageId":"87sg29n9mt.fsf@evledraar.gmail.com","threadId":"55739","inReplyTo":"YKWggLGDhTOY+lcy@google.com","subject":"Re: RFC: error codes on exit","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-05-26T08:21:41Z","receivedAt":"2021-05-26T09:10:25Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 19 2021, Jonathan Nieder wrote:\n\n> Hi,\n>\n> (Danger, jrn is wading into error handling again...)\n>\n> At $DAYJOB we are setting up some alerting for some bot fleets and\n> developer workstations, using trace2 as the data source.  Having\n> trace2 has been great --- combined with gradual weekly rollouts of\n> \"next\", it helps us to understand quickly when a change is creating a\n> regression for users, which hopefully improves the quality of Git for\n> everyone.\n>\n> One kind of signal we haven't been able to make good use of is error\n> rates.  The problem is that a die() call can be an indication of\n>\n>  a. the user asked to do something that isn't sensible, and we kindly\n>     rebuked the user\n>\n>  b. we contacted a server, and the server was not happy with our\n>     request\n>\n>  c. the local Git repository is corrupt\n>\n>  d. we ran out of resources (e.g., disk space)\n>\n>  e. we encountered an internal error in handling the user's\n>     legitimate request\n> [...]\n> Further down the line I can imagine making use of git_error_code\n> elsewhere for e.g. some limited retries of the corresponding\n> transaction when we fail to lock a file.\n>\n> Thoughts?  Good idea?  Bad idea?\n\nHaving read the thread at large (and some of this is a more general\nresponse) a few points, not against or as a retort to this, just related\nthoughts, complimentary suggestions etc:\n\n 1. As shown in my f6d25d78789 (api docs: document that BUG() emits a\n    trace2 error event, 2021-04-13) all of BUG/die/error/warning just\n    emit \"error\" under trace2.\n\n    It seems to me a good place to start with this effort would be for\n    someone to split that up. It requires changing the trace2 schema,\n    but it can be done in some backwards compatible way. Perhaps event:\n    error, error_type: [bug,die,error,warning] ?\n\n 1.5. Split up error_errno() from error() for trace2 purposes? This gets\n      you partway to your \"d\".\n\n 2. Similarly we need to log the correct line numbers for\n    die/error/warning. They need to be a macro/function like BUG() /\n    BUG_fl().\n\n 3. You can then key error events/frequencies on the \"fmt\".\n\n 4. To the extent tha #3 isn't true on client machines due to i18n we\n    could change the API in a backwards-compatible way from\n    e.g. error(_(\"string\") to error(_N(\"string\")). We'd then always\n    transmit the C locale \"fmt\".\n\nBasically I wonder if a more granular approach with just better logging\nof information we have now (but lose in trace2) + maybe some split-up of\nthe current functions, e.g. having a user_error() distinct from\nrepository_error() or whatever wouldn't get us most/all of the way to\nthis.\n\n> Further down the line I can imagine making use of git_error_code\n> elsewhere for e.g. some limited retries of the corresponding\n> transaction when we fail to lock a file.\n\nMaybe, but that seems highly problem-dependant, and not e.g. something\nwhere we'd like to just do a blind retry in one of our own porcelain\ntools if a plumbing one failed with a \"had an issue, retries might work\"\ncode.\n"}]}