{"thread":{"id":"31459","subject":"approxidate parsing for bad time units","startedAt":"2012-09-06T16:24:02Z","lastAt":"2012-09-10T21:19:11Z","messageCount":6,"participants":["Jeffrey Middleton","Junio C Hamano","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"198458","messageId":"CAFE6XRFgQa10vTWXfxRG53W6K4U=VGqpK5sQwH7xp9GfKd=2Uw@mail.gmail.com","threadId":"31459","inReplyTo":null,"subject":"approxidate parsing for bad time units","fromName":"Jeffrey Middleton","fromEmail":"jefromi@gmail.com","sentAt":"2012-09-06T16:24:02Z","receivedAt":"2012-09-06T16:24:02Z","isPatch":false,"sender":{"key":"jefromi@gmail.com","avatar":null},"body":"In telling someone what date formats git accepts, and how to verify it\nunderstands, I noticed this weirdness:\n\n$ export TEST_DATE_NOW=`date -u +%s --date='September 10'`;\n./test-date approxidate now; for i in `seq 1 10`; do ./test-date\napproxidate \"$i frobbles ago\"; done\nnow -> 2012-09-10 00:00:00 +0000\n1 frobbles ago -> 2012-09-02 00:00:00 +0000\n2 frobbles ago -> 2012-09-03 00:00:00 +0000\n3 frobbles ago -> 2012-09-04 00:00:00 +0000\n4 frobbles ago -> 2012-09-05 00:00:00 +0000\n5 frobbles ago -> 2012-09-06 00:00:00 +0000\n6 frobbles ago -> 2012-09-07 00:00:00 +0000\n7 frobbles ago -> 2012-09-08 00:00:00 +0000\n8 frobbles ago -> 2012-09-09 00:00:00 +0000\n9 frobbles ago -> 2012-09-10 00:00:00 +0000\n10 frobbles ago -> 2012-09-11 00:00:00 +0000\n\nWhich gets more concerning once you realize the same thing happens no\nmatter what fake unit of time you use... including things like \"yaers\"\nand \"moths\". Perhaps approxidate could be a little stricter?\n\nThanks,\nJeffrey\n"},{"id":"198476","messageId":"7vehme3n49.fsf@alter.siamese.dyndns.org","threadId":"31459","inReplyTo":"CAFE6XRFgQa10vTWXfxRG53W6K4U=VGqpK5sQwH7xp9GfKd=2Uw@mail.gmail.com","subject":"Re: approxidate parsing for bad time units","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-09-06T20:36:06Z","receivedAt":"2012-09-06T20:36:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeffrey Middleton <jefromi@gmail.com> writes:\n\n> In telling someone what date formats git accepts, and how to verify it\n> understands, I noticed this weirdness:\n>\n> $ export TEST_DATE_NOW=`date -u +%s --date='September 10'`;\n> ./test-date approxidate now; for i in `seq 1 10`; do ./test-date\n> approxidate \"$i frobbles ago\"; done\n> now -> 2012-09-10 00:00:00 +0000\n> 1 frobbles ago -> 2012-09-02 00:00:00 +0000\n> ...\n> 10 frobbles ago -> 2012-09-11 00:00:00 +0000\n>\n> Which gets more concerning once you realize the same thing happens no\n> matter what fake unit of time you use... including things like \"yaers\"\n> and \"moths\". Perhaps approxidate could be a little stricter?\n\n\"Could be stricter\", perhaps.\n\nDo we care deeply?  I doubt it, and for a good reason.  The fuzzy\nparsing is primarily [*1*] for humans getting interactive results\nwho are expected to be able to notice when the fuzziness went far\noff.\n\nAs long as we have ways for scripts and humans to feed its input in\na more strict and unambiguous way [*2*], it does not hurt anybody if\nthe fuzzy parser ignored crufts that it does not understand.\n\n\n[Footnotes]\n\n*1* ... and of course some coding fun and easter egg values. Think\nof it as our own Eliza or Zork parser ;-).\n\n*2* And of course we do.\n"},{"id":"198478","messageId":"CAFE6XREG5-gwjzvyP9r_hfyY3bWSV2=Bjv9ZbXkejXQRoqYERA@mail.gmail.com","threadId":"31459","inReplyTo":"7vehme3n49.fsf@alter.siamese.dyndns.org","subject":"Re: approxidate parsing for bad time units","fromName":"Jeffrey Middleton","fromEmail":"jefromi@gmail.com","sentAt":"2012-09-06T21:01:30Z","receivedAt":"2012-09-06T21:01:30Z","isPatch":false,"sender":{"key":"jefromi@gmail.com","avatar":null},"body":"I'm generally very happy with the fuzzy parsing. It's a great feature\nthat is designed to and in general does save users a lot of time and\nthought. In this case I don't think it does. The problems are:\n(1) It's not ignoring things it can't understand, it's silently\ninterpreting them in a useless way. I'm pretty sure that \"n units ago\"\nis equivalent to \"the same time of day on the last day of the previous\nmonth, plus n days.\"\n(2) Though in some cases it's really obvious, in others it's quite\npossible not to notice, e.g. if `git rev-list --since=5.dyas.ago` is\nsilently the same as `git rev-list --since=4.days.ago`.\n\nSo I do think it's worth improving. (Yes, I know, send patches; I'll\nthink about it.)\n\n\nOn Thu, Sep 6, 2012 at 1:36 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Jeffrey Middleton <jefromi@gmail.com> writes:\n>\n>> In telling someone what date formats git accepts, and how to verify it\n>> understands, I noticed this weirdness:\n>>\n>> $ export TEST_DATE_NOW=`date -u +%s --date='September 10'`;\n>> ./test-date approxidate now; for i in `seq 1 10`; do ./test-date\n>> approxidate \"$i frobbles ago\"; done\n>> now -> 2012-09-10 00:00:00 +0000\n>> 1 frobbles ago -> 2012-09-02 00:00:00 +0000\n>> ...\n>> 10 frobbles ago -> 2012-09-11 00:00:00 +0000\n>>\n>> Which gets more concerning once you realize the same thing happens no\n>> matter what fake unit of time you use... including things like \"yaers\"\n>> and \"moths\". Perhaps approxidate could be a little stricter?\n>\n> \"Could be stricter\", perhaps.\n>\n> Do we care deeply?  I doubt it, and for a good reason.  The fuzzy\n> parsing is primarily [*1*] for humans getting interactive results\n> who are expected to be able to notice when the fuzziness went far\n> off.\n>\n> As long as we have ways for scripts and humans to feed its input in\n> a more strict and unambiguous way [*2*], it does not hurt anybody if\n> the fuzzy parser ignored crufts that it does not understand.\n>\n>\n> [Footnotes]\n>\n> *1* ... and of course some coding fun and easter egg values. Think\n> of it as our own Eliza or Zork parser ;-).\n>\n> *2* And of course we do.\n"},{"id":"198525","messageId":"20120907135452.GA1290@sigill.intra.peff.net","threadId":"31459","inReplyTo":"CAFE6XREG5-gwjzvyP9r_hfyY3bWSV2=Bjv9ZbXkejXQRoqYERA@mail.gmail.com","subject":"Re: approxidate parsing for bad time units","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-09-07T13:54:52Z","receivedAt":"2012-09-07T13:54:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 06, 2012 at 02:01:30PM -0700, Jeffrey Middleton wrote:\n\n> I'm generally very happy with the fuzzy parsing. It's a great feature\n> that is designed to and in general does save users a lot of time and\n> thought. In this case I don't think it does. The problems are:\n> (1) It's not ignoring things it can't understand, it's silently\n> interpreting them in a useless way.\n\nRight, but we would then need to come up with a list of things it _does_\nunderstand. So right now I can say \"6 June\" or \"6th of June\" or even \"6\nde June\", and it works because we just ignore the cruft in the middle.\n\nSo I think you'd need to either whitelist what everybody is typing, or\nblacklist some common typos (or convince people to be stricter in what\nthey type).\n\n> So I do think it's worth improving. (Yes, I know, send patches; I'll\n> think about it.)\n\nYou read my mind. :)\n\n-Peff\n"},{"id":"198728","messageId":"CAFE6XRHmX6TjGu7Jte_KW82nYX7ZUw6imO1ktbUcYpNbc6ZBsA@mail.gmail.com","threadId":"31459","inReplyTo":"20120907135452.GA1290@sigill.intra.peff.net","subject":"Re: approxidate parsing for bad time units","fromName":"Jeffrey Middleton","fromEmail":"jefromi@gmail.com","sentAt":"2012-09-10T21:07:02Z","receivedAt":"2012-09-10T21:07:02Z","isPatch":false,"sender":{"key":"jefromi@gmail.com","avatar":null},"body":"As you mentioned, parsing \"n ... [month]\", and even \"...n...\" (e.g.\n\"the 3rd\") as the nth day of a month is great, but in this case, I\nthink \"n ... ago\" is a pretty strong sign that that's not the intended\nbehavior.\n\nMy first thought was just to make it an error if the string ends in\n\"ago\" but the date is parsed as a day of the month. You don't actually\nhave to come up with any typos to blacklist, just keep the \"ago\" from\nbeing silently ignored. I suspect \"n units ago\" is by far the most\ncommon use of the approxidate parsing in the wild, since it's\ndocumented and has been popularized online. So throwing an error just\nin that case would save essentially everyone. I hadn't even realized\nit worked without \"ago\" until I looked at the code.\n\nIf that doesn't sound like a good plan, then yes, I agree, it'd be\ntricky to catch it in the general case without breaking things.\n(Levenshtein distance to the target strings instead of exact matching,\nI guess, so that it could say \"did you mean...\" like for misspelled\ncommands.)\n\nOn Fri, Sep 7, 2012 at 6:54 AM, Jeff King <peff@peff.net> wrote:\n>\n> On Thu, Sep 06, 2012 at 02:01:30PM -0700, Jeffrey Middleton wrote:\n>\n> > I'm generally very happy with the fuzzy parsing. It's a great feature\n> > that is designed to and in general does save users a lot of time and\n> > thought. In this case I don't think it does. The problems are:\n> > (1) It's not ignoring things it can't understand, it's silently\n> > interpreting them in a useless way.\n>\n> Right, but we would then need to come up with a list of things it _does_\n> understand. So right now I can say \"6 June\" or \"6th of June\" or even \"6\n> de June\", and it works because we just ignore the cruft in the middle.\n>\n> So I think you'd need to either whitelist what everybody is typing, or\n> blacklist some common typos (or convince people to be stricter in what\n> they type).\n>\n> > So I do think it's worth improving. (Yes, I know, send patches; I'll\n> > think about it.)\n>\n> You read my mind. :)\n>\n> -Peff\n"},{"id":"198736","messageId":"20120910211911.GA1537@sigill.intra.peff.net","threadId":"31459","inReplyTo":"CAFE6XRHmX6TjGu7Jte_KW82nYX7ZUw6imO1ktbUcYpNbc6ZBsA@mail.gmail.com","subject":"Re: approxidate parsing for bad time units","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-09-10T21:19:11Z","receivedAt":"2012-09-10T21:19:11Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 10, 2012 at 02:07:02PM -0700, Jeffrey Middleton wrote:\n\n> As you mentioned, parsing \"n ... [month]\", and even \"...n...\" (e.g.\n> \"the 3rd\") as the nth day of a month is great, but in this case, I\n> think \"n ... ago\" is a pretty strong sign that that's not the intended\n> behavior.\n\nYeah, agreed. We are really talking about two distinct cases:\n\n  1. An absolute date (\"the 3rd of June\", \"last tuesday\") whose exact\n     location may need to be inferred from the context of the current\n     date.\n\n  2. A relative unit difference from the current time (\"7 days ago\")\n\nHowever, I'm not sure that the word \"ago\" is always present when\nchoosing the latter. For example, you can say \"7 days\" and approxidate\nwill treat it like \"7 days ago\". Nor is it simply using a unit like\n\"days\". You can even say \"7 tuesdays\" to go backwards that many Tuesdays\n(e.g., the 24th of July from today).\n\nSo you can use \"ago\" as a sign that you are definitely in case (2), but\ncannot assume that its absence means you are in case (1).\n\nThat means we can catch \"3 dasy ago\" as nonsensical, but not \"3 dasy\",\nas the latter simply looks like \"the 3rd\" from approxidate's\nperspective. Still, something is better than nothing, and it means if\nyou are careful to always say \"ago\", you can catch some errors (of\ncourse, you might typo \"ago\"... :) ).\n\n> My first thought was just to make it an error if the string ends in\n> \"ago\" but the date is parsed as a day of the month. You don't actually\n> have to come up with any typos to blacklist, just keep the \"ago\" from\n> being silently ignored. I suspect \"n units ago\" is by far the most\n> common use of the approxidate parsing in the wild, since it's\n> documented and has been popularized online. So throwing an error just\n> in that case would save essentially everyone. I hadn't even realized\n> it worked without \"ago\" until I looked at the code.\n\nYeah, I think that would work, and would provide some safety. And it\nshouldn't be too hard to implement.\n\n> If that doesn't sound like a good plan, then yes, I agree, it'd be\n> tricky to catch it in the general case without breaking things.\n> (Levenshtein distance to the target strings instead of exact matching,\n> I guess, so that it could say \"did you mean...\" like for misspelled\n> commands.)\n\nGross. :)\n\n-Peff\n"}]}