{"thread":{"id":"44156","subject":"Changing the default for \"core.abbrev\"?","startedAt":"2016-09-26T01:39:41Z","lastAt":"2019-02-06T18:36:34Z","messageCount":111,"participants":["Linus Torvalds","Junio C Hamano","Jeff King","Matthieu Moy","Christian Couder","Jacob Keller","SZEDER Gábor","Lukas Fleischer","Johannes Sixt","Jakub Narębski","Kyle J. McKay","Mike Hommey","Ævar Arnfjörð Bjarmason"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"302549","messageId":"CA+55aFy0_pwtFOYS1Tmnxipw9ZkRNCQHmoYyegO00pjMiZQfbg@mail.gmail.com","threadId":"44156","inReplyTo":null,"subject":"Changing the default for \"core.abbrev\"?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-26T01:39:35Z","receivedAt":"2016-09-26T01:39:41Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"The default value for commit abbreviation (environment.c: 19) is seven:\n\n    int minimum_abbrev = 4, default_abbrev = 7;\n\nwhich back in the dark early days of git was fairly reasonable.\n\nIt's probably still a perfectly fine default for lots of projects,\nsince 7 hex digits is a few hundred million unique values, and you\nwon't really start to get very many collisions in that until you get\ncloser to a million objects.\n\nThe kernel, these days, is at roughly 5 million objects, and while the\nseven hex digits are still often enough for uniqueness (and git will\nalways add digits *until* it is unique), it's long been at the point\nwhere I tell people to do\n\n    git config --global core.abbrev 12\n\nbecause even though git will extend the seven hex digits until the\nobject name is unique, that only reflects the *current* situation in\nthe repository. With 5 million objects and a very healthy growth rate,\na 7-8 hex digit number that is unique today is not necessarily unique\na month or two from now, and then it gets annoying when a commit\nmessage has a short git ID that is no longer unique when you go back\nand try to figure out what went wrong in that commit.\n\nI can just keep reminding kernel maintainers and developers to update\ntheir git config, but maybe it would be a good idea to just admit that\nthe defaults picked in 2005 weren't necessarily the best ones\npossible, and those could be bumped up a bit?\n\nI think I mentioned this some time ago, and it's not a huge deal, but\nI thought I'd just mention it again because it came up again today for\nme..\n\nThanks,\n\n              Linus\n"},{"id":"302550","messageId":"xmqq37knwcf4.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFy0_pwtFOYS1Tmnxipw9ZkRNCQHmoYyegO00pjMiZQfbg@mail.gmail.com","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T03:46:39Z","receivedAt":"2016-09-26T03:46:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> I can just keep reminding kernel maintainers and developers to update\n> their git config, but maybe it would be a good idea to just admit that\n> the defaults picked in 2005 weren't necessarily the best ones\n> possible, and those could be bumped up a bit?\n>\n> I think I mentioned this some time ago, and it's not a huge deal, but\n> I thought I'd just mention it again because it came up again today for\n> me..\n\nI am not quite sure how good any new default would be, though.  Just\nlike any timeout is not long enough for somebody, growing projects\nwill eventually hit whatever abbreviation length they start with.\n\nEven if we bump it to 12 for everybody, majority of projects at\nGitHub would probably be just wasting 5 more hexdigits in addition\nto whatever they are already wasting.  The kernel folks will keep\nhaving the problem of having harder time looking up objects referred\nto by ancient commits no matter what the new default is anyway, and\nthen they will again regret we didn't bump it to 16 in year 2016 in\nseveral decades; by that time both of us are probably retired so it\nmay no longer be our problems, though ;-)\n\nI am not opposed to bump the default to 12 or whatever, but I\nsuspect any lengthening today may need to be accompanied by a tool\nsupport that finds the set of objects that are reachable from a\ncommit whose names begin with non-unique abbreviations that appear\nin the commit log message. Assuming that it is very hard to refer to\nfuture objects in the log message you write today, such a tool may\nfind a single object that used to be the unique instance of that\nabbreviation back then, and with reachability bitmap support, it may\nnot be too expensive to run.\n\n"},{"id":"302552","messageId":"20160926043442.3pz7ccawdcsn2kzb@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqq37knwcf4.fsf@gitster.mtv.corp.google.com","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T04:34:42Z","receivedAt":"2016-09-26T04:34:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Sep 25, 2016 at 08:46:39PM -0700, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > I can just keep reminding kernel maintainers and developers to update\n> > their git config, but maybe it would be a good idea to just admit that\n> > the defaults picked in 2005 weren't necessarily the best ones\n> > possible, and those could be bumped up a bit?\n> >\n> > I think I mentioned this some time ago, and it's not a huge deal, but\n> > I thought I'd just mention it again because it came up again today for\n> > me..\n> \n> I am not quite sure how good any new default would be, though.  Just\n> like any timeout is not long enough for somebody, growing projects\n> will eventually hit whatever abbreviation length they start with.\n\nI actually think \"12\" might be sane for a long time. That's 48 bits of\nsha1, so we'd expect a 50% change of a _single_ collision at 2^24, or 16\nmillion.  The biggest repository I know about (in number of objects) is\nthe one holding all of the objects for all of the forks of\ntorvalds/linux on GitHub. It's at about 15 million objects.\n\nWhich _seems_ close, but remember that's the size where we expect to see\na single collision. They don't become common until much later (I didn't\ncompute an exact number, but Linus's 16x sounds about right). I know\nthat the growth of the kernel isn't really linear, but I think the need\nto bump to \"13\" might not just be decades, but possibly a century or\nmore.\n\nSo 12 seems reasonable, and the only downside for it (or for \"13\", for\nthat matter) is a few extra bytes. I dunno, maybe people will really\nhate that, but I have a feeling these are mostly cut-and-pasted anyway.\n\n> I am not opposed to bump the default to 12 or whatever, but I\n> suspect any lengthening today may need to be accompanied by a tool\n> support that finds the set of objects that are reachable from a\n> commit whose names begin with non-unique abbreviations that appear\n> in the commit log message. Assuming that it is very hard to refer to\n> future objects in the log message you write today, such a tool may\n> find a single object that used to be the unique instance of that\n> abbreviation back then, and with reachability bitmap support, it may\n> not be too expensive to run.\n\nI had a similar thought, but I think it's not just reachability. You\nmight refer to a short sha1 on an alternate branch that isn't reachable\nfrom you. Or you may even use the short sha1 in an email message or a\nbug tracker. So I think the extra context you want is probably a\ntimestamp: at time t, what was a reasonable guess for this sha1?\n\nThat's easy to answer for commits and tags (cull the ones that are too\nnew), but harder for blobs and trees (you'd want to know the earliest\ncommit which contains them).\n\nAn easier (but less automatic) tool would be to improve our error\nmessage for the ambiguous case, and actually report details of the\ncandidates. I'm working up a patch now.\n\n-Peff\n"},{"id":"302553","messageId":"CAPc5daV1YJaEqH5eZCej3nkg8htHVDWQu0V0uoC4gVmPYpDL9Q@mail.gmail.com","threadId":"44156","inReplyTo":"20160926043442.3pz7ccawdcsn2kzb@sigill.intra.peff.net","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T04:45:18Z","receivedAt":"2016-09-26T04:45:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Sun, Sep 25, 2016 at 9:34 PM, Jeff King <peff@peff.net> wrote:\n>\n> An easier (but less automatic) tool would be to improve our error\n> message for the ambiguous case, and actually report details of the\n> candidates. I'm working up a patch now.\n\nThat sounds like a fun little lunch-break project. Thanks.\n"},{"id":"302555","messageId":"vpq37kntbjj.fsf@anie.imag.fr","threadId":"44156","inReplyTo":"xmqq37knwcf4.fsf@gitster.mtv.corp.google.com","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-09-26T06:33:52Z","receivedAt":"2016-09-26T06:34:08Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I am not opposed to bump the default to 12 or whatever, but I\n> suspect any lengthening today may need to be accompanied by a tool\n> support that finds the set of objects that are reachable from a\n> commit whose names begin with non-unique abbreviations that appear\n> in the commit log message.\n\nSomething much simpler would be to set core.abbrev at clone time,\ndepending on the size of the project just cloned. So, when cloning a\nhello-world, we'd keep the 7 but when cloning a big project we'd get a\nlarger value.\n\nThis doesn't cover the case of someone growing his own project without\ncloning, and isn't as clever as actually looking for colision, but it\nwould probably provide a sane default in 99% cases, and wouldn't be\nworse than hardcoding 7 in the 1% remaining cases.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"302557","messageId":"CAP8UFD3UDm_CT2_4YYeQ+-gdC2kJPp7v0k5hZOB4oALUzR-JPg@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFy0_pwtFOYS1Tmnxipw9ZkRNCQHmoYyegO00pjMiZQfbg@mail.gmail.com","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2016-09-26T07:13:32Z","receivedAt":"2016-09-26T07:16:22Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Mon, Sep 26, 2016 at 3:39 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> The kernel, these days, is at roughly 5 million objects, and while the\n> seven hex digits are still often enough for uniqueness (and git will\n> always add digits *until* it is unique), it's long been at the point\n> where I tell people to do\n>\n>     git config --global core.abbrev 12\n>\n> because even though git will extend the seven hex digits until the\n> object name is unique, that only reflects the *current* situation in\n> the repository. With 5 million objects and a very healthy growth rate,\n> a 7-8 hex digit number that is unique today is not necessarily unique\n> a month or two from now, and then it gets annoying when a commit\n> message has a short git ID that is no longer unique when you go back\n> and try to figure out what went wrong in that commit.\n\nAEvar sent a patch recently\n(https://public-inbox.org/git/20160921114428.28664-3-avarab@gmail.com/)\nto have gitweb link to \"git describe\"'d commits in log messages, and\nthis makes me wonder if it woudn't be better for the kernel to also\nuse the output of a command like `git describe --verylong` or `git\ndescribe --long=12` instead of a regular git ID in commit messages.\n"},{"id":"302560","messageId":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","threadId":"44156","inReplyTo":"CAPc5daV1YJaEqH5eZCej3nkg8htHVDWQu0V0uoC4gVmPYpDL9Q@mail.gmail.com","subject":"[PATCH 0/10] helping people resolve ambiguous sha1s","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T11:57:20Z","receivedAt":"2016-09-26T11:57:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Sep 25, 2016 at 09:45:18PM -0700, Junio C Hamano wrote:\n\n> On Sun, Sep 25, 2016 at 9:34 PM, Jeff King <peff@peff.net> wrote:\n> >\n> > An easier (but less automatic) tool would be to improve our error\n> > message for the ambiguous case, and actually report details of the\n> > candidates. I'm working up a patch now.\n> \n> That sounds like a fun little lunch-break project. Thanks.\n\nThat's what I thought, but it turned out to be quite involved. :)\n\nI started by trying to teach get_short_sha1() to remember all of the\ncandidates it sees, but it turns out to be surprisingly complicated. I\ndid have something working, but I scrapped it in favor of just looking\nat the object database again. It's the error code path, so it's OK to be\nslower (especially if it keeps the non-error code path much simpler).\n\nBut then being the diligent programmer that I am, I added a tests.\nAnd that failed because of an unrelated bug. Fixing that revealed\nanother bug. And so on.\n\nThe good news is that I think I've finally cleared up all of the\nlong-standing bugs where git will print the same error message twice.\nThose have been annoying me for yours (and apparently others[1]).\n\nPatches 2-4 and 9 are all bugfixes. Patch 10 is the interesting part.\nThe rest are just cleanups and refactoring.\n\n  [01/10]: get_sha1: detect buggy calls with multiple disambiguators\n  [02/10]: get_sha1: avoid repeating ourselves via ONLY_TO_DIE\n  [03/10]: get_sha1: propagate flags to child functions\n  [04/10]: get_short_sha1: peel tags when looking for treeish\n  [05/10]: get_short_sha1: refactor init of disambiguation code\n  [06/10]: get_short_sha1: NUL-terminate hex prefix\n  [07/10]: get_short_sha1: mark ambiguity error for translation\n  [08/10]: sha1_array: let callbacks interrupt iteration\n  [09/10]: for_each_abbrev: drop duplicate objects\n  [10/10]: get_short_sha1: list ambiguous objects on error\n\nOf course this is all totally orthogonal to Linus's original question. I\nhope it will make things more pleasant when somebody does end up having\nto look up a too-short sha1, but it's probably still a good idea to\nbump the default.\n\n-Peff\n\n[1] http://public-inbox.org/git/504B91B7.1000406@avtalion.name/\n"},{"id":"302561","messageId":"20160926115901.txmbr4e6xzwyfpmo@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 01/10] get_sha1: detect buggy calls with multiple disambiguators","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T11:59:01Z","receivedAt":"2016-09-26T11:59:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"The get_sha1() family of functions takes a flags field, but\nsome of the flags are mutually exclusive. In particular, we\ncan only handle one disambiguating function, and the flags\nquietly override each other. Let's instead detect these as\nprogramming bugs.\n\nTechnically some of the flags are supersets of the others,\nso treating COMMITTISH|TREEISH as just COMMITTISH is not\nwrong, but it's a good sign the caller is confused. And\ncertainly asking for BLOB|TREE does not work.\n\nWe can do the check easily with some bit-twiddling, and as a\nbonus, the bit-mask of disambiguators will come in handy in\na future patch.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n cache.h     | 5 +++++\n sha1_name.c | 9 +++++++++\n 2 files changed, 14 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex d0494c8..7bd78ca 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1203,6 +1203,11 @@ struct object_context {\n #define GET_SHA1_FOLLOW_SYMLINKS 0100\n #define GET_SHA1_ONLY_TO_DIE    04000\n \n+#define GET_SHA1_DISAMBIGUATORS \\\n+\t(GET_SHA1_COMMIT | GET_SHA1_COMMITTISH | \\\n+\tGET_SHA1_TREE | GET_SHA1_TREEISH | \\\n+\tGET_SHA1_BLOB)\n+\n extern int get_sha1(const char *str, unsigned char *sha1);\n extern int get_sha1_commit(const char *str, unsigned char *sha1);\n extern int get_sha1_committish(const char *str, unsigned char *sha1);\ndiff --git a/sha1_name.c b/sha1_name.c\nindex faf873c..f9812ff 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -310,6 +310,11 @@ static int prepare_prefixes(const char *name, int len,\n \treturn 0;\n }\n \n+static int multiple_bits_set(unsigned flags)\n+{\n+\treturn !!(flags & (flags - 1));\n+}\n+\n static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\t\t  unsigned flags)\n {\n@@ -327,6 +332,10 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \tprepare_alt_odb();\n \n \tmemset(&ds, 0, sizeof(ds));\n+\n+\tif (multiple_bits_set(flags & GET_SHA1_DISAMBIGUATORS))\n+\t\tdie(\"BUG: multiple get_short_sha1 disambiguator flags\");\n+\n \tif (flags & GET_SHA1_COMMIT)\n \t\tds.fn = disambiguate_commit_only;\n \telse if (flags & GET_SHA1_COMMITTISH)\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302562","messageId":"20160926115915.7lo6tr2mvs5irza2@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 02/10] get_sha1: avoid repeating ourselves via ONLY_TO_DIE","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T11:59:15Z","receivedAt":"2016-09-26T11:59:34Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"When the revision code cannot parse an argument like\n\"HEAD:foo\", it will call maybe_die_on_misspelt_object_name(),\nwhich re-runs get_sha1() with an extra ONLY_TO_DIE flag. We\nthen spend more effort to generate a better error message.\n\nUnfortunately, a side effect is that our second call may\nrepeat the same error messages from the original get_sha1()\ncall. You can see this with:\n\n  $ git show 0017\n  error: short SHA1 0017 is ambiguous.\n  error: short SHA1 0017 is ambiguous.\n  fatal: ambiguous argument '0017': unknown revision or path not in the working tree.\n  Use '--' to separate paths from revisions, like this:\n  'git <command> [<revision>...] -- [<file>...]'\n\nwhere the second \"error:\" line comes from the ONLY_TO_DIE\ncall.\n\nTo fix this, we can make ONLY_TO_DIE imply QUIETLY. This is\na little odd, because the whole point of ONLY_TO_DIE is to\noutput error messages. But what we want to do is tell the\nrest of the get_sha1() code (particularly get_sha1_1()) that\nthe _regular_ messages should be quiet, but the only-to-die\nones should not.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c                         | 3 +++\n t/t1512-rev-parse-disambiguation.sh | 6 ++++++\n 2 files changed, 9 insertions(+)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex f9812ff..fe05ba0 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -1391,6 +1391,9 @@ static int get_sha1_with_context_1(const char *name,\n \tconst char *cp;\n \tint only_to_die = flags & GET_SHA1_ONLY_TO_DIE;\n \n+\tif (only_to_die)\n+\t\tflags |= GET_SHA1_QUIETLY;\n+\n \tmemset(oc, 0, sizeof(*oc));\n \toc->mode = S_IFINVALID;\n \tret = get_sha1_1(name, namelen, sha1, flags);\ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex e221167..16f9709 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -291,4 +291,10 @@ test_expect_success 'ambiguous short sha1 ref' '\n \tgrep \"refname.*${REF}.*ambiguous\" err\n '\n \n+test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated' '\n+\ttest_must_fail git rev-parse 00000 2>stderr &&\n+\tgrep \"is ambiguous\" stderr >errors &&\n+\ttest_line_count = 1 errors\n+'\n+\n test_done\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302563","messageId":"20160926115940.yyedwgz4duxerm4i@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 03/10] get_sha1: propagate flags to child functions","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T11:59:41Z","receivedAt":"2016-09-26T11:59:47Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"The get_sha1() function is actually implementation by many\nsub-functions, but we do not always pass our flags around to\nall of those functions. As a result, we may forget that our\ncaller asked us to resolve with GET_SHA1_QUIETLY and output\nmessages. The two triggerable cases are:\n\n  1. Resolving treeish:path will resolve the \"treeish\"\n     portion using GET_SHA1_TREEISH, dropping all other\n     flags.\n\n  2. The peel_onion() function did not take flags at all\n     but recurses to get_sha1_1(), which does.\n\nThe solution for both is to bitwise-OR their new flags with\nthe existing ones (after dropping any mutually exclusive\ndisambiguation flags).\n\nThis bug can trigger with \"git rev-parse --quiet\", which\nasks for quiet resolution. But it can also happen in a more\nvanilla code path when we do a follow-up ONLY_TO_DIE\ninvocation of get_sha1(), and that's what the tests check.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c                         | 18 ++++++++++++------\n t/t1512-rev-parse-disambiguation.sh | 14 +++++++++++++-\n 2 files changed, 25 insertions(+), 7 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex fe05ba0..38e51d9 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -686,12 +686,12 @@ struct object *peel_to_type(const char *name, int namelen,\n \t}\n }\n \n-static int peel_onion(const char *name, int len, unsigned char *sha1)\n+static int peel_onion(const char *name, int len, unsigned char *sha1,\n+\t\t      unsigned lookup_flags)\n {\n \tunsigned char outer[20];\n \tconst char *sp;\n \tunsigned int expected_type = 0;\n-\tunsigned lookup_flags = 0;\n \tstruct object *o;\n \n \t/*\n@@ -731,10 +731,11 @@ static int peel_onion(const char *name, int len, unsigned char *sha1)\n \telse\n \t\treturn -1;\n \n+\tlookup_flags &= ~GET_SHA1_DISAMBIGUATORS;\n \tif (expected_type == OBJ_COMMIT)\n-\t\tlookup_flags = GET_SHA1_COMMITTISH;\n+\t\tlookup_flags |= GET_SHA1_COMMITTISH;\n \telse if (expected_type == OBJ_TREE)\n-\t\tlookup_flags = GET_SHA1_TREEISH;\n+\t\tlookup_flags |= GET_SHA1_TREEISH;\n \n \tif (get_sha1_1(name, sp - name - 2, outer, lookup_flags))\n \t\treturn -1;\n@@ -835,7 +836,7 @@ static int get_sha1_1(const char *name, int len, unsigned char *sha1, unsigned l\n \t\treturn get_nth_ancestor(name, len1, sha1, num);\n \t}\n \n-\tret = peel_onion(name, len, sha1);\n+\tret = peel_onion(name, len, sha1, lookup_flags);\n \tif (!ret)\n \t\treturn 0;\n \n@@ -1470,7 +1471,12 @@ static int get_sha1_with_context_1(const char *name,\n \tif (*cp == ':') {\n \t\tunsigned char tree_sha1[20];\n \t\tint len = cp - name;\n-\t\tif (!get_sha1_1(name, len, tree_sha1, GET_SHA1_TREEISH)) {\n+\t\tunsigned sub_flags = flags;\n+\n+\t\tsub_flags &= ~GET_SHA1_DISAMBIGUATORS;\n+\t\tsub_flags |= GET_SHA1_TREEISH;\n+\n+\t\tif (!get_sha1_1(name, len, tree_sha1, sub_flags)) {\n \t\t\tconst char *filename = cp+1;\n \t\t\tchar *new_filename = NULL;\n \ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex 16f9709..30e0b80 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -291,10 +291,22 @@ test_expect_success 'ambiguous short sha1 ref' '\n \tgrep \"refname.*${REF}.*ambiguous\" err\n '\n \n-test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated' '\n+test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated (raw)' '\n \ttest_must_fail git rev-parse 00000 2>stderr &&\n \tgrep \"is ambiguous\" stderr >errors &&\n \ttest_line_count = 1 errors\n '\n \n+test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated (treeish)' '\n+\ttest_must_fail git rev-parse 00000:foo 2>stderr &&\n+\tgrep \"is ambiguous\" stderr >errors &&\n+\ttest_line_count = 1 errors\n+'\n+\n+test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated (peel)' '\n+\ttest_must_fail git rev-parse 00000^{commit} 2>stderr &&\n+\tgrep \"is ambiguous\" stderr >errors &&\n+\ttest_line_count = 1 errors\n+'\n+\n test_done\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302564","messageId":"20160926115947.hksmtkqp3i4tfftx@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 04/10] get_short_sha1: peel tags when looking for treeish","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T11:59:48Z","receivedAt":"2016-09-26T11:59:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"The treeish disambiguation function tries to peel tags, but\nit does so by calling:\n\n  deref_tag(lookup_object(sha1), ...);\n\nThis will only work if we have previously looked at the tag\nand created a \"struct tag\" for it. Since parsing revision\narguments typically happens before anything else, this is\nusually not the case, and we would fail to peel the tag (we\nare lucky that deref_tag() gracefully handles the NULL and\ndoes not segfault).\n\nInstead, we can use parse_object(). Note that this is the\nsame fix done by 94d75d1 (get_short_sha1(): correctly\ndisambiguate type-limited abbreviation, 2013-07-01), but\nthat commit fixed only the committish disambiguator, and\nleft the bug in the treeish one.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c                         | 2 +-\n t/t1512-rev-parse-disambiguation.sh | 7 +++++++\n 2 files changed, 8 insertions(+), 1 deletion(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 38e51d9..432a308 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -269,7 +269,7 @@ static int disambiguate_treeish_only(const unsigned char *sha1, void *cb_data_un\n \t\treturn 0;\n \n \t/* We need to do this the hard way... */\n-\tobj = deref_tag(lookup_object(sha1), NULL, 0);\n+\tobj = deref_tag(parse_object(sha1), NULL, 0);\n \tif (obj && (obj->type == OBJ_TREE || obj->type == OBJ_COMMIT))\n \t\treturn 1;\n \treturn 0;\ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex 30e0b80..dfd3567 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -264,6 +264,13 @@ test_expect_success 'ambiguous commit-ish' '\n \ttest_must_fail git log 000000000...\n '\n \n+# There are three objects with this prefix: a blob, a tree, and a tag. We know\n+# the blob will not pass as a treeish, but the tree and tag should (and thus\n+# cause an error).\n+test_expect_success 'ambiguous tags peel to treeish' '\n+\ttest_must_fail git rev-parse 0000000000f^{tree}\n+'\n+\n test_expect_success 'rev-parse --disambiguate' '\n \t# The test creates 16 objects that share the prefix and two\n \t# commits created by commit-tree in earlier tests share a\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302565","messageId":"20160926120003.oixhtotisw5xnvh4@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 05/10] get_short_sha1: refactor init of disambiguation code","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:04Z","receivedAt":"2016-09-26T12:00:14Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"The disambiguation machinery has two callers: get_short_sha1\nand for_each_abbrev. Both need to repeat much of the same\nsetup: declaring buffers, sanity-checking lengths, preparing\nthe prefixes, etc.  Let's pull that into a single init\nfunction so we can avoid repeating ourselves.\n\nPulling the buffers into the \"struct disambiguate_state\"\nisn't strictly necessary, but it does make things simpler\nfor the callers, who no longer have to worry about sizing\nthem correctly (i.e., it's an implicit requirement that\nthe caller provide 20- and 40-byte buffers).\n\nAnd while we're touching this code, we can convert any\nmagic-number sizes to the more modern GIT_SHA1_* constants.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c | 79 +++++++++++++++++++++++++++----------------------------------\n 1 file changed, 35 insertions(+), 44 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 432a308..79eb1ee 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -13,9 +13,13 @@ static int get_sha1_oneline(const char *, unsigned char *, struct commit_list *)\n typedef int (*disambiguate_hint_fn)(const unsigned char *, void *);\n \n struct disambiguate_state {\n+\tint len; /* length of prefix in hex chars */\n+\tchar hex_pfx[GIT_SHA1_HEXSZ];\n+\tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n+\n \tdisambiguate_hint_fn fn;\n \tvoid *cb_data;\n-\tunsigned char candidate[20];\n+\tunsigned char candidate[GIT_SHA1_RAWSZ];\n \tunsigned candidate_exists:1;\n \tunsigned candidate_checked:1;\n \tunsigned candidate_ok:1;\n@@ -72,10 +76,10 @@ static void update_candidates(struct disambiguate_state *ds, const unsigned char\n \t/* otherwise, current can be discarded and candidate is still good */\n }\n \n-static void find_short_object_filename(int len, const char *hex_pfx, struct disambiguate_state *ds)\n+static void find_short_object_filename(struct disambiguate_state *ds)\n {\n \tstruct alternate_object_database *alt;\n-\tchar hex[40];\n+\tchar hex[GIT_SHA1_HEXSZ];\n \tstatic struct alternate_object_database *fakeent;\n \n \tif (!fakeent) {\n@@ -95,7 +99,7 @@ static void find_short_object_filename(int len, const char *hex_pfx, struct disa\n \t}\n \tfakeent->next = alt_odb_list;\n \n-\txsnprintf(hex, sizeof(hex), \"%.2s\", hex_pfx);\n+\txsnprintf(hex, sizeof(hex), \"%.2s\", ds->hex_pfx);\n \tfor (alt = fakeent; alt && !ds->ambiguous; alt = alt->next) {\n \t\tstruct dirent *de;\n \t\tDIR *dir;\n@@ -103,7 +107,7 @@ static void find_short_object_filename(int len, const char *hex_pfx, struct disa\n \t\t * every alt_odb struct has 42 extra bytes after the base\n \t\t * for exactly this purpose\n \t\t */\n-\t\txsnprintf(alt->name, 42, \"%.2s/\", hex_pfx);\n+\t\txsnprintf(alt->name, 42, \"%.2s/\", ds->hex_pfx);\n \t\tdir = opendir(alt->base);\n \t\tif (!dir)\n \t\t\tcontinue;\n@@ -113,7 +117,7 @@ static void find_short_object_filename(int len, const char *hex_pfx, struct disa\n \n \t\t\tif (strlen(de->d_name) != 38)\n \t\t\t\tcontinue;\n-\t\t\tif (memcmp(de->d_name, hex_pfx + 2, len - 2))\n+\t\t\tif (memcmp(de->d_name, ds->hex_pfx + 2, ds->len - 2))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(hex + 2, de->d_name, 38);\n \t\t\tif (!get_sha1_hex(hex, sha1))\n@@ -138,9 +142,7 @@ static int match_sha(unsigned len, const unsigned char *a, const unsigned char *\n \treturn 1;\n }\n \n-static void unique_in_pack(int len,\n-\t\t\t  const unsigned char *bin_pfx,\n-\t\t\t   struct packed_git *p,\n+static void unique_in_pack(struct packed_git *p,\n \t\t\t   struct disambiguate_state *ds)\n {\n \tuint32_t num, last, i, first = 0;\n@@ -155,7 +157,7 @@ static void unique_in_pack(int len,\n \t\tint cmp;\n \n \t\tcurrent = nth_packed_object_sha1(p, mid);\n-\t\tcmp = hashcmp(bin_pfx, current);\n+\t\tcmp = hashcmp(ds->bin_pfx, current);\n \t\tif (!cmp) {\n \t\t\tfirst = mid;\n \t\t\tbreak;\n@@ -174,20 +176,19 @@ static void unique_in_pack(int len,\n \t */\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tcurrent = nth_packed_object_sha1(p, i);\n-\t\tif (!match_sha(len, bin_pfx, current))\n+\t\tif (!match_sha(ds->len, ds->bin_pfx, current))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, current);\n \t}\n }\n \n-static void find_short_packed_object(int len, const unsigned char *bin_pfx,\n-\t\t\t\t     struct disambiguate_state *ds)\n+static void find_short_packed_object(struct disambiguate_state *ds)\n {\n \tstruct packed_git *p;\n \n \tprepare_packed_git();\n \tfor (p = packed_git; p && !ds->ambiguous; p = p->next)\n-\t\tunique_in_pack(len, bin_pfx, p, ds);\n+\t\tunique_in_pack(p, ds);\n }\n \n #define SHORT_NAME_NOT_FOUND (-1)\n@@ -281,14 +282,17 @@ static int disambiguate_blob_only(const unsigned char *sha1, void *cb_data_unuse\n \treturn kind == OBJ_BLOB;\n }\n \n-static int prepare_prefixes(const char *name, int len,\n-\t\t\t    unsigned char *bin_pfx,\n-\t\t\t    char *hex_pfx)\n+static int init_object_disambiguation(const char *name, int len,\n+\t\t\t\t      struct disambiguate_state *ds)\n {\n \tint i;\n \n-\thashclr(bin_pfx);\n-\tmemset(hex_pfx, 'x', 40);\n+\tif (len < MINIMUM_ABBREV || len > GIT_SHA1_HEXSZ)\n+\t\treturn -1;\n+\n+\tmemset(ds, 0, sizeof(*ds));\n+\tmemset(ds->hex_pfx, 'x', GIT_SHA1_HEXSZ);\n+\n \tfor (i = 0; i < len ;i++) {\n \t\tunsigned char c = name[i];\n \t\tunsigned char val;\n@@ -302,11 +306,14 @@ static int prepare_prefixes(const char *name, int len,\n \t\t}\n \t\telse\n \t\t\treturn -1;\n-\t\thex_pfx[i] = c;\n+\t\tds->hex_pfx[i] = c;\n \t\tif (!(i & 1))\n \t\t\tval <<= 4;\n-\t\tbin_pfx[i >> 1] |= val;\n+\t\tds->bin_pfx[i >> 1] |= val;\n \t}\n+\n+\tds->len = len;\n+\tprepare_alt_odb();\n \treturn 0;\n }\n \n@@ -319,20 +326,12 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\t\t  unsigned flags)\n {\n \tint status;\n-\tchar hex_pfx[40];\n-\tunsigned char bin_pfx[20];\n \tstruct disambiguate_state ds;\n \tint quietly = !!(flags & GET_SHA1_QUIETLY);\n \n-\tif (len < MINIMUM_ABBREV || len > 40)\n-\t\treturn -1;\n-\tif (prepare_prefixes(name, len, bin_pfx, hex_pfx) < 0)\n+\tif (init_object_disambiguation(name, len, &ds) < 0)\n \t\treturn -1;\n \n-\tprepare_alt_odb();\n-\n-\tmemset(&ds, 0, sizeof(ds));\n-\n \tif (multiple_bits_set(flags & GET_SHA1_DISAMBIGUATORS))\n \t\tdie(\"BUG: multiple get_short_sha1 disambiguator flags\");\n \n@@ -347,36 +346,28 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \telse if (flags & GET_SHA1_BLOB)\n \t\tds.fn = disambiguate_blob_only;\n \n-\tfind_short_object_filename(len, hex_pfx, &ds);\n-\tfind_short_packed_object(len, bin_pfx, &ds);\n+\tfind_short_object_filename(&ds);\n+\tfind_short_packed_object(&ds);\n \tstatus = finish_object_disambiguation(&ds, sha1);\n \n \tif (!quietly && (status == SHORT_NAME_AMBIGUOUS))\n-\t\treturn error(\"short SHA1 %.*s is ambiguous.\", len, hex_pfx);\n+\t\treturn error(\"short SHA1 %.*s is ambiguous.\", ds.len, ds.hex_pfx);\n \treturn status;\n }\n \n int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n {\n-\tchar hex_pfx[40];\n-\tunsigned char bin_pfx[20];\n \tstruct disambiguate_state ds;\n-\tint len = strlen(prefix);\n \n-\tif (len < MINIMUM_ABBREV || len > 40)\n+\tif (init_object_disambiguation(prefix, strlen(prefix), &ds) < 0)\n \t\treturn -1;\n-\tif (prepare_prefixes(prefix, len, bin_pfx, hex_pfx) < 0)\n-\t\treturn -1;\n-\n-\tprepare_alt_odb();\n \n-\tmemset(&ds, 0, sizeof(ds));\n \tds.always_call_fn = 1;\n \tds.cb_data = cb_data;\n \tds.fn = fn;\n \n-\tfind_short_object_filename(len, hex_pfx, &ds);\n-\tfind_short_packed_object(len, bin_pfx, &ds);\n+\tfind_short_object_filename(&ds);\n+\tfind_short_packed_object(&ds);\n \treturn ds.ambiguous;\n }\n \n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302566","messageId":"20160926120007.eswpfrzs2ed66d2o@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 06/10] get_short_sha1: NUL-terminate hex prefix","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:07Z","receivedAt":"2016-09-26T12:00:22Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We store the hex prefix in a 40-byte buffer with the prefix\nitself followed by 40-minus-len \"x\" characters. These x's\nserve no purpose, and the lack of NUL termination makes the\nprefix string annoying to use. Let's just terminate it.\n\nNote that this is in contrast to the binary prefix, which\n_must_ be zero-padded, because we look at the whole thing\nduring a binary search to find the first potential match in\neach pack index. The loose-object hex search cannot use the\nsame trick because it has to do a linear walk through the\nunsorted results of readdir() (and even if it could, you'd\nwant zeroes instead of x's).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 79eb1ee..549ef3f 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -14,7 +14,7 @@ typedef int (*disambiguate_hint_fn)(const unsigned char *, void *);\n \n struct disambiguate_state {\n \tint len; /* length of prefix in hex chars */\n-\tchar hex_pfx[GIT_SHA1_HEXSZ];\n+\tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n \n \tdisambiguate_hint_fn fn;\n@@ -291,7 +291,6 @@ static int init_object_disambiguation(const char *name, int len,\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n-\tmemset(ds->hex_pfx, 'x', GIT_SHA1_HEXSZ);\n \n \tfor (i = 0; i < len ;i++) {\n \t\tunsigned char c = name[i];\n@@ -313,6 +312,7 @@ static int init_object_disambiguation(const char *name, int len,\n \t}\n \n \tds->len = len;\n+\tds->hex_pfx[len] = '\\0';\n \tprepare_alt_odb();\n \treturn 0;\n }\n@@ -351,7 +351,7 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \tstatus = finish_object_disambiguation(&ds, sha1);\n \n \tif (!quietly && (status == SHORT_NAME_AMBIGUOUS))\n-\t\treturn error(\"short SHA1 %.*s is ambiguous.\", ds.len, ds.hex_pfx);\n+\t\treturn error(\"short SHA1 %s is ambiguous.\", ds.hex_pfx);\n \treturn status;\n }\n \n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302567","messageId":"20160926120014.6gvp22awmrkjgmoc@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 07/10] get_short_sha1: mark ambiguity error for translation","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:14Z","receivedAt":"2016-09-26T12:00:26Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This is a human-readable message, and there's no reason it\nshould not be translated. While we're at it, let's drop the\nperiod from the end, which is not our usual style.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 549ef3f..d4c7e26 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -351,7 +351,7 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \tstatus = finish_object_disambiguation(&ds, sha1);\n \n \tif (!quietly && (status == SHORT_NAME_AMBIGUOUS))\n-\t\treturn error(\"short SHA1 %s is ambiguous.\", ds.hex_pfx);\n+\t\treturn error(_(\"short SHA1 %s is ambiguous\"), ds.hex_pfx);\n \treturn status;\n }\n \n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302568","messageId":"20160926120029.zuyasa3hph372beb@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 08/10] sha1_array: let callbacks interrupt iteration","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:29Z","receivedAt":"2016-09-26T12:00:39Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"The callbacks for iterating a sha1_array must have a void\nreturn.  This is unlike our usual for_each semantics, where\na callback may interrupt iteration and have its value\npropagated. Let's switch it to the usual form, which will\nenable its use in more places (e.g., where we are replacing\nan existing iteration with a different data structure).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Documentation/technical/api-sha1-array.txt | 8 ++++++--\n builtin/cat-file.c                         | 3 ++-\n builtin/receive-pack.c                     | 3 ++-\n sha1-array.c                               | 8 ++++++--\n sha1-array.h                               | 8 ++++----\n submodule.c                                | 3 ++-\n t/helper/test-sha1-array.c                 | 3 ++-\n 7 files changed, 24 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/technical/api-sha1-array.txt b/Documentation/technical/api-sha1-array.txt\nindex 3e75497..dcc5294 100644\n--- a/Documentation/technical/api-sha1-array.txt\n+++ b/Documentation/technical/api-sha1-array.txt\n@@ -38,16 +38,20 @@ Functions\n `sha1_array_for_each_unique`::\n \tEfficiently iterate over each unique element of the list,\n \texecuting the callback function for each one. If the array is\n-\tnot sorted, this function has the side effect of sorting it.\n+\tnot sorted, this function has the side effect of sorting it. If\n+\tthe callback returns a non-zero value, the iteration ends\n+\timmediately and the callback's return is propagated; otherwise,\n+\t0 is returned.\n \n Examples\n --------\n \n -----------------------------------------\n-void print_callback(const unsigned char sha1[20],\n+int print_callback(const unsigned char sha1[20],\n \t\t    void *data)\n {\n \tprintf(\"%s\\n\", sha1_to_hex(sha1));\n+\treturn 0; /* always continue */\n }\n \n void some_func(void)\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 94e67eb..cca97a8 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -401,11 +401,12 @@ struct object_cb_data {\n \tstruct expand_data *expand;\n };\n \n-static void batch_object_cb(const unsigned char sha1[20], void *vdata)\n+static int batch_object_cb(const unsigned char sha1[20], void *vdata)\n {\n \tstruct object_cb_data *data = vdata;\n \thashcpy(data->expand->oid.hash, sha1);\n \tbatch_object_write(NULL, data->opt, data->expand);\n+\treturn 0;\n }\n \n static int batch_loose_object(const unsigned char *sha1,\ndiff --git a/builtin/receive-pack.c b/builtin/receive-pack.c\nindex 896b16f..f7cd180 100644\n--- a/builtin/receive-pack.c\n+++ b/builtin/receive-pack.c\n@@ -268,9 +268,10 @@ static int show_ref_cb(const char *path_full, const struct object_id *oid,\n \treturn 0;\n }\n \n-static void show_one_alternate_sha1(const unsigned char sha1[20], void *unused)\n+static int show_one_alternate_sha1(const unsigned char sha1[20], void *unused)\n {\n \tshow_ref(\".have\", sha1);\n+\treturn 0;\n }\n \n static void collect_one_alternate_ref(const struct ref *ref, void *data)\ndiff --git a/sha1-array.c b/sha1-array.c\nindex 6f4a224..af1d7d5 100644\n--- a/sha1-array.c\n+++ b/sha1-array.c\n@@ -42,7 +42,7 @@ void sha1_array_clear(struct sha1_array *array)\n \tarray->sorted = 0;\n }\n \n-void sha1_array_for_each_unique(struct sha1_array *array,\n+int sha1_array_for_each_unique(struct sha1_array *array,\n \t\t\t\tfor_each_sha1_fn fn,\n \t\t\t\tvoid *data)\n {\n@@ -52,8 +52,12 @@ void sha1_array_for_each_unique(struct sha1_array *array,\n \t\tsha1_array_sort(array);\n \n \tfor (i = 0; i < array->nr; i++) {\n+\t\tint ret;\n \t\tif (i > 0 && !hashcmp(array->sha1[i], array->sha1[i-1]))\n \t\t\tcontinue;\n-\t\tfn(array->sha1[i], data);\n+\t\tret = fn(array->sha1[i], data);\n+\t\tif (ret)\n+\t\t\treturn ret;\n \t}\n+\treturn 0;\n }\ndiff --git a/sha1-array.h b/sha1-array.h\nindex 72bb33b..b3230be 100644\n--- a/sha1-array.h\n+++ b/sha1-array.h\n@@ -14,10 +14,10 @@ void sha1_array_append(struct sha1_array *array, const unsigned char *sha1);\n int sha1_array_lookup(struct sha1_array *array, const unsigned char *sha1);\n void sha1_array_clear(struct sha1_array *array);\n \n-typedef void (*for_each_sha1_fn)(const unsigned char sha1[20],\n-\t\t\t\t void *data);\n-void sha1_array_for_each_unique(struct sha1_array *array,\n-\t\t\t\tfor_each_sha1_fn fn,\n+typedef int (*for_each_sha1_fn)(const unsigned char sha1[20],\n \t\t\t\tvoid *data);\n+int sha1_array_for_each_unique(struct sha1_array *array,\n+\t\t\t       for_each_sha1_fn fn,\n+\t\t\t       void *data);\n \n #endif /* SHA1_ARRAY_H */\ndiff --git a/submodule.c b/submodule.c\nindex 0ef2ff4..aba94dd 100644\n--- a/submodule.c\n+++ b/submodule.c\n@@ -728,9 +728,10 @@ void check_for_new_submodule_commits(unsigned char new_sha1[20])\n \tsha1_array_append(&ref_tips_after_fetch, new_sha1);\n }\n \n-static void add_sha1_to_argv(const unsigned char sha1[20], void *data)\n+static int add_sha1_to_argv(const unsigned char sha1[20], void *data)\n {\n \targv_array_push(data, sha1_to_hex(sha1));\n+\treturn 0;\n }\n \n static void calculate_changed_submodule_paths(void)\ndiff --git a/t/helper/test-sha1-array.c b/t/helper/test-sha1-array.c\nindex 09f7790..f7a53c4 100644\n--- a/t/helper/test-sha1-array.c\n+++ b/t/helper/test-sha1-array.c\n@@ -1,9 +1,10 @@\n #include \"cache.h\"\n #include \"sha1-array.h\"\n \n-static void print_sha1(const unsigned char sha1[20], void *data)\n+static int print_sha1(const unsigned char sha1[20], void *data)\n {\n \tputs(sha1_to_hex(sha1));\n+\treturn 0;\n }\n \n int cmd_main(int argc, const char **argv)\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302569","messageId":"20160926120032.7zl65iam3hgk5t62@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 09/10] for_each_abbrev: drop duplicate objects","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:33Z","receivedAt":"2016-09-26T12:00:43Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"If an object appears multiple times in the object database\n(e.g., in both loose and packed form, or in two separate\npacks), the disambiguation machinery may see it more than\nonce. The get_short_sha1() function handles this already,\nbut for_each_abbrev() blindly fires the callback for each\ninstance it finds.\n\nWe can fix this by collecting the output in a sha1 array and\nde-duplicating it.  As a bonus, the sort done for the\nde-duplication means that our output will be stable,\nregardless of the order in which the objects are found.\n\nNote that the old code normalized the callback's output to\n0/1 to store in the 1-bit ds->ambiguous flag (which both\nhalted the iteration and was returned from the\nfor_each_abbrev function). Now that we are using sha1_array,\nwe can return the real value. In practice, it doesn't matter\nas the sole caller only ever returns 0.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c                         | 19 +++++++++++++++----\n t/t1512-rev-parse-disambiguation.sh |  7 +++++++\n 2 files changed, 22 insertions(+), 4 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex d4c7e26..f7403d7 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -7,6 +7,7 @@\n #include \"refs.h\"\n #include \"remote.h\"\n #include \"dir.h\"\n+#include \"sha1-array.h\"\n \n static int get_sha1_oneline(const char *, unsigned char *, struct commit_list *);\n \n@@ -355,20 +356,30 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \treturn status;\n }\n \n+static int collect_ambiguous(const unsigned char *sha1, void *data)\n+{\n+\tsha1_array_append(data, sha1);\n+\treturn 0;\n+}\n+\n int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n {\n+\tstruct sha1_array collect = SHA1_ARRAY_INIT;\n \tstruct disambiguate_state ds;\n+\tint ret;\n \n \tif (init_object_disambiguation(prefix, strlen(prefix), &ds) < 0)\n \t\treturn -1;\n \n \tds.always_call_fn = 1;\n-\tds.cb_data = cb_data;\n-\tds.fn = fn;\n-\n+\tds.fn = collect_ambiguous;\n+\tds.cb_data = &collect;\n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\n-\treturn ds.ambiguous;\n+\n+\tret = sha1_array_for_each_unique(&collect, fn, cb_data);\n+\tsha1_array_clear(&collect);\n+\treturn ret;\n }\n \n int find_unique_abbrev_r(char *hex, const unsigned char *sha1, int len)\ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex dfd3567..1d8f550 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -280,6 +280,13 @@ test_expect_success 'rev-parse --disambiguate' '\n \ttest \"$(sed -e \"s/^\\(.........\\).*/\\1/\" actual | sort -u)\" = 000000000\n '\n \n+test_expect_success 'rev-parse --disambiguate drops duplicates' '\n+\tgit rev-parse --disambiguate=000000000 >expect &&\n+\tgit pack-objects .git/objects/pack/pack <expect &&\n+\tgit rev-parse --disambiguate=000000000 >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_expect_success 'ambiguous 40-hex ref' '\n \tTREE=$(git mktree </dev/null) &&\n \tREF=$(git rev-parse HEAD) &&\n-- \n2.10.0.492.g14f803f\n\n"},{"id":"302570","messageId":"20160926120036.mqs435a36njeihq6@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115720.p2yb22lcq37gboon@sigill.intra.peff.net","subject":"[PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:00:36Z","receivedAt":"2016-09-26T12:00:47Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"When the user gives us an ambiguous short sha1, we print an\nerror and refuse to resolve it. In some cases, the next step\nis for them to feed us more characters (e.g., if they were\nretyping or cut-and-pasting from a full sha1). But in other\ncases, that might be all they have. For example, an old\ncommit message may have used a 7-character hex that was\nunique at the time, but is now ambiguous.  Git doesn't\nprovide any information about the ambiguous objects it\nfound, so it's hard for the user to find out which one they\nprobably meant.\n\nThis patch teaches get_short_sha1() to list the sha1s of the\nobjects it found, along with a few bits of information that\nmay help the user decide which one they meant. Here's what\nit looks like on git.git:\n\n  $ git rev-parse b2e1\n  error: short SHA1 b2e1 is ambiguous\n  hint: The candidates are:\n  hint:   b2e1196 tag v2.8.0-rc1\n  hint:   b2e11d1 tree\n  hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n  hint:   b2e1759 blob\n  hint:   b2e18954 blob\n  hint:   b2e1895c blob\n  fatal: ambiguous argument 'b2e1': unknown revision or path not in the working tree.\n  Use '--' to separate paths from revisions, like this:\n  'git <command> [<revision>...] -- [<file>...]'\n\nWe show the tagname for tags, and the date and subject for\ncommits. For trees and blobs, in theory we could dig in the\nhistory to find the paths at which they were present. But\nthat's very expensive (on the order of 30s for the kernel),\nand it's not likely to be all that helpful. Most short\nreferences are to commits, so the useful information is\ntypically going to be that the object in question _isn't_ a\ncommit. So it's silly to spend a lot of CPU preemptively\ndigging up the path; the user can do it themselves if they\nreally need to.\n\nAnd of course it's somewhat ironic that we abbreviate the\nsha1s in the disambiguation hint. But full sha1s would cause\nannoying line wrapping for the commit lines, and presumably\nthe user is going to just re-issue their command immediately\nwith the corrected sha1.\n\nWe also restrict the list to those that match any\ndisambiguation hint. E.g.:\n\n  $ git rev-parse b2e1:foo\n  error: short SHA1 b2e1 is ambiguous\n  hint: The candidates are:\n  hint:   b2e1196 tag v2.8.0-rc1\n  hint:   b2e11d1 tree\n  hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n  fatal: Invalid object name 'b2e1'.\n\ndoes not bother reporting the blobs, because they cannot\nwork as a treeish.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_name.c                         | 50 +++++++++++++++++++++++++++++++++++--\n t/t1512-rev-parse-disambiguation.sh | 24 ++++++++++++++++++\n 2 files changed, 72 insertions(+), 2 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex f7403d7..35d943d 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -318,6 +318,38 @@ static int init_object_disambiguation(const char *name, int len,\n \treturn 0;\n }\n \n+static int show_ambiguous_object(const unsigned char *sha1, void *data)\n+{\n+\tconst struct disambiguate_state *ds = data;\n+\tstruct strbuf desc = STRBUF_INIT;\n+\tint type;\n+\n+\tif (ds->fn && !ds->fn(sha1, ds->cb_data))\n+\t\treturn 0;\n+\n+\ttype = sha1_object_info(sha1, NULL);\n+\tif (type == OBJ_COMMIT) {\n+\t\tstruct commit *commit = lookup_commit(sha1);\n+\t\tif (commit) {\n+\t\t\tstruct pretty_print_context pp = {0};\n+\t\t\tpp.date_mode.type = DATE_SHORT;\n+\t\t\tformat_commit_message(commit, \" %ad - %s\", &desc, &pp);\n+\t\t}\n+\t} else if (type == OBJ_TAG) {\n+\t\tstruct tag *tag = lookup_tag(sha1);\n+\t\tif (!parse_tag(tag) && tag->tag)\n+\t\t\tstrbuf_addf(&desc, \" %s\", tag->tag);\n+\t}\n+\n+\tadvise(\"  %s %s%s\",\n+\t       find_unique_abbrev(sha1, DEFAULT_ABBREV),\n+\t       typename(type) ? typename(type) : \"unknown type\",\n+\t       desc.buf);\n+\n+\tstrbuf_release(&desc);\n+\treturn 0;\n+}\n+\n static int multiple_bits_set(unsigned flags)\n {\n \treturn !!(flags & (flags - 1));\n@@ -351,8 +383,22 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \tfind_short_packed_object(&ds);\n \tstatus = finish_object_disambiguation(&ds, sha1);\n \n-\tif (!quietly && (status == SHORT_NAME_AMBIGUOUS))\n-\t\treturn error(_(\"short SHA1 %s is ambiguous\"), ds.hex_pfx);\n+\tif (!quietly && (status == SHORT_NAME_AMBIGUOUS)) {\n+\t\terror(_(\"short SHA1 %s is ambiguous\"), ds.hex_pfx);\n+\n+\t\t/*\n+\t\t * We may still have ambiguity if we simply saw a series of\n+\t\t * candidates that did not satisfy our hint function. In\n+\t\t * that case, we still want to show them, so disable the hint\n+\t\t * function entirely.\n+\t\t */\n+\t\tif (!ds.ambiguous)\n+\t\t\tds.fn = NULL;\n+\n+\t\tadvise(_(\"The candidates are:\"));\n+\t\tfor_each_abbrev(ds.hex_pfx, show_ambiguous_object, &ds);\n+\t}\n+\n \treturn status;\n }\n \ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex 1d8f550..c5447ef 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -323,4 +323,28 @@ test_expect_success C_LOCALE_OUTPUT 'ambiguity errors are not repeated (peel)' '\n \ttest_line_count = 1 errors\n '\n \n+test_expect_success C_LOCALE_OUTPUT 'ambiguity hints' '\n+\ttest_must_fail git rev-parse 000000000 2>stderr &&\n+\tgrep ^hint: stderr >hints &&\n+\t# 16 candidates, plus one intro line\n+\ttest_line_count = 17 hints\n+'\n+\n+test_expect_success C_LOCALE_OUTPUT 'ambiguity hints respect type' '\n+\ttest_must_fail git rev-parse 000000000^{commit} 2>stderr &&\n+\tgrep ^hint: stderr >hints &&\n+\t# 5 commits, 1 tag (which is a commitish), plus intro line\n+\ttest_line_count = 7 hints\n+'\n+\n+test_expect_success C_LOCALE_OUTPUT 'failed type-selector still shows hint' '\n+\t# these two blobs share the same prefix \"ee3d\", but neither\n+\t# will pass for a commit\n+\techo 851 | git hash-object --stdin -w &&\n+\techo 872 | git hash-object --stdin -w &&\n+\ttest_must_fail git rev-parse ee3d^{commit} 2>stderr &&\n+\tgrep ^hint: stderr >hints &&\n+\ttest_line_count = 3 hints\n+'\n+\n test_done\n-- \n2.10.0.492.g14f803f\n"},{"id":"302571","messageId":"20160926120942.tjaxhcwwz3lyxy25@sigill.intra.peff.net","threadId":"44156","inReplyTo":"vpq37kntbjj.fsf@anie.imag.fr","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:09:43Z","receivedAt":"2016-09-26T12:09:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 08:33:52AM +0200, Matthieu Moy wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > I am not opposed to bump the default to 12 or whatever, but I\n> > suspect any lengthening today may need to be accompanied by a tool\n> > support that finds the set of objects that are reachable from a\n> > commit whose names begin with non-unique abbreviations that appear\n> > in the commit log message.\n> \n> Something much simpler would be to set core.abbrev at clone time,\n> depending on the size of the project just cloned. So, when cloning a\n> hello-world, we'd keep the 7 but when cloning a big project we'd get a\n> larger value.\n> \n> This doesn't cover the case of someone growing his own project without\n> cloning, and isn't as clever as actually looking for colision, but it\n> would probably provide a sane default in 99% cases, and wouldn't be\n> worse than hardcoding 7 in the 1% remaining cases.\n\nI think we could easily make this even more dynamic, and just base the\nminimum for DEFAULT_ABBREV on the number of objects _currently_ in the\nrepository, plus some safety factor. We could do this cheaply by just\ncounting the number of objects in the packs (which we get for free when\nwe open their pack index). That misses loose objects, but if you have 4\nmillion loose objects you have bigger problems than abbreviation\nlengths, I think.\n\nOTOH, any scheme that looks at the current repository size will\neventually grow outdated. The safety factor depends on how fast your\nrepository grows, and how big you expect it to eventually get. Such a\ndefault might still have been using 7-character abbreviations on\nlinux.git in 2006, and we'd be stuck with them now.\n\nThe idea of a 12-character default is basically that we'd expect decades\nor more for even the largest projects to get there, so you err on the\nside of future-proofing.\n\n-Peff\n"},{"id":"302572","messageId":"20160926121104.nonvu7h3shygspi6@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160926115947.hksmtkqp3i4tfftx@sigill.intra.peff.net","subject":"Re: [PATCH 04/10] get_short_sha1: peel tags when looking for treeish","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T12:11:04Z","receivedAt":"2016-09-26T12:11:14Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 07:59:48AM -0400, Jeff King wrote:\n\n> Subject: Re: [PATCH 04/10] get_short_sha1: peel tags when looking for treeish\n>\n> The treeish disambiguation function tries to peel tags, but\n> it does so by calling:\n\nProbably the subject should be \"parse tags when...\". We already try to\npeel, we just don't do it right.\n\n-Peff\n"},{"id":"302583","messageId":"CA+55aFyfvvqq1c=hZcuL-yPavp2tjzx8r3bFJnMY7DAE7YcB=Q@mail.gmail.com","threadId":"44156","inReplyTo":"20160926120036.mqs435a36njeihq6@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-26T16:36:23Z","receivedAt":"2016-09-26T16:36:28Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Sep 26, 2016 at 5:00 AM, Jeff King <peff@peff.net> wrote:\n>\n> This patch teaches get_short_sha1() to list the sha1s of the\n> objects it found, along with a few bits of information that\n> may help the user decide which one they meant.\n\nThis looks very good to me, but I wonder if it couldn't be even more aggressive.\n\nIn particular, the only hashes that most people ever use in short form\nare commit hashes. Those are the ones you'd use in normal human\ninteractions to point to something happening.\n\nSo when the disambiguation notices that there is ambiguity, but there\nis only _one_ commit, maybe it should just have an aggressive mode\nthat says \"use that as if it wasn't ambiguous\".\n\nAnd then have an explicit command (or flag) to do disambiguation for\nwhen you explicitly want it.\n\nRationale: you'd never care about short forms for tags. You'd just use\nthe tag name. And while blob ID's certainly show up in short form in\ndiff output (in the \"index\" line), very few people will use them. And\ntree hashes are basically never seen outside of any plumbing commands\nand then seldom in shortened form.\n\nSo I think it would make sense to default to a mode that just picks\nthe commit hash if there is only one such hash. Sure, some command\nmight want a \"treeish\", but a commit is still more likely than a tree\nor a tag.\n\nBut regardless, this series looks like a good thing.\n\n                        Linus\n"},{"id":"302584","messageId":"xmqqbmzavcqx.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926115901.txmbr4e6xzwyfpmo@sigill.intra.peff.net","subject":"Re: [PATCH 01/10] get_sha1: detect buggy calls with multiple disambiguators","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T16:37:10Z","receivedAt":"2016-09-26T16:37:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> The get_sha1() family of functions takes a flags field, but\n> some of the flags are mutually exclusive. In particular, we\n> can only handle one disambiguating function, and the flags\n> quietly override each other. Let's instead detect these as\n> programming bugs.\n>\n> Technically some of the flags are supersets of the others,\n> so treating COMMITTISH|TREEISH as just COMMITTISH is not\n> wrong, but it's a good sign the caller is confused. And\n> certainly asking for BLOB|TREE does not work.\n>\n> We can do the check easily with some bit-twiddling, and as a\n> bonus, the bit-mask of disambiguators will come in handy in\n> a future patch.\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n\nOther than your reinvention of HAS_MULTI_BITS(), which has been with\nus since db7244bd (\"parse-options new features.\", 2007-11-07), this\nlooks like a reasonable thing to do.\n\n;-)\n\n>  cache.h     | 5 +++++\n>  sha1_name.c | 9 +++++++++\n>  2 files changed, 14 insertions(+)\n>\n> diff --git a/cache.h b/cache.h\n> index d0494c8..7bd78ca 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1203,6 +1203,11 @@ struct object_context {\n>  #define GET_SHA1_FOLLOW_SYMLINKS 0100\n>  #define GET_SHA1_ONLY_TO_DIE    04000\n>  \n> +#define GET_SHA1_DISAMBIGUATORS \\\n> +\t(GET_SHA1_COMMIT | GET_SHA1_COMMITTISH | \\\n> +\tGET_SHA1_TREE | GET_SHA1_TREEISH | \\\n> +\tGET_SHA1_BLOB)\n> +\n>  extern int get_sha1(const char *str, unsigned char *sha1);\n>  extern int get_sha1_commit(const char *str, unsigned char *sha1);\n>  extern int get_sha1_committish(const char *str, unsigned char *sha1);\n> diff --git a/sha1_name.c b/sha1_name.c\n> index faf873c..f9812ff 100644\n> --- a/sha1_name.c\n> +++ b/sha1_name.c\n> @@ -310,6 +310,11 @@ static int prepare_prefixes(const char *name, int len,\n>  \treturn 0;\n>  }\n>  \n> +static int multiple_bits_set(unsigned flags)\n> +{\n> +\treturn !!(flags & (flags - 1));\n> +}\n> +\n>  static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n>  \t\t\t  unsigned flags)\n>  {\n> @@ -327,6 +332,10 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n>  \tprepare_alt_odb();\n>  \n>  \tmemset(&ds, 0, sizeof(ds));\n> +\n> +\tif (multiple_bits_set(flags & GET_SHA1_DISAMBIGUATORS))\n> +\t\tdie(\"BUG: multiple get_short_sha1 disambiguator flags\");\n> +\n>  \tif (flags & GET_SHA1_COMMIT)\n>  \t\tds.fn = disambiguate_commit_only;\n>  \telse if (flags & GET_SHA1_COMMITTISH)\n"},{"id":"302585","messageId":"xmqq7f9yvbwn.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926115947.hksmtkqp3i4tfftx@sigill.intra.peff.net","subject":"Re: [PATCH 04/10] get_short_sha1: peel tags when looking for treeish","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T16:55:20Z","receivedAt":"2016-09-26T16:55:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> The treeish disambiguation function tries to peel tags, but\n> it does so by calling:\n>\n>   deref_tag(lookup_object(sha1), ...);\n>\n> This will only work if we have previously looked at the tag\n> and created a \"struct tag\" for it. Since parsing revision\n> arguments typically happens before anything else, this is\n> usually not the case, and we would fail to peel the tag (we\n> are lucky that deref_tag() gracefully handles the NULL and\n> does not segfault).\n\nMakes perfect sense.\n\n> Instead, we can use parse_object(). Note that this is the\n> same fix done by 94d75d1 (get_short_sha1(): correctly\n> disambiguate type-limited abbreviation, 2013-07-01), but\n> that commit fixed only the committish disambiguator, and\n> left the bug in the treeish one.\n\nCan you share your secret tool you use to find this kind of thing?\nYes, the patch from that commit does look very similar to what we\nsee in this patch, but I'd love to see \"I am fixing an incorrect\ncall to lookup-object by replacing it with parse-object; has there\nbeen a similar fix?\" automated ;-)\n\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n>  sha1_name.c                         | 2 +-\n>  t/t1512-rev-parse-disambiguation.sh | 7 +++++++\n>  2 files changed, 8 insertions(+), 1 deletion(-)\n>\n> diff --git a/sha1_name.c b/sha1_name.c\n> index 38e51d9..432a308 100644\n> --- a/sha1_name.c\n> +++ b/sha1_name.c\n> @@ -269,7 +269,7 @@ static int disambiguate_treeish_only(const unsigned char *sha1, void *cb_data_un\n>  \t\treturn 0;\n>  \n>  \t/* We need to do this the hard way... */\n> -\tobj = deref_tag(lookup_object(sha1), NULL, 0);\n> +\tobj = deref_tag(parse_object(sha1), NULL, 0);\n>  \tif (obj && (obj->type == OBJ_TREE || obj->type == OBJ_COMMIT))\n>  \t\treturn 1;\n>  \treturn 0;\n> diff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\n> index 30e0b80..dfd3567 100755\n> --- a/t/t1512-rev-parse-disambiguation.sh\n> +++ b/t/t1512-rev-parse-disambiguation.sh\n> @@ -264,6 +264,13 @@ test_expect_success 'ambiguous commit-ish' '\n>  \ttest_must_fail git log 000000000...\n>  '\n>  \n> +# There are three objects with this prefix: a blob, a tree, and a tag. We know\n> +# the blob will not pass as a treeish, but the tree and tag should (and thus\n> +# cause an error).\n> +test_expect_success 'ambiguous tags peel to treeish' '\n> +\ttest_must_fail git rev-parse 0000000000f^{tree}\n> +'\n> +\n>  test_expect_success 'rev-parse --disambiguate' '\n>  \t# The test creates 16 objects that share the prefix and two\n>  \t# commits created by commit-tree in earlier tests share a\n"},{"id":"302587","messageId":"xmqq37kmvb6x.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926120007.eswpfrzs2ed66d2o@sigill.intra.peff.net","subject":"Re: [PATCH 06/10] get_short_sha1: NUL-terminate hex prefix","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T17:10:46Z","receivedAt":"2016-09-26T17:10:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> We store the hex prefix in a 40-byte buffer with the prefix\n> itself followed by 40-minus-len \"x\" characters. These x's\n> serve no purpose, and the lack of NUL termination makes the\n> prefix string annoying to use. Let's just terminate it.\n\n> Note that this is in contrast to the binary prefix, which\n> _must_ be zero-padded, because we look at the whole thing\n> during a binary search to find the first potential match in\n> each pack index. \n\nMakes sense.\n\n> The loose-object hex search cannot use the\n> same trick because it has to do a linear walk through the\n> unsorted results of readdir() (and even if it could, you'd\n> want zeroes instead of x's).\n\nOK.\n\n>  struct disambiguate_state {\n>  \tint len; /* length of prefix in hex chars */\n> -\tchar hex_pfx[GIT_SHA1_HEXSZ];\n> +\tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n>  \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n>  \n>  \tdisambiguate_hint_fn fn;\n> @@ -291,7 +291,6 @@ static int init_object_disambiguation(const char *name, int len,\n>  \t\treturn -1;\n>  \n>  \tmemset(ds, 0, sizeof(*ds));\n> -\tmemset(ds->hex_pfx, 'x', GIT_SHA1_HEXSZ);\n\nAs the whole thing is cleared here...\n\n>  \n>  \tfor (i = 0; i < len ;i++) {\n>  \t\tunsigned char c = name[i];\n> @@ -313,6 +312,7 @@ static int init_object_disambiguation(const char *name, int len,\n>  \t}\n>  \n>  \tds->len = len;\n> +\tds->hex_pfx[len] = '\\0';\n\n... do we even need this one?  It would not hurt, though.\n\n> @@ -351,7 +351,7 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n>  \tstatus = finish_object_disambiguation(&ds, sha1);\n>  \n>  \tif (!quietly && (status == SHORT_NAME_AMBIGUOUS))\n> -\t\treturn error(\"short SHA1 %.*s is ambiguous.\", ds.len, ds.hex_pfx);\n> +\t\treturn error(\"short SHA1 %s is ambiguous.\", ds.hex_pfx);\n\nMakes sense.\n\nThanks.\n"},{"id":"302589","messageId":"20160926172111.sve6tsse2figcved@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqbmzavcqx.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 01/10] get_sha1: detect buggy calls with multiple disambiguators","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T17:21:12Z","receivedAt":"2016-09-26T17:21:25Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 09:37:10AM -0700, Junio C Hamano wrote:\n\n> > We can do the check easily with some bit-twiddling, and as a\n> > bonus, the bit-mask of disambiguators will come in handy in\n> > a future patch.\n> >\n> > Signed-off-by: Jeff King <peff@peff.net>\n> > ---\n> \n> Other than your reinvention of HAS_MULTI_BITS(), which has been with\n> us since db7244bd (\"parse-options new features.\", 2007-11-07), this\n> looks like a reasonable thing to do.\n\nHeh, I _thought_ we had something like that but couldn't find it. I\ngrepped for \"[^&]& .*-\", which does match it, but stupidly did it only\nin '*.c'. Definitely it should use the existing macro instead.\n\n-Peff\n"},{"id":"302591","messageId":"20160926172350.ikqfnanrrj5oepmq@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqq7f9yvbwn.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 04/10] get_short_sha1: peel tags when looking for treeish","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T17:23:50Z","receivedAt":"2016-09-26T17:23:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 09:55:20AM -0700, Junio C Hamano wrote:\n\n> > Instead, we can use parse_object(). Note that this is the\n> > same fix done by 94d75d1 (get_short_sha1(): correctly\n> > disambiguate type-limited abbreviation, 2013-07-01), but\n> > that commit fixed only the committish disambiguator, and\n> > left the bug in the treeish one.\n> \n> Can you share your secret tool you use to find this kind of thing?\n> Yes, the patch from that commit does look very similar to what we\n> see in this patch, but I'd love to see \"I am fixing an incorrect\n> call to lookup-object by replacing it with parse-object; has there\n> been a similar fix?\" automated ;-)\n\nI wish there was an answer besides \"persistence and patience\". I was\njust finishing up the commit message for the final patch, and noticed\nthat the tag was not present in the second example output, which happens\nto use the tree-ish syntax. And I noticed it was doubly weird that the\nsame bug did not show up in the test scripts, which look for\ncommittishes. That made me peek at the implementation, and from there it\nwas an easy `git blame` away.\n\n-Peff\n"},{"id":"302592","messageId":"20160926172516.frftagyt6aycp75q@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqq37kmvb6x.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 06/10] get_short_sha1: NUL-terminate hex prefix","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T17:25:16Z","receivedAt":"2016-09-26T17:25:23Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 10:10:46AM -0700, Junio C Hamano wrote:\n\n> >  struct disambiguate_state {\n> >  \tint len; /* length of prefix in hex chars */\n> > -\tchar hex_pfx[GIT_SHA1_HEXSZ];\n> > +\tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n> >  \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n> >  \n> >  \tdisambiguate_hint_fn fn;\n> > @@ -291,7 +291,6 @@ static int init_object_disambiguation(const char *name, int len,\n> >  \t\treturn -1;\n> >  \n> >  \tmemset(ds, 0, sizeof(*ds));\n> > -\tmemset(ds->hex_pfx, 'x', GIT_SHA1_HEXSZ);\n> \n> As the whole thing is cleared here...\n> \n> >  \n> >  \tfor (i = 0; i < len ;i++) {\n> >  \t\tunsigned char c = name[i];\n> > @@ -313,6 +312,7 @@ static int init_object_disambiguation(const char *name, int len,\n> >  \t}\n> >  \n> >  \tds->len = len;\n> > +\tds->hex_pfx[len] = '\\0';\n> \n> ... do we even need this one?  It would not hurt, though.\n\nSharp eyes. I noticed that while writing it, but wondered if anybody\nelse would. :)\n\nI left the second one in to make the intention more explicit, and so\nreaders did not have to worry that the NULs were overwritten in the\nloop. I'd be OK with it either way, though.\n\n-Peff\n"},{"id":"302593","messageId":"xmqqwphytvp3.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926120036.mqs435a36njeihq6@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T17:30:48Z","receivedAt":"2016-09-26T17:30:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> We also restrict the list to those that match any\n> disambiguation hint. E.g.:\n>\n>   $ git rev-parse b2e1:foo\n>   error: short SHA1 b2e1 is ambiguous\n>   hint: The candidates are:\n>   hint:   b2e1196 tag v2.8.0-rc1\n>   hint:   b2e11d1 tree\n>   hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n>   fatal: Invalid object name 'b2e1'.\n>\n> does not bother reporting the blobs, because they cannot\n> work as a treeish.\n\nThat's a nice touch, and it even comes free--how wonderful.\n\nIt somehow felt strange to have an expensive (compared to no-op,\nanyway) loop whose only externally visible effect is to call\nadvise(), but there does not appear to be a way to even disable this\nadvise() output, so it probably is OK, I guess.\n\n>  \n> +test_expect_success C_LOCALE_OUTPUT 'ambiguity hints' '\n> +\ttest_must_fail git rev-parse 000000000 2>stderr &&\n> +\tgrep ^hint: stderr >hints &&\n> +\t# 16 candidates, plus one intro line\n> +\ttest_line_count = 17 hints\n> +'\n> +\n> +test_expect_success C_LOCALE_OUTPUT 'ambiguity hints respect type' '\n> +\ttest_must_fail git rev-parse 000000000^{commit} 2>stderr &&\n> +\tgrep ^hint: stderr >hints &&\n> +\t# 5 commits, 1 tag (which is a commitish), plus intro line\n> +\ttest_line_count = 7 hints\n> +'\n> +\n> +test_expect_success C_LOCALE_OUTPUT 'failed type-selector still shows hint' '\n> +\t# these two blobs share the same prefix \"ee3d\", but neither\n> +\t# will pass for a commit\n> +\techo 851 | git hash-object --stdin -w &&\n> +\techo 872 | git hash-object --stdin -w &&\n> +\ttest_must_fail git rev-parse ee3d^{commit} 2>stderr &&\n> +\tgrep ^hint: stderr >hints &&\n> +\ttest_line_count = 3 hints\n> +'\n> +\n>  test_done\n"},{"id":"302595","messageId":"20160926173413.prp3wevf6kkksy7c@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqwphytvp3.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-26T17:34:13Z","receivedAt":"2016-09-26T17:34:20Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 10:30:48AM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > We also restrict the list to those that match any\n> > disambiguation hint. E.g.:\n> >\n> >   $ git rev-parse b2e1:foo\n> >   error: short SHA1 b2e1 is ambiguous\n> >   hint: The candidates are:\n> >   hint:   b2e1196 tag v2.8.0-rc1\n> >   hint:   b2e11d1 tree\n> >   hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n> >   fatal: Invalid object name 'b2e1'.\n> >\n> > does not bother reporting the blobs, because they cannot\n> > work as a treeish.\n> \n> That's a nice touch, and it even comes free--how wonderful.\n> \n> It somehow felt strange to have an expensive (compared to no-op,\n> anyway) loop whose only externally visible effect is to call\n> advise(), but there does not appear to be a way to even disable this\n> advise() output, so it probably is OK, I guess.\n\nRight, advise() always has an effect. But that reminds me.  I wasn't\nsure if we should attach an advice.* config to this. If we do, then the\nright place to put the conditional is right after the error() call in\nget_short_sha1().\n\nSince it's attached to an error path, I'm guessing nobody will be too\nupset about it, so my inclination was to wait and let somebody add the\nconditional advice code if they're bothered.\n\n-Peff\n"},{"id":"302598","messageId":"xmqqk2dytvet.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926172516.frftagyt6aycp75q@sigill.intra.peff.net","subject":"Re: [PATCH 06/10] get_short_sha1: NUL-terminate hex prefix","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T17:36:58Z","receivedAt":"2016-09-26T17:37:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I left the second one in to make the intention more explicit, and so\n> readers did not have to worry that the NULs were overwritten in the\n> loop. I'd be OK with it either way, though.\n\nYes, I agree with that it is a good thing to make our intention more\nexplicit and I am perfectly fine with leaving it as-is.\n\nThanks.\n\n"},{"id":"302600","messageId":"xmqqfuomtvb5.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926173413.prp3wevf6kkksy7c@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T17:39:10Z","receivedAt":"2016-09-26T17:39:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Since it's attached to an error path, I'm guessing nobody will be too\n> upset about it, so my inclination was to wait and let somebody add the\n> conditional advice code if they're bothered.\n\nFair enough.  At that point of getting an error message, the only\nthing they can do is to start wondering what object the person who\ngave the now-non-unique abbrevation to them, so I suspect this is\none of the \"advice\" messages that can always be there.\n"},{"id":"302603","messageId":"xmqq7f9yturq.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160926172111.sve6tsse2figcved@sigill.intra.peff.net","subject":"Re: [PATCH 01/10] get_sha1: detect buggy calls with multiple disambiguators","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-26T17:50:49Z","receivedAt":"2016-09-26T17:50:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> Other than your reinvention of HAS_MULTI_BITS(), which has been with\n>> us since db7244bd (\"parse-options new features.\", 2007-11-07), this\n>> looks like a reasonable thing to do.\n>\n> Heh, I _thought_ we had something like that but couldn't find it. I\n> grepped for \"[^&]& .*-\", which does match it, but stupidly did it only\n> in '*.c'. Definitely it should use the existing macro instead.\n\nOK, I'll queue this on top to be squashed, so no need to resend only\nfor this one.\n\n sha1_name.c | 7 +------\n 1 file changed, 1 insertion(+), 6 deletions(-)\n\ndiff --git a/sha1_name.c b/sha1_name.c\nindex f9812ff..0ff83a9 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -310,11 +310,6 @@ static int prepare_prefixes(const char *name, int len,\n \treturn 0;\n }\n \n-static int multiple_bits_set(unsigned flags)\n-{\n-\treturn !!(flags & (flags - 1));\n-}\n-\n static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\t\t  unsigned flags)\n {\n@@ -333,7 +328,7 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \n \tmemset(&ds, 0, sizeof(ds));\n \n-\tif (multiple_bits_set(flags & GET_SHA1_DISAMBIGUATORS))\n+\tif (HAS_MULTI_BITS(flags & GET_SHA1_DISAMBIGUATORS))\n \t\tdie(\"BUG: multiple get_short_sha1 disambiguator flags\");\n \n \tif (flags & GET_SHA1_COMMIT)\n"},{"id":"302677","messageId":"CA+P7+xpQQpZECE=XFjPbv5x8rYvY9UY4JyeTMFkSwFqjHmO_vA@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFyfvvqq1c=hZcuL-yPavp2tjzx8r3bFJnMY7DAE7YcB=Q@mail.gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2016-09-27T05:42:26Z","receivedAt":"2016-09-27T05:42:52Z","isPatch":true,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Mon, Sep 26, 2016 at 9:36 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> This looks very good to me, but I wonder if it couldn't be even more aggressive.\n>\n> In particular, the only hashes that most people ever use in short form\n> are commit hashes. Those are the ones you'd use in normal human\n> interactions to point to something happening.\n>\n> So when the disambiguation notices that there is ambiguity, but there\n> is only _one_ commit, maybe it should just have an aggressive mode\n> that says \"use that as if it wasn't ambiguous\".\n>\n> And then have an explicit command (or flag) to do disambiguation for\n> when you explicitly want it.\n>\n> Rationale: you'd never care about short forms for tags. You'd just use\n> the tag name. And while blob ID's certainly show up in short form in\n> diff output (in the \"index\" line), very few people will use them. And\n> tree hashes are basically never seen outside of any plumbing commands\n> and then seldom in shortened form.\n>\n> So I think it would make sense to default to a mode that just picks\n> the commit hash if there is only one such hash. Sure, some command\n> might want a \"treeish\", but a commit is still more likely than a tree\n> or a tag.\n>\n\nI'd think we would want to phase this in over a few releases if we do\nthis? Maybe at least sort commits first in the list so that they are\nfaster to spot.\n\nI am trying to think of what problems we'd cause by having the\nbehavior be this aggressive...\n\nThanks,\nJake\n\n> But regardless, this series looks like a good thing.\n>\n>                         Linus\n"},{"id":"302698","messageId":"20160927123801.3bpdg3hap3kzzfmv@sigill.intra.peff.net","threadId":"44156","inReplyTo":"CA+55aFyfvvqq1c=hZcuL-yPavp2tjzx8r3bFJnMY7DAE7YcB=Q@mail.gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-27T12:38:01Z","receivedAt":"2016-09-27T12:38:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 26, 2016 at 09:36:23AM -0700, Linus Torvalds wrote:\n\n> On Mon, Sep 26, 2016 at 5:00 AM, Jeff King <peff@peff.net> wrote:\n> >\n> > This patch teaches get_short_sha1() to list the sha1s of the\n> > objects it found, along with a few bits of information that\n> > may help the user decide which one they meant.\n> \n> This looks very good to me, but I wonder if it couldn't be even more\n> aggressive.\n> \n> In particular, the only hashes that most people ever use in short form\n> are commit hashes. Those are the ones you'd use in normal human\n> interactions to point to something happening.\n> \n> So when the disambiguation notices that there is ambiguity, but there\n> is only _one_ commit, maybe it should just have an aggressive mode\n> that says \"use that as if it wasn't ambiguous\".\n\nYou can basically get that by using \"1234^{commit}\" all the time, as\nthat turns on the committish disambiguator function (though it's not\nquite the same, as it would pick a tag, too; you really want the\ncommit-only disambiguator). But presumably you'd want it on all the\ntime. See the patch below, which lets you do:\n\n  git config --global core.disambiguate commit\n\nand I think should do what you want. I'm up in the air on whether it is\na good idea or not, but then I do not usually run into ambiguous sha1s.\n\n> And then have an explicit command (or flag) to do disambiguation for\n> when you explicitly want it.\n\nIn my patch you can tweak the config variable off, though it might make\nsense to also have some per-short-sha1 syntax.\n\n> Rationale: you'd never care about short forms for tags. You'd just use\n> the tag name. And while blob ID's certainly show up in short form in\n> diff output (in the \"index\" line), very few people will use them. And\n> tree hashes are basically never seen outside of any plumbing commands\n> and then seldom in shortened form.\n\nI think I do sometimes \"git show $blob_sha1\" based on a diff index line.\nOTOH, I don't think of though as \"long-term\" references. I'm usually\ntrying to apply the patch at the time, so it's fairly fresh (it's true\nthat the short-sha1 may have been generated on the sender's side, who\nhas fewer objects, but I doubt that's a big problem in general; the real\nissue is that it was unique at one point, and isn't a few years later).\n\nBut more importantly, any fallback like this should take a backseat to\ncontext provided by the rest of git. So for instance, the index-building\nin \"am -3\" uses the blob disambiguator, and should continue to do so\n(and does with my patch).\n\n> So I think it would make sense to default to a mode that just picks\n> the commit hash if there is only one such hash. Sure, some command\n> might want a \"treeish\", but a commit is still more likely than a tree\n> or a tag.\n\nBy the same rule I just mentioned above, if you use the short sha1 in a\ntreeish context, it will look for any treeish (so \"1234:foo\" would\ncontinue to look for any treeish, not just a commit). So that might not\nbe as desirable, but I think it does make sense (and of course it will\nstill tell you immediately what the options are, and you can decide what\nto do).\n\n-- >8 --\nSubject: [PATCH] get_short_sha1: make default disambiguation configurable\n\nWhen we find ambiguous short sha1s, we may get a\ndisambiguation rule from our caller's context. But if we\ndon't, we fall back to treating all sha1s the same, even\nthough most projects will tend to refer only to commits by\ntheir short sha1s.\n\nThis patch introduces a configuration option that lets the\nuser pick a different fallback (e.g., only commits). It's\npossible that we may want to make this the default, but it's\na good idea to start as a config option for two reasons:\n\n  1. It lets people experiment with this and see if it's a\n     good idea (i.e., the \"tend to\" above is an assumption;\n     we don't really know if this will break some obscure\n     cases).\n\n  2. Even if we do flip the default, it gives people an\n     escape hatch if it causes problems (you can sometimes\n     override it by asking for \"1234^{tree}\", but not all\n     combinations are possible).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n cache.h                             |  2 ++\n config.c                            |  3 +++\n sha1_name.c                         | 32 ++++++++++++++++++++++++++++++++\n t/t1512-rev-parse-disambiguation.sh | 14 ++++++++++++++\n 4 files changed, 51 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex 5df0f33..b9583c4 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1224,6 +1224,8 @@ extern int get_oid(const char *str, struct object_id *oid);\n typedef int each_abbrev_fn(const unsigned char *sha1, void *);\n extern int for_each_abbrev(const char *prefix, each_abbrev_fn, void *);\n \n+extern int set_disambiguate_hint_config(const char *var, const char *value);\n+\n /*\n  * Try to read a SHA1 in hexadecimal format from the 40 characters\n  * starting at hex.  Write the 20-byte result to sha1 in binary form.\ndiff --git a/config.c b/config.c\nindex 1e4b617..83fdecb 100644\n--- a/config.c\n+++ b/config.c\n@@ -841,6 +841,9 @@ static int git_default_core_config(const char *var, const char *value)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.disambiguate\"))\n+\t\treturn set_disambiguate_hint_config(var, value);\n+\n \tif (!strcmp(var, \"core.loosecompression\")) {\n \t\tint level = git_config_int(var, value);\n \t\tif (level == -1)\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 0513f14..3b647fd 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -283,6 +283,36 @@ static int disambiguate_blob_only(const unsigned char *sha1, void *cb_data_unuse\n \treturn kind == OBJ_BLOB;\n }\n \n+static disambiguate_hint_fn default_disambiguate_hint;\n+\n+int set_disambiguate_hint_config(const char *var, const char *value)\n+{\n+\tstatic const struct {\n+\t\tconst char *name;\n+\t\tdisambiguate_hint_fn fn;\n+\t} hints[] = {\n+\t\t{ \"none\", NULL },\n+\t\t{ \"commit\", disambiguate_commit_only },\n+\t\t{ \"committish\", disambiguate_committish_only },\n+\t\t{ \"tree\", disambiguate_tree_only },\n+\t\t{ \"treeish\", disambiguate_treeish_only },\n+\t\t{ \"blob\", disambiguate_blob_only }\n+\t};\n+\tint i;\n+\n+\tif (!value)\n+\t\treturn config_error_nonbool(var);\n+\n+\tfor (i = 0; i < ARRAY_SIZE(hints); i++) {\n+\t\tif (!strcasecmp(value, hints[i].name)) {\n+\t\t\tdefault_disambiguate_hint = hints[i].fn;\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\n+\treturn error(\"unknown hint type for '%s': %s\", var, value);\n+}\n+\n static int init_object_disambiguation(const char *name, int len,\n \t\t\t\t      struct disambiguate_state *ds)\n {\n@@ -373,6 +403,8 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\tds.fn = disambiguate_treeish_only;\n \telse if (flags & GET_SHA1_BLOB)\n \t\tds.fn = disambiguate_blob_only;\n+\telse\n+\t\tds.fn = default_disambiguate_hint;\n \n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex c5447ef..7c659eb 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -347,4 +347,18 @@ test_expect_success C_LOCALE_OUTPUT 'failed type-selector still shows hint' '\n \ttest_line_count = 3 hints\n '\n \n+test_expect_success 'core.disambiguate config can prefer types' '\n+\t# ambiguous between tree and tag\n+\tsha1=0000000000f &&\n+\ttest_must_fail git rev-parse $sha1 &&\n+\tgit rev-parse $sha1^{commit} &&\n+\tgit -c core.disambiguate=committish rev-parse $sha1\n+'\n+\n+test_expect_success 'core.disambiguate does not override context' '\n+\t# treeish ambiguous between tag and tree\n+\ttest_must_fail \\\n+\t\tgit -c core.disambiguate=committish rev-parse $sha1^{tree}\n+'\n+\n test_done\n-- \n2.10.0.564.g318c4ae\n\n"},{"id":"302854","messageId":"20160928233047.14313-4-gitster@pobox.com","threadId":"44156","inReplyTo":"20160928233047.14313-1-gitster@pobox.com","subject":"[PATCH 3/4] worktree: honor configuration variables","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-28T23:30:46Z","receivedAt":"2016-09-28T23:34:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The command accesses default_abbrev (defined in environment.c and is\nupdated via core.abbrev configuration), but never makes any call to\ngit_config().  The output from \"worktree list\" ignores the abbrev\nsetting for this reason.\n\nMake a call to git_config() to read the default set of configuration\nvariables at the beginning of the command.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * This is already queued separately from this series.\n\n builtin/worktree.c | 2 ++\n 1 file changed, 2 insertions(+)\n\ndiff --git a/builtin/worktree.c b/builtin/worktree.c\nindex 6dcf7bd..5c4854d 100644\n--- a/builtin/worktree.c\n+++ b/builtin/worktree.c\n@@ -528,6 +528,8 @@ int cmd_worktree(int ac, const char **av, const char *prefix)\n \t\tOPT_END()\n \t};\n \n+\tgit_config(git_default_config, NULL);\n+\n \tif (ac < 2)\n \t\tusage_with_options(worktree_usage, options);\n \tif (!prefix)\n-- \n2.10.0-584-gc9e068c\n\n"},{"id":"302855","messageId":"20160928233047.14313-5-gitster@pobox.com","threadId":"44156","inReplyTo":"20160928233047.14313-1-gitster@pobox.com","subject":"[PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-28T23:30:47Z","receivedAt":"2016-09-28T23:34:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"As Peff said, responding in a thread started by Linus's suggestion\nto raise the default abbreviation to 12 hexdigits:\n\n    I actually think \"12\" might be sane for a long time. That's 48 bits of\n    sha1, so we'd expect a 50% change of a _single_ collision at 2^24, or 16\n    million.  The biggest repository I know about (in number of objects) is\n    the one holding all of the objects for all of the forks of\n    torvalds/linux on GitHub. It's at about 15 million objects.\n\n    Which _seems_ close, but remember that's the size where we expect to see\n    a single collision. They don't become common until much later (I didn't\n    compute an exact number, but Linus's 16x sounds about right). I know\n    that the growth of the kernel isn't really linear, but I think the need\n    to bump to \"13\" might not just be decades, but possibly a century or\n    more.\n\n    So 12 seems reasonable, and the only downside for it (or for \"13\", for\n    that matter) is a few extra bytes. I dunno, maybe people will really\n    hate that, but I have a feeling these are mostly cut-and-pasted anyway.\n\nAnd this does exactly that.\n\nKeep the tests working by explicitly asking for the old 7 hexdigits\nsetting in the fake system-wide configuration file used for tests.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n environment.c        | 2 +-\n t/gitconfig-for-test | 3 +++\n 2 files changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/environment.c b/environment.c\nindex ca72464..25daddb 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -16,7 +16,7 @@ int trust_executable_bit = 1;\n int trust_ctime = 1;\n int check_stat = 1;\n int has_symlinks = 1;\n-int minimum_abbrev = 4, default_abbrev = 7;\n+int minimum_abbrev = 4, default_abbrev = 12;\n int ignore_case;\n int assume_unchanged;\n int prefer_symlink_refs;\ndiff --git a/t/gitconfig-for-test b/t/gitconfig-for-test\nindex 4598885..8c28442 100644\n--- a/t/gitconfig-for-test\n+++ b/t/gitconfig-for-test\n@@ -4,3 +4,6 @@\n ;; [user]\n ;;\tname = A U Thor\n ;;\temail = author@example.com\n+\n+[core]\n+\tabbrev = 7\n-- \n2.10.0-584-gc9e068c\n\n"},{"id":"302856","messageId":"20160928233047.14313-1-gitster@pobox.com","threadId":"44156","inReplyTo":"CA+55aFy0_pwtFOYS1Tmnxipw9ZkRNCQHmoYyegO00pjMiZQfbg@mail.gmail.com","subject":"[PATCH 0/4] raising core.abbrev default to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-28T23:30:43Z","receivedAt":"2016-09-28T23:35:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Per request/suggestion by Linus. \n\nThis took far more effort to keep the existing tests working than\nthe actual change.\n\nJunio C Hamano (4):\n  config: allow customizing /etc/gitconfig location\n  t13xx: do not assume system config is empty\n  worktree: honor configuration variables\n  core.abbrev: raise the default abbreviation to 12 hexdigits\n\n builtin/worktree.c     |  2 ++\n cache.h                |  1 +\n config.c               |  2 ++\n environment.c          |  2 +-\n t/gitconfig-for-test   |  9 +++++++++\n t/t1300-repo-config.sh | 39 ++++++++++++++++++++++++++++-----------\n t/t1308-config-set.sh  |  1 +\n t/test-lib.sh          |  4 ++--\n 8 files changed, 46 insertions(+), 14 deletions(-)\n create mode 100644 t/gitconfig-for-test\n\n-- \n2.10.0-584-gc9e068c\n\n"},{"id":"302857","messageId":"20160928233047.14313-3-gitster@pobox.com","threadId":"44156","inReplyTo":"20160928233047.14313-1-gitster@pobox.com","subject":"[PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-28T23:30:45Z","receivedAt":"2016-09-28T23:35:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Most parts of these two tests want to read from the local\nconfiguration file they prepare and make sure expected names and\nvalues appear with \"git config --list\".\n\nOnce we add custom configuration items that we want to affect the\ntests with globally to t/gitconfig-for-test file, these will start\nseeing the contents from there and break.  Clarify with --local that\nthey only care about the contents from their local configuration.\n\nThe tests for show-origin codepath in \"git config\" however cannot be\ntweaked with \"--local\" etc., because they wants to read also from\n$HOME/.gitconfig and make sure what comes from where.  Disable\nreading from the system-wide config with GIT_CONFIG_NOSYSTEM=1 for\nthese tests.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/t1300-repo-config.sh | 24 +++++++++++++-----------\n t/t1308-config-set.sh  |  1 +\n 2 files changed, 14 insertions(+), 11 deletions(-)\n\ndiff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\nindex 1184f43..b998568 100755\n--- a/t/t1300-repo-config.sh\n+++ b/t/t1300-repo-config.sh\n@@ -341,13 +341,11 @@ version.1.2.3eX.alpha=beta\n EOF\n \n test_expect_success 'working --list' '\n-\tgit config --list > output &&\n+\tgit config --local --list > output &&\n \ttest_cmp expect output\n '\n-cat > expect << EOF\n-EOF\n-\n-test_expect_success '--list without repo produces empty output' '\n+test_expect_success '--list without repo shows only from the global' '\n+\tgit config --system --list >expect &&\n \tgit --git-dir=nonexistent config --list >output &&\n \ttest_cmp expect output\n '\n@@ -360,7 +358,7 @@ version.1.2.3eX.alpha\n EOF\n \n test_expect_success '--name-only --list' '\n-\tgit config --name-only --list >output &&\n+\tgit config --local --name-only --list >output &&\n \ttest_cmp expect output\n '\n \n@@ -370,7 +368,7 @@ nextsection.nonewline wow2 for me\n EOF\n \n test_expect_success '--get-regexp' '\n-\tgit config --get-regexp in >output &&\n+\tgit config --local --get-regexp in >output &&\n \ttest_cmp expect output\n '\n \n@@ -380,7 +378,7 @@ nextsection.nonewline\n EOF\n \n test_expect_success '--name-only --get-regexp' '\n-\tgit config --name-only --get-regexp in >output &&\n+\tgit config --local --name-only --get-regexp in >output &&\n \ttest_cmp expect output\n '\n \n@@ -391,7 +389,7 @@ EOF\n \n test_expect_success '--add' '\n \tgit config --add nextsection.nonewline \"wow4 for you\" &&\n-\tgit config --get-all nextsection.nonewline > output &&\n+\tgit config --local --get-all nextsection.nonewline > output &&\n \ttest_cmp expect output\n '\n \n@@ -935,7 +933,7 @@ section.quotecont=cont;inued\n EOF\n \n test_expect_success 'value continued on next line' '\n-\tgit config --list > result &&\n+\tgit config --local --list > result &&\n \ttest_cmp result expect\n '\n \n@@ -959,7 +957,7 @@ Qsection.sub=section.val4\n Qsection.sub=section.val5Q\n EOF\n test_expect_success '--null --list' '\n-\tgit config --null --list >result.raw &&\n+\tgit config --null --local --list >result.raw &&\n \tnul_to_q <result.raw >result &&\n \techo >>result &&\n \ttest_cmp expect result\n@@ -1264,6 +1262,7 @@ test_expect_success '--show-origin with --list' '\n \t\tfile:.git/../include/relative.include\tuser.relative=include\n \t\tcommand line:\tuser.cmdline=true\n \tEOF\n+\tGIT_CONFIG_NOSYSTEM=1 \\\n \tgit -c user.cmdline=true config --list --show-origin >output &&\n \ttest_cmp expect output\n '\n@@ -1281,6 +1280,7 @@ test_expect_success '--show-origin with --list --null' '\n \t\tincludeQcommand line:Quser.cmdline\n \t\ttrueQ\n \tEOF\n+\tGIT_CONFIG_NOSYSTEM=1 \\\n \tgit -c user.cmdline=true config --null --list --show-origin >output.raw &&\n \tnul_to_q <output.raw >output &&\n \t# The here-doc above adds a newline that the --null output would not\n@@ -1304,6 +1304,7 @@ test_expect_success '--show-origin with --get-regexp' '\n \t\tfile:$HOME/.gitconfig\tuser.global true\n \t\tfile:.git/config\tuser.local true\n \tEOF\n+\tGIT_CONFIG_NOSYSTEM=1 \\\n \tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n \ttest_cmp expect output\n '\n@@ -1312,6 +1313,7 @@ test_expect_success '--show-origin getting a single key' '\n \tcat >expect <<-\\EOF &&\n \t\tfile:.git/config\tlocal\n \tEOF\n+\tGIT_CONFIG_NOSYSTEM=1 \\\n \tgit config --show-origin user.override >output &&\n \ttest_cmp expect output\n '\ndiff --git a/t/t1308-config-set.sh b/t/t1308-config-set.sh\nindex 7655c94..5d5adb1 100755\n--- a/t/t1308-config-set.sh\n+++ b/t/t1308-config-set.sh\n@@ -260,6 +260,7 @@ test_expect_success 'iteration shows correct origins' '\n \tname=\n \tscope=cmdline\n \tEOF\n+\tGIT_CONFIG_NOSYSTEM=1 \\\n \tGIT_CONFIG_PARAMETERS=$cmdline_config test-config iterate >actual &&\n \ttest_cmp expect actual\n '\n-- \n2.10.0-584-gc9e068c\n\n"},{"id":"302858","messageId":"20160928233047.14313-2-gitster@pobox.com","threadId":"44156","inReplyTo":"20160928233047.14313-1-gitster@pobox.com","subject":"[PATCH 1/4] config: allow customizing /etc/gitconfig location","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-28T23:30:44Z","receivedAt":"2016-09-28T23:35:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"With a new environment variable GIT_ETC_GITCONFIG, the users can\nspecify a file that is used instead of /etc/gitconfig to read (and\nwrite) the system-wide configuration.\n\nEarlier, we introduced GIT_CONFIG_NOSYSTEM environment variable\nab88c363 (\"allow suppressing of global and system config\",\n2008-02-06), primarily to protect our tests from random set of\nconfiguration variables the system administrators would put in their\n/etc/gitconfig file.  We can replace the use of this mechanism in\nour tests by pointing GIT_ETC_GITCONFIG at our own instead.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * The next step is to add \"[core]abbrev=7\" to this file and update\n   default_abbrev to 12 in environment.c and see what breaks.  I\n   suspect that \"git worktree list\" would break without my recent\n   patch.  I also know some tests expect \"git config -l\" to show\n   only values they set to their local configuration, which would\n   need to be corrected.  We'll see them in next steps.\n\n cache.h                |  1 +\n config.c               |  2 ++\n t/gitconfig-for-test   |  6 ++++++\n t/t1300-repo-config.sh | 15 +++++++++++++++\n t/test-lib.sh          |  4 ++--\n 5 files changed, 26 insertions(+), 2 deletions(-)\n create mode 100644 t/gitconfig-for-test\n\ndiff --git a/cache.h b/cache.h\nindex b0dae4b..81a07bf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -408,6 +408,7 @@ static inline enum object_type object_type(unsigned int mode)\n #define GIT_NAMESPACE_ENVIRONMENT \"GIT_NAMESPACE\"\n #define GIT_WORK_TREE_ENVIRONMENT \"GIT_WORK_TREE\"\n #define GIT_PREFIX_ENVIRONMENT \"GIT_PREFIX\"\n+#define GIT_ETC_GITCONFIG_ENVIRONMENT \"GIT_ETC_GITCONFIG\"\n #define DEFAULT_GIT_DIR_ENVIRONMENT \".git\"\n #define DB_ENVIRONMENT \"GIT_OBJECT_DIRECTORY\"\n #define INDEX_ENVIRONMENT \"GIT_INDEX_FILE\"\ndiff --git a/config.c b/config.c\nindex 0dfed68..124699b 100644\n--- a/config.c\n+++ b/config.c\n@@ -1253,6 +1253,8 @@ const char *git_etc_gitconfig(void)\n {\n \tstatic const char *system_wide;\n \tif (!system_wide)\n+\t\tsystem_wide = getenv(GIT_ETC_GITCONFIG_ENVIRONMENT);\n+\tif (!system_wide)\n \t\tsystem_wide = system_path(ETC_GITCONFIG);\n \treturn system_wide;\n }\ndiff --git a/t/gitconfig-for-test b/t/gitconfig-for-test\nnew file mode 100644\nindex 0000000..4598885\n--- /dev/null\n+++ b/t/gitconfig-for-test\n@@ -0,0 +1,6 @@\n+;; This file is used as if it were /etc/gitconfig while running the\n+;; test scripts in this directory.\n+;;\n+;; [user]\n+;;\tname = A U Thor\n+;;\temail = author@example.com\ndiff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\nindex 923bfc5..1184f43 100755\n--- a/t/t1300-repo-config.sh\n+++ b/t/t1300-repo-config.sh\n@@ -1372,4 +1372,19 @@ test_expect_success !MINGW '--show-origin blob ref' '\n \ttest_cmp expect output\n '\n \n+test_expect_success 'system-wide configuration' '\n+\tsystem=\"$TRASH_DIRECTORY/system-wide\" &&\n+\t>\"$system\" &&\n+\tgit config -f \"$system\" --add frotz.nitfol xyzzy &&\n+\n+\tgit config -f \"$system\" frotz.nitfol >expect &&\n+\tGIT_ETC_GITCONFIG=\"$system\" \\\n+\tgit config --system frotz.nitfol >actual &&\n+\n+\tGIT_ETC_GITCONFIG=\"$system\" \\\n+\tgit config --system --replace-all frotz.nitfol blorb &&\n+\techo blorb >expect &&\n+\tGIT_ETC_GITCONFIG=\"$system\" git config --system frotz.nitfol >actual\n+'\n+\n test_done\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex ac56512..6803212 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -851,9 +851,9 @@ else # normal case, use ../bin-wrappers only unless $with_dashes:\n \tfi\n fi\n GIT_TEMPLATE_DIR=\"$GIT_BUILD_DIR\"/templates/blt\n-GIT_CONFIG_NOSYSTEM=1\n+GIT_ETC_GITCONFIG=\"$GIT_BUILD_DIR/t/gitconfig-for-test\"\n GIT_ATTR_NOSYSTEM=1\n-export PATH GIT_EXEC_PATH GIT_TEMPLATE_DIR GIT_CONFIG_NOSYSTEM GIT_ATTR_NOSYSTEM\n+export PATH GIT_EXEC_PATH GIT_TEMPLATE_DIR GIT_ETC_GITCONFIG GIT_ATTR_NOSYSTEM\n \n if test -z \"$GIT_TEST_CMP\"\n then\n-- \n2.10.0-584-gc9e068c\n\n"},{"id":"302859","messageId":"20160929024400.22605-1-szeder@ira.uka.de","threadId":"44156","inReplyTo":"20160928233047.14313-5-gitster@pobox.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"SZEDER Gábor","fromEmail":"szeder@ira.uka.de","sentAt":"2016-09-29T02:44:00Z","receivedAt":"2016-09-29T02:47:17Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"> As Peff said, responding in a thread started by Linus's suggestion\n> to raise the default abbreviation to 12 hexdigits:\n> \n>     I actually think \"12\" might be sane for a long time. That's 48 bits of\n>     sha1, so we'd expect a 50% change of a _single_ collision at 2^24, or 16\n\ns/change/chance/\n\nI know it's quoted, but still.\n\n>     million.  The biggest repository I know about (in number of objects) is\n>     the one holding all of the objects for all of the forks of\n>     torvalds/linux on GitHub. It's at about 15 million objects.\n> \n>     Which _seems_ close, but remember that's the size where we expect to see\n>     a single collision. They don't become common until much later (I didn't\n>     compute an exact number, but Linus's 16x sounds about right). I know\n>     that the growth of the kernel isn't really linear, but I think the need\n>     to bump to \"13\" might not just be decades, but possibly a century or\n>     more.\n> \n>     So 12 seems reasonable, and the only downside for it (or for \"13\", for\n>     that matter) is a few extra bytes. I dunno, maybe people will really\n>     hate that, but I have a feeling these are mostly cut-and-pasted anyway.\n\nI for one raise my hand in protest...\n\n\"few extra bytes\" is not the only downside, and it's not at all about\nhow many characters are copy-and-pasted.  In my opinion it's much more\nimportant that this change wastes 5 columns worth of valuable screen\nreal estate e.g. for 'git blame' or 'git log --oneline' in projects\nthat don't need it and certainly won't ever need it.\n\nSure, users working on smaller repos are free to reset core.abbrev to\nits original value.  I don't have any numbers, of course, but I\nsuspect that there are many more smaller repos out there that this\nchange will affect disadvantageously, than there are large repos for\nwhich it's beneficial.\n\n\n> And this does exactly that.\n> \n> Keep the tests working by explicitly asking for the old 7 hexdigits\n> setting in the fake system-wide configuration file used for tests.\n> \n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n"},{"id":"302860","messageId":"147512682170.7989.11263315726387435240@typhoon","threadId":"44156","inReplyTo":"20160929024400.22605-1-szeder@ira.uka.de","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Lukas Fleischer","fromEmail":"lfleischer@lfos.de","sentAt":"2016-09-29T05:27:01Z","receivedAt":"2016-09-29T05:33:50Z","isPatch":true,"sender":{"key":"lfleischer@lfos.de","avatar":"https://avatars.githubusercontent.com/u/5530842?v=4"},"body":"On Thu, 29 Sep 2016 at 04:44:00, SZEDER Gábor wrote:\n> I for one raise my hand in protest...\n> \n> \"few extra bytes\" is not the only downside, and it's not at all about\n> how many characters are copy-and-pasted.  In my opinion it's much more\n> important that this change wastes 5 columns worth of valuable screen\n> real estate e.g. for 'git blame' or 'git log --oneline' in projects\n> that don't need it and certainly won't ever need it.\n> \n> Sure, users working on smaller repos are free to reset core.abbrev to\n> its original value.  I don't have any numbers, of course, but I\n> suspect that there are many more smaller repos out there that this\n> change will affect disadvantageously, than there are large repos for\n> which it's beneficial.\n\nI know this suggestion comes a bit late but would it make sense to let\nthe repository owner overwrite the core.abbrev setting?\n\nOne possible way to implement this would be adding .gitconfig support to\nrepositories with a very limited set of whitelisted variables allowed in\nthere (could be core.abbrev only to begin with). Or some entirely\nseparate mechanism like .gitignore.\n\nWith such a mechanism, we could keep the default of 7 which works fine\nfor most projects. Linus could bump the default to 12 for linux.git. If\nsome users are not happy with that, they can still overwrite it in their\nlocal Git config. Anybody starting a project could change the initial\nvalue to a suitable value in one of the first commits -- provided they\nalready have an idea how much the project will grow. That way, hashes\nwill be \"long enough\" even for early commits, before any heuristics\ncould guess that the project would become large.\n\nOpinions?\n\nRegards,\nLukas\n"},{"id":"302862","messageId":"ae9dbf3b-4190-8145-a59f-0d578067032a@kdbg.org","threadId":"44156","inReplyTo":"20160928233047.14313-5-gitster@pobox.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2016-09-29T05:58:49Z","receivedAt":"2016-09-29T05:59:00Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 29.09.2016 um 01:30 schrieb Junio C Hamano:\n> As Peff said, responding in a thread started by Linus's suggestion\n> to raise the default abbreviation to 12 hexdigits:\n\nThis is waayy too large for a new default. The vast majority of \nrepositories is smallish. For those, the long sequences of hex digits \nare an uglification that is almost unbearable.\n\nI know that kernel developers are important, but their importance has \nlong been outnumbered by the anonymous and silent masses of users.\n\nPersonally, I use 8 digits just because it is a \"rounder\" number than 7, \nbut in all of my repositories 7 would still work just as well.\n\n-- Hannes\n\n"},{"id":"302870","messageId":"20160929090108.hf2jzfcvbcsfaxw7@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160928233047.14313-3-gitster@pobox.com","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T09:01:09Z","receivedAt":"2016-09-29T09:01:39Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 28, 2016 at 04:30:45PM -0700, Junio C Hamano wrote:\n\n> The tests for show-origin codepath in \"git config\" however cannot be\n> tweaked with \"--local\" etc., because they wants to read also from\n> $HOME/.gitconfig and make sure what comes from where.  Disable\n> reading from the system-wide config with GIT_CONFIG_NOSYSTEM=1 for\n> these tests.\n\nI think anytime you would use GIT_CONFIG_NOSYSTEM over --local, it is an\nindication that the test is trying to check how multiple sources\ninteract. And the right thing to do for them is to set GIT_ETC_GITCONFIG\nto some known quantity. We just couldn't do that before, so we skipped\nit.\n\nIOW, something like the patch below (on top of yours). Note that the\ncommands that are doing a \"--get\" and not a \"--list\" don't actually seem\nto need either (because they are getting the values out of the local\nfile anyway), so we could drop the setting of GIT_ETC_GITCONFIG from\nthem entirely.\n\ndiff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\nindex b998568..d2476a8 100755\n--- a/t/t1300-repo-config.sh\n+++ b/t/t1300-repo-config.sh\n@@ -1234,6 +1234,11 @@ test_expect_success 'set up --show-origin tests' '\n \t\t[user]\n \t\t\trelative = include\n \tEOF\n+\tcat >\"$HOME\"/etc-gitconfig <<-\\EOF &&\n+\t\t[user]\n+\t\t\tsystem = true\n+\t\t\toverride = system\n+\tEOF\n \tcat >\"$HOME\"/.gitconfig <<-EOF &&\n \t\t[user]\n \t\t\tglobal = true\n@@ -1252,6 +1257,8 @@ test_expect_success 'set up --show-origin tests' '\n \n test_expect_success '--show-origin with --list' '\n \tcat >expect <<-EOF &&\n+\t\tfile:$HOME/etc-gitconfig\tuser.system=true\n+\t\tfile:$HOME/etc-gitconfig\tuser.override=system\n \t\tfile:$HOME/.gitconfig\tuser.global=true\n \t\tfile:$HOME/.gitconfig\tuser.override=global\n \t\tfile:$HOME/.gitconfig\tinclude.path=$INCLUDE_DIR/absolute.include\n@@ -1262,14 +1269,16 @@ test_expect_success '--show-origin with --list' '\n \t\tfile:.git/../include/relative.include\tuser.relative=include\n \t\tcommand line:\tuser.cmdline=true\n \tEOF\n-\tGIT_CONFIG_NOSYSTEM=1 \\\n+\tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n \tgit -c user.cmdline=true config --list --show-origin >output &&\n \ttest_cmp expect output\n '\n \n test_expect_success '--show-origin with --list --null' '\n \tcat >expect <<-EOF &&\n-\t\tfile:$HOME/.gitconfigQuser.global\n+\t\tfile:$HOME/etc-gitconfigQuser.system\n+\t\ttrueQfile:$HOME/etc-gitconfigQuser.override\n+\t\tsystemQfile:$HOME/.gitconfigQuser.global\n \t\ttrueQfile:$HOME/.gitconfigQuser.override\n \t\tglobalQfile:$HOME/.gitconfigQinclude.path\n \t\t$INCLUDE_DIR/absolute.includeQfile:$INCLUDE_DIR/absolute.includeQuser.absolute\n@@ -1280,7 +1289,7 @@ test_expect_success '--show-origin with --list --null' '\n \t\tincludeQcommand line:Quser.cmdline\n \t\ttrueQ\n \tEOF\n-\tGIT_CONFIG_NOSYSTEM=1 \\\n+\tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n \tgit -c user.cmdline=true config --null --list --show-origin >output.raw &&\n \tnul_to_q <output.raw >output &&\n \t# The here-doc above adds a newline that the --null output would not\n@@ -1304,7 +1313,7 @@ test_expect_success '--show-origin with --get-regexp' '\n \t\tfile:$HOME/.gitconfig\tuser.global true\n \t\tfile:.git/config\tuser.local true\n \tEOF\n-\tGIT_CONFIG_NOSYSTEM=1 \\\n+\tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n \tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n \ttest_cmp expect output\n '\n@@ -1313,7 +1322,7 @@ test_expect_success '--show-origin getting a single key' '\n \tcat >expect <<-\\EOF &&\n \t\tfile:.git/config\tlocal\n \tEOF\n-\tGIT_CONFIG_NOSYSTEM=1 \\\n+\tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n \tgit config --show-origin user.override >output &&\n \ttest_cmp expect output\n '\n"},{"id":"302871","messageId":"20160929091509.2n4mdrevwxechqol@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160929024400.22605-1-szeder@ira.uka.de","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T09:15:10Z","receivedAt":"2016-09-29T09:15:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 04:44:00AM +0200, SZEDER Gábor wrote:\n\n> >     So 12 seems reasonable, and the only downside for it (or for \"13\", for\n> >     that matter) is a few extra bytes. I dunno, maybe people will really\n> >     hate that, but I have a feeling these are mostly cut-and-pasted anyway.\n> \n> I for one raise my hand in protest...\n> \n> \"few extra bytes\" is not the only downside, and it's not at all about\n> how many characters are copy-and-pasted.  In my opinion it's much more\n> important that this change wastes 5 columns worth of valuable screen\n> real estate e.g. for 'git blame' or 'git log --oneline' in projects\n> that don't need it and certainly won't ever need it.\n\nTrue. The core of the issue is that we really only care about this\nminimum length when _storing_ an abbreviation, but we don't know when\nthe user is just looking at it in the moment, and when they are going to\nstick it in a commit message, email, or bug tracker.\n\nIn an ideal world, anybody who was about to store it would run \"git\ndescribe\" or something to come up with some canonical reference format.\nAnd we could just bump the default minimum there. Personally, I almost\nexclusively cite commits as the output of:\n\n  git log -1 --pretty='tformat:%h (%s, %ad)' --date=short\n\nand I'd be fine to stick \"--abbrev=12\" in there for future-proofing. But\nI don't know what the kernel or other projects do.\n\nI'd also be curious to know if the patch I sent in [1] to more\naggressively prefer commits would make this less of an issue, and people\nwouldn't care as much about using longer hashes in the first place. So\none option is to merge that (and possibly even make it the default) and\nsee if people still care in 6 months.\n\n-Peff\n\n[1] http://public-inbox.org/git/20160927123801.3bpdg3hap3kzzfmv@sigill.intra.peff.net/\n"},{"id":"302872","messageId":"20160929092204.eod2cvtrqg5whu6h@sigill.intra.peff.net","threadId":"44156","inReplyTo":"147512682170.7989.11263315726387435240@typhoon","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T09:22:04Z","receivedAt":"2016-09-29T09:22:26Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 07:27:01AM +0200, Lukas Fleischer wrote:\n\n> > Sure, users working on smaller repos are free to reset core.abbrev to\n> > its original value.  I don't have any numbers, of course, but I\n> > suspect that there are many more smaller repos out there that this\n> > change will affect disadvantageously, than there are large repos for\n> > which it's beneficial.\n> \n> I know this suggestion comes a bit late but would it make sense to let\n> the repository owner overwrite the core.abbrev setting?\n> \n> One possible way to implement this would be adding .gitconfig support to\n> repositories with a very limited set of whitelisted variables allowed in\n> there (could be core.abbrev only to begin with). Or some entirely\n> separate mechanism like .gitignore.\n\nThe suggestion for versioned repository-level config comes up from time\nto time; you can find other instances in the list archive. Usually the\nbiggest issue is that usually nobody comes up with a good example of\nsomething that the project would actually want to set. Setting\n\"core.abbrev\" at least seems plausible.\n\nThough...\n\n> With such a mechanism, we could keep the default of 7 which works fine\n> for most projects. Linus could bump the default to 12 for linux.git. If\n> some users are not happy with that, they can still overwrite it in their\n> local Git config. Anybody starting a project could change the initial\n> value to a suitable value in one of the first commits -- provided they\n> already have an idea how much the project will grow. That way, hashes\n> will be \"long enough\" even for early commits, before any heuristics\n> could guess that the project would become large.\n\nI wonder if in practice we would do just as well to size default_abbrev\ndynamically based on the number of objects. That doesn't help projects\nwhich are just starting, but will eventually grow gigantic.  But I doubt\nthat most projects would have the foresight to preemptively set\ncore.abbrev. And that would at least reduce the impact as the project\n_does_ get big.\n\n-Peff\n"},{"id":"302873","messageId":"20160929092556.3kkmdd6uprss76zx@sigill.intra.peff.net","threadId":"44156","inReplyTo":"20160928233047.14313-5-gitster@pobox.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T09:25:56Z","receivedAt":"2016-09-29T09:26:04Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 28, 2016 at 04:30:47PM -0700, Junio C Hamano wrote:\n\n> As Peff said, responding in a thread started by Linus's suggestion\n> to raise the default abbreviation to 12 hexdigits:\n> \n>     I actually think \"12\" might be sane for a long time. That's 48 bits of\n>     sha1, so we'd expect a 50% change of a _single_ collision at 2^24, or 16\n>     million.  The biggest repository I know about (in number of objects) is\n>     the one holding all of the objects for all of the forks of\n>     torvalds/linux on GitHub. It's at about 15 million objects.\n> \n>     Which _seems_ close, but remember that's the size where we expect to see\n>     a single collision. They don't become common until much later (I didn't\n>     compute an exact number, but Linus's 16x sounds about right). I know\n>     that the growth of the kernel isn't really linear, but I think the need\n>     to bump to \"13\" might not just be decades, but possibly a century or\n>     more.\n> \n>     So 12 seems reasonable, and the only downside for it (or for \"13\", for\n>     that matter) is a few extra bytes. I dunno, maybe people will really\n>     hate that, but I have a feeling these are mostly cut-and-pasted anyway.\n\nI am not sure my quote is a good rationale for this bump. It was meant\nto be a rationale that \"12\" is big enough, but the \"I dunno\" at the end\nkind of glosses over the downsides.\n\n-Peff\n"},{"id":"302875","messageId":"f239b2eb-d122-9c4b-187b-fbd40a94bcf4@gmail.com","threadId":"44156","inReplyTo":"20160928233047.14313-2-gitster@pobox.com","subject":"Re: [PATCH 1/4] config: allow customizing /etc/gitconfig location","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2016-09-29T09:53:10Z","receivedAt":"2016-09-29T09:53:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 29.09.2016 o 01:30, Junio C Hamano pisze:\n> With a new environment variable GIT_ETC_GITCONFIG, the users can\n> specify a file that is used instead of /etc/gitconfig to read (and\n> write) the system-wide configuration.\n\nWhy it is named GIT_ETC_GITCONFIG (which is Unix-ism), and not\nGIT_CONFIG_SYSTEM / GIT_CONFIG_SYSTEM_PATH, that is something\nOS-neutral?\n\n-- \nJakub Narębski\n\n"},{"id":"302876","messageId":"vpq7f9v9g5l.fsf@anie.imag.fr","threadId":"44156","inReplyTo":"20160929091509.2n4mdrevwxechqol@sigill.intra.peff.net","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-09-29T10:03:34Z","receivedAt":"2016-09-29T10:04:05Z","isPatch":true,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Thu, Sep 29, 2016 at 04:44:00AM +0200, SZEDER Gábor wrote:\n>\n>> >     So 12 seems reasonable, and the only downside for it (or for \"13\", for\n>> >     that matter) is a few extra bytes. I dunno, maybe people will really\n>> >     hate that, but I have a feeling these are mostly cut-and-pasted anyway.\n>> \n>> I for one raise my hand in protest...\n>> \n>> \"few extra bytes\" is not the only downside, and it's not at all about\n>> how many characters are copy-and-pasted.  In my opinion it's much more\n>> important that this change wastes 5 columns worth of valuable screen\n>> real estate e.g. for 'git blame' or 'git log --oneline' in projects\n>> that don't need it and certainly won't ever need it.\n>\n> True. The core of the issue is that we really only care about this\n> minimum length when _storing_ an abbreviation, but we don't know when\n> the user is just looking at it in the moment, and when they are going to\n> stick it in a commit message, email, or bug tracker.\n\nPerhaps a compromise would be to adapt the length to the size of the\nproject _and_ keep a huge margin. So, essentially, we'd have small\nprojects stick to the 7 characters, and very quickly bump to 12.\n\nSo, for a fast-growing project, there would be a short window at the\nbeginning of the project where people could cut-and-past short hashes.\nOTOH, small projects could keep these few columns of screen real-estate.\n\nThat said, I can certainly live without these 5 columns, don't take my\nmessage as an objection to setting to 12 right away.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"302879","messageId":"2242637D-4C3B-4AF2-8BE4-823B3E1745D5@gmail.com","threadId":"44156","inReplyTo":"20160926120036.mqs435a36njeihq6@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Kyle J. McKay","fromEmail":"mackyle@gmail.com","sentAt":"2016-09-29T11:46:19Z","receivedAt":"2016-09-29T11:46:39Z","isPatch":true,"sender":{"key":"mackyle@gmail.com","avatar":"https://avatars.githubusercontent.com/u/813346?v=4"},"body":"On Sep 26, 2016, at 05:00, Jeff King wrote:\n\n>  $ git rev-parse b2e1\n>  error: short SHA1 b2e1 is ambiguous\n>  hint: The candidates are:\n>  hint:   b2e1196 tag v2.8.0-rc1\n>  hint:   b2e11d1 tree\n>  hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit- \n> options'\n>  hint:   b2e1759 blob\n>  hint:   b2e18954 blob\n>  hint:   b2e1895c blob\n>  fatal: ambiguous argument 'b2e1': unknown revision or path not in  \n> the working tree.\n>  Use '--' to separate paths from revisions, like this:\n>  'git <command> [<revision>...] -- [<file>...]'\n\nThis hint: information is excellent.  There needs to be a way to show  \nit on demand.\n\n$ git rev-parse --disambiguate=b2e1\nb2e11962c5e6a9c81aa712c751c83a743fd4f384\nb2e11d1bb40c5f81a2f4e37b9f9a60ec7474eeab\nb2e163272c01aca4aee4684f5c683ba341c1953d\nb2e18954c03ff502053cb74d142faab7d2a8dacb\nb2e1895ca92ec2037349d88b945ba64ebf16d62d\n\nNot nearly so helpful, but the operation of --disambiguate cannot be  \nchanged without breaking current scripts.\n\nCan your excellent \"hint:\" output above be attached to the -- \ndisambiguate option somehow, please.  Something like this perhaps:\n\n$ git rev-parse --disambiguate-list=b2e1\nb2e1196 tag v2.8.0-rc1\nb2e11d1 tree\nb2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\nb2e1759 blob\nb2e18954 blob\nb2e1895c blob\n\nAny option name will do, --disambiguate-verbose, --disambiguate- \nextended, --disambiguate-long, --disambiguate-log, --disambiguate- \nhelp, --disambiguate-show-me-something-useful-to-humans-not-scripts ...\n\n--Kyle\n"},{"id":"302881","messageId":"20160929145210.Horde.g9CV6T-OpWCeozsbBpTood-@webmail.informatik.kit.edu","threadId":"44156","inReplyTo":"20160929091509.2n4mdrevwxechqol@sigill.intra.peff.net","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"SZEDER Gábor","fromEmail":"szeder@ira.uka.de","sentAt":"2016-09-29T12:52:10Z","receivedAt":"2016-09-29T12:52:27Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"\nQuoting Jeff King <peff@peff.net>:\n\n> On Thu, Sep 29, 2016 at 04:44:00AM +0200, SZEDER Gábor wrote:\n>\n>> >     So 12 seems reasonable, and the only downside for it (or for \"13\", for\n>> >     that matter) is a few extra bytes. I dunno, maybe people will really\n>> >     hate that, but I have a feeling these are mostly  \n>> cut-and-pasted anyway.\n>>\n>> I for one raise my hand in protest...\n>>\n>> \"few extra bytes\" is not the only downside, and it's not at all about\n>> how many characters are copy-and-pasted.  In my opinion it's much more\n>> important that this change wastes 5 columns worth of valuable screen\n>> real estate e.g. for 'git blame' or 'git log --oneline' in projects\n>> that don't need it and certainly won't ever need it.\n>\n> True. The core of the issue is that we really only care about this\n> minimum length when _storing_ an abbreviation, but we don't know when\n> the user is just looking at it in the moment, and when they are going to\n> stick it in a commit message, email, or bug tracker.\n>\n> In an ideal world, anybody who was about to store it would run \"git\n> describe\" or something to come up with some canonical reference format.\n> And we could just bump the default minimum there. Personally, I almost\n> exclusively cite commits as the output of:\n>\n>   git log -1 --pretty='tformat:%h (%s, %ad)' --date=short\n\nInteresting, I have a pretty format alias that looks almost like this,\nexcept that I carry a patch locally allowing me to say %as for short\ndate format :)\n\nWhat I sometimes wished for is a pretty format specifier for 'git\ndescribe --contains', which would make it convenient to cite commits\nlike this: v0.99~954 (Initial revision of \"git\", the information manager\nfrom hell, 2005-04-07).  It's better than the abbreviated object name,\nbecause it will stay unique, assuming that the chosen tag is never\ndeleted, and it carries extra information for humans (the first release\ncontaining the referenced commit), while the abbreviated object name is\ncompletely meaningless.\n\nThe obvious drawback that makes it a non-solution for the problem at\nhand is that this format can only refer to commits that are reachable\nfrom a tag and can't be used for commits that are descendants of the\nmost recent tag, e.g. when fixing a bug introduced after the last\nrelease.  Oh, and the user has to fetch the tag first to be able to\nmake sense of such a reference.\n\n> and I'd be fine to stick \"--abbrev=12\" in there for future-proofing. But\n> I don't know what the kernel or other projects do.\n>\n> I'd also be curious to know if the patch I sent in [1] to more\n> aggressively prefer commits would make this less of an issue, and people\n> wouldn't care as much about using longer hashes in the first place. So\n> one option is to merge that (and possibly even make it the default) and\n> see if people still care in 6 months.\n>\n> -Peff\n>\n> [1]  \n> http://public-inbox.org/git/20160927123801.3bpdg3hap3kzzfmv@sigill.intra.peff.net/\n\n\n"},{"id":"302883","messageId":"80D70A8B-0EB6-4118-8193-8D113621C57D@gmail.com","threadId":"44156","inReplyTo":"xmqq37knwcf4.fsf@gitster.mtv.corp.google.com","subject":"Re: Changing the default for \"core.abbrev\"?","fromName":"Kyle J. McKay","fromEmail":"mackyle@gmail.com","sentAt":"2016-09-29T13:01:44Z","receivedAt":"2016-09-29T13:01:56Z","isPatch":false,"sender":{"key":"mackyle@gmail.com","avatar":"https://avatars.githubusercontent.com/u/813346?v=4"},"body":"On Sep 25, 2016, at 18:39, Linus Torvalds wrote:\n\n> The kernel, these days, is at roughly 5 million objects, and while the\n> seven hex digits are still often enough for uniqueness (and git will\n> always add digits *until* it is unique), it's long been at the point\n> where I tell people to do\n>\n>    git config --global core.abbrev 12\n>\n> because even though git will extend the seven hex digits until the\n> object name is unique, that only reflects the *current* situation in\n> the repository. With 5 million objects and a very healthy growth rate,\n> a 7-8 hex digit number that is unique today is not necessarily unique\n> a month or two from now, and then it gets annoying when a commit\n> message has a short git ID that is no longer unique when you go back\n> and try to figure out what went wrong in that commit.\n\nOn Sep 25, 2016, at 20:46, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> I can just keep reminding kernel maintainers and developers to update\n>> their git config, but maybe it would be a good idea to just admit  \n>> that\n>> the defaults picked in 2005 weren't necessarily the best ones\n>> possible, and those could be bumped up a bit?\n>\n> I am not quite sure how good any new default would be, though.  Just\n> like any timeout is not long enough for somebody, growing projects\n> will eventually hit whatever abbreviation length they start with.\n\nThis made me curious what the situation is really like.  So I crunched  \nsome data.\n\nUsing a recent clone of $korg/torvalds/linux:\n\n$ git rev-parse --verify d597639e203\nerror: short SHA1 d597639e203 is ambiguous.\nfatal: Needed a single revision\n\nSo the kernel already has 11-character \"short\" SHA1s that are  \nambiguous.  Is a core.abbrev setting of 12 really good enough?\n\nHere are the stats on the kernel's repository:\n\nAmbiguous length 11 (but not at length 12) info:\n   prefixes:       2\n                   0 (with 1 or more commit disambiguations)\n\nAmbiguous length 10 (but not at length 11) info:\n   prefixes:      12\n                   3 (with 1 or more commit disambiguations)\n                   0 (with 2 or more commit disambiguations)\n\nAmbiguous length 9 (but not at length 10) info:\n   prefixes:     186\n                  43 (with 1 or more commit disambiguations)\n                   1 (with 2 or more commit disambiguations)\n                   0 (with 3 or more disambiguations)\n\nAmbiguous length 8 (but not at length 9) info:\n   prefixes:    2723\n                 651 (with 1 or more commit disambiguations)\n                  40 (with 2 or more commit disambiguations)\n                   1 (with 3 or more disambiguations)\n   maxambig:       3 (there is 1 of them)\n\nAmbiguous length 7 (but not at length 8) info:\n   prefixes:   41864\n                9842 (with 1 or more commit disambiguations)\n                 680 (with 2 or more commit disambiguations)\n                 299 (with 3 or more disambiguations)\n   maxambig:       3 (there are 299 of them)\n\nThe \"maxambig\" value is the maximum number of disambiguations for any  \nsingle prefix at that prefix length.  So for prefixes of length 7  \nthere are 299 that disambiguate into 3 objects.\n\nJust out of curiosity, generating stats on the Git repository gives:\n\nAmbiguous length 8 (but not at length 9) info:\n   prefixes:       7\n                   3 (with 1 or more commit disambiguations)\n                   2 (with 2 or more commit disambiguations)\n                   0 (with 3 or more disambiguations)\n\nAmbiguous length 7 (but not at length 8) info:\n   prefixes:      87\n                  36 (with 1 or more commit disambiguations)\n                   3 (with 2 or more commit disambiguations)\n                   0 (with 3 or more disambiguations)\n\nRunning the stats on $github/gitster/git produces some ambiguous  \nlength 9 prefixes (one of which contains a commit disambiguation).\n\n--Kyle\n"},{"id":"302884","messageId":"841D4FC2-9673-486A-8D94-8967188CCC60@gmail.com","threadId":"44156","inReplyTo":"CA+55aFyfvvqq1c=hZcuL-yPavp2tjzx8r3bFJnMY7DAE7YcB=Q@mail.gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Kyle J. McKay","fromEmail":"mackyle@gmail.com","sentAt":"2016-09-29T13:01:51Z","receivedAt":"2016-09-29T13:02:05Z","isPatch":true,"sender":{"key":"mackyle@gmail.com","avatar":"https://avatars.githubusercontent.com/u/813346?v=4"},"body":"On Sep 26, 2016, at 09:36, Linus Torvalds wrote:\n\n> On Mon, Sep 26, 2016 at 5:00 AM, Jeff King <peff@peff.net> wrote:\n>>\n>> This patch teaches get_short_sha1() to list the sha1s of the\n>> objects it found, along with a few bits of information that\n>> may help the user decide which one they meant.\n>\n> This looks very good to me, but I wonder if it couldn't be even more  \n> aggressive.\n>\n> In particular, the only hashes that most people ever use in short form\n> are commit hashes. Those are the ones you'd use in normal human\n> interactions to point to something happening.\n>\n> So when the disambiguation notices that there is ambiguity, but there\n> is only _one_ commit, maybe it should just have an aggressive mode\n> that says \"use that as if it wasn't ambiguous\".\n\nIf you have this:\n\nfaa23ec9b437812ce2fc9a5b3d59418d672debc1 refs/heads/ambig\n7f40afe646fa3f8a0f361b6f567d8f7d7a184c10 refs/tags/ambig\n\nand you do this:\n\n$ git rev-parse ambig\nwarning: refname 'ambig' is ambiguous.\n7f40afe646fa3f8a0f361b6f567d8f7d7a184c10\n\nGit automatically prefers the tag over the branch, but it does spit  \nout a warning.\n\n> And then have an explicit command (or flag) to do disambiguation for\n> when you explicitly want it.\n\nI think you don't even need that.  Git already does disambiguation for  \nref names, picks one and spits out a warning.\n\nWhy not do the same for short hash names when it makes sense?\n\n> Rationale: you'd never care about short forms for tags. You'd just use\n> the tag name. And while blob ID's certainly show up in short form in\n> diff output (in the \"index\" line), very few people will use them. And\n> tree hashes are basically never seen outside of any plumbing commands\n> and then seldom in shortened form.\n>\n> So I think it would make sense to default to a mode that just picks\n> the commit hash if there is only one such hash. Sure, some command\n> might want a \"treeish\", but a commit is still more likely than a tree\n> or a tag.\n>\n> But regardless, this series looks like a good thing.\n\nI like it too.\n\nBut perhaps it makes sense to actually pick one if there's only one  \ndisambiguation of the type you're looking for.\n\nFor example given:\n\n235234a blob\n2352347 tag\n235234f tree\n2352340 commit\n\nIf you are doing \"git cat-file blob 235234\" it should pick the blob  \nand spit out a warning (and similarly for other cat-file types).  But  \n\"git cat-file -p 235234\" would give the fatal error with the  \ndisambiguation hints because it wants type \"any\".\n\nIf you are doing \"git show 235234\" it should pick the tag (if it peels  \nto a committish) because Git has already set a precedent of preferring  \ntags over commits when it disambiguates ref names and otherwise pick  \nthe commit.\n\nLets consider this approach using the stats for the Linux kernel:\n\n> Ambiguous prefix length 7 counts:\n>   prefixes:   44733\n>    objects:   89766\n>\n> Ambiguous length 11 (but not at length 12) info:\n>   prefixes:       2\n>                   0 (with 1 or more commit disambiguations)\n>\n> Ambiguous length 10 (but not at length 11) info:\n>   prefixes:      12\n>                   3 (with 1 or more commit disambiguations)\n>                   0 (with 2 or more commit disambiguations)\n>\n> Ambiguous length 9 (but not at length 10) info:\n>   prefixes:     186\n>                  43 (with 1 or more commit disambiguations)\n>                   1 (with 2 or more commit disambiguations)\n>\n> Ambiguous length 8 (but not at length 9) info:\n>   prefixes:    2723\n>                 651 (with 1 or more commit disambiguations)\n>                  40 (with 2 or more commit disambiguations)\n>\n> Ambiguous length 7 (but not at length 8) info:\n>   prefixes:   41864\n>                9842 (with 1 or more commit disambiguations)\n>                 680 (with 2 or more commit disambiguations)\n\nOf the 44733 ambiguous length 7 prefixes, only about 10539 of them  \ndisambiguate into one or more commit objects.\n\nBut if we apply the \"spit a warning and prefer a commit object if  \nthere's only one and you're looking for a committish\" rule, that drops  \nthe number from 10539 to about 721.  In other words, only about 7% of  \nthe previously ambiguous short commit SHA1 prefixes would continue to  \nbe ambiguous at length 7.  In fact it almost makes a prefix length of  \n9 good enough, there's just the one at length 9 that disambiguates  \ninto more than one commit (45f014c52).\n\n--Kyle\n"},{"id":"302885","messageId":"20160929130322.562ng4t2ktk6qzok@sigill.intra.peff.net","threadId":"44156","inReplyTo":"2242637D-4C3B-4AF2-8BE4-823B3E1745D5@gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T13:03:22Z","receivedAt":"2016-09-29T13:03:31Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 04:46:19AM -0700, Kyle J. McKay wrote:\n\n> This hint: information is excellent.  There needs to be a way to show it on\n> demand.\n> \n> $ git rev-parse --disambiguate=b2e1\n> b2e11962c5e6a9c81aa712c751c83a743fd4f384\n> b2e11d1bb40c5f81a2f4e37b9f9a60ec7474eeab\n> b2e163272c01aca4aee4684f5c683ba341c1953d\n> b2e18954c03ff502053cb74d142faab7d2a8dacb\n> b2e1895ca92ec2037349d88b945ba64ebf16d62d\n> \n> Not nearly so helpful, but the operation of --disambiguate cannot be changed\n> without breaking current scripts.\n> \n> Can your excellent \"hint:\" output above be attached to the --disambiguate\n> option somehow, please.  Something like this perhaps:\n> \n> $ git rev-parse --disambiguate-list=b2e1\n> b2e1196 tag v2.8.0-rc1\n> b2e11d1 tree\n> b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n> b2e1759 blob\n> b2e18954 blob\n> b2e1895c blob\n\nI think the \"right\" way to do this is pipe the list of sha1s into\nanother git commit which can format them however you want.\nUnfortunately, there isn't a single command that does a great job:\n\n  - \"cat-file --batch-check\" can show you the sha1 and type, but it\n    won't abbreviate sha1s, and it won't show you commit/tag information\n\n  - \"log --stdin --no-walk\" will format the commit however you like, but\n    skips the trees and blobs entirely, and the tag can only be seen via\n    \"%d\"\n\n  - \"for-each-ref\" has flexible formatting, too, but wants to format\n    refs, not objects (and doesn't read from stdin).\n\nIMHO that is a sign that our formatting tools aren't as good as they\ncould be (I think the right tool is cat-file, but it should be able to\ndo all of the formatting that the other commands can do).\n\nOf course if you really just want human-readable output, then:\n\n  $ git cat-file -e b2e1\n  error: short SHA1 b2e1 is ambiguous\n  hint: The candidates are:\n  hint:   b2e1196 tag v2.8.0-rc1\n  hint:   b2e11d1 tree\n  hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n  hint:   b2e1759 blob\n  hint:   b2e18954 blob\n  hint:   b2e1895c blob\n  fatal: Not a valid object name b2e1\n\nis pretty easy.\n\nThat being said, I don't mind if somebody wanted to do a rev-parse\noption on top of my series. The formatting code is already split into\nits own function.\n\n-Peff\n"},{"id":"302886","messageId":"20160929132425.of7m5t4tsqcb6bbk@sigill.intra.peff.net","threadId":"44156","inReplyTo":"841D4FC2-9673-486A-8D94-8967188CCC60@gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T13:24:25Z","receivedAt":"2016-09-29T13:24:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 06:01:51AM -0700, Kyle J. McKay wrote:\n\n> But perhaps it makes sense to actually pick one if there's only one\n> disambiguation of the type you're looking for.\n> \n> For example given:\n> \n> 235234a blob\n> 2352347 tag\n> 235234f tree\n> 2352340 commit\n> \n> If you are doing \"git cat-file blob 235234\" it should pick the blob and spit\n> out a warning (and similarly for other cat-file types).  But \"git cat-file\n> -p 235234\" would give the fatal error with the disambiguation hints because\n> it wants type \"any\".\n\nThat code is already there; it's just a matter of whether git has enough\ninformation to know the context. E.g. (in git.git):\n\n  $ git show b2e11\n  error: short SHA1 b2e11 is ambiguous\n  hint: The candidates are:\n  hint:   b2e1196 tag v2.8.0-rc1\n  hint:   b2e11d1 tree\n  ...\n\n  $ git log b2e11\n  commit ab5d01a29eb7380ceab070f0807c2939849c44bc (tag: v2.8.0-rc1)\n  ...\n\nThe \"show\" command can show anything, but \"log\" really wants\ncommittishes, so it's able to disambiguate. It looks like cat-file never\nlearned to feed its context, but it's probably something like this:\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 94e67eb..ecbb959 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -56,12 +56,22 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \tstruct object_info oi = {NULL};\n \tstruct strbuf sb = STRBUF_INIT;\n \tunsigned flags = LOOKUP_REPLACE_OBJECT;\n+\tunsigned sha1_flags = 0;\n \tconst char *path = force_path;\n \n \tif (unknown_type)\n \t\tflags |= LOOKUP_UNKNOWN_OBJECT;\n \n-\tif (get_sha1_with_context(obj_name, 0, oid.hash, &obj_context))\n+\tif (exp_type) {\n+\t\tif (!strcmp(exp_type, \"commit\"))\n+\t\t\tsha1_flags |= GET_SHA1_COMMITTISH;\n+\t\telse if(!strcmp(exp_type, \"tree\"))\n+\t\t\tsha1_flags |= GET_SHA1_TREEISH;\n+\t\telse if(!strcmp(exp_type, \"blob\"))\n+\t\t\tsha1_flags |= GET_SHA1_BLOB;\n+\t}\n+\n+\tif (get_sha1_with_context(obj_name, sha1_flags, oid.hash, &obj_context))\n \t\tdie(\"Not a valid object name %s\", obj_name);\n \n \tif (!path)\n\n> If you are doing \"git show 235234\" it should pick the tag (if it peels to a\n> committish) because Git has already set a precedent of preferring tags over\n> commits when it disambiguates ref names and otherwise pick the commit.\n\nI'm not convinced that picking the tag is actually helpful in this case;\nI agree with Linus that feeding something to \"git show\" almost always\nwants to choose the commit.\n\nI also don't think tag ambiguity in short sha1s is all that interesting.\nThere are a tiny number of tag objects. Most of your collisions are\ngoing to be with trees or blobs, which should generally outnumber\ncommits by a factor of 5-10, though it depends on your workflow (git.git\ndoes not have a deep tree, so it's only a factor of 4).\n\nAnd if you just want to choose a committish over trees and blobs, well,\nthen; I invite you to check out the core.disambiguate patch I sent\nelsewhere in the thread. :)\n\n-Peff\n"},{"id":"302888","messageId":"2FECD796-7B92-41BB-A0AF-57650FF7E78D@gmail.com","threadId":"44156","inReplyTo":"20160929132425.of7m5t4tsqcb6bbk@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Kyle J. McKay","fromEmail":"mackyle@gmail.com","sentAt":"2016-09-29T14:36:27Z","receivedAt":"2016-09-29T14:42:00Z","isPatch":true,"sender":{"key":"mackyle@gmail.com","avatar":"https://avatars.githubusercontent.com/u/813346?v=4"},"body":"On Sep 29, 2016, at 06:24, Jeff King wrote:\n\n>> If you are doing \"git show 235234\" it should pick the tag (if it  \n>> peels to a\n>> committish) because Git has already set a precedent of preferring  \n>> tags over\n>> commits when it disambiguates ref names and otherwise pick the  \n>> commit.\n>\n> I'm not convinced that picking the tag is actually helpful in this  \n> case;\n> I agree with Linus that feeding something to \"git show\" almost always\n> wants to choose the commit.\n\nSince \"git show\" peels tags you end up seeing the commit it refers to  \n(assuming it's a committish tag).\n\n> I also don't think tag ambiguity in short sha1s is all that  \n> interesting.\n\nThe Linux repository has this:\n\n    901069c:\n       901069c71415a76d commit iwlagn: change Copyright to 2011\n       901069c5c5b15532 tag    (v2.6.38-rc4) Linux 2.6.38-rc4\n\nSince that tag peels to a commit, it seems like it would be incorrect  \nto pick the commit over the tag when you're looking for a committish.\n\nEither 901069c should resolve to the tag (which gets peeled to the  \ncommit) or it should error out with the hint messages.\n\nThe Git repository has this:\n\n    c512b03:\n       c512b035556eff4d commit Merge branch 'rc/maint-reflog-msg-for- \nforced\n       c512b0344196931a tag    (v0.99.9a) GIT 0.99.9a\n\nSo perhaps it's a little bit more interesting than it first appears.  :)\n\n--Kyle\n"},{"id":"302889","messageId":"20160929145520.dgyj57df4tyqrl4y@sigill.intra.peff.net","threadId":"44156","inReplyTo":"2FECD796-7B92-41BB-A0AF-57650FF7E78D@gmail.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T14:55:20Z","receivedAt":"2016-09-29T14:55:28Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 07:36:27AM -0700, Kyle J. McKay wrote:\n\n> On Sep 29, 2016, at 06:24, Jeff King wrote:\n> \n> > > If you are doing \"git show 235234\" it should pick the tag (if it\n> > > peels to a\n> > > committish) because Git has already set a precedent of preferring\n> > > tags over\n> > > commits when it disambiguates ref names and otherwise pick the\n> > > commit.\n> > \n> > I'm not convinced that picking the tag is actually helpful in this case;\n> > I agree with Linus that feeding something to \"git show\" almost always\n> > wants to choose the commit.\n> \n> Since \"git show\" peels tags you end up seeing the commit it refers to\n> (assuming it's a committish tag).\n\nYes, but it's almost certainly _not_ the commit you meant. From your\nexample:\n\n>    c512b03:\n>       c512b035556eff4d commit Merge branch 'rc/maint-reflog-msg-for-forced\n>       c512b0344196931a tag    (v0.99.9a) GIT 0.99.9a\n\nIf I'm looking for the commit c512b03, then it almost certainly isn't\nv0.99.9a. That tag's commit is e634aec. Or another way of thinking about\nit: you want to guess what the _writer_ of the note meant. Why would\nsomebody write \"c512b03\" when they could have written \"v0.99.9a\"? And\nthey certainly would not have written it if they meant \"e634aec\". :)\n\n> > I also don't think tag ambiguity in short sha1s is all that interesting.\n> \n> The Linux repository has this:\n> \n>    901069c:\n>       901069c71415a76d commit iwlagn: change Copyright to 2011\n>       901069c5c5b15532 tag    (v2.6.38-rc4) Linux 2.6.38-rc4\n\nSure, I'm not surprised there's a collision. But I'd expect those to be\na tiny fraction of collisions. Here's the breakdown of object types in\nmy clone of linux.git:\n\n  $ git cat-file --batch-all-objects --batch-check='%(objecttype)' |\n    sort | uniq -c\n  1421198 blob\n   618073 commit\n      479 tag\n  2877913 tree\n\nThat's a hundredth of a percent tag objects.  The chance that you have\n_a_ 7-hex collision with a tag is relatively high. But the chance that\nany given collision involves a tag is rather small.\n\n-Peff\n"},{"id":"302895","messageId":"xmqqbmz6hbdk.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929130322.562ng4t2ktk6qzok@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T17:19:35Z","receivedAt":"2016-09-29T17:19:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> $ git rev-parse --disambiguate-list=b2e1\n>> b2e1196 tag v2.8.0-rc1\n>> b2e11d1 tree\n>> b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n>> b2e1759 blob\n>> b2e18954 blob\n>> b2e1895c blob\n>\n> I think the \"right\" way to do this is pipe the list of sha1s into\n> another git commit which can format them however you want.\n> Unfortunately, there isn't a single command that does a great job:\n>\n>   - \"cat-file --batch-check\" can show you the sha1 and type, but it\n>     won't abbreviate sha1s, and it won't show you commit/tag information\n>\n>   - \"log --stdin --no-walk\" will format the commit however you like, but\n>     skips the trees and blobs entirely, and the tag can only be seen via\n>     \"%d\"\n>\n>   - \"for-each-ref\" has flexible formatting, too, but wants to format\n>     refs, not objects (and doesn't read from stdin).\n\n    - \"name-rev\" is used to give \"describe --contains\", and can read\n      from its standard input, but has no format customization.\n      Another downside of it is that it only wants to see\n      committishes.\n\n> IMHO that is a sign that our formatting tools aren't as good as they\n> could be (I think the right tool is cat-file, but it should be able to\n> do all of the formatting that the other commands can do).\n>\n> Of course if you really just want human-readable output, then:\n>\n>   $ git cat-file -e b2e1\n>   error: short SHA1 b2e1 is ambiguous\n>   hint: The candidates are:\n>   hint:   b2e1196 tag v2.8.0-rc1\n>   hint:   b2e11d1 tree\n>   hint:   b2e1632 commit 2007-11-14 - Merge branch 'bs/maint-commit-options'\n>   hint:   b2e1759 blob\n>   hint:   b2e18954 blob\n>   hint:   b2e1895c blob\n>   fatal: Not a valid object name b2e1\n>\n> is pretty easy.\n\nYes.  I think adding this to rev-parse that is meant for machines is\nprobably a mistake, as this \"hint\" machinery's output will become\neven more human friendly over time as we gain experience.\n\n - If the hypothetical \"--disambiguate-list\" option wants to produce\n   machine parseable output for scripts, it would mean its output\n   (and whatgver the reading script can do based on its output for\n   humans) will become less useful for humans over time.\n\n - If the hypothetical \"--disambiguate-list\" option only wants to\n   replicate the human readable output that is designed to be\n   improved over time and expects its output _not_ to be interpreted\n   by scripts but merely be relayed, then why aren't these scripts\n   just invoking the commands that already gives the \"hint:\" output\n   and showing that directly to humans in the first place?\n\n> That being said, I don't mind if somebody wanted to do a rev-parse\n> option on top of my series. The formatting code is already split into\n> its own function.\n\nSo let's not go there.\n"},{"id":"302896","messageId":"xmqq7f9uhbc4.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"f239b2eb-d122-9c4b-187b-fbd40a94bcf4@gmail.com","subject":"Re: [PATCH 1/4] config: allow customizing /etc/gitconfig location","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T17:20:27Z","receivedAt":"2016-09-29T17:20:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narębski <jnareb@gmail.com> writes:\n\n> W dniu 29.09.2016 o 01:30, Junio C Hamano pisze:\n>> With a new environment variable GIT_ETC_GITCONFIG, the users can\n>> specify a file that is used instead of /etc/gitconfig to read (and\n>> write) the system-wide configuration.\n>\n> Why it is named GIT_ETC_GITCONFIG (which is Unix-ism), and not\n> GIT_CONFIG_SYSTEM / GIT_CONFIG_SYSTEM_PATH, that is something\n> OS-neutral?\n\nIsn't \"environment variable\" something that came from POSIX world?\n"},{"id":"302900","messageId":"vpqmviqk3b5.fsf@anie.imag.fr","threadId":"44156","inReplyTo":"xmqq7f9uhbc4.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 1/4] config: allow customizing /etc/gitconfig location","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-09-29T17:45:34Z","receivedAt":"2016-09-29T17:45:52Z","isPatch":true,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Jakub Narębski <jnareb@gmail.com> writes:\n>\n>> W dniu 29.09.2016 o 01:30, Junio C Hamano pisze:\n>>> With a new environment variable GIT_ETC_GITCONFIG, the users can\n>>> specify a file that is used instead of /etc/gitconfig to read (and\n>>> write) the system-wide configuration.\n>>\n>> Why it is named GIT_ETC_GITCONFIG (which is Unix-ism), and not\n>> GIT_CONFIG_SYSTEM / GIT_CONFIG_SYSTEM_PATH, that is something\n>> OS-neutral?\n>\n> Isn't \"environment variable\" something that came from POSIX world?\n\nI don't know who invented the concept, but environment variables have\nbeen there in the windows world since it exists I think (it existed in\nMS-DOS).\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"302904","messageId":"xmqqmviqfuoh.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"ae9dbf3b-4190-8145-a59f-0d578067032a@kdbg.org","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T18:05:34Z","receivedAt":"2016-09-29T18:05:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j6t@kdbg.org> writes:\n\n> Am 29.09.2016 um 01:30 schrieb Junio C Hamano:\n>> As Peff said, responding in a thread started by Linus's suggestion\n>> to raise the default abbreviation to 12 hexdigits:\n>\n> This is waayy too large for a new default. The vast majority of\n> repositories is smallish. For those, the long sequences of hex digits\n> are an uglification that is almost unbearable.\n>\n> I know that kernel developers are important, but their importance has\n> long been outnumbered by the anonymous and silent masses of users.\n>\n> Personally, I use 8 digits just because it is a \"rounder\" number than\n> 7, but in all of my repositories 7 would still work just as well.\n\nYes, \"git log --oneline\" looks somewhat different and strange for\nme, too ;-)\n\nI am sure I'll get used to it if I keep using it, but I suspect that\nI'd be irritated as I find myself typing 'q' more and more often to\n\"less -S\" that is automatically invoked when I do \"git log --oneline\nmaster..\" to see what commits are on my current topic branch.\n"},{"id":"302906","messageId":"xmqqintefuau.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929090108.hf2jzfcvbcsfaxw7@sigill.intra.peff.net","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T18:13:45Z","receivedAt":"2016-09-29T18:14:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I think anytime you would use GIT_CONFIG_NOSYSTEM over --local, it is an\n> indication that the test is trying to check how multiple sources\n> interact. And the right thing to do for them is to set GIT_ETC_GITCONFIG\n> to some known quantity. We just couldn't do that before, so we skipped\n> it.  IOW, something like the patch below (on top of yours).\n\nOK, that way we can make sure that \"multiple sources\" operations do\nlook at the system-wide stuff.\n\n> Note that the\n> commands that are doing a \"--get\" and not a \"--list\" don't actually seem\n> to need either (because they are getting the values out of the local\n> file anyway), so we could drop the setting of GIT_ETC_GITCONFIG from\n> them entirely.\n\n\"either\" meaning \"we do not need to add --local and we do not need\nGIT_CONFIG_NOSYSTEM\"?\n\n"},{"id":"302908","messageId":"20160929182621.lobihscwl7amtu7s@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqintefuau.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T18:26:21Z","receivedAt":"2016-09-29T18:27:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 11:13:45AM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > I think anytime you would use GIT_CONFIG_NOSYSTEM over --local, it is an\n> > indication that the test is trying to check how multiple sources\n> > interact. And the right thing to do for them is to set GIT_ETC_GITCONFIG\n> > to some known quantity. We just couldn't do that before, so we skipped\n> > it.  IOW, something like the patch below (on top of yours).\n> \n> OK, that way we can make sure that \"multiple sources\" operations do\n> look at the system-wide stuff.\n\nExactly.\n\n> > Note that the\n> > commands that are doing a \"--get\" and not a \"--list\" don't actually seem\n> > to need either (because they are getting the values out of the local\n> > file anyway), so we could drop the setting of GIT_ETC_GITCONFIG from\n> > them entirely.\n> \n> \"either\" meaning \"we do not need to add --local and we do not need\n> GIT_CONFIG_NOSYSTEM\"?\n\nYes. I didn't test it with your core.abbrev patch 4/4, but I _didn't_\nhave to touch their expected output after pointing them at a non-empty\netc-gitconfig file in the trash directory. Which implies to me they\ndon't care either way (which makes sense; they are asking for a specific\nkey which is supposed to be found in one of the other files).\n\n-Peff\n"},{"id":"302911","messageId":"CA+55aFyYWWpz+9+KKf=9y3vBrEDyy-5h6J3boiitGE7Zb=uL-Q@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqmviqfuoh.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-29T18:37:41Z","receivedAt":"2016-09-29T18:37:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 11:05 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Yes, \"git log --oneline\" looks somewhat different and strange for\n> me, too ;-)\n\nI'm playing with an early patch to make the default more dynamic.\nLet's see how well it works in practice, but it looks fairly\npromising. Let me test a bit more and send out an RFC patch..\n\n              Linus\n"},{"id":"302916","messageId":"CA+55aFwbCNiF0nDppZ5SuRcZwc9kNvKYzgyd_bR8Ut8XRW_p4Q@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFyYWWpz+9+KKf=9y3vBrEDyy-5h6J3boiitGE7Zb=uL-Q@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-29T18:55:46Z","receivedAt":"2016-09-29T18:55:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 11:37 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> I'm playing with an early patch to make the default more dynamic.\n> Let's see how well it works in practice, but it looks fairly\n> promising. Let me test a bit more and send out an RFC patch..\n\nOk, this is *very* rough, and it doesn't actuall pass all the tests,\nand I didn't even try to look at why. But it passes the trivial\nsmell-test, and in particular it actually makes mathematical sense...\n\nI think the patch can speak for itself, but the basic core is this\nsection in get_short_sha1():\n\n  +       if (len < 16 && !status && (flags & GET_SHA1_AUTOMATIC)) {\n  +               unsigned int expect_collision = 1 << (len * 2);\n  +               if (ds.nrobjects > expect_collision)\n  +                       return SHORT_NAME_AMBIGUOUS;\n  +       }\n\nbasically, what it says is that we will consider a sha1 ambiguous even\nif it was *technically* unique (that's the '!status' part of the test)\nif:\n\n - the length was 15 or less\n\n*and*\n\n - the number of objects we have is larger than the expected point\nwhere statistically we should start to expect to get one collision.\n\nThat \"expect_collision\" math is actually very simple: each hex\ncharacter adds four bits of range, but since we expect collisions at\nthe square root of the maximum number of objects, we shift by just two\nbits per hex digits instead.\n\nThe rest of the patch is a trivial change to just initialize the\ndefault short size to -1, and consider that to mean \"enable the\nautomatic size checking with a minimum of 7\". And the trivial code to\nestimate the number of objects (which ignores duplicates between packs\netc _entirely_).\n\nFor the kernel, just the *math* right now actually gives 12\ncharacters. For current git it actually seems to say that 8 is the\ncorrect number. For small projects, you'll still see 7.\n\nANYWAY. This patch is on top of Jeff's patches in 'pu' (I think those\nare great regardless of this patch!), and as mentioned, it fails some\ntests. I suspect that the failures might be due to the abbrev_default\nbeing -1, and some other code finds that surprising now. But as\nmentioned, I didn't really even look at it.\n\nWhat do you think? It's actually a fairly simple patch and I really do\nthink it makes sense and it seems to just DTRT automatically.\n\n              Linus\n\n\n cache.h       |  1 +\n environment.c |  2 +-\n sha1_name.c   | 21 ++++++++++++++++++++-\n 3 files changed, 22 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 6e33f2f..d2da6d1 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1207,6 +1207,7 @@ struct object_context {\n #define GET_SHA1_TREEISH          020\n #define GET_SHA1_BLOB             040\n #define GET_SHA1_FOLLOW_SYMLINKS 0100\n+#define GET_SHA1_AUTOMATIC\t 0200\n #define GET_SHA1_ONLY_TO_DIE    04000\n \n #define GET_SHA1_DISAMBIGUATORS \\\ndiff --git a/environment.c b/environment.c\nindex c1442df..fd6681e 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -16,7 +16,7 @@ int trust_executable_bit = 1;\n int trust_ctime = 1;\n int check_stat = 1;\n int has_symlinks = 1;\n-int minimum_abbrev = 4, default_abbrev = 7;\n+int minimum_abbrev = 4, default_abbrev = -1;\n int ignore_case;\n int assume_unchanged;\n int prefer_symlink_refs;\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 3b647fd..8791ff3 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -15,6 +15,7 @@ typedef int (*disambiguate_hint_fn)(const unsigned char *, void *);\n \n struct disambiguate_state {\n \tint len; /* length of prefix in hex chars */\n+\tunsigned int nrobjects;\n \tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n \n@@ -118,6 +119,12 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \n \t\t\tif (strlen(de->d_name) != 38)\n \t\t\t\tcontinue;\n+\n+\t\t\t// We only look at the one subdirectory, and we assume\n+\t\t\t// each subdirectory is roughly similar, so each object\n+\t\t\t// we find probably has 255 other objects in the other\n+\t\t\t// fan-out directories\n+\t\t\tds->nrobjects += 256;\n \t\t\tif (memcmp(de->d_name, ds->hex_pfx + 2, ds->len - 2))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(hex + 2, de->d_name, 38);\n@@ -151,6 +158,7 @@ static void unique_in_pack(struct packed_git *p,\n \n \topen_pack_index(p);\n \tnum = p->num_objects;\n+\tds->nrobjects += num;\n \tlast = num;\n \twhile (first < last) {\n \t\tuint32_t mid = (first + last) / 2;\n@@ -426,6 +434,12 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\tfor_each_abbrev(ds.hex_pfx, show_ambiguous_object, &ds);\n \t}\n \n+\tif (len < 16 && !status && (flags & GET_SHA1_AUTOMATIC)) {\n+\t\tunsigned int expect_collision = 1 << (len * 2);\n+\t\tif (ds.nrobjects > expect_collision)\n+\t\t\treturn SHORT_NAME_AMBIGUOUS;\n+\t}\n+\n \treturn status;\n }\n \n@@ -458,14 +472,19 @@ int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n int find_unique_abbrev_r(char *hex, const unsigned char *sha1, int len)\n {\n \tint status, exists;\n+\tint flags = GET_SHA1_QUIETLY;\n \n+\tif (len < 0) {\n+\t\tflags |= GET_SHA1_AUTOMATIC;\n+\t\tlen = 7;\n+\t}\n \tsha1_to_hex_r(hex, sha1);\n \tif (len == 40 || !len)\n \t\treturn 40;\n \texists = has_sha1_file(sha1);\n \twhile (len < 40) {\n \t\tunsigned char sha1_ret[20];\n-\t\tstatus = get_short_sha1(hex, len, sha1_ret, GET_SHA1_QUIETLY);\n+\t\tstatus = get_short_sha1(hex, len, sha1_ret, flags);\n \t\tif (exists\n \t\t    ? !status\n \t\t    : status == SHORT_NAME_NOT_FOUND) {\n"},{"id":"302917","messageId":"xmqqa8eqfsap.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929182621.lobihscwl7amtu7s@sigill.intra.peff.net","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T18:57:02Z","receivedAt":"2016-09-29T18:57:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> \"either\" meaning \"we do not need to add --local and we do not need\n>> GIT_CONFIG_NOSYSTEM\"?\n>\n> Yes. I didn't test it with your core.abbrev patch 4/4, but I _didn't_\n> have to touch their expected output after pointing them at a non-empty\n> etc-gitconfig file in the trash directory. Which implies to me they\n> don't care either way (which makes sense; they are asking for a specific\n> key which is supposed to be found in one of the other files).\n\nThere is a bit of problem here, though.\n\n * If we make t1300 point at its own system-wide config, it will be\n   in control of its contents, so \"find this key\" will find only it\n   wants to find (or we found a regression).\n\n * But then if it ever does something that depends on the default\n   value of core.abbrev (or whatever we'd tweak in response to the\n   next suggestion by Linus ;-), we cannot really allow it to do\n   so.  We'd want t/gitconfig-for-test to be the single place that\n   we can tweak these things, but we'll have to know t1300 uses its\n   own and need to make the same change there, too.\n\nSo, I dunno.\n"},{"id":"302918","messageId":"xmqq60pefrvc.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929182621.lobihscwl7amtu7s@sigill.intra.peff.net","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T19:06:15Z","receivedAt":"2016-09-29T19:06:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Thu, Sep 29, 2016 at 11:13:45AM -0700, Junio C Hamano wrote:\n>\n>> Jeff King <peff@peff.net> writes:\n>> \n>> > I think anytime you would use GIT_CONFIG_NOSYSTEM over --local, it is an\n>> > indication that the test is trying to check how multiple sources\n>> > interact. And the right thing to do for them is to set GIT_ETC_GITCONFIG\n>> > to some known quantity. We just couldn't do that before, so we skipped\n>> > it.  IOW, something like the patch below (on top of yours).\n>> \n>> OK, that way we can make sure that \"multiple sources\" operations do\n>> look at the system-wide stuff.\n>\n> Exactly.\n\nI think it deserves a separate patch and the result is more\nunderstandable.  I've queued this for now (on top of a revised 1/4\nthat uses GIT_CONFIG_SYSTEM_PATH instead).\n\n-- >8 --\nFrom: Jeff King <peff@peff.net>\nDate: Thu, 29 Sep 2016 11:29:10 -0700\nSubject: [PATCH] t1300: check also system-wide configuration file in\n --show-origin tests\n\nBecause we used to run our tests with GIT_CONFIG_NOSYSTEM, these did\nnot test that the system-wide configuration file is also read and\nshown as one of the origins.  Create a custom/fake system-wide\nconfiguration file and make sure it appears in the output, using the\nnewly introduced GIT_CONFIG_SYSTEM_PATH mechanism.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/t1300-repo-config.sh | 15 ++++++++++++++-\n 1 file changed, 14 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\nindex 0543b62227bf..aa25577709c5 100755\n--- a/t/t1300-repo-config.sh\n+++ b/t/t1300-repo-config.sh\n@@ -1236,6 +1236,11 @@ test_expect_success 'set up --show-origin tests' '\n \t\t[user]\n \t\t\trelative = include\n \tEOF\n+\tcat >\"$HOME\"/etc-gitconfig <<-\\EOF &&\n+\t\t[user]\n+\t\t\tsystem = true\n+\t\t\toverride = system\n+\tEOF\n \tcat >\"$HOME\"/.gitconfig <<-EOF &&\n \t\t[user]\n \t\t\tglobal = true\n@@ -1254,6 +1259,8 @@ test_expect_success 'set up --show-origin tests' '\n \n test_expect_success '--show-origin with --list' '\n \tcat >expect <<-EOF &&\n+\t\tfile:$HOME/etc-gitconfig\tuser.system=true\n+\t\tfile:$HOME/etc-gitconfig\tuser.override=system\n \t\tfile:$HOME/.gitconfig\tuser.global=true\n \t\tfile:$HOME/.gitconfig\tuser.override=global\n \t\tfile:$HOME/.gitconfig\tinclude.path=$INCLUDE_DIR/absolute.include\n@@ -1264,13 +1271,16 @@ test_expect_success '--show-origin with --list' '\n \t\tfile:.git/../include/relative.include\tuser.relative=include\n \t\tcommand line:\tuser.cmdline=true\n \tEOF\n+\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n \tgit -c user.cmdline=true config --list --show-origin >output &&\n \ttest_cmp expect output\n '\n \n test_expect_success '--show-origin with --list --null' '\n \tcat >expect <<-EOF &&\n-\t\tfile:$HOME/.gitconfigQuser.global\n+\t\tfile:$HOME/etc-gitconfigQuser.system\n+\t\ttrueQfile:$HOME/etc-gitconfigQuser.override\n+\t\tsystemQfile:$HOME/.gitconfigQuser.global\n \t\ttrueQfile:$HOME/.gitconfigQuser.override\n \t\tglobalQfile:$HOME/.gitconfigQinclude.path\n \t\t$INCLUDE_DIR/absolute.includeQfile:$INCLUDE_DIR/absolute.includeQuser.absolute\n@@ -1281,6 +1291,7 @@ test_expect_success '--show-origin with --list --null' '\n \t\tincludeQcommand line:Quser.cmdline\n \t\ttrueQ\n \tEOF\n+\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n \tgit -c user.cmdline=true config --null --list --show-origin >output.raw &&\n \tnul_to_q <output.raw >output &&\n \t# The here-doc above adds a newline that the --null output would not\n@@ -1304,6 +1315,7 @@ test_expect_success '--show-origin with --get-regexp' '\n \t\tfile:$HOME/.gitconfig\tuser.global true\n \t\tfile:.git/config\tuser.local true\n \tEOF\n+\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n \tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n \ttest_cmp expect output\n '\n@@ -1312,6 +1324,7 @@ test_expect_success '--show-origin getting a single key' '\n \tcat >expect <<-\\EOF &&\n \t\tfile:.git/config\tlocal\n \tEOF\n+\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n \tgit config --show-origin user.override >output &&\n \ttest_cmp expect output\n '\n-- \n2.10.0-589-g5adf4e1\n\n"},{"id":"302919","messageId":"CA+55aFx9Utm9yDZceks+5q9c8ydc2QMYshWwJ0G0GHWWLwSsXQ@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFwbCNiF0nDppZ5SuRcZwc9kNvKYzgyd_bR8Ut8XRW_p4Q@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-29T19:06:23Z","receivedAt":"2016-09-29T19:06:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 11:55 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> For the kernel, just the *math* right now actually gives 12\n> characters. For current git it actually seems to say that 8 is the\n> correct number. For small projects, you'll still see 7.\n\nSorry, the git number is 9, not 8. The reason is that git has roughly\n212k objects, and 9 hex digits gets expected collisions at about 256k\nobjects.\n\nSo the logic means that we'll see 7 hex digits for projects with less\nthan 16k objects, 8 hex digits if there are less than 64k objects, and\n9 hex digits for projects like git that currently have fewer than 256k\nobjects.\n\nBut git itself might not be *that* far from going to 10 hex digits\nwith my patch.\n\nThe kernel uses 12 he digits because the collision math says that's\nthe right thing for a project with between 4M and 16M objects (with\nthe kernel being at 5M).\n\nSo on the whole the patch really does seem to just do the right thing\nautomatically.\n\n              Linus\n"},{"id":"302920","messageId":"20160929191609.maxggcli76472t4g@sigill.intra.peff.net","threadId":"44156","inReplyTo":"CA+55aFwbCNiF0nDppZ5SuRcZwc9kNvKYzgyd_bR8Ut8XRW_p4Q@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T19:16:09Z","receivedAt":"2016-09-29T19:16:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 11:55:46AM -0700, Linus Torvalds wrote:\n\n> I think the patch can speak for itself, but the basic core is this\n> section in get_short_sha1():\n> \n>   +       if (len < 16 && !status && (flags & GET_SHA1_AUTOMATIC)) {\n>   +               unsigned int expect_collision = 1 << (len * 2);\n>   +               if (ds.nrobjects > expect_collision)\n>   +                       return SHORT_NAME_AMBIGUOUS;\n>   +       }\n\nHmm. So at length 7, we expect collisions at 2^14, which is 16384. That\nseems really low. I mean, by the birthday paradox that's where expect\na 50% chance of a collision. But that's a single collision. We\ndefinitely don't expect them to be common at that size.\n\nSo I suspect this could be a bit looser. The real number we care about\nis probably something like \"there is probability 'p' of a collision when\nwe add a new object\", but I'm not sure what that 'p' would be. Or\nperhaps \"we accept collisions in 'n' percent of objects\". But again, I\ndon't know that 'n'.\n\nI dunno. I suppose being overly conservative with this number leaves\nroom for growth. Repositories generally get bigger, not smaller. :)\n\n> What do you think? It's actually a fairly simple patch and I really do\n> think it makes sense and it seems to just DTRT automatically.\n\nI like the general idea.\n\nAs far as the implementation, I was surprised to see it touch\nget_short_sha1() at all. That's, after all, for lookups, and we would\nnever want to require more characters on the reading side.\n\nI see you worked around it with a flag so that this behavior only kicks\nin when called via find_unique_abbrev(). But if you look at the caller:\n\n> @@ -458,14 +472,19 @@ int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n>  int find_unique_abbrev_r(char *hex, const unsigned char *sha1, int len)\n>  {\n>  \tint status, exists;\n> +\tint flags = GET_SHA1_QUIETLY;\n>  \n> +\tif (len < 0) {\n> +\t\tflags |= GET_SHA1_AUTOMATIC;\n> +\t\tlen = 7;\n> +\t}\n>  \tsha1_to_hex_r(hex, sha1);\n>  \tif (len == 40 || !len)\n>  \t\treturn 40;\n>  \texists = has_sha1_file(sha1);\n>  \twhile (len < 40) {\n>  \t\tunsigned char sha1_ret[20];\n> -\t\tstatus = get_short_sha1(hex, len, sha1_ret, GET_SHA1_QUIETLY);\n> +\t\tstatus = get_short_sha1(hex, len, sha1_ret, flags);\n>  \t\tif (exists\n>  \t\t    ? !status\n>  \t\t    : status == SHORT_NAME_NOT_FOUND) {\n\nYou can see that we're going to do more work than we would otherwise\nneed to. Because we start at 7, and ask get_short_sha1() \"is this unique\nenough?\", and looping. But if we _know_ we won't accept any answer\nshorter than some N based on the number of objects in the repository,\nthen we should start at that N.\n\nIOW, something like:\n\n  if (len < 0)\n\tlen = ceil(log_base_2(repository_object_count()));\n\nhere, and then you don't have to touch get_short_sha1() at all.\n\nI suspect you pushed it down into get_short_sha1() because it kind-of\ndoes the repository_object_count() step for \"free\" as it's looking at\nthe object anyway. But that step is really not very expensive. And I'd\neven say you could just ignore loose objects entirely, and treat them\nlike a rounding error (the way that duplicate objects in packs are\ntreated).\n\nThat leaves you with just an O(# of packs) loop over a linked list. You\ncould even just keep a global object count up to date in\nadd_packed_git(), and then it's O(1).\n\n-Peff\n"},{"id":"302921","messageId":"20160929191857.lxcgzf2cg5zfjkrq@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqa8eqfsap.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T19:18:57Z","receivedAt":"2016-09-29T19:19:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 11:57:02AM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> >> \"either\" meaning \"we do not need to add --local and we do not need\n> >> GIT_CONFIG_NOSYSTEM\"?\n> >\n> > Yes. I didn't test it with your core.abbrev patch 4/4, but I _didn't_\n> > have to touch their expected output after pointing them at a non-empty\n> > etc-gitconfig file in the trash directory. Which implies to me they\n> > don't care either way (which makes sense; they are asking for a specific\n> > key which is supposed to be found in one of the other files).\n> \n> There is a bit of problem here, though.\n> \n>  * If we make t1300 point at its own system-wide config, it will be\n>    in control of its contents, so \"find this key\" will find only it\n>    wants to find (or we found a regression).\n> \n>  * But then if it ever does something that depends on the default\n>    value of core.abbrev (or whatever we'd tweak in response to the\n>    next suggestion by Linus ;-), we cannot really allow it to do\n>    so.  We'd want t/gitconfig-for-test to be the single place that\n>    we can tweak these things, but we'll have to know t1300 uses its\n>    own and need to make the same change there, too.\n\nRight, but I think that's fine. Tests that care deeply about the\ncontents of etc-gitconfig are unlikely to care about core.abbrev. And in\nthe off chance that they do, then the worst case is...they get updated\nto handle core.abbrev (either passing a command line option, or just\nputting core.abbrev in their test file).\n\nI just don't see it being a problem. Adding core.abbrev for the whole\ntest suite is just about not having a big flag day where we change all\nthe tests. Changing one or two tests (and again, I'd be surprised if we\neven have to do that) doesn't seem like a big deal.\n\n-Peff\n"},{"id":"302925","messageId":"20160929192613.o6q2fqp3mjntz2l6@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqq60pefrvc.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T19:26:14Z","receivedAt":"2016-09-29T19:26:20Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 12:06:15PM -0700, Junio C Hamano wrote:\n\n> I think it deserves a separate patch and the result is more\n> understandable.  I've queued this for now (on top of a revised 1/4\n> that uses GIT_CONFIG_SYSTEM_PATH instead).\n\nThanks, makes sense (and I like the new variable name better, by the\nway).\n\n> -- >8 --\n> From: Jeff King <peff@peff.net>\n> Date: Thu, 29 Sep 2016 11:29:10 -0700\n> Subject: [PATCH] t1300: check also system-wide configuration file in\n>  --show-origin tests\n> \n> Because we used to run our tests with GIT_CONFIG_NOSYSTEM, these did\n> not test that the system-wide configuration file is also read and\n> shown as one of the origins.  Create a custom/fake system-wide\n> configuration file and make sure it appears in the output, using the\n> newly introduced GIT_CONFIG_SYSTEM_PATH mechanism.\n> \n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n\nGood description.\n\nSigned-off-by: Jeff King <peff@peff.net>\n\nof course.\n\n> @@ -1304,6 +1315,7 @@ test_expect_success '--show-origin with --get-regexp' '\n>  \t\tfile:$HOME/.gitconfig\tuser.global true\n>  \t\tfile:.git/config\tuser.local true\n>  \tEOF\n> +\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n>  \tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n>  \ttest_cmp expect output\n>  '\n\nThis is one is trying to do a multi-file lookup, but we couldn't look in\nthe system config before. But to naturally extend it, it ought to look\nlike this on top:\n\ndiff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\nindex d2476a8..4dd5ce3 100755\n--- a/t/t1300-repo-config.sh\n+++ b/t/t1300-repo-config.sh\n@@ -1310,11 +1310,12 @@ test_expect_success '--show-origin with single file' '\n \n test_expect_success '--show-origin with --get-regexp' '\n \tcat >expect <<-EOF &&\n+\t\tfile:$HOME/etc-gitconfig\tuser.system true\n \t\tfile:$HOME/.gitconfig\tuser.global true\n \t\tfile:.git/config\tuser.local true\n \tEOF\n \tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n-\tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n+\tgit config --show-origin --get-regexp \"user\\.[g|l|s].*\" >output &&\n \ttest_cmp expect output\n '\n \n> @@ -1312,6 +1324,7 @@ test_expect_success '--show-origin getting a single key' '\n>  \tcat >expect <<-\\EOF &&\n>  \t\tfile:.git/config\tlocal\n>  \tEOF\n> +\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n>  \tgit config --show-origin user.override >output &&\n>  \ttest_cmp expect output\n>  '\n\nAnd I was tempted to say this one should not need to care, but I guess\nit is testing that we correctly read the override from the local config\nover the global one. So likewise, it is good to check that we also\noverride the system config (it does not effect the \"expect\" output, but\nthat does not mean it is not enhancing the test).\n\n-Peff\n"},{"id":"302927","messageId":"CA+55aFxNVbvyERNc_xEhrtfTVMGz3hkeAx1nv9vW+dhJwCpp6g@mail.gmail.com","threadId":"44156","inReplyTo":"20160929191609.maxggcli76472t4g@sigill.intra.peff.net","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-29T19:40:43Z","receivedAt":"2016-09-29T19:40:49Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 12:16 PM, Jeff King <peff@peff.net> wrote:\n>\n> Hmm. So at length 7, we expect collisions at 2^14, which is 16384. That\n> seems really low. I mean, by the birthday paradox that's where expect\n> a 50% chance of a collision. But that's a single collision. We\n> definitely don't expect them to be common at that size.\n>\n> So I suspect this could be a bit looser.\n\nSo I have to admit that I was surprised by how quickly it actually\ndecided that 7 isn't enough. In fact, the reason I initially said that\ngit used 8 digits was that I didn't count very closely, and just\nverified that it was more than the default 7.\n\nBut quite frankly, I think the math is correct, and part of that is\nthat the logic is all about not just the current state, but the\n\"reasonably near future\".\n\nSo it is indeed fairly aggressive, and the moment you have more\nobjects than the \"we'd expect to probably see  _one_ collision\" it\ngrows the size. But looking at the kernel situation, that really is\nwhat we'd want, because the whole problem with the existing code is\nthat it only takes the *current* situation into account. That's what\nwe want to get away from. We want git to pick a number that is sane\nfrom a standpoint of \"this project is still growing\".\n\nAnd git _already_ has commits that are ambiguous in 8 hex digits and\nneed 9. Yes, it's rare today, but the reason I'm telling kernel\ndevelopers to use 12 is because while a size-11 collision is very rare\ntoday, it does actually happen, and we want o pick a value where it is\nrare enough that even in the near future it's not going to be a big\ndeal.\n\nDon't get me wrong: collisions aren't fatal. So it's not like we have\nto absolutely avoid them, and I really like your patch series exactly\nbecause it makes collisions even less of a deal (particularly since I\nexpect people will not upgrade immediately, so we'll continue to see\neven new 7-hex-digit short forms even in the kernel). So it's a\nbalance of making the hex string long enough that it's simply not a\nbig worry.\n\nSo I'm sure it *could* be looser, but I actually also really suspect\nthat git truly *should* use a 9-digit abbreviation rather than 8 (and\n7 is definitely starting to be borderline, I think).\n\n> As far as the implementation, I was surprised to see it touch\n> get_short_sha1() at all. That's, after all, for lookups, and we would\n> never want to require more characters on the reading side.\n\nHeh. The implementation is crap. It was literally a \"how can I make\nthe smallest possible patch\" implementation. I was finishing it off\nwhile at a talk by Nicolas Pitre at Linaro Connect where I am right\nnow.\n\nSo I agree - it does extra work just because that's where it all\nslotted in with minimal effort.\n\nAt a minimum, once it finds a good new default, it should just memoize\nthat. So a minimal fix to the \"it's stupldly recalculating things over\nrand over again\" would be to just set \"default_abbrev\" to the value it\nfinds acceptable after the first time it finds something, so that it\ndoesn't end up looping _again_ in the future.\n\nBut you could easily also just instead have it do something like\n\n      if (default_abbrev < 0)\n            default_abbrev = initialize_abbrev();\n\nat startup time if \"abbrev_commit\" is set, and just do it once and for\nall rather rthan the odd loping behavior.\n\nI really just wanted to see how well the concept worked, and I was\nhappy to see that it gave what I thought were the \"correct\" numbers.\nAnd the loop was salready there ...\n\n            Linus\n"},{"id":"302928","messageId":"xmqq1t02fq6p.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFx9Utm9yDZceks+5q9c8ydc2QMYshWwJ0G0GHWWLwSsXQ@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T19:42:38Z","receivedAt":"2016-09-29T19:42:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, Sep 29, 2016 at 11:55 AM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n>>\n>> For the kernel, just the *math* right now actually gives 12\n>> characters. For current git it actually seems to say that 8 is the\n>> correct number. For small projects, you'll still see 7.\n>\n> Sorry, the git number is 9, not 8. The reason is that git has roughly\n> 212k objects, and 9 hex digits gets expected collisions at about 256k\n> objects.\n>\n> So the logic means that we'll see 7 hex digits for projects with less\n> than 16k objects, 8 hex digits if there are less than 64k objects, and\n> 9 hex digits for projects like git that currently have fewer than 256k\n> objects.\n\nWhew.  I was wondering where my brain went wrong, as I knew we have\n200k objects and 8 hexdigits means 1<<16 = 64k which is way too\nshort.\n"},{"id":"302929","messageId":"xmqqwphuebhd.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFxNVbvyERNc_xEhrtfTVMGz3hkeAx1nv9vW+dhJwCpp6g@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T19:45:34Z","receivedAt":"2016-09-29T19:45:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> But you could easily also just instead have it do something like\n>\n>       if (default_abbrev < 0)\n>             default_abbrev = initialize_abbrev();\n>\n> at startup time if \"abbrev_commit\" is set, and just do it once and for\n> all rather rthan the odd loping behavior.\n\nI think that is a reasonable way to go.\n\n#define DEFAULT_ABBREV get_default_abbrev()\n\nwould help.\n"},{"id":"302930","messageId":"xmqqshsieax6.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929191857.lxcgzf2cg5zfjkrq@sigill.intra.peff.net","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T19:57:41Z","receivedAt":"2016-09-29T19:57:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I just don't see it being a problem. Adding core.abbrev for the whole\n> test suite is just about not having a big flag day where we change all\n> the tests. Changing one or two tests (and again, I'd be surprised if we\n> even have to do that) doesn't seem like a big deal.\n\nI've already wasted several hours whipping t1300 into shape, because\nit was done in not so forward-looking future-proofed way.  I am not\nworried about core.abbrev but I am worried more about the next thing\nthat requires us to add an entry to t/gitconfig-for-test.  Adding a\ncorresponding entry to retain the old default for that new config to\ntwo places may not be a big deal, but it still makes me feel a bit\nuneasy.\n\nIn any case, I suspect that Linus's \"auto\" thing may still need the\ncustom system config with t1300 clean-up to pass the test, even\nthough I suspect it would compute that 7 is enough for most of the\ntiny repositories our tests use, so I'll polish this a bit more\nwhile waiting for that discussion to settle.\n\nThanks.\n\n\n\n\n"},{"id":"302943","messageId":"xmqqoa36e7v8.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160929192613.o6q2fqp3mjntz2l6@sigill.intra.peff.net","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T21:03:39Z","receivedAt":"2016-09-29T21:03:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Good description.\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n>\n> of course.\n>\n>> @@ -1304,6 +1315,7 @@ test_expect_success '--show-origin with --get-regexp' '\n>>  \t\tfile:$HOME/.gitconfig\tuser.global true\n>>  \t\tfile:.git/config\tuser.local true\n>>  \tEOF\n>> +\tGIT_CONFIG_SYSTEM_PATH=$HOME/etc-gitconfig \\\n>>  \tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n>>  \ttest_cmp expect output\n>>  '\n>\n> This is one is trying to do a multi-file lookup, but we couldn't look in\n> the system config before. But to naturally extend it, it ought to look\n> like this on top:\n>\n> diff --git a/t/t1300-repo-config.sh b/t/t1300-repo-config.sh\n> index d2476a8..4dd5ce3 100755\n> --- a/t/t1300-repo-config.sh\n> +++ b/t/t1300-repo-config.sh\n> @@ -1310,11 +1310,12 @@ test_expect_success '--show-origin with single file' '\n>  \n>  test_expect_success '--show-origin with --get-regexp' '\n>  \tcat >expect <<-EOF &&\n> +\t\tfile:$HOME/etc-gitconfig\tuser.system true\n>  \t\tfile:$HOME/.gitconfig\tuser.global true\n>  \t\tfile:.git/config\tuser.local true\n>  \tEOF\n>  \tGIT_ETC_GITCONFIG=$HOME/etc-gitconfig \\\n> -\tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n> +\tgit config --show-origin --get-regexp \"user\\.[g|l|s].*\" >output &&\n>  \ttest_cmp expect output\n>  '\n\nMakes sense modulo you inherited useless vertical bars from the\noriginal.  I'll squash something like that in but without || ;-)\n\nThanks.\n"},{"id":"302944","messageId":"20160929210839.77ikl33g44w2mpey@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqoa36e7v8.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 2/4] t13xx: do not assume system config is empty","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-29T21:08:39Z","receivedAt":"2016-09-29T21:08:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 02:03:39PM -0700, Junio C Hamano wrote:\n\n> > -\tgit config --show-origin --get-regexp \"user\\.[g|l].*\" >output &&\n> > +\tgit config --show-origin --get-regexp \"user\\.[g|l|s].*\" >output &&\n> >  \ttest_cmp expect output\n> >  '\n> \n> Makes sense modulo you inherited useless vertical bars from the\n> original.  I'll squash something like that in but without || ;-)\n\nHeh, I glossed over that completely. Thanks.\n\n-Peff\n"},{"id":"302957","messageId":"CA+55aFyVEQ+8TBBUm5KG9APtd9wy8cp_mRO=3nj12DXZNLAC9A@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqwphuebhd.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-29T21:53:41Z","receivedAt":"2016-09-29T21:53:48Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 12:45 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> I think that is a reasonable way to go.\n>\n> #define DEFAULT_ABBREV get_default_abbrev()\n>\n> would help.\n\nSo something like this that replaces the previous patch?\n\nSomebody should really double-check my heuristics, to see that I did\nthe pack counting etc right.  It doesn't do alternate loose file\ncounting at all, and maybe it could matter.  The advantage of the\nprevious patch was that it got the object counting right almost\nautomatically, this actually has its own new object counting code and\nmaybe I screwed it up.\n\n                Linus\n\n\n cache.h       |  3 ++-\n environment.c |  2 +-\n sha1_file.c   | 43 +++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 46 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 6e33f2f28..a022e1bd2 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1186,8 +1186,9 @@ static inline int hex2chr(const char *s)\n }\n \n /* Convert to/from hex/sha1 representation */\n+extern int get_default_abbrev(void);\n #define MINIMUM_ABBREV minimum_abbrev\n-#define DEFAULT_ABBREV default_abbrev\n+#define DEFAULT_ABBREV get_default_abbrev()\n \n struct object_context {\n \tunsigned char tree[20];\ndiff --git a/environment.c b/environment.c\nindex c1442df9a..fd6681e46 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -16,7 +16,7 @@ int trust_executable_bit = 1;\n int trust_ctime = 1;\n int check_stat = 1;\n int has_symlinks = 1;\n-int minimum_abbrev = 4, default_abbrev = 7;\n+int minimum_abbrev = 4, default_abbrev = -1;\n int ignore_case;\n int assume_unchanged;\n int prefer_symlink_refs;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex ca149a607..28ba04b65 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -3720,3 +3720,46 @@ int for_each_packed_object(each_packed_object_fn cb, void *data, unsigned flags)\n \t}\n \treturn r ? r : pack_errors;\n }\n+\n+static int init_default_abbrev(void)\n+{\n+\tunsigned long count = 0;\n+\tstruct packed_git *p;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tDIR *dir;\n+\tchar *name;\n+\tint ret;\n+\n+\tprepare_packed_git();\n+\tfor (p = packed_git; p; p = p->next) {\n+\t\tif (open_pack_index(p))\n+\t\t\tcontinue;\n+\t\tcount += p->num_objects;\n+\t}\n+\n+\tstrbuf_addstr(&buf, get_object_directory());\n+\tstrbuf_addstr(&buf, \"/42/\");\n+\tname = strbuf_detach(&buf, NULL);\n+\tdir = opendir(name);\n+\tfree(name);\n+\tif (dir) {\n+\t\tstruct dirent *de;\n+\t\twhile ((de = readdir(dir)) != NULL) {\n+\t\t\tcount += 256;\n+\t\t}\n+\t\tclosedir(dir);\n+\t}\n+\tfor (ret = 7; ret < 15; ret++) {\n+\t\tunsigned long expect_collision = 1ul << (ret * 2);\n+\t\tif (count < expect_collision)\n+\t\t\tbreak;\n+\t}\n+\treturn ret;\n+}\n+\n+int get_default_abbrev(void)\n+{\n+\tif (default_abbrev < 0)\n+\t\tdefault_abbrev = init_default_abbrev();\n+\treturn default_abbrev;\n+}\n"},{"id":"302960","messageId":"xmqqbmz6cna5.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFyVEQ+8TBBUm5KG9APtd9wy8cp_mRO=3nj12DXZNLAC9A@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T23:13:38Z","receivedAt":"2016-09-29T23:13:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Somebody should really double-check my heuristics, to see that I did\n> the pack counting etc right.  It doesn't do alternate loose file\n> counting at all, and maybe it could matter.  The advantage of the\n> previous patch was that it got the object counting right almost\n> automatically, this actually has its own new object counting code and\n> maybe I screwed it up.\n\nOne thing that worries me is if we are ready to start accessing the\nobject store in all codepaths when we ask for DEFAULT_ABBREV.  The\nworries are twofold:\n\n (1) Do we do the right thing if object store is not available to\n     us?  Some commands can be run outside repository, and if our\n     call to prepare_packed_git() or loose object iteration barfed\n     in some way, that would introduce a regression.\n\n (2) Is calling prepare_packed_git() too early interfere with how\n     the commands expect its own prepare_packed_git() work?  That\n     is, if a command has this sequence, \"ask DEFAULT_ABBREV,\n     arrange things, and then call prepare_packed_git()\", and the\n     existing \"arrange things\" step had something that causes a new\n     pack to become eligible to be read by prepare_packed_git(),\n     like adding to the list of alternate object stores, its own\n     prepare_packed_git() will now become a no-op.\n\nI browsed through \"tig grep DEFAULT_ABBREV \\*.c\" and it seems that\nin majority of the hits, we not just are ready to start accessing,\nbut already have an object or two, which must have come from an\nalready open object store, so they are OK.  Especially the ones that\nuse it as the last argument to find_unique_abbrev() are OK as we are\nabout to open the object store to do the computation.\n\nThere are very early ones in the program startup sequence in the\nfollowing functions, but I do not think of a reason why our new and\nearly call to prepare_packed_git() might be problematic, given that\nall of them require us to have an access to the repository (i.e.\nthis change cannot introduce a regression where a command used to\nwork outside a repository but barf when prepare_packed_git() is\ncalled early):\n\n - builtin/describe.c\n - builtin/rev-list.c\n - builtin/rev-parse.c\n\nI thought that the one in diff.c might be problematic when the \"git\ndiff\" command is run outside a repository with the \"--no-index\"\noption, but it appears that init_default_abbrev() seems to be OK\nwhen run outside a repository.\n\nThere is one in parse-options-cb.c that is used to parse the --abbrev\ncommand line option.  This might cause a cosmetic problem but when\nthe user is asking for an abbreviation, it is expected that we will\nhave an access to the object store anyway, so it may be OK.\n\nI am sorry that none of the above is about your math ;-)  I suck at\nmath so I won't comment.\n\n"},{"id":"302961","messageId":"xmqq7f9ucmyk.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"xmqqbmz6cna5.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-29T23:20:35Z","receivedAt":"2016-09-29T23:20:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> The advantage of the\n>> previous patch was that it got the object counting right almost\n>> automatically, this actually has its own new object counting code and\n>> maybe I screwed it up.\n\nI guess another advantage of your original approach was that it\ndelayed the counting to the very last minute, so the things that\nworried me in my previous response were automatically made\nnon-issues.\n"},{"id":"302967","messageId":"CA+55aFysvNc4p_nFcV=edctCizJBJtDjFOHJa-YYgVZQgBZfiA@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqbmz6cna5.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T00:20:47Z","receivedAt":"2016-09-30T00:20:54Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 4:13 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> One thing that worries me is if we are ready to start accessing the\n> object store in all codepaths when we ask for DEFAULT_ABBREV.\n\nYes. That was my main worry too. I also looked at just doing an explicit\n\n     if (abbrev_commit && default_abbrev < 0)\n          default_abbrev = get_default_abbrev();\n\nand in many ways that would be nicer exactly because the point where\nthis happens is then explicit, instead of being hidden behind that\nmacro that may end up being done in random places.\n\nBut it wasn't entirely obvious which all paths would need that\ninitialization either, so on the whole it was very much a \"six of one,\nhalf a dozen of the other\" thing.\n\nAs you say, my original patch had neither of those issues. It just\nstupidly re-did the loop over and over, and maybe the right thing to\ndo is to have that original code, but just short-circuit the \"over and\nover\" behavior by just resetting default_abbrev to the value we do\nfind.\n\n              Linus\n"},{"id":"302968","messageId":"CA+55aFyXxQSygO-gqevLZDjuggOaHs7HsRO=P6GhpC3GStqwvQ@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFysvNc4p_nFcV=edctCizJBJtDjFOHJa-YYgVZQgBZfiA@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T00:28:55Z","receivedAt":"2016-09-30T00:29:00Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 5:20 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> As you say, my original patch had neither of those issues.\n\nTo be fair, my original patch had a different worry that I didn't\nbother with: what if one of the _other_ callers of \"get_short_sha1()\"\npassed in -1 to it.  I only handled the -1 case in th eone path care\nabout in that first RFC for testing. So I'm *not* suggesting you\nshould apply my first version,, It has issues too.\n\nLet me see if I can massage my first hacky RFC test-patch into\nsomething more reliable.\n\n              Linus\n"},{"id":"302969","messageId":"CA+55aFxsfxvDQqi2M3TUVvAHUx3Qm1hHQ4DMyzXzN6V2v7o-3A@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFyXxQSygO-gqevLZDjuggOaHs7HsRO=P6GhpC3GStqwvQ@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T00:57:49Z","receivedAt":"2016-09-30T00:57:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 5:28 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> To be fair, my original patch had a different worry that I didn't\n> bother with: what if one of the _other_ callers of \"get_short_sha1()\"\n> passed in -1 to it.  I only handled the -1 case in th eone path care\n> about in that first RFC for testing. So I'm *not* suggesting you\n> should apply my first version,, It has issues too.\n\nActually, all the other cases seem to be \"parse a SHA1 with a known\nlength\", so they really don't have a negative length.  So this seems\nok, and is easier to verify than the \"what all contexts might use\nDEFAULT_ABBREV\" thing. There's only a few callers, and it's a static\nfunction so it's easy to check it locally in sha1_name.c.\n\n               Linus\n"},{"id":"302970","messageId":"CA+55aFzjTB0peMDPoPA6JyeUy90x=Lh4qdfiYLNf6RQU3ey9Hg@mail.gmail.com","threadId":"44156","inReplyTo":"20160930005638.almd66ralshknoxa@glandium.org","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T01:01:40Z","receivedAt":"2016-09-30T01:01:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 5:56 PM, Mike Hommey <mh@glandium.org> wrote:\n>\n> OTOH, how often does one refer to trees or blobs with abbreviated sha1s?\n> Most of the time, you'd use abbreviated sha1s for commits. And the number\n> of commits in git and the kernel repositories are much lower than the\n> number of overall objects.\n\nSee that whole other discussion about this. I agree. If we only ever\nworried about just commits, the abbreviation length wouldn't need to\nbe grown nearly as aggressively. The current default would still be\nwrong for the kernel, but it wouldn't be as noticeably wrong, and\nupdating it to 8 or 9 would be fine.\n\nThat said, people argued against that too. We *do* end up having\nabbreviated SHA1's for blobs in the diff index. When I said that _I_\nneer use it, somebody piped up to say that they do.\n\nSo I'd rather just keep the existing semantics (a hash is a hash is a\nhash), and just abbreviate at a sufficient point that we don't have to\nworry too much about disambiguating further by object type.\n\n      Linus\n"},{"id":"302971","messageId":"CA+55aFyHn0Q-qPq4dPEJ7X_4jf5UbsVw2vE-4LoWYbPn6gS10g@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFxsfxvDQqi2M3TUVvAHUx3Qm1hHQ4DMyzXzN6V2v7o-3A@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T01:18:03Z","receivedAt":"2016-09-30T01:18:10Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 5:57 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> Actually, all the other cases seem to be \"parse a SHA1 with a known\n> length\", so they really don't have a negative length.  So this seems\n> ok, and is easier to verify than the \"what all contexts might use\n> DEFAULT_ABBREV\" thing. There's only a few callers, and it's a static\n> function so it's easy to check it locally in sha1_name.c.\n\nHere's my original patch with just a tiny change that instead of\nstarting the automatic guessing at 7 each time, it starts at\n\"default_automatic_abbrev\", which is initialized to 7.\n\nThe difference is that if we decide that \"oh, that was too small, need\nto repeat\", we also update that \"default_automatic_abbrev\" value, so\nthat we won't start at the number that we now know was too small.\n\nSo it still loops over the abbrev values, but now it only loops a\ncouple of times.\n\nI actually verified the performance impact by doing\n\n      time git rev-list --abbrev-commit HEAD > /dev/null\n\non the kernel git tree, and it does actually matter. With my original\npatch, we wasted a noticeable amount of time on just the extra\nlooping, with this it's down to the same performance as just doing it\nonce at init time (it's about 12s vs 9s on my laptop).\n\nSo this patch may actually be \"production ready\" apart from the fact\nthat some tests still fail (at least t2027-worktree-list.sh) because\nof different short SHA1 cases.\n\n                     Linus\n\n\n cache.h       |  1 +\n environment.c |  2 +-\n sha1_name.c   | 26 +++++++++++++++++++++++++-\n 3 files changed, 27 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 6e33f2f28..d2da6d186 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1207,6 +1207,7 @@ struct object_context {\n #define GET_SHA1_TREEISH          020\n #define GET_SHA1_BLOB             040\n #define GET_SHA1_FOLLOW_SYMLINKS 0100\n+#define GET_SHA1_AUTOMATIC\t 0200\n #define GET_SHA1_ONLY_TO_DIE    04000\n \n #define GET_SHA1_DISAMBIGUATORS \\\ndiff --git a/environment.c b/environment.c\nindex c1442df9a..fd6681e46 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -16,7 +16,7 @@ int trust_executable_bit = 1;\n int trust_ctime = 1;\n int check_stat = 1;\n int has_symlinks = 1;\n-int minimum_abbrev = 4, default_abbrev = 7;\n+int minimum_abbrev = 4, default_abbrev = -1;\n int ignore_case;\n int assume_unchanged;\n int prefer_symlink_refs;\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 3b647fd7c..1003c96ea 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -15,6 +15,7 @@ typedef int (*disambiguate_hint_fn)(const unsigned char *, void *);\n \n struct disambiguate_state {\n \tint len; /* length of prefix in hex chars */\n+\tunsigned int nrobjects;\n \tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n \n@@ -118,6 +119,12 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \n \t\t\tif (strlen(de->d_name) != 38)\n \t\t\t\tcontinue;\n+\n+\t\t\t// We only look at the one subdirectory, and we assume\n+\t\t\t// each subdirectory is roughly similar, so each object\n+\t\t\t// we find probably has 255 other objects in the other\n+\t\t\t// fan-out directories\n+\t\t\tds->nrobjects += 256;\n \t\t\tif (memcmp(de->d_name, ds->hex_pfx + 2, ds->len - 2))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(hex + 2, de->d_name, 38);\n@@ -151,6 +158,7 @@ static void unique_in_pack(struct packed_git *p,\n \n \topen_pack_index(p);\n \tnum = p->num_objects;\n+\tds->nrobjects += num;\n \tlast = num;\n \twhile (first < last) {\n \t\tuint32_t mid = (first + last) / 2;\n@@ -380,6 +388,9 @@ static int show_ambiguous_object(const unsigned char *sha1, void *data)\n \treturn 0;\n }\n \n+// Why seven? That's our historical default before the automatic abbreviation\n+static int default_automatic_abbrev = 7;\n+\n static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\t\t  unsigned flags)\n {\n@@ -426,6 +437,14 @@ static int get_short_sha1(const char *name, int len, unsigned char *sha1,\n \t\tfor_each_abbrev(ds.hex_pfx, show_ambiguous_object, &ds);\n \t}\n \n+\tif (len < 16 && !status && (flags & GET_SHA1_AUTOMATIC)) {\n+\t\tunsigned int expect_collision = 1 << (len * 2);\n+\t\tif (ds.nrobjects > expect_collision) {\n+\t\t\tdefault_automatic_abbrev = len+1;\n+\t\t\treturn SHORT_NAME_AMBIGUOUS;\n+\t\t}\n+\t}\n+\n \treturn status;\n }\n \n@@ -458,14 +477,19 @@ int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n int find_unique_abbrev_r(char *hex, const unsigned char *sha1, int len)\n {\n \tint status, exists;\n+\tint flags = GET_SHA1_QUIETLY;\n \n+\tif (len < 0) {\n+\t\tflags |= GET_SHA1_AUTOMATIC;\n+\t\tlen = default_automatic_abbrev;\n+\t}\n \tsha1_to_hex_r(hex, sha1);\n \tif (len == 40 || !len)\n \t\treturn 40;\n \texists = has_sha1_file(sha1);\n \twhile (len < 40) {\n \t\tunsigned char sha1_ret[20];\n-\t\tstatus = get_short_sha1(hex, len, sha1_ret, GET_SHA1_QUIETLY);\n+\t\tstatus = get_short_sha1(hex, len, sha1_ret, flags);\n \t\tif (exists\n \t\t    ? !status\n \t\t    : status == SHORT_NAME_NOT_FOUND) {\n"},{"id":"302972","messageId":"20160930005638.almd66ralshknoxa@glandium.org","threadId":"44156","inReplyTo":"CA+55aFx9Utm9yDZceks+5q9c8ydc2QMYshWwJ0G0GHWWLwSsXQ@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2016-09-30T00:56:38Z","receivedAt":"2016-09-30T01:29:10Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Thu, Sep 29, 2016 at 12:06:23PM -0700, Linus Torvalds wrote:\n> On Thu, Sep 29, 2016 at 11:55 AM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> >\n> > For the kernel, just the *math* right now actually gives 12\n> > characters. For current git it actually seems to say that 8 is the\n> > correct number. For small projects, you'll still see 7.\n> \n> Sorry, the git number is 9, not 8. The reason is that git has roughly\n> 212k objects, and 9 hex digits gets expected collisions at about 256k\n> objects.\n> \n> So the logic means that we'll see 7 hex digits for projects with less\n> than 16k objects, 8 hex digits if there are less than 64k objects, and\n> 9 hex digits for projects like git that currently have fewer than 256k\n> objects.\n> \n> But git itself might not be *that* far from going to 10 hex digits\n> with my patch.\n> \n> The kernel uses 12 he digits because the collision math says that's\n> the right thing for a project with between 4M and 16M objects (with\n> the kernel being at 5M).\n\nOTOH, how often does one refer to trees or blobs with abbreviated sha1s?\nMost of the time, you'd use abbreviated sha1s for commits. And the number\nof commits in git and the kernel repositories are much lower than the\nnumber of overall objects.\n\nrev-list --all --count on the git repo gives me 46790. On the kernel, it\ngives 618078.\n\nNow, the interesting thing is looking at the *actual* collisions in\nthose spaces.\n\nAt 9 digits, there's only one commit collision in the kernel repo:\n  45f014c5264f5e68ef0e51b36f4ef5ede3d18397\n  45f014c52eef022873b19d6a20eb0ec9668f2b09\n\nAnd two commit collisions at 8 digits in the git repo:\n  1536dd9c1df0b7167b139f6666080cc4774ef63f\n  1536dd9c61b5582cf079999057cb715dd6dc6620\n\n  2e6e3e82ee36b3e1bec1db8db24817270080424e\n  2e6e3e829f3759823d70e7af511bc04cd05ad0af\n\nAt 7 digits, there are 5 actual commit collisions in the git repo and\n718 in the kernel repo only one of those collisions involve more than 2\ncommits.\n\nMike\n"},{"id":"302973","messageId":"xmqqtwcyavou.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFyHn0Q-qPq4dPEJ7X_4jf5UbsVw2vE-4LoWYbPn6gS10g@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T03:54:57Z","receivedAt":"2016-09-30T03:55:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> So this patch may actually be \"production ready\" apart from the fact\n> that some tests still fail (at least t2027-worktree-list.sh) because\n> of different short SHA1 cases.\n\nt2027 has at least two problems.\n\n * \"git worktree\" does not read the core.abbrev configuration,\n   without a recent fix in jc/worktree-config, i.e. d49028e6\n   (\"worktree: honor configuration variables\", 2016-09-26).\n\n * The script uses \"git rev-parse --short HEAD\"; I suspect that it\n   says \"ah, default_abbrev is -1 and minimum_abbrev is 4, so let's\n   try abbreviating to 4 hexdigits\".\n\nThe first failure in t3203 seems to come from the same issue in\n\"rev-parse --short\".\n"},{"id":"302974","messageId":"xmqqoa36auyx.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"xmqqtwcyavou.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T04:10:30Z","receivedAt":"2016-09-30T04:11:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> So this patch may actually be \"production ready\" apart from the fact\n>> that some tests still fail (at least t2027-worktree-list.sh) because\n>> of different short SHA1 cases.\n>\n> t2027 has at least two problems.\n>\n>  * \"git worktree\" does not read the core.abbrev configuration,\n>    without a recent fix in jc/worktree-config, i.e. d49028e6\n>    (\"worktree: honor configuration variables\", 2016-09-26).\n>\n>  * The script uses \"git rev-parse --short HEAD\"; I suspect that it\n>    says \"ah, default_abbrev is -1 and minimum_abbrev is 4, so let's\n>    try abbreviating to 4 hexdigits\".\n>\n> The first failure in t3203 seems to come from the same issue in\n> \"rev-parse --short\".\n\nA quick and dirty fix for it may look like this.\n\nWe leave the variable abbrev to DEFAULT_ABBREV and let\nfind_unique_abbrev() react to \"eh, -1? I need to do the\nauto-scaling\".  \"git diff-tree --abbrev\" seems to have a similar\nproblem, and the fix is the same.\n\nThere still are breakages seen in t5510 and t5526 that are about the\nverbose output of \"git fetch\".  I'll stop digging at this point\ntonight, and welcome others who look into it ;-)\n\n builtin/rev-parse.c | 14 ++++++++------\n diff.c              |  2 +-\n 2 files changed, 9 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 76cf05e2ad..f8c8c6c22e 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -642,13 +642,15 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t    starts_with(arg, \"--short=\")) {\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n-\t\t\t\tabbrev = DEFAULT_ABBREV;\n-\t\t\t\tif (arg[7] == '=')\n+\t\t\t\tif (arg[7] != '=') {\n+\t\t\t\t\tabbrev = DEFAULT_ABBREV;\n+\t\t\t\t} else {\n \t\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n-\t\t\t\tif (abbrev < MINIMUM_ABBREV)\n-\t\t\t\t\tabbrev = MINIMUM_ABBREV;\n-\t\t\t\telse if (40 <= abbrev)\n-\t\t\t\t\tabbrev = 40;\n+\t\t\t\t\tif (abbrev < MINIMUM_ABBREV)\n+\t\t\t\t\t\tabbrev = MINIMUM_ABBREV;\n+\t\t\t\t\telse if (40 <= abbrev)\n+\t\t\t\t\t\tabbrev = 40;\n+\t\t\t\t}\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (!strcmp(arg, \"--sq\")) {\ndiff --git a/diff.c b/diff.c\nindex c6da383c56..cefc13eb8e 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3399,7 +3399,7 @@ void diff_setup_done(struct diff_options *options)\n \t\t\t */\n \t\t\tread_cache();\n \t}\n-\tif (options->abbrev <= 0 || 40 < options->abbrev)\n+\tif (40 < options->abbrev)\n \t\toptions->abbrev = 40; /* full */\n \n \t/*\n"},{"id":"302975","messageId":"CA+55aFwsWxJqaK3mCZsjF6qcRmBKD0ffM9iW-1A6rWn05H5DFw@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqtwcyavou.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T04:11:11Z","receivedAt":"2016-09-30T04:11:23Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 8:54 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n>  * The script uses \"git rev-parse --short HEAD\"; I suspect that it\n>    says \"ah, default_abbrev is -1 and minimum_abbrev is 4, so let's\n>    try abbreviating to 4 hexdigits\".\n\nAhh, right you are. The logic there is\n\n                                abbrev = DEFAULT_ABBREV;\n                                if (arg[7] == '=')\n                                        abbrev = strtoul(arg + 8, NULL, 10);\n                                if (abbrev < MINIMUM_ABBREV)\n                                        abbrev = MINIMUM_ABBREV;\n                                ....\n\nwhich now does something different than what it used to do because\nDEFAULT_ABBREV is -1.\n\nPutting the \"sanity-check the abbrev range\" tests inside the \"if()\"\nstatement that does strtoul() should fix it. Let me test...\n\n[ short time passes ]\n\nYup. Incremental patch for that single issue attached.  I made it do\nan early \"continue\" instead of adding another level on indentation.\n\n                 Linus\n\n\n builtin/rev-parse.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 4da1f1da2..cfb0f1510 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -671,8 +671,9 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n \t\t\t\tabbrev = DEFAULT_ABBREV;\n-\t\t\t\tif (arg[7] == '=')\n-\t\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n+\t\t\t\tif (!arg[7])\n+\t\t\t\t\tcontinue;\n+\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n \t\t\t\tif (abbrev < MINIMUM_ABBREV)\n \t\t\t\t\tabbrev = MINIMUM_ABBREV;\n \t\t\t\telse if (40 <= abbrev)\n"},{"id":"302976","messageId":"CA+55aFw4=tGQZd0QO_8Zzs0AqPCpew_Wvnwft-JP2OzFbask8w@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqoa36auyx.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T04:18:06Z","receivedAt":"2016-09-30T04:18:12Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 9:10 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> A quick and dirty fix for it may look like this.\n\nCrossed emails.\n\nIndeed, I just solved the builtin/rev-parse.c thing slightly differently.\n\nAnd you found another failure in the diff code similarly not liking\nthe negative DEFAULT_ABBREV.  There are probably other things like\nthat.\n\n              Linus\n"},{"id":"302977","messageId":"xmqqeg42au5w.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"xmqqoa36auyx.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T04:27:55Z","receivedAt":"2016-09-30T04:28:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> There still are breakages seen in t5510 and t5526 that are about the\n> verbose output of \"git fetch\".  I'll stop digging at this point\n> tonight, and welcome others who look into it ;-)\n\nOK, just before I leave the keyboard for the night...\n\n-- >8 --\nFrom: Junio C Hamano <gitster@pobox.com>\nDate: Thu, 29 Sep 2016 21:19:20 -0700\nSubject: [PATCH] abbrev: adjust to the new world order\n\nThe default_abbrev used to be a concrete value usable as the default\nabbreviation length.  The code that sets custom abbreviation length,\nin response to command line argument, often did something like:\n\n\tif (skip_prefix(arg, \"--abbrev=\", &arg))\n\t\tabbrev = atoi(arg);\n\telse if (!strcmp(\"--abbrev\", &arg))\n\t\tabbrev = DEFAULT_ABBREV;\n\t/* make the value sane */\n\tif (abbrev < 0 || 40 < abbrev)\n\t\tabbrev = ... some sane value ...\n\nThe new world order however is that the default_abbrev is a negative\nvalue that signals find_unique_abbrev() that it needs to dynamically\nfind out a good default value.  We shouldn't coerce a negative value\ninto a random positive value like the above sample code.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n builtin/rev-parse.c | 5 +++--\n diff.c              | 2 +-\n 2 files changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 76cf05e2ad..17cbfabdde 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -643,8 +643,9 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n \t\t\t\tabbrev = DEFAULT_ABBREV;\n-\t\t\t\tif (arg[7] == '=')\n-\t\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n+\t\t\t\tif (!arg[7])\n+\t\t\t\t\tcontinue;\n+\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n \t\t\t\tif (abbrev < MINIMUM_ABBREV)\n \t\t\t\t\tabbrev = MINIMUM_ABBREV;\n \t\t\t\telse if (40 <= abbrev)\ndiff --git a/diff.c b/diff.c\nindex c6da383c56..cefc13eb8e 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3399,7 +3399,7 @@ void diff_setup_done(struct diff_options *options)\n \t\t\t */\n \t\t\tread_cache();\n \t}\n-\tif (options->abbrev <= 0 || 40 < options->abbrev)\n+\tif (40 < options->abbrev)\n \t\toptions->abbrev = 40; /* full */\n \n \t/*\n-- \n2.10.0-612-g22341905f2\n\n"},{"id":"302978","messageId":"CA+55aFz4uEzVw4yc+2X4=UPaRewxBkO9gbfXAPQ96kauQ883Zw@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFw4=tGQZd0QO_8Zzs0AqPCpew_Wvnwft-JP2OzFbask8w@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T04:29:49Z","receivedAt":"2016-09-30T04:29:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Sep 29, 2016 at 9:18 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> There are probably other things like that.\n\nt5510-fetch.sh fails oddly, looks like the output is off by one character.\n\n   not ok 77 - fetch aligned output\n\nIt has a magic \"cut -c 22-\" that expects the output at a specific\nplace, and now it's at column 21 instead of column 22. Strange test,\nbut it still seems to be aligned, just in a different column.\n\nBut clearly something changed.\n\n             Linus\n"},{"id":"302979","messageId":"xmqqa8eqatt1.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"xmqqeg42au5w.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T04:35:38Z","receivedAt":"2016-09-30T04:35:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> There still are breakages seen in t5510 and t5526 that are about the\n>> verbose output of \"git fetch\".  I'll stop digging at this point\n>> tonight, and welcome others who look into it ;-)\n>\n> OK, just before I leave the keyboard for the night...\n>\n> -- >8 --\n> From: Junio C Hamano <gitster@pobox.com>\n> Date: Thu, 29 Sep 2016 21:19:20 -0700\n> Subject: [PATCH] abbrev: adjust to the new world order\n\nTo those who are following from sidelines, this builds on Linus's\nthird iteration patch (which is based on his first patch), applied\non Peff's \"give disambiguation help when giving an ambiguity error\"\nseries.  I didn't merge the work-in-progress going back and forth\nbetween Linus and I tonight to any of the integration branches, but\nit is available as lt/abbrev-auto-2 branch of the \"broken down\"\nrepository, i.e.\n\n    git://github.com/gitster/git.git lt/abbrev-auto-2\n\n"},{"id":"302980","messageId":"CA+P7+xr4ZNCCJkS0=yR-FNu+MrL60YX-+Wsz9L_5LCNhnY_d=A@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqbmz6hbdk.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 10/10] get_short_sha1: list ambiguous objects on error","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2016-09-30T05:51:20Z","receivedAt":"2016-09-30T05:51:46Z","isPatch":true,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Thu, Sep 29, 2016 at 10:19 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> Jeff King <peff@peff.net> writes:\n>>   - \"cat-file --batch-check\" can show you the sha1 and type, but it\n>>     won't abbreviate sha1s, and it won't show you commit/tag information\n>>\n>>   - \"log --stdin --no-walk\" will format the commit however you like, but\n>>     skips the trees and blobs entirely, and the tag can only be seen via\n>>     \"%d\"\n>>\n>>   - \"for-each-ref\" has flexible formatting, too, but wants to format\n>>     refs, not objects (and doesn't read from stdin).\n>\n>     - \"name-rev\" is used to give \"describe --contains\", and can read\n>       from its standard input, but has no format customization.\n>       Another downside of it is that it only wants to see\n>       committishes.\n>\n\nSome tool which reads standard input and can be formatted would be\nnice. Extending name-rev with the same format options as for-each-ref\nwould be nice.\n\nThanks,\nJake\n"},{"id":"302982","messageId":"20160930074708.pthqg4ttl5rtpy3i@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqqbmz6cna5.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-30T07:47:08Z","receivedAt":"2016-09-30T07:47:17Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 04:13:38PM -0700, Junio C Hamano wrote:\n\n> There are very early ones in the program startup sequence in the\n> following functions, but I do not think of a reason why our new and\n> early call to prepare_packed_git() might be problematic, given that\n> all of them require us to have an access to the repository (i.e.\n> this change cannot introduce a regression where a command used to\n> work outside a repository but barf when prepare_packed_git() is\n> called early):\n> \n>  - builtin/describe.c\n>  - builtin/rev-list.c\n>  - builtin/rev-parse.c\n> \n> I thought that the one in diff.c might be problematic when the \"git\n> diff\" command is run outside a repository with the \"--no-index\"\n> option, but it appears that init_default_abbrev() seems to be OK\n> when run outside a repository.\n\nActually, \"diff --no-index\" is currently buggy in this regard. In the\nfollowup series to jk/setup-sequence-update (which I mentioned but\nhaven't posted yet), I teach get_object_dir() not to blindly default to\n\".git\", and found that \"diff --no-index\" is perfectly happy to look in\n\".git/objects\" for find_unique_abbrev(), even when we know there's no\nrepository (or it has an unknown vintage).\n\nI fixed it there by just using the default abbrev value for out-of-repo\ndiffs, and skip calling find_unique_abbrev() at all. That would here,\ntoo.\n\nBut if we add object-store initialization at other times, it's a\npotential conflict. IMHO this should stay inside find_unique_abbrev(),\nwhere we know we already must look at the object store.\n\n-Peff\n"},{"id":"302983","messageId":"20160930080658.lyi7aovvazjmy346@sigill.intra.peff.net","threadId":"44156","inReplyTo":"CA+55aFyHn0Q-qPq4dPEJ7X_4jf5UbsVw2vE-4LoWYbPn6gS10g@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-30T08:06:58Z","receivedAt":"2016-09-30T08:07:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 29, 2016 at 06:18:03PM -0700, Linus Torvalds wrote:\n\n> On Thu, Sep 29, 2016 at 5:57 PM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> >\n> > Actually, all the other cases seem to be \"parse a SHA1 with a known\n> > length\", so they really don't have a negative length.  So this seems\n> > ok, and is easier to verify than the \"what all contexts might use\n> > DEFAULT_ABBREV\" thing. There's only a few callers, and it's a static\n> > function so it's easy to check it locally in sha1_name.c.\n> \n> Here's my original patch with just a tiny change that instead of\n> starting the automatic guessing at 7 each time, it starts at\n> \"default_automatic_abbrev\", which is initialized to 7.\n> \n> The difference is that if we decide that \"oh, that was too small, need\n> to repeat\", we also update that \"default_automatic_abbrev\" value, so\n> that we won't start at the number that we now know was too small.\n> \n> So it still loops over the abbrev values, but now it only loops a\n> couple of times.\n> \n> I actually verified the performance impact by doing\n> \n>       time git rev-list --abbrev-commit HEAD > /dev/null\n> \n> on the kernel git tree, and it does actually matter. With my original\n> patch, we wasted a noticeable amount of time on just the extra\n> looping, with this it's down to the same performance as just doing it\n> once at init time (it's about 12s vs 9s on my laptop).\n\nI agree that this deals with the performance concerns by caching the\ndefault_abbrev_len and starting there. I still think it's unnecessarily\ninvasive to touch get_short_sha1() at all, which is otherwise only a\nreading function.\n\nSo IMHO, the best combination is the init_default_abbrev() you posted in\n[1], but initialized at the top of find_unique_abbrev(). And cached\nthere, obviously, in a similar way.\n\n-Peff\n\n[1] http://public-inbox.org/git/CA+55aFyVEQ+8TBBUm5KG9APtd9wy8cp_mRO=3nj12DXZNLAC9A@mail.gmail.com/\n"},{"id":"302993","messageId":"CA+55aFxW1S6FNUh8YjSXkfC8=F5dka1rY-As6PWfG2rqmrsXXA@mail.gmail.com","threadId":"44156","inReplyTo":"20160930080658.lyi7aovvazjmy346@sigill.intra.peff.net","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T17:54:16Z","receivedAt":"2016-09-30T17:54:22Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Sep 30, 2016 at 1:06 AM, Jeff King <peff@peff.net> wrote:\n>\n> I agree that this deals with the performance concerns by caching the\n> default_abbrev_len and starting there. I still think it's unnecessarily\n> invasive to touch get_short_sha1() at all, which is otherwise only a\n> reading function.\n\nSo the reason that d oesn't work is that the \"disambiguate_state\" data\nwhere we keep the number of objects is only visible within\nget_short_sha1().\n\nSo outside that function, you don't have any sane way to figure out\nhow many objects. So then you have to do the extra counting function..\n\n> So IMHO, the best combination is the init_default_abbrev() you posted in\n> [1], but initialized at the top of find_unique_abbrev(). And cached\n> there, obviously, in a similar way.\n\nThat's certainly possible, but I'm really not happy with how the\ncounting function looks.  And nobody actually stood up to say \"yeah,\nthat gets alternate loose objects right\" or \"if you have tons of those\nalternate loose objects you have other issues anyway\". I think\nsomebody would have to \"own\" that counting function, the advantage of\njust putting it into disambiguate_state is that we just get the\ncounting for free..\n\n                         Linus\n"},{"id":"302995","messageId":"xmqqr3819sqx.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"20160930080658.lyi7aovvazjmy346@sigill.intra.peff.net","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T17:56:06Z","receivedAt":"2016-09-30T17:56:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I agree that this deals with the performance concerns by caching the\n> default_abbrev_len and starting there. I still think it's unnecessarily\n> invasive to touch get_short_sha1() at all, which is otherwise only a\n> reading function.\n>\n> So IMHO, the best combination is the init_default_abbrev() you posted in\n> [1], but initialized at the top of find_unique_abbrev(). And cached\n> there, obviously, in a similar way.\n\nHmm. I am undecided; both approaches look OK to me.\n\n"},{"id":"302996","messageId":"20160930180504.p6ueuxj6g7qzcbwq@sigill.intra.peff.net","threadId":"44156","inReplyTo":"CA+55aFxW1S6FNUh8YjSXkfC8=F5dka1rY-As6PWfG2rqmrsXXA@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-09-30T18:05:04Z","receivedAt":"2016-09-30T18:05:12Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 30, 2016 at 10:54:16AM -0700, Linus Torvalds wrote:\n\n> On Fri, Sep 30, 2016 at 1:06 AM, Jeff King <peff@peff.net> wrote:\n> >\n> > I agree that this deals with the performance concerns by caching the\n> > default_abbrev_len and starting there. I still think it's unnecessarily\n> > invasive to touch get_short_sha1() at all, which is otherwise only a\n> > reading function.\n> \n> So the reason that d oesn't work is that the \"disambiguate_state\" data\n> where we keep the number of objects is only visible within\n> get_short_sha1().\n> \n> So outside that function, you don't have any sane way to figure out\n> how many objects. So then you have to do the extra counting function..\n\nRight. I think you should do the extra counting function. It's a few\nmore lines, but the design is way less tangled.\n\n> > So IMHO, the best combination is the init_default_abbrev() you posted in\n> > [1], but initialized at the top of find_unique_abbrev(). And cached\n> > there, obviously, in a similar way.\n> \n> That's certainly possible, but I'm really not happy with how the\n> counting function looks.  And nobody actually stood up to say \"yeah,\n> that gets alternate loose objects right\" or \"if you have tons of those\n> alternate loose objects you have other issues anyway\". I think\n> somebody would have to \"own\" that counting function, the advantage of\n> just putting it into disambiguate_state is that we just get the\n> counting for free..\n\nI don't think you _need_ get the alternate loose objects right. In fact,\nI don't think you need to care about loose objects at all. For the\nscales we're talking about, they're a rounding error. I would have done\nit like this:\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 65deaf9..1845502 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1382,6 +1382,32 @@ static void prepare_packed_git_one(char *objdir, int local)\n \tstrbuf_release(&path);\n }\n \n+static int approximate_object_count_valid;\n+\n+/*\n+ * Give a fast, rough count of the number of objects in the repository. This\n+ * ignores loose objects completely. If you have a lot of them, then either\n+ * you should repack because your performance will be awful, or they are\n+ * all unreachable objects about to be pruned, in which case they're not really\n+ * interesting as a measure of repo size in the first place.\n+ */\n+unsigned long approximate_object_count(void)\n+{\n+\tstatic unsigned long count;\n+\tif (!approximate_object_count_valid) {\n+\t\tstruct packed_git *p;\n+\n+\t\tprepare_packed_git();\n+\t\tcount = 0;\n+\t\tfor (p = packed_git; p; p = p->next) {\n+\t\t\tif (open_pack_index(p))\n+\t\t\t\tcontinue;\n+\t\t\tcount += p->num_objects;\n+\t\t}\n+\t}\n+\treturn count;\n+}\n+\n static void *get_next_packed_git(const void *p)\n {\n \treturn ((const struct packed_git *)p)->next;\n@@ -1456,6 +1482,7 @@ void prepare_packed_git(void)\n \n void reprepare_packed_git(void)\n {\n+\tapproximate_object_count_valid = 0;\n \tprepare_packed_git_run_once = 0;\n \tprepare_packed_git();\n }\n"},{"id":"302998","messageId":"CA+55aFxyF=xX84AXr8MG14MRHwdrQw00PBM20UfqBdidaeqdMg@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFxW1S6FNUh8YjSXkfC8=F5dka1rY-As6PWfG2rqmrsXXA@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T18:21:08Z","receivedAt":"2016-09-30T18:21:14Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Sep 30, 2016 at 10:54 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n>> So IMHO, the best combination is the init_default_abbrev() you posted in\n>> [1], but initialized at the top of find_unique_abbrev(). And cached\n>> there, obviously, in a similar way.\n>\n> That's certainly possible, but I'm really not happy with how the\n> counting function looks.  And nobody actually stood up to say \"yeah,\n> that gets alternate loose objects right\" or \"if you have tons of those\n> alternate loose objects you have other issues anyway\". I think\n> somebody would have to \"own\" that counting function, the advantage of\n> just putting it into disambiguate_state is that we just get the\n> counting for free..\n\nSide note: maybe we can mix the two approaches, and keep the counting\nin the disambiguation state, and just make the counting function do\n\n    init_object_disambiguation();\n    find_short_object_filename(&ds);\n    find_short_packed_object(&ds);\n    finish_object_disambiguation(&ds, sha1);\n\nand then just use \"ds.nrobjects\". So the counting would still be done\nby the disambiguation code, it just woudln't be in get_short_sha1().\n\nSo here's another version that takes that approach. And if somebody\n(hint hint) wants to do the counting differently, they can perhaps\nsend an incremental patch to do that.\n\n(This patch also contains the few setup issues Junio found with the\nnew \"default_abbrev is negative\" model)\n\n              Linus\n\n\n builtin/rev-parse.c |  5 +++--\n diff.c              |  2 +-\n environment.c       |  2 +-\n sha1_name.c         | 39 ++++++++++++++++++++++++++++++++++++++-\n 4 files changed, 43 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 4da1f1da2..cfb0f1510 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -671,8 +671,9 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n \t\t\t\tabbrev = DEFAULT_ABBREV;\n-\t\t\t\tif (arg[7] == '=')\n-\t\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n+\t\t\t\tif (!arg[7])\n+\t\t\t\t\tcontinue;\n+\t\t\t\tabbrev = strtoul(arg + 8, NULL, 10);\n \t\t\t\tif (abbrev < MINIMUM_ABBREV)\n \t\t\t\t\tabbrev = MINIMUM_ABBREV;\n \t\t\t\telse if (40 <= abbrev)\ndiff --git a/diff.c b/diff.c\nindex 59920747d..c6d445915 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3421,7 +3421,7 @@ void diff_setup_done(struct diff_options *options)\n \t\t\t */\n \t\t\tread_cache();\n \t}\n-\tif (options->abbrev <= 0 || 40 < options->abbrev)\n+\tif (options->abbrev > 40)\n \t\toptions->abbrev = 40; /* full */\n \n \t/*\ndiff --git a/environment.c b/environment.c\nindex c1442df9a..fd6681e46 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -16,7 +16,7 @@ int trust_executable_bit = 1;\n int trust_ctime = 1;\n int check_stat = 1;\n int has_symlinks = 1;\n-int minimum_abbrev = 4, default_abbrev = 7;\n+int minimum_abbrev = 4, default_abbrev = -1;\n int ignore_case;\n int assume_unchanged;\n int prefer_symlink_refs;\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 3b647fd7c..684b36dba 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -15,6 +15,7 @@ typedef int (*disambiguate_hint_fn)(const unsigned char *, void *);\n \n struct disambiguate_state {\n \tint len; /* length of prefix in hex chars */\n+\tunsigned int nrobjects;\n \tchar hex_pfx[GIT_SHA1_HEXSZ + 1];\n \tunsigned char bin_pfx[GIT_SHA1_RAWSZ];\n \n@@ -118,6 +119,12 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \n \t\t\tif (strlen(de->d_name) != 38)\n \t\t\t\tcontinue;\n+\n+\t\t\t// We only look at the one subdirectory, and we assume\n+\t\t\t// each subdirectory is roughly similar, so each object\n+\t\t\t// we find probably has 255 other objects in the other\n+\t\t\t// fan-out directories\n+\t\t\tds->nrobjects += 256;\n \t\t\tif (memcmp(de->d_name, ds->hex_pfx + 2, ds->len - 2))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(hex + 2, de->d_name, 38);\n@@ -151,6 +158,7 @@ static void unique_in_pack(struct packed_git *p,\n \n \topen_pack_index(p);\n \tnum = p->num_objects;\n+\tds->nrobjects += num;\n \tlast = num;\n \twhile (first < last) {\n \t\tuint32_t mid = (first + last) / 2;\n@@ -455,17 +463,46 @@ int for_each_abbrev(const char *prefix, each_abbrev_fn fn, void *cb_data)\n \treturn ret;\n }\n \n+static int get_automatic_abbrev(const char *hex)\n+{\n+\tstatic int len;\n+\tstruct disambiguate_state ds;\n+\n+\tif (init_object_disambiguation(hex, 7, &ds) < 0)\n+\t\treturn 7;\n+\n+\tfind_short_object_filename(&ds);\n+\tfind_short_packed_object(&ds);\n+\n+\tfor (len = 7; len < 16; len++) {\n+\t\tunsigned int expect_collision = 1 << (len * 2);\n+\t\tif (ds.nrobjects < expect_collision)\n+\t\t\tbreak;\n+\t}\n+\treturn len;\n+}\n+\n int find_unique_abbrev_r(char *hex, const unsigned char *sha1, int len)\n {\n \tint status, exists;\n+\tint flags = GET_SHA1_QUIETLY;\n \n \tsha1_to_hex_r(hex, sha1);\n \tif (len == 40 || !len)\n \t\treturn 40;\n+\n+\tif (len < 0) {\n+\t\tstatic int automatic_abbrev = -1;\n+\n+\t\tif (automatic_abbrev < 0)\n+\t\t\tautomatic_abbrev = get_automatic_abbrev(hex);\n+\t\tlen = automatic_abbrev;\n+\t}\n+\n \texists = has_sha1_file(sha1);\n \twhile (len < 40) {\n \t\tunsigned char sha1_ret[20];\n-\t\tstatus = get_short_sha1(hex, len, sha1_ret, GET_SHA1_QUIETLY);\n+\t\tstatus = get_short_sha1(hex, len, sha1_ret, flags);\n \t\tif (exists\n \t\t    ? !status\n \t\t    : status == SHORT_NAME_NOT_FOUND) {\n"},{"id":"303000","messageId":"xmqqmvip9qo7.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"xmqqeg42au5w.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T18:40:56Z","receivedAt":"2016-09-30T18:41:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> From: Junio C Hamano <gitster@pobox.com>\n> Date: Thu, 29 Sep 2016 21:19:20 -0700\n> Subject: [PATCH] abbrev: adjust to the new world order\n>\n> The default_abbrev used to be a concrete value usable as the default\n> abbreviation length.  The code that sets custom abbreviation length,\n> in response to command line argument, often did something like:\n>\n> \tif (skip_prefix(arg, \"--abbrev=\", &arg))\n> \t\tabbrev = atoi(arg);\n> \telse if (!strcmp(\"--abbrev\", &arg))\n> \t\tabbrev = DEFAULT_ABBREV;\n> \t/* make the value sane */\n> \tif (abbrev < 0 || 40 < abbrev)\n> \t\tabbrev = ... some sane value ...\n>\n> The new world order however is that the default_abbrev is a negative\n> value that signals find_unique_abbrev() that it needs to dynamically\n> find out a good default value.  We shouldn't coerce a negative value\n> into a random positive value like the above sample code.\n>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n\nThere is another instance buried deep in an obscure macro.  A\nminimum fix may look like this, but I really hope somebody else\nfinds a better approach.  Peff alluded to \"when it is still -1\nsubstituting it with a reasonable value like 7\" in a separate\nthread, and we probably would want a way to allow accessing that\n\"reasonable value like 7\" without triggering auto sizing logic\ntoo early.\n\nWith this and the patch in the message I am responding to, your\npatch from the last night seems to pass all the tests for me.\n\n transport.h | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/transport.h b/transport.h\nindex 6fe3485325..8a96e22bb0 100644\n--- a/transport.h\n+++ b/transport.h\n@@ -142,7 +142,7 @@ struct transport {\n #define TRANSPORT_PUSH_ATOMIC 8192\n #define TRANSPORT_PUSH_OPTIONS 16384\n \n-#define TRANSPORT_SUMMARY_WIDTH (2 * DEFAULT_ABBREV + 3)\n+#define TRANSPORT_SUMMARY_WIDTH (2 * (DEFAULT_ABBREV < 0 ? 7 : DEFAULT_ABBREV) + 3)\n #define TRANSPORT_SUMMARY(x) (int)(TRANSPORT_SUMMARY_WIDTH + strlen(x) - gettext_width(x)), (x)\n \n /* Returns a transport suitable for the url */\n"},{"id":"303001","messageId":"CA+55aFyDqYCBCvw0MjZ8fNhaVbRSjsSXNDH--unYkoeJNwVcVg@mail.gmail.com","threadId":"44156","inReplyTo":"xmqqmvip9qo7.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2016-09-30T18:51:15Z","receivedAt":"2016-09-30T18:51:37Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Sep 30, 2016 at 11:40 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> There is another instance buried deep in an obscure macro.  A\n> minimum fix may look like this, but I really hope somebody else\n> finds a better approach.\n\nHeh. Yeah, that's just ugly. I assume this is why the odd git fetch\npretty-printing test was off by one column..\n\nConsidering that TRANSPORT_SUMMARY and TRANSPORT_SUMMARY_WIDTH are\nboth used in exactly one place each, I'd suggest getting rid of that\ncrazy macro, and just expanding it in those places to avoid these\nkinds of crazy \"hiding variables inside complex defines thning\".\n\nAnd maybe just deciding to hardcode TRANSPORT_SUMMARY_WIDTH to 17\n(which was it's original default value and presumably is what the test\nis effectively hardcoded for too), and avoiding that complexity\nentirely.\n\n                Linus\n"},{"id":"303003","messageId":"xmqqintd9prc.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFyDqYCBCvw0MjZ8fNhaVbRSjsSXNDH--unYkoeJNwVcVg@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T19:00:39Z","receivedAt":"2016-09-30T19:00:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Considering that TRANSPORT_SUMMARY and TRANSPORT_SUMMARY_WIDTH are\n> both used in exactly one place each, I'd suggest getting rid of that\n> crazy macro, and just expanding it in those places to avoid these\n> kinds of crazy \"hiding variables inside complex defines thning\".\n>\n> And maybe just deciding to hardcode TRANSPORT_SUMMARY_WIDTH to 17\n> (which was it's original default value and presumably is what the test\n> is effectively hardcoded for too), and avoiding that complexity\n> entirely.\n\nFor all fairness, when the WIDTH thing was introduced, there were\ntwo places that needed reference it at f1863d0d16 (\"refactor\nduplicated code in builtin-send-pack.c and transport.c\",\n2010-02-16).  But that is no longer the case, and it makes sense to\nhardcode it as 17 (or something derived from a symbolic constant\nthat gives the new \"default to default\").\n\nWhat TRANSPORT_SUMMARY() does is even more crazy and it really\nshouldn't be exposed as a public interface.  Let's move it to its\nsingle calling place.\n"},{"id":"303014","messageId":"CACBZZX4p8KEHnkDQ-c2La-1rkgiBs47T+dCNOfpji3tKs3YhVA@mail.gmail.com","threadId":"44156","inReplyTo":"CA+55aFzjTB0peMDPoPA6JyeUy90x=Lh4qdfiYLNf6RQU3ey9Hg@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2016-09-30T19:41:48Z","receivedAt":"2016-09-30T19:42:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Fri, Sep 30, 2016 at 3:01 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> On Thu, Sep 29, 2016 at 5:56 PM, Mike Hommey <mh@glandium.org> wrote:\n>>\n>> OTOH, how often does one refer to trees or blobs with abbreviated sha1s?\n>> Most of the time, you'd use abbreviated sha1s for commits. And the number\n>> of commits in git and the kernel repositories are much lower than the\n>> number of overall objects.\n>\n> See that whole other discussion about this. I agree. If we only ever\n> worried about just commits, the abbreviation length wouldn't need to\n> be grown nearly as aggressively. The current default would still be\n> wrong for the kernel, but it wouldn't be as noticeably wrong, and\n> updating it to 8 or 9 would be fine.\n>\n> That said, people argued against that too. We *do* end up having\n> abbreviated SHA1's for blobs in the diff index. When I said that _I_\n> neer use it, somebody piped up to say that they do.\n>\n> So I'd rather just keep the existing semantics (a hash is a hash is a\n> hash), and just abbreviate at a sufficient point that we don't have to\n> worry too much about disambiguating further by object type.\n\nI work on a repo that's around the size of linux.git in every way\n(commits, objects etc.), and growing twice as fast.\n\nSo I also see 8 or 9 digit abbreviations on a daily basis, even with\nthe current defaults core.abbrev, but I still think growing it so\naggressively is the wrong thing to do.\n\nThe fact that we have a core.abbrev option at all and nobody's talking\nabout getting rid of it entirely means we all acknowledge the UX\nconvenience of short SHA1s.\n\nI don't think it's a good idea for such UX options to have defaults\nthat really only make sense for repositories at the very far end of\nthe bell curve, which is the case with linux.git and the repo I work\non.\n\nEither way you're going to waste somebody's time. I think it's a\nbetter trade-off that some kernel dev occasionally has to look at\nPeff's new disambiguation output, than have the wast hordes of\neveryday Git users have less screen real estate, need to recite longer\nsha1s over the phone during outages (people do that), and any number\nof other every day use cases.\n\nI think if anything we should be talking about making the default\nshorter & then have some clever auto-scaling by repository size as has\nbeen discussed in this thread to deal with the repositories at the far\nend of the bell curve.\n"},{"id":"303015","messageId":"xmqq4m4x9myd.fsf@gitster.mtv.corp.google.com","threadId":"44156","inReplyTo":"CA+55aFxyF=xX84AXr8MG14MRHwdrQw00PBM20UfqBdidaeqdMg@mail.gmail.com","subject":"Re: [PATCH 4/4] core.abbrev: raise the default abbreviation to 12 hexdigits","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-09-30T20:01:14Z","receivedAt":"2016-09-30T20:01:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Fri, Sep 30, 2016 at 10:54 AM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n>>\n>>> So IMHO, the best combination is the init_default_abbrev() you posted in\n>>> [1], but initialized at the top of find_unique_abbrev(). And cached\n>>> there, obviously, in a similar way.\n>>\n>> That's certainly possible, but I'm really not happy with how the\n>> counting function looks.  And nobody actually stood up to say \"yeah,\n>> that gets alternate loose objects right\" or \"if you have tons of those\n>> alternate loose objects you have other issues anyway\". I think\n>> somebody would have to \"own\" that counting function, the advantage of\n>> just putting it into disambiguate_state is that we just get the\n>> counting for free..\n>\n> Side note: maybe we can mix the two approaches, and keep the counting\n> in the disambiguation state, and just make the counting function do\n>\n>     init_object_disambiguation();\n>     find_short_object_filename(&ds);\n>     find_short_packed_object(&ds);\n>     finish_object_disambiguation(&ds, sha1);\n>\n> and then just use \"ds.nrobjects\". So the counting would still be done\n> by the disambiguation code, it just woudln't be in get_short_sha1().\n>\n> So here's another version that takes that approach. And if somebody\n> (hint hint) wants to do the counting differently, they can perhaps\n> send an incremental patch to do that.\n>\n> (This patch also contains the few setup issues Junio found with the\n> new \"default_abbrev is negative\" model)\n\nSorry, but I do not quite see the point in the difference between\nthis one and your original that had a hook in get_short_sha1(), as\nit seemed to me that Peff's objection was about the counting done in\nfind_short_object_filename() and find_short_packed_object(), which\nis (understandably) still here.\n\n\n\n"},{"id":"368461","messageId":"20190204161217.20047-1-avarab@gmail.com","threadId":"44156","inReplyTo":"20160926043442.3pz7ccawdcsn2kzb@sigill.intra.peff.net","subject":"[RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-02-04T16:12:17Z","receivedAt":"2019-02-04T16:12:33Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The algorithm we use to pick the default abbreviation length as a\nfunction of the approximate number of objects is described in the\ncommit message for e6c587c733 (\"abbrev: auto size the default\nabbreviation\", 2016-09-30), as well as in and downthread of [1], but\nit hasn't been documented.\n\nLet's do that, and while we're at it explicitly test for when the\ncurrent implementation will \"roll over\" up to values of 2^32-1 (the\nmaximum portable \"unsigned long\" value).\n\n1. https://public-inbox.org/git/20160926043442.3pz7ccawdcsn2kzb@sigill.intra.peff.net/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\nThis is a patch from the middle of a series I'm currently working on\nre-rolling. See\nhttps://public-inbox.org/git/20180608224136.20220-1-avarab@gmail.com/\n\nWhat I'd like to get here is commentary on the phrasing and accuracy\nof the doc patch I'm adding here.\n\nThis patch assumes that we have a abbrev_length_for_object_count()\nfunction, which I've added in an eariler unpublished patch. It just\nexposes the length picking algorithm found in find_unique_abbrev_r().\n\n Documentation/config/core.txt       | 17 +++++++\n builtin/rev-parse.c                 |  8 ++++\n t/t1512-rev-parse-disambiguation.sh | 74 +++++++++++++++++++++++++++++\n 3 files changed, 99 insertions(+)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 185857a13f..2175761833 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -599,6 +599,23 @@ core.abbrev::\n \tabbreviated object names to stay unique for some time.\n \tThe minimum length is 4.\n +\n+The algorithm to pick the the current abbreviation length is\n+considered an implementation detail, and might be changed in the\n+future. Since Git version 2.11, the length has been configured to\n+auto-scale based on the estimated number of objects in the\n+repository. We pick a length such that if all objects in the\n+repository were abbreviated, we'd have a 50% chance of a *single*\n+collision.\n++\n+For example, with 2^14-1 is the last object count at which we'll pick\n+a short length of \"7\", and will roll over to \"8\" once we have one more\n+object at 2^14. Since each hexdigit we add (4 bits) allows us to have\n+four times (2 bits) as many objects in the repository, we'll roll over\n+to a length of \"9\" at 2^16 objects, \"10\" at 2^18 etc. We'll never\n+automatically pick a length less than \"7\", which effectively hardcodes\n+2^12 as the minimum number of objects in a repository we'll consider\n+when choosing the abbreviation length.\n++\n This can also be set to relative values such as `+2` or `-2`, which\n means to add or subtract N characters from the SHA-1 that Git would\n otherwise print, this allows for producing more future-proof SHA-1s\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex d0d751a009..e7bf4375a2 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -773,6 +773,14 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\t\treturn 1;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (opt_with_value(arg, \"--abbrev-len\", &arg)) {\n+\t\t\t\tunsigned long v;\n+\t\t\t\tif (!git_parse_ulong(arg, &v))\n+\t\t\t\t\treturn 1;\n+\t\t\t\tint len = abbrev_length_for_object_count(v);\n+\t\t\t\tprintf(\"%d\\n\", len);\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tif (!strcmp(arg, \"--bisect\")) {\n \t\t\t\tfor_each_fullref_in(\"refs/bisect/bad\", show_reference, NULL, 0);\n \t\t\t\tfor_each_fullref_in(\"refs/bisect/good\", anti_reference, NULL, 0);\ndiff --git a/t/t1512-rev-parse-disambiguation.sh b/t/t1512-rev-parse-disambiguation.sh\nindex 265a6972fc..0e97888a44 100755\n--- a/t/t1512-rev-parse-disambiguation.sh\n+++ b/t/t1512-rev-parse-disambiguation.sh\n@@ -450,4 +450,78 @@ test_expect_success C_LOCALE_OUTPUT 'ambiguous commits are printed by type first\n \tdone\n '\n \n+test_expect_success 'abbreviation length at 2^N-1 and 2^N' '\n+\tpow_2_min=$(git rev-parse --abbrev-len=3) &&\n+\tpow_2_eql=$(git rev-parse --abbrev-len=4) &&\n+\tpow_4_min=$(git rev-parse --abbrev-len=15) &&\n+\tpow_4_eql=$(git rev-parse --abbrev-len=16) &&\n+\tpow_6_min=$(git rev-parse --abbrev-len=63) &&\n+\tpow_6_eql=$(git rev-parse --abbrev-len=64) &&\n+\tpow_8_min=$(git rev-parse --abbrev-len=255) &&\n+\tpow_8_eql=$(git rev-parse --abbrev-len=256) &&\n+\tpow_10_min=$(git rev-parse --abbrev-len=1023) &&\n+\tpow_10_eql=$(git rev-parse --abbrev-len=1024) &&\n+\tpow_12_min=$(git rev-parse --abbrev-len=4095) &&\n+\tpow_12_eql=$(git rev-parse --abbrev-len=4096) &&\n+\tpow_14_min=$(git rev-parse --abbrev-len=16383) &&\n+\tpow_14_eql=$(git rev-parse --abbrev-len=16384) &&\n+\tpow_16_min=$(git rev-parse --abbrev-len=65535) &&\n+\tpow_16_eql=$(git rev-parse --abbrev-len=65536) &&\n+\tpow_18_min=$(git rev-parse --abbrev-len=262143) &&\n+\tpow_18_eql=$(git rev-parse --abbrev-len=262144) &&\n+\tpow_20_min=$(git rev-parse --abbrev-len=1048575) &&\n+\tpow_20_eql=$(git rev-parse --abbrev-len=1048576) &&\n+\tpow_22_min=$(git rev-parse --abbrev-len=4194303) &&\n+\tpow_22_eql=$(git rev-parse --abbrev-len=4194304) &&\n+\tpow_24_min=$(git rev-parse --abbrev-len=16777215) &&\n+\tpow_24_eql=$(git rev-parse --abbrev-len=16777216) &&\n+\tpow_26_min=$(git rev-parse --abbrev-len=67108863) &&\n+\tpow_26_eql=$(git rev-parse --abbrev-len=67108864) &&\n+\tpow_28_min=$(git rev-parse --abbrev-len=268435455) &&\n+\tpow_28_eql=$(git rev-parse --abbrev-len=268435456) &&\n+\tpow_30_min=$(git rev-parse --abbrev-len=1073741823) &&\n+\tpow_30_eql=$(git rev-parse --abbrev-len=1073741824) &&\n+\tpow_32_min=$(git rev-parse --abbrev-len=4294967295) &&\n+\n+\tcat >actual <<-EOF &&\n+\t2 = $pow_2_min $pow_2_eql\n+\t4 = $pow_4_min $pow_4_eql\n+\t6 = $pow_6_min $pow_6_eql\n+\t8 = $pow_8_min $pow_8_eql\n+\t10 = $pow_10_min $pow_10_eql\n+\t12 = $pow_12_min $pow_12_eql\n+\t14 = $pow_14_min $pow_14_eql\n+\t16 = $pow_16_min $pow_16_eql\n+\t18 = $pow_18_min $pow_18_eql\n+\t20 = $pow_20_min $pow_20_eql\n+\t22 = $pow_22_min $pow_22_eql\n+\t24 = $pow_24_min $pow_24_eql\n+\t26 = $pow_26_min $pow_26_eql\n+\t28 = $pow_28_min $pow_28_eql\n+\t30 = $pow_30_min $pow_30_eql\n+\t32 = 16\n+\tEOF\n+\n+\tcat >expected <<-\\EOF &&\n+\t2 = 7 7\n+\t4 = 7 7\n+\t6 = 7 7\n+\t8 = 7 7\n+\t10 = 7 7\n+\t12 = 7 7\n+\t14 = 7 8\n+\t16 = 8 9\n+\t18 = 9 10\n+\t20 = 10 11\n+\t22 = 11 12\n+\t24 = 12 13\n+\t26 = 13 14\n+\t28 = 14 15\n+\t30 = 15 16\n+\t32 = 16\n+\tEOF\n+\n+\ttest_cmp expected actual\n+'\n+\n test_done\n-- \n2.20.1.611.gfbb209baf1\n\n"},{"id":"368476","messageId":"xmqq7eefv02i.fsf@gitster-ct.c.googlers.com","threadId":"44156","inReplyTo":"20190204161217.20047-1-avarab@gmail.com","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-04T19:13:41Z","receivedAt":"2019-02-04T19:13:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> +The algorithm to pick the the current abbreviation length is\n> +considered an implementation detail, and might be changed in the\n> +future. Since Git version 2.11, the length has been configured to\n> +auto-scale based on the estimated number of objects in the\n> +repository. We pick a length such that if all objects in the\n> +repository were abbreviated, we'd have a 50% chance of a *single*\n> +collision.\n\nCorrect and reads well.\n\n> +For example, with 2^14-1 is the last object count at which we'll pick\n> +a short length of \"7\", and will roll over to \"8\" once we have one more\n> +object at 2^14. Since each hexdigit we add (4 bits) allows us to have\n> +four times (2 bits) as many objects in the repository\n\nSomething is missing at this point in the sentence. \n\n\t\"without raising the chance of a single collision higher\"\n\nor something like that.\n\n> , we'll roll over\n> +to a length of \"9\" at 2^16 objects, \"10\" at 2^18 etc.\n\nCorrect and reads well.\n\n> We'll never\n> +automatically pick a length less than \"7\", which effectively hardcodes\n> +2^12 as the minimum number of objects in a repository we'll consider\n> +when choosing the abbreviation length.\n\nThis may be technicaly correct, but to me, it seems to place stress\non the wrong side of the equation.  Since nobody would find \"Ah, so\nI can create up to 2^12 objects without fearing that my abbreviated\nobject name would become longer than 7\", I do not see much point in\nsaying \"hardcoded floor for the number of objects\".\n\nOn the other hand, saying that 7 is the hardcoded floor for the\nabbreviation length does make sense, as those adept at math after\nreading the paragraph up to this point would wonder why their tiny\nrepository still uses 7 hexdigits, which is way too many to ensure\nthe low collision rate for the size of their toy repository.\n\n\tWe do not use abbreviation shorter than 7 hexdigits by default,\n\tso a small repository with less than 2^12 objects may have even\n\tsmaller chance than 50% to have a single collision.\n\nmay be an improvement.\n\n"},{"id":"368482","messageId":"xmqq36p3uxq2.fsf@gitster-ct.c.googlers.com","threadId":"44156","inReplyTo":"20190204161217.20047-1-avarab@gmail.com","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-04T20:04:21Z","receivedAt":"2019-02-04T20:04:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> @@ -773,6 +773,14 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n>  \t\t\t\t\treturn 1;\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n> +\t\t\tif (opt_with_value(arg, \"--abbrev-len\", &arg)) {\n> +\t\t\t\tunsigned long v;\n> +\t\t\t\tif (!git_parse_ulong(arg, &v))\n> +\t\t\t\t\treturn 1;\n> +\t\t\t\tint len = abbrev_length_for_object_count(v);\n> +\t\t\t\tprintf(\"%d\\n\", len);\n> +\t\t\t\tcontinue;\n> +\t\t\t}\n\nInstead of exposing this pretty-much \"test-only\" feature as a new\noption to t/helper/test-tool, I think it is OK, if not even better,\nto have it in rev-parse proper like this patch does.\n\nI however have a mildly strong suspition that people would expect\n\"rev-parse --abbrev-len=<num>\" to be a synonym of \"--short=<num>\"\n\nAs this is pretty-much a test-only option, perhaps going longer but\nmore descriptive would make sense?  \n\n\tgit rev-parse --compute-abbrev-length-for <object-count>\n\nmay be an overkill, but something along those lines.\n\nOh by the way, the code has decl-after-stmt, and perhaps len needs\nto be of type \"const int\" ;-)\n\n"},{"id":"368492","messageId":"87tvhjkzhz.fsf@evledraar.gmail.com","threadId":"44156","inReplyTo":"xmqq36p3uxq2.fsf@gitster-ct.c.googlers.com","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-02-04T21:36:08Z","receivedAt":"2019-02-04T21:36:13Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Feb 04 2019, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n>> @@ -773,6 +773,14 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n>>  \t\t\t\t\treturn 1;\n>>  \t\t\t\tcontinue;\n>>  \t\t\t}\n>> +\t\t\tif (opt_with_value(arg, \"--abbrev-len\", &arg)) {\n>> +\t\t\t\tunsigned long v;\n>> +\t\t\t\tif (!git_parse_ulong(arg, &v))\n>> +\t\t\t\t\treturn 1;\n>> +\t\t\t\tint len = abbrev_length_for_object_count(v);\n>> +\t\t\t\tprintf(\"%d\\n\", len);\n>> +\t\t\t\tcontinue;\n>> +\t\t\t}\n>\n> Instead of exposing this pretty-much \"test-only\" feature as a new\n> option to t/helper/test-tool, I think it is OK, if not even better,\n> to have it in rev-parse proper like this patch does.\n\nWhile I mainly added this code so I could prove the docs correct with a\ntest for both myself & others, I think having this exposed is probably\nuseful.\n\nI've seen more than once some feature of a web frontend for git where\nthere's both access to aggregate statistics (number of commits or\nobjects), and SHA-1 shortening going on, but the latter is just done via\nsubstr().\n\nRight now we have nothing directly exposed to answer \"what length would\ngit pick\", you can of course e.g. \"log --abbrev\" a single commit, but if\nthat commit happens to be more ambiguous than most you'll get the right\nanswer.\n\n\n> I however have a mildly strong suspition that people would expect\n> \"rev-parse --abbrev-len=<num>\" to be a synonym of \"--short=<num>\"\n>\n> As this is pretty-much a test-only option, perhaps going longer but\n> more descriptive would make sense?\n>\n> \tgit rev-parse --compute-abbrev-length-for <object-count>\n>\n> may be an overkill, but something along those lines.\n\nYeah I think that's better. This is so rare that it's better to be\nverbose.\n\n> Oh by the way, the code has decl-after-stmt, and perhaps len needs\n> to be of type \"const int\" ;-)\n\nThanks. Will fix.\n"},{"id":"368497","messageId":"20190204233235.GB2366@sigill.intra.peff.net","threadId":"44156","inReplyTo":"xmqq36p3uxq2.fsf@gitster-ct.c.googlers.com","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-02-04T23:32:36Z","receivedAt":"2019-02-04T23:32:42Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 04, 2019 at 12:04:21PM -0800, Junio C Hamano wrote:\n\n> Instead of exposing this pretty-much \"test-only\" feature as a new\n> option to t/helper/test-tool, I think it is OK, if not even better,\n> to have it in rev-parse proper like this patch does.\n> \n> I however have a mildly strong suspition that people would expect\n> \"rev-parse --abbrev-len=<num>\" to be a synonym of \"--short=<num>\"\n> \n> As this is pretty-much a test-only option, perhaps going longer but\n> more descriptive would make sense?  \n> \n> \tgit rev-parse --compute-abbrev-length-for <object-count>\n> \n> may be an overkill, but something along those lines.\n\nYou could even default <object-count> to the number of objects in the\nrepository. Which implies that perhaps the best spot is the command\nwhere we already count the number of objects, git-count-objects.\n\n-Peff\n"},{"id":"368499","messageId":"87r2cnkta8.fsf@evledraar.gmail.com","threadId":"44156","inReplyTo":"20190204233235.GB2366@sigill.intra.peff.net","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-02-04T23:50:23Z","receivedAt":"2019-02-04T23:50:28Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Feb 05 2019, Jeff King wrote:\n\n> On Mon, Feb 04, 2019 at 12:04:21PM -0800, Junio C Hamano wrote:\n>\n>> Instead of exposing this pretty-much \"test-only\" feature as a new\n>> option to t/helper/test-tool, I think it is OK, if not even better,\n>> to have it in rev-parse proper like this patch does.\n>>\n>> I however have a mildly strong suspition that people would expect\n>> \"rev-parse --abbrev-len=<num>\" to be a synonym of \"--short=<num>\"\n>>\n>> As this is pretty-much a test-only option, perhaps going longer but\n>> more descriptive would make sense?\n>>\n>> \tgit rev-parse --compute-abbrev-length-for <object-count>\n>>\n>> may be an overkill, but something along those lines.\n>\n> You could even default <object-count> to the number of objects in the\n> repository. Which implies that perhaps the best spot is the command\n> where we already count the number of objects, git-count-objects.\n\nThat's documented as reporting loose objects by default, although it has\na full report with -v.\n\nMaybe rev-parse isn't the right place, I just picked it because it seems\nto be the general utility belt for stuff that doesn't fit elsewhere.\n\nBut putting it in git-count-objects seems like a bit more of a stretch\ngiven the above.\n"},{"id":"368620","messageId":"20190206182950.GB10231@sigill.intra.peff.net","threadId":"44156","inReplyTo":"87r2cnkta8.fsf@evledraar.gmail.com","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-02-06T18:29:51Z","receivedAt":"2019-02-06T18:29:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 05, 2019 at 12:50:23AM +0100, Ævar Arnfjörð Bjarmason wrote:\n\n> >> As this is pretty-much a test-only option, perhaps going longer but\n> >> more descriptive would make sense?\n> >>\n> >> \tgit rev-parse --compute-abbrev-length-for <object-count>\n> >>\n> >> may be an overkill, but something along those lines.\n> >\n> > You could even default <object-count> to the number of objects in the\n> > repository. Which implies that perhaps the best spot is the command\n> > where we already count the number of objects, git-count-objects.\n> \n> That's documented as reporting loose objects by default, although it has\n> a full report with -v.\n\nTrue, though I think that's mostly for historical reasons. It _could_ be\npart of the full report, like:\n\n  $ git count-objects -v\n  ...\n  abbrev-len: 12\n\nbut from your test-script usage, I'd expect you'd want to be able to\nfeed a fake count to it, like:\n\n  git count-objects --compute-abbrev-len=1234\n\nor something (of course you _could_ also make a repository with N\nobjects, but that's a lot more expensive).\n\n> Maybe rev-parse isn't the right place, I just picked it because it seems\n> to be the general utility belt for stuff that doesn't fit elsewhere.\n> \n> But putting it in git-count-objects seems like a bit more of a stretch\n> given the above.\n\nI dunno. It seems like less of a stretch to me, but it is true that\nrev-parse is already a kitchen sink repository. I can live with it\neither way.\n\n-Peff\n"},{"id":"368622","messageId":"87ftt0lq6q.fsf@evledraar.gmail.com","threadId":"44156","inReplyTo":"20190206182950.GB10231@sigill.intra.peff.net","subject":"Re: [RFC/PATCH] core.abbrev doc: document and test the abbreviation length","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-02-06T18:36:29Z","receivedAt":"2019-02-06T18:36:34Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Feb 06 2019, Jeff King wrote:\n\n> On Tue, Feb 05, 2019 at 12:50:23AM +0100, Ævar Arnfjörð Bjarmason wrote:\n>\n>> >> As this is pretty-much a test-only option, perhaps going longer but\n>> >> more descriptive would make sense?\n>> >>\n>> >> \tgit rev-parse --compute-abbrev-length-for <object-count>\n>> >>\n>> >> may be an overkill, but something along those lines.\n>> >\n>> > You could even default <object-count> to the number of objects in the\n>> > repository. Which implies that perhaps the best spot is the command\n>> > where we already count the number of objects, git-count-objects.\n>>\n>> That's documented as reporting loose objects by default, although it has\n>> a full report with -v.\n>\n> True, though I think that's mostly for historical reasons. It _could_ be\n> part of the full report, like:\n>\n>   $ git count-objects -v\n>   ...\n>   abbrev-len: 12\n>\n> but from your test-script usage, I'd expect you'd want to be able to\n> feed a fake count to it, like:\n>\n>   git count-objects --compute-abbrev-len=1234\n\nYeah for just reporting it count-objects makes more sense. I think I'll\nadd it there...\n\n> or something (of course you _could_ also make a repository with N\n> objects, but that's a lot more expensive).\n\n...but yes, for the test script & to export the info I'd like to have\nthe \"what's the abbrev length for a repo with N objects\" option, which\nwould be for rev-parse.\n\n>> Maybe rev-parse isn't the right place, I just picked it because it seems\n>> to be the general utility belt for stuff that doesn't fit elsewhere.\n>>\n>> But putting it in git-count-objects seems like a bit more of a stretch\n>> given the above.\n>\n> I dunno. It seems like less of a stretch to me, but it is true that\n> rev-parse is already a kitchen sink repository. I can live with it\n> either way.\n>\n> -Peff\n"}]}