{"thread":{"id":"36970","subject":"[PATCH v5 0/2] cleanup duplicate name_compare() functions","startedAt":"2014-06-20T02:06:42Z","lastAt":"2014-06-20T17:15:27Z","messageCount":5,"participants":["Jeremiah Mahler","Junio C Hamano"],"isPatch":true,"patchVersion":5,"patchTotal":2},"messages":[{"id":"244734","messageId":"1403230004-11034-1-git-send-email-jmmahler@gmail.com","threadId":"36970","inReplyTo":null,"subject":"[PATCH v5 0/2] cleanup duplicate name_compare() functions","fromName":"Jeremiah Mahler","fromEmail":"jmmahler@gmail.com","sentAt":"2014-06-20T02:06:42Z","receivedAt":"2014-06-20T02:06:42Z","isPatch":true,"sender":{"key":"jmmahler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/154028?v=4"},"body":"Version 5 of the patch series to cleanup the duplicate name_compare()\nfunctions.\n\n  - name-hash.c had a call to cache_name_compare() but it required that\n    the lengths were equal.  Since cache_name_compare() is equivalent to\n    memcmp() when the lengths are equal, replace it with memcmp().\n    This avoids renaming cache_name_compare() to name_compare() in a\n    later patch.\n\n  - Cleanup of log message by Junio C Humano.\n\n\nJeremiah Mahler (2):\n  name-hash.c: replace cache_name_compare() with memcmp()\n  cleanup duplicate name_compare() functions\n\n cache.h        |  2 +-\n dir.c          |  3 +--\n name-hash.c    |  2 +-\n read-cache.c   | 23 +++++++++++++----------\n tree-walk.c    | 10 ----------\n unpack-trees.c | 11 -----------\n 6 files changed, 16 insertions(+), 35 deletions(-)\n\n-- \n2.0.0.694.g5736dad\n"},{"id":"244735","messageId":"1403230004-11034-2-git-send-email-jmmahler@gmail.com","threadId":"36970","inReplyTo":"1403230004-11034-1-git-send-email-jmmahler@gmail.com","subject":"[PATCH v5 1/2] name-hash.c: replace cache_name_compare() with memcmp()","fromName":"Jeremiah Mahler","fromEmail":"jmmahler@gmail.com","sentAt":"2014-06-20T02:06:43Z","receivedAt":"2014-06-20T02:06:43Z","isPatch":true,"sender":{"key":"jmmahler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/154028?v=4"},"body":"When cache_name_compare() is used on counted strings of the same\nlength, it is equivalent to a memcmp().  Since the one use of\ncache_name_compare() in name-hash.c requires that the lengths are\nequal, just replace it with memcmp().\n\nSigned-off-by: Jeremiah Mahler <jmmahler@gmail.com>\n---\n name-hash.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/name-hash.c b/name-hash.c\nindex be7c4ae..63cc188 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -179,7 +179,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \t * Always do exact compare, even if we want a case-ignoring comparison;\n \t * we do the quick exact one first, because it will be the common case.\n \t */\n-\tif (len == namelen && !cache_name_compare(name, namelen, ce->name, len))\n+\tif (len == namelen && !memcmp(name, ce->name, len))\n \t\treturn 1;\n \n \tif (!icase)\n-- \n2.0.0.694.g5736dad\n"},{"id":"244736","messageId":"1403230004-11034-3-git-send-email-jmmahler@gmail.com","threadId":"36970","inReplyTo":"1403230004-11034-1-git-send-email-jmmahler@gmail.com","subject":"[PATCH v5 2/2] cleanup duplicate name_compare() functions","fromName":"Jeremiah Mahler","fromEmail":"jmmahler@gmail.com","sentAt":"2014-06-20T02:06:44Z","receivedAt":"2014-06-20T02:06:44Z","isPatch":true,"sender":{"key":"jmmahler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/154028?v=4"},"body":"We often represent our strings as a counted string, i.e. a pair of the\npointer to the beginning of the string and its length, and the string\nmay not be NUL terminated to that length.\n\nTo compare a pair of such counted strings, unpack-trees.c and\nread-cache.c implement their own name_compare() functions identically.\nIn addition, the cache_name_compare() function in read-cache.c is nearly\nidentical.  The only difference is when one string is the prefix of the\nother string, in which case the former returns -1/+1 to show which one\nis longer and the latter returns the difference of the lengths to show\nthe same information.\n\nUnify these three functions by using the implementation from\ncache_name_compare().  This does not make any difference to the existing\nand future callers, as they must be paying attention only to the sign of\nthe returned value (and not the magnitude) because the original\nimplementations of these two functions return values returned by\nmemcmp(3) when the one string is not a prefix of the other string, and\nthe only thing memcmp(3) guarantees its callers is the sign of the\nreturned value, not the magnitude.\n\nSigned-off-by: Jeremiah Mahler <jmmahler@gmail.com>\n---\n cache.h        |  2 +-\n dir.c          |  3 +--\n read-cache.c   | 23 +++++++++++++----------\n tree-walk.c    | 10 ----------\n unpack-trees.c | 11 -----------\n 5 files changed, 15 insertions(+), 34 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex c498a30..e3205fe 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1027,7 +1027,7 @@ extern int validate_headref(const char *ref);\n \n extern int base_name_compare(const char *name1, int len1, int mode1, const char *name2, int len2, int mode2);\n extern int df_name_compare(const char *name1, int len1, int mode1, const char *name2, int len2, int mode2);\n-extern int cache_name_compare(const char *name1, int len1, const char *name2, int len2);\n+extern int name_compare(const char *name1, size_t len1, const char *name2, size_t len2);\n extern int cache_name_stage_compare(const char *name1, int len1, int stage1, const char *name2, int len2, int stage2);\n \n extern void *read_object_with_reference(const unsigned char *sha1,\ndiff --git a/dir.c b/dir.c\nindex 797805d..e65888d 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -1354,8 +1354,7 @@ static int cmp_name(const void *p1, const void *p2)\n \tconst struct dir_entry *e1 = *(const struct dir_entry **)p1;\n \tconst struct dir_entry *e2 = *(const struct dir_entry **)p2;\n \n-\treturn cache_name_compare(e1->name, e1->len,\n-\t\t\t\t  e2->name, e2->len);\n+\treturn name_compare(e1->name, e1->len, e2->name, e2->len);\n }\n \n static struct path_simplify *create_simplify(const char **pathspec)\ndiff --git a/read-cache.c b/read-cache.c\nindex 9f56d76..158241d 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -434,18 +434,26 @@ int df_name_compare(const char *name1, int len1, int mode1,\n \treturn c1 - c2;\n }\n \n-int cache_name_stage_compare(const char *name1, int len1, int stage1, const char *name2, int len2, int stage2)\n+int name_compare(const char *name1, size_t len1, const char *name2, size_t len2)\n {\n-\tint len = len1 < len2 ? len1 : len2;\n-\tint cmp;\n-\n-\tcmp = memcmp(name1, name2, len);\n+\tsize_t min_len = (len1 < len2) ? len1 : len2;\n+\tint cmp = memcmp(name1, name2, min_len);\n \tif (cmp)\n \t\treturn cmp;\n \tif (len1 < len2)\n \t\treturn -1;\n \tif (len1 > len2)\n \t\treturn 1;\n+\treturn 0;\n+}\n+\n+int cache_name_stage_compare(const char *name1, int len1, int stage1, const char *name2, int len2, int stage2)\n+{\n+\tint cmp;\n+\n+\tcmp = name_compare(name1, len1, name2, len2);\n+\tif (cmp)\n+\t\treturn cmp;\n \n \tif (stage1 < stage2)\n \t\treturn -1;\n@@ -454,11 +462,6 @@ int cache_name_stage_compare(const char *name1, int len1, int stage1, const char\n \treturn 0;\n }\n \n-int cache_name_compare(const char *name1, int len1, const char *name2, int len2)\n-{\n-\treturn cache_name_stage_compare(name1, len1, 0, name2, len2, 0);\n-}\n-\n static int index_name_stage_pos(const struct index_state *istate, const char *name, int namelen, int stage)\n {\n \tint first, last;\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 4dc86c7..5dd9a71 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -144,16 +144,6 @@ struct tree_desc_x {\n \tstruct tree_desc_skip *skip;\n };\n \n-static int name_compare(const char *a, int a_len,\n-\t\t\tconst char *b, int b_len)\n-{\n-\tint len = (a_len < b_len) ? a_len : b_len;\n-\tint cmp = memcmp(a, b, len);\n-\tif (cmp)\n-\t\treturn cmp;\n-\treturn (a_len - b_len);\n-}\n-\n static int check_entry_match(const char *a, int a_len, const char *b, int b_len)\n {\n \t/*\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 4a9cdf2..c4a97ca 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -629,17 +629,6 @@ static int unpack_failed(struct unpack_trees_options *o, const char *message)\n \treturn -1;\n }\n \n-/* NEEDSWORK: give this a better name and share with tree-walk.c */\n-static int name_compare(const char *a, int a_len,\n-\t\t\tconst char *b, int b_len)\n-{\n-\tint len = (a_len < b_len) ? a_len : b_len;\n-\tint cmp = memcmp(a, b, len);\n-\tif (cmp)\n-\t\treturn cmp;\n-\treturn (a_len - b_len);\n-}\n-\n /*\n  * The tree traversal is looking at name p.  If we have a matching entry,\n  * return it.  If name p is a directory in the index, do not return\n-- \n2.0.0.694.g5736dad\n"},{"id":"244984","messageId":"xmqqvbro7k54.fsf@gitster.dls.corp.google.com","threadId":"36970","inReplyTo":"1403230004-11034-2-git-send-email-jmmahler@gmail.com","subject":"Re: [PATCH v5 1/2] name-hash.c: replace cache_name_compare() with memcmp()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-06-20T17:10:56Z","receivedAt":"2014-06-20T17:10:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeremiah Mahler <jmmahler@gmail.com> writes:\n\n> When cache_name_compare() is used on counted strings of the same\n> length, it is equivalent to a memcmp().  Since the one use of\n> cache_name_compare() in name-hash.c requires that the lengths are\n> equal, just replace it with memcmp().\n\nI do not think it is not \"requires that the lengths are equal\"; it\nmerely is a premature optimization, as it wants to catch only the\ncase where the names are the same.\n\nYour patch is not wrong per-se, but with the above justification of\nyours, I would actually have expected to see it updated to use\n!cache_name_compare() and then later !name_compare().  That way, if\nit ever turns out that giving name_compare() semantics specific to\n\"name\" (as opposed to just byte-for-byte comparison given by\nmemcmp(3)) is a good idea, we will use that comparison with\nsemantics specific to \"name\"s here, without having to change it from\nmemcmp().\n\nHaving said all that, I think we see a more correct justification\nfor this change in the pre-context of the patch.  We want the exact\ncomparison, without any funky \"name\"-specific semantics in the\nquick-and-exact case.\n\nI've queued it like this (no need to reroll).\nThanks.\n\n    name-hash.c: replace cache_name_compare() with memcmp(3)\n    \n    The same_name() private function wants a quick-and-exact check to\n    see if they two names are byte-for-byte identical first and then\n    fall back to the slow path.  Use memcmp(3) for the former to make it\n    clear that we do not want any \"name\" specific comparison.\n    \n    Signed-off-by: Jeremiah Mahler <jmmahler@gmail.com>\n    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n\ndiff --git a/name-hash.c b/name-hash.c\nindex 97444d0..49fd508 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -179,7 +179,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \t * Always do exact compare, even if we want a case-ignoring comparison;\n \t * we do the quick exact one first, because it will be the common case.\n \t */\n-\tif (len == namelen && !cache_name_compare(name, namelen, ce->name, len))\n+\tif (len == namelen && !memcmp(name, ce->name, len))\n \t\treturn 1;\n \n \tif (!icase)\n"},{"id":"244985","messageId":"xmqqpphw7k4y.fsf@gitster.dls.corp.google.com","threadId":"36970","inReplyTo":"1403230004-11034-3-git-send-email-jmmahler@gmail.com","subject":"Re: [PATCH v5 2/2] cleanup duplicate name_compare() functions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-06-20T17:15:27Z","receivedAt":"2014-06-20T17:15:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeremiah Mahler <jmmahler@gmail.com> writes:\n\n> We often represent our strings as a counted string, i.e. a pair of the\n> pointer to the beginning of the string and its length, and the string\n> may not be NUL terminated to that length.\n>\n> To compare a pair of such counted strings, unpack-trees.c and\n> read-cache.c implement their own name_compare() functions identically.\n> In addition, the cache_name_compare() function in read-cache.c is nearly\n> identical.  The only difference is when one string is the prefix of the\n> other string, in which case the former returns -1/+1 to show which one\n> is longer and the latter returns the difference of the lengths to show\n> the same information.\n\nI think I got the former/latter swapped by mistake when I wrote this\n(two name_compare() give us the difference, and cache_name_compare()\ngives -1/+1); I'll spell their names out when I queue this patch.\n\nThanks.\n"}]}