git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] git gc: Speed it up by 18% via faster hash comparisons

From
Erik Faye-Lund <kusmabite@gmail.com>
Date
Apr 28, 2011, 09:17 UTC
Message-ID
<BANLkTik_2sHZ0OTgQeHpRnpmNsAmT=sAcA@mail.gmail.com>
In-Reply-To
<20110428062717.GA952@elte.hu>
2011/4/28 Ingo Molnar <mingo@elte.hu>:
Show 16 quoted lines
>
> * Junio C Hamano <gitster@pobox.com> wrote:
>
>> Ingo Molnar <mingo@elte.hu> writes:
>>
>> > +static inline int hashcmp(const unsigned char *sha1, const unsigned char *sha2)
>> >  {
>> > -   return !memcmp(sha1, null_sha1, 20);
>> > +   int i;
>> > +
>> > +   for (i = 0; i < 20; i++, sha1++, sha2++) {
>> > +           if (*sha1 != *sha2) {
>> > +                   if (*sha1 < *sha2)
>> > +                           return -1;
>> > +                   return +1;
>> > +           }
Why not just:
if (*sha1 != *sha2)
        return *sha2 - *sha1;

memcmp isn't guaranteed to return onlt the values -1, 0, +1, it can return any value, just as long as it's sign of a non-zero return express the relationship between the first mis-matching byte.

Show 7 quoted lines
>> > +   }
>> > +
>> > +   return 0;
>>
>> This is very unfortunate, as it is so trivially correct and we shouldn't
>> have to do it.  If the compiler does not use a good inlined memcmp(), this
>> patch may fly, but I fear it may hurt other compilers, no?

If the common case is that the hashes are random (as the assumption in this patch is), then this patch should give very close to ideal performance, no? A good memcmp might be faster when there's a match; but how often do we really compare SHA-1's that are identical?

I see your worry, but if the assumption is correct, I doubt it'd turn out to be a real problem. ~99.6% of the time we'd early-out on the first byte, which is ideal.

If comparing identical SHA-1's are important, perhaps just having a early-out just on the first byte and then doing memcmp is a good solution (similar to what Jonathan Nieder proposed, but without alignment problems)?

> Secondly, the combined speedup of the cached case with my two patches appears
> to be more than 30% on my testbox so it's a very nifty win from two relatively
> simple changes.

That speed-up was on ONE test vector, no? There are a lot of other uses of hash-comparisons in Git, did you measure those?

Show 13 quoted lines
>> > +static inline int is_null_sha1(const unsigned char *sha1)
>> >  {
>> > -   return memcmp(sha1, sha2, 20);
>> > +   const unsigned long long *sha1_64 = (void *)sha1;
>> > +   const unsigned int *sha1_32 = (void *)sha1;
>>
>> Can everybody do unaligned accesses just fine?
>
> I have added some quick debug code and none of the sha1 pointers (in my
> admittedly very limited) testing showed misaligned pointers on 64-bit systems.
>
> On 32-bit systems the pointer might be 32-bit aligned only - the patch below
> implements the function 32-bit comparisons.

That's simply wrong. Unsigned char arrays can and will be unaligned, and this causes exceptions on most architectures (x86 is pretty much the exception here). While some systems for these architectures support unaligned reads from the exception handler, others doesn't. So this patch is pretty much guaranteed to cause a crash in some setups.

Previous: Ingo MolnarNext: Ingo Molnar
Message 15 of 45 in “git gc: Speed it up by 18% via faster hash comparisons”
  1. git gc: Speed it up by 18% via faster hash comparisonsIngo Molnar, Apr 27, 2011
  2. Ingo MolnarApr 27, 2011
  3. Jonathan NiederApr 27, 2011
  4. Ingo MolnarApr 28, 2011
  5. Jonathan NiederApr 28, 2011
  6. Ingo MolnarApr 28, 2011
  7. Dmitry PotapovApr 28, 2011
  8. Junio C HamanoApr 27, 2011
  9. Ralf BaechleApr 28, 2011
  10. Bernhard R. LinkApr 28, 2011
  11. Andreas EricssonApr 28, 2011
  12. Erik Faye-LundApr 28, 2011
  13. H. Peter AnvinApr 28, 2011
  14. Ingo MolnarApr 28, 2011
  15. Erik Faye-LundApr 28, 2011
  16. Ingo MolnarApr 28, 2011
  17. Ingo MolnarApr 28, 2011
  18. Erik Faye-LundApr 28, 2011
  19. Pekka EnbergApr 28, 2011
  20. Erik Faye-LundApr 28, 2011
  21. Pekka EnbergApr 28, 2011
  22. Erik Faye-LundApr 28, 2011
  23. Pekka EnbergApr 28, 2011
  24. Jonathan NiederApr 28, 2011
  25. Erik Faye-LundApr 28, 2011
  26. Ingo MolnarApr 28, 2011
  27. Ingo MolnarApr 28, 2011
  28. Erik Faye-LundApr 28, 2011
  29. Ingo MolnarApr 28, 2011
  30. Alex RiesenApr 29, 2011
  31. H. Peter AnvinApr 29, 2011
  32. Tor ArntsenApr 28, 2011
  33. H. Peter AnvinApr 28, 2011
  34. Andreas EricssonApr 28, 2011
  35. Erik Faye-LundApr 28, 2011
  36. Ingo MolnarApr 28, 2011
  37. Nguyen Thai Ngoc DuyApr 28, 2011
  38. Erik Faye-LundApr 28, 2011
  39. Junio C HamanoApr 28, 2011
  40. Dmitry PotapovApr 28, 2011
  41. Dmitry PotapovApr 28, 2011
  42. Ingo MolnarApr 28, 2011
  43. Dmitry PotapovApr 28, 2011
  44. Ingo MolnarApr 28, 2011
  45. Ingo MolnarApr 28, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.