git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v7 1/1] refs.c: SSE4.2 optimizations for check_refname_component

From
David Turner <dturner@twopensource.com>
Date
Jun 15, 2014, 05:53 UTC
Message-ID
<1402811620.5629.77.camel@stross>
In-Reply-To
<20140614152209.GA14125@domone.podge>
On Sat, 2014-06-14 at 17:22 +0200, Ondřej Bílka wrote:
Show 21 quoted lines
> On Thu, Jun 05, 2014 at 07:56:15PM -0400, David Turner wrote:
> > Optimize check_refname_component using SSE4.2, where available.
> > 
> > git rev-parse HEAD is a good test-case for this, since it does almost
> > nothing except parse refs.  For one particular repo with about 60k
> > refs, almost all packed, the timings are:
> > 
> > Look up table: 29 ms
> > SSE4.2:        25 ms
> > 
> > This is about a 15% improvement.
> > 
> > The configure.ac changes include code from the GNU C Library written
> > by Joseph S. Myers <joseph at codesourcery dot com>.
> > 
> > Only supports GCC and Clang at present, because C interfaces to the
> > cpuid instruction are not well-standardized.
> >
> Still a SSE4.2 is not that useful, in most cases SSE2 is faster. Here I
> think that difference will not be that big when correctly implemented.
> That will avoid a runtime checks.

Surprisingly to me, this is true! At least, on my machine. Sadly, the only way to make it avoid a runtime check is to exclude 32-bit machines (or to make the option non-default, which I would prefer not to do).

> For parallelisation you need to take extra step and paralelize whole
> check than going component-by-component.
Good idea.
> For detecting sequences a faster way is construct bitmasks with SSE2 so
> you could combine these. It avoids needing special casing on 16-byte
> boundaries.
That does seem to be faster.
> Below is untested implementation where you could add a bad character
> check with SSE4.2 which would speed it up. Are refs mostly
> alphanumerical? If so we could speed this up by paralelized alnum check
> and handling other characters in slower path.

Twitter's are almost entirely in [-._/a-zA-Z0-9] -- there are only a handful of exceptions. So, a method that has some bycatch outside of this range is just as fast as the SSE4.2 bad character check (but somewhat more code).

Previous: Ondřej BílkaNext: Junio C Hamano
Message 4 of 12 in “refs.c: SSE4.2 optimizations for check_refname_component”
  1. 0/1 refs.c: SSE4.2 optimizations for check_refname_componentDavid Turner, Jun 5, 2014
  2. 1/1 refs.c: SSE4.2 optimizations for check_refname_componentDavid Turner, Jun 5, 2014
  3. Ondřej BílkaJun 14, 2014
  4. David TurnerJun 15, 2014
  5. Junio C HamanoJun 9, 2014
  6. David TurnerJun 9, 2014
  7. Junio C HamanoJun 9, 2014
  8. Johannes SixtJun 10, 2014
  9. Junio C HamanoJun 10, 2014
  10. David TurnerJun 13, 2014
  11. Torsten BögershausenJun 13, 2014
  12. Philip OakleyJun 14, 2014

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.