git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4] Threaded grep

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Jan 26, 2010, 16:44 UTC
Message-ID
<alpine.LFD.2.00.1001260836520.3574@localhost.localdomain>
In-Reply-To
<4B5F1894.4070509@googlemail.com>
On Tue, 26 Jan 2010, Benjamin Kramer wrote:
> 
> BSD and glibc have an "REG_STARTEND" flag to do that. I made a small
> PoC patch to use it if it's available but it didn't give any significant
> speedup on my system.
Goodie.  It's noticeable for me. This is what I reported earlier:
Show 12 quoted lines
> > $ /usr/bin/time git grep void
> 
> Before:
> 
>         real    0m1.144s
>         user    0m0.988s
>         sys     0m0.148s
> 
> After:
>         real    0m0.290s
>         user    0m1.732s
>         sys     0m0.232s
and with your patch I get
	real	0m0.239s
	user	0m1.392s
	sys	0m0.276s
and the profile shows no strlen in it:
    57.12%      git  libc-2.11.1.so                 [.] re_search_internal
     5.59%      git  [kernel]                       [k] copy_user_generic_string
     4.09%      git  [kernel]                       [k] _raw_spin_lock
     2.57%      git  [kernel]                       [k] intel_pmu_enable_all
     2.46%      git  [kernel]                       [k] __d_lookup
     1.94%      git  libc-2.11.1.so                 [.] re_string_reconstruct
     1.87%      git  [kernel]                       [k] kmem_cache_alloc
     1.68%      git  libc-2.11.1.so                 [.] _int_free
     1.53%      git  [kernel]                       [k] find_get_page
     1.43%      git  [kernel]                       [k] update_curr
     1.27%      git  libc-2.11.1.so                 [.] __GI___libc_malloc
     1.17%      git  [kernel]                       [k] _atomic_dec_and_lock
     1.00%      git  libc-2.11.1.so                 [.] __GI_memcpy

Side note: the tailing end of the profiles aren't very stable, probably because the grep executes so quickly and in so many threads, so the functions in the one-percent range will move up and down the list depending on just exactly where we happened to get profile hits. Similarly, the raw_spin_lock numbers vary.

But the big picture is stable, and that 57% number (and the nonlock copy_user_generic_string) is consistent. And your patch definitely helped both actual performance and is visible in the profile: re_search_internal went from ~52% to ~57%.

So ack on that patch. Looks like a good thing to do, and with the #ifdef, it looks like it should just automatically DTRT based on regexec implementation.

		Linus
Previous: Benjamin KramerNext: Linus Torvalds
Message 6 of 12 in “Threaded grep”
  1. Threaded grepFredrik Kuivinen, Jan 25, 2010
  2. Linus TorvaldsJan 25, 2010
  3. Fredrik KuivinenJan 26, 2010
  4. Linus TorvaldsJan 26, 2010
  5. Benjamin KramerJan 26, 2010
  6. Linus TorvaldsJan 26, 2010
  7. Linus TorvaldsJan 26, 2010
  8. Mike HommeyJan 26, 2010
  9. grep: use REG_STARTEND (if available) to speed up regexecBenjamin Kramer, Jan 26, 2010
  10. Junio C HamanoJan 26, 2010
  11. Fredrik KuivinenJan 26, 2010
  12. Junio C HamanoJan 26, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.