git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4] Threaded grep

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Jan 25, 2010, 23:59 UTC
Message-ID
<alpine.LFD.2.00.1001251542100.3574@localhost.localdomain>
In-Reply-To
<20100125225139.GA3048@fredrik-laptop>
On Mon, 25 Jan 2010, Fredrik Kuivinen wrote:
Show 5 quoted lines
> 
> The results below are best of five runs in the Linux repository (on a
> box with two cores).
> 
> git grep qwerty
Before:
	real	0m0.531s
	user	0m0.412s
	sys	0m0.112s
After:
	real	0m0.151s
	user	0m0.720s
	sys	0m0.272s
> $ /usr/bin/time git grep void
Before:
	real	0m1.144s
	user	0m0.988s
	sys	0m0.148s
After:
	real	0m0.290s
	user	0m1.732s
	sys	0m0.232s
So it's helping a lot (~3.5x and ~3.9x) on this 4-core HT setup. 

I don't seem to ever get more than a 4x speedup, so my guess is that HT simply isn't able to do much of anything with this load.

The profile for the threaded case says:
    51.73%      git  libc-2.11.1.so                 [.] re_search_internal
    11.47%      git  [kernel]                       [k] copy_user_generic_string
     2.90%      git  libc-2.11.1.so                 [.] __strlen_sse2
     2.66%      git  [kernel]                       [k] link_path_walk
     2.55%      git  [kernel]                       [k] intel_pmu_enable_all
     2.40%      git  [kernel]                       [k] __d_lookup
     1.71%      git  libc-2.11.1.so                 [.] __GI___libc_malloc
     1.55%      git  [kernel]                       [k] _raw_spin_lock
     1.43%      git  [kernel]                       [k] sys_futex
     1.30%      git  libc-2.11.1.so                 [.] __cfree
     1.28%      git  [kernel]                       [k] intel_pmu_disable_all
     1.25%      git  libc-2.11.1.so                 [.] __GI_memchr
     1.14%      git  libc-2.11.1.so                 [.] _int_malloc
     1.02%      git  [kernel]                       [k] effective_load

and the only thing that makes me go "eh?" there is the strlen(). Why is that so hot? But locking doesn't seem to be the biggest issue, and in general I think this is all pretty good. The 'effective_load' thing is the scheduler, so there's certainly some context switching going on, probably still due to excessive synchronization, but it's equally clear that that is certainly not a dominant factor.

One potentially interesting data point is that if I make NR_THREADS be 16, performance goes down, and I get more locking overhead. So NR_THREADS of 8 works well on this machine.

So ack from me. The patch looks reasonably clean too, at least for something as complex as a multi-threaded grep.

One worry is, of course, whether all regex() implementations are thread-safe. Maybe there are broken libraries that have hidden global state in them?

			Linus
Previous: Fredrik KuivinenNext: Fredrik Kuivinen
Message 2 of 12 in “Threaded grep”
  1. Threaded grepFredrik Kuivinen, Jan 25, 2010
  2. Linus TorvaldsJan 25, 2010
  3. Fredrik KuivinenJan 26, 2010
  4. Linus TorvaldsJan 26, 2010
  5. Benjamin KramerJan 26, 2010
  6. Linus TorvaldsJan 26, 2010
  7. Linus TorvaldsJan 26, 2010
  8. Mike HommeyJan 26, 2010
  9. grep: use REG_STARTEND (if available) to speed up regexecBenjamin Kramer, Jan 26, 2010
  10. Junio C HamanoJan 26, 2010
  11. Fredrik KuivinenJan 26, 2010
  12. Junio C HamanoJan 26, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.