Search Capability Leaderboard
Latest update
September 25, 2026
What's the best model for agentic search?
Large language models rely on web search to find information that isn't available in the model weights, but not all models benefit equally. We use standard and proprietary benchmarks to evaluate how well top models orchestrate searches and synthesize results into correct answers.
Latest update September 25, 2026
Insights
Highest Search Intelligence Score (75.4). Great for tasks that require complex reasoning and maximum accuracy.
Most cost-efficient model ($33.1 per 1K tasks). Great for most queries, at 29× lower cost than the top model.
Largest lift from search (+43.8). Most improved when paired with search.
# Search Capability Leaderboard
Updated September 25, 2026
What's the best model for agentic search?
Large language models rely on web search to find information that isn't available in the model weights, but not all models benefit equally. We use standard and proprietary benchmarks to evaluate how well top models orchestrate searches and synthesize results into correct answers. See the Methodology section below.
Insights
- Claude Opus 5.5 (Gold, Search Intelligence): Highest Search Intelligence Score (75.4). Great for tasks that require complex reasoning and maximum accuracy.
- GPT-6 Luna (Gold, Search Efficiency): Most cost-efficient model ($33.1 per 1K tasks). Great for most queries, at 29× lower cost than the top model.
- Claude Sonnet 5: Largest lift from search (+43.8). Most improved when paired with search.
## Search Intelligence Leaderboard
| Rank | Model | Score with Search | Score without Search | Lift | Cost per 1K tasks | Time per task |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 75.4 | 44.2 | +31.2 | $949 | 512s |
| 2 | Claude Fable 5.1 | 72.7 | 42.4 | +30.3 | $1,655 | 512s |
| 3 | GPT-6 Astra | 70.8 | 43.0 | +27.9 | $401 | 83.7s |
| 4 | Claude Opus 5 | 70.0 | 37.2 | +32.8 | $1,012 | 342s |
| 5 | GPT-5.6 Sol | 67.7 | 41.4 | +26.3 | $269 | 107s |
| 6 | Gemini 3.7 Flash | 66.8 | 39.4 | +27.4 | $130 | 229s |
| 7 | GPT-6 Sol | 66.6 | 37.9 | +28.7 | $183 | 260s |
| 8 | Claude Sonnet 5 | 66.6 | 22.8 | +43.8 | $688 | 846s |
| 9 | Muse Spark 1.3 | 66.1 | 30.3 | +35.8 | $142 | 274s |
| 10 | Gemini 3.8 Flash | 65.1 | 39.0 | +26.1 | $177 | 317s |
| 11 | Kimi K3 | 64.2 | 30.6 | +33.6 | $266 | 432s |
| 12 | DeepSeek V4.1 Flash | 62.7 | 27.0 | +35.7 | $35.8 | 549s |
| 13 | GLM 5.3 | 62.7 | 21.9 | +40.8 | $78.5 | 308s |
| 14 | GPT-6 Luna | 61.9 | 28.0 | +33.9 | $33.1 | 391s |
| 15 | GPT-5.6 Luna | 60.7 | 26.9 | +33.9 | $36.2 | 98.0s |
| 16 | DeepSeek V4 Flash (0731) | 59.2 | 24.0 | +35.2 | $13.6 | 292s |
| 17 | Gemini 3 Flash | 58.6 | 34.1 | +24.5 | $114 | 276s |
| 18 | DeepSeek V4 Pro | 58.4 | 30.5 | +27.9 | $110 | 383s |
| 19 | Hunyuan 3 | 58.2 | 22.6 | +35.6 | $28.4 | 363s |
| 20 | GLM-5.2 | 54.3 | 18.4 | +35.9 | $66.6 | 301s |
| 21 | DeepSeek V4 Pro (0813) | 53.5 | 33.1 | +20.4 | $234 | 394s |
| 22 | DeepSeek V4 Flash | 53.1 | 19.5 | +33.7 | $24.4 | 358s |
| 23 | Nemotron 3 Ultra 550B | 53.0 | 15.3 | +37.7 | $113 | 328s |
| 24 | MiniMax M3 | 52.6 | 23.1 | +29.5 | $41.1 | 291s |
| 25 | MiMo v2.5 | 49.5 | 13.4 | +36.1 | $23.5 | 894s |
| 26 | Laguna S 2.1 | 39.5 | 12.3 | +27.2 | $21.4 | 459s |
| 27 | Nemotron 3.5 Lightning | 37.9 | 9.8 | +28.0 | $12.1 | 176s |
## Search Efficiency Leaderboard
Most cost-efficient models scoring at or above the median Search Intelligence Score (61.9), cheapest first. Costs are per 1,000 tasks.
| Rank | Model | Cost per 1K tasks | Score with Search |
|---|---|---|---|
| 1 | GPT-6 Luna | $33.1 | 61.9 |
| 2 | DeepSeek V4.1 Flash | $35.8 | 62.7 |
| 3 | GLM 5.3 | $78.5 | 62.7 |
| 4 | Gemini 3.7 Flash | $130 | 66.8 |
| 5 | Muse Spark 1.3 | $142 | 66.1 |
| 6 | Gemini 3.8 Flash | $177 | 65.1 |
| 7 | GPT-6 Sol | $183 | 66.6 |
| 8 | Kimi K3 | $266 | 64.2 |
| 9 | GPT-5.6 Sol | $269 | 67.7 |
| 10 | GPT-6 Astra | $401 | 70.8 |
| 11 | Claude Sonnet 5 | $688 | 66.6 |
| 12 | Claude Opus 5.5 | $949 | 75.4 |
| 13 | Claude Opus 5 | $1,012 | 70.0 |
| 14 | Claude Fable 5.1 | $1,655 | 72.7 |
Search Intelligence Leaderboard
10 of 27 models
- 1Opus 5.575.4Cost $949
- 2Fable 5.172.7Cost $1,655
- 3GPT-6 Astra70.8Cost $401
- 4Opus 570.0Cost $1,012
- 5GPT-5.6 Sol67.7Cost $269
- 6Gemini 3.766.8Cost $130
- 7GPT-6 Sol66.6Cost $183
- 8Sonnet 566.6Cost $688
- 9Muse Spark 1.366.1Cost $142
- 10Gemini 3.865.1Cost $177
| 1 | Opus 5.5Claude Opus 5.5 | 75.4 | $949 |
| 2 | Fable 5.1Claude Fable 5.1 | 72.7 | $1,655 |
| 3 | GPT-6 AstraGPT-6 Astra | 70.8 | $401 |
| 4 | Opus 5Claude Opus 5 | 70.0 | $1,012 |
| 5 | GPT-5.6 SolGPT-5.6 Sol | 67.7 | $269 |
| 6 | Gemini 3.7Gemini 3.7 Flash | 66.8 | $130 |
| 7 | GPT-6 SolGPT-6 Sol | 66.6 | $183 |
| 8 | Sonnet 5Claude Sonnet 5 | 66.6 | $688 |
| 9 | Muse Spark 1.3Muse Spark 1.3 | 66.1 | $142 |
| 10 | Gemini 3.8Gemini 3.8 Flash | 65.1 | $177 |
Search Efficiency Leaderboard
10 of 14 · 27 models
- 1GPT-6 Luna$33.1Score 61.9
- 2DS Flash 4.1$35.8Score 62.7
- 3GLM 5.3$78.5Score 62.7
- 4Gemini 3.7$130Score 66.8
- 5Muse Spark 1.3$142Score 66.1
- 6Gemini 3.8$177Score 65.1
- 7GPT-6 Sol$183Score 66.6
- 8Kimi K3$266Score 64.2
- 9GPT-5.6 Sol$269Score 67.7
- 10GPT-6 Astra$401Score 70.8
| 1 | GPT-6 LunaGPT-6 Luna | $33.1 | 61.9 |
| 2 | DS Flash 4.1DeepSeek V4.1 Flash | $35.8 | 62.7 |
| 3 | GLM 5.3GLM 5.3 | $78.5 | 62.7 |
| 4 | Gemini 3.7Gemini 3.7 Flash | $130 | 66.8 |
| 5 | Muse Spark 1.3Muse Spark 1.3 | $142 | 66.1 |
| 6 | Gemini 3.8Gemini 3.8 Flash | $177 | 65.1 |
| 7 | GPT-6 SolGPT-6 Sol | $183 | 66.6 |
| 8 | Kimi K3Kimi K3 | $266 | 64.2 |
| 9 | GPT-5.6 SolGPT-5.6 Sol | $269 | 67.7 |
| 10 | GPT-6 AstraGPT-6 Astra | $401 | 70.8 |
Intelligence
Search Intelligence vs. Cost
Search Intelligence Score vs. cost to complete a task (per 1,000 tasks). Cost uses a logarithmic x-axis; up and to the left is better.
Pareto frontier
Search Intelligence Score
Search Intelligence Score, highest first.
Cost
Cost per Task
Cost per 1,000 tasks with Parallel Search, cheapest first, split across costs associated with model inference (token generation), search, and page extraction.
Cost of Search vs. Lift
Compares the percentage of total cost that goes to search (and extract) tools versus the lift in intelligence. A smaller percentage can reflect higher inference spend. Tap a point to label it.
Models
Methodology
Our goal here is to measure how effective models are at using search. For simplicity, we combine publicly accepted benchmark datasets with our own WISER benchmark into a composite score. This is not meant to reflect the comprehensive evaluations Parallel runs internally. We use the following evaluation datasets:
- DSQADeepSearchQA. Google's benchmark of multi-step research questions.
- HLEHumanity's Last Exam. Expert-written questions that need web-grounded reasoning.
- WISERWISER. Parallel's own benchmark of real-world business research queries.
Example questions
Shortened examples from public benchmark materials. The scored sample may differ.
DSQA
Research across sources
Using Macrotrends for NVIDIA’s stock performance and Worldometer for U.S. GDP, identify years from 2020 through 2023 when annual stock gains exceeded 125% and GDP growth exceeded 2.5%.
Find both data series, align the years, and return every year that meets both thresholds.
HLE
Specialist knowledge
In hummingbirds, how many tendon pairs does the sesamoid bone within the m. depressor caudae insertion support?
Interpret specialist anatomy terms and establish the precise count from relevant evidence.
WISER
Business research
For Salesforce’s FY2024, calculate the share of subscription and support revenue from Tableau’s reporting segment. Compare it with Salesforce’s 2023 CRM market share.
Identify the reporting segment, calculate its revenue share, and compare it with a separate market estimate.
Process
Each model answers the same 100 questions per benchmark, with and without Parallel Search Fast mode and Extract. Tool budgets are fixed; code execution is disabled. Once the tool budget is exhausted, we request a final answer with tools disabled. Reasoning settings and token limits may vary by model to reflect provider recommendations, when provided.
DSQA answers are extracted without access to the reference answer, then graded. HLE and WISER use correct-or-incorrect grading.
- Search Intelligence Score
DSQA F1, HLE accuracy, and WISER accuracy, weighted equally on a 0–100 scale. F1 allows partial credit for incomplete answers and penalizes incorrect items. Scores use all 100 questions per benchmark; failed and pending tasks count as zero.
- Lift
Lift measures the difference between a model's score with search and its score without. A lift of +35 means search added 35 points.
- Search Efficiency Leaderboard
Models at or above the median score, ranked by cost per 1,000 tasks. Costs include inference and estimated Search and Extract usage, excluding grading. Cost records are incomplete and may omit recovery attempts.
- Time per Task
Average seconds from question to final answer across available records, with equal weight for each benchmark.
- Pareto Frontier
A model is on the frontier when no other model is both cheaper and higher scoring.
## Methodology
Our goal here is to measure how effective models are at using search. For simplicity, we combine publicly accepted benchmark datasets with our own WISER benchmark into a composite score. This is not meant to reflect the comprehensive evaluations Parallel runs internally. We use the following evaluation datasets: DeepSearchQA (DSQA): Google's benchmark of multi-step research questions. Humanity's Last Exam (HLE): Expert-written questions that need web-grounded reasoning. WISER (WISER): Parallel's own benchmark of real-world business research queries.
Example questions
Shortened examples from public benchmark materials. The scored sample may differ.
DSQA
Research across sources
Using Macrotrends for NVIDIA’s stock performance and Worldometer for U.S. GDP, identify years from 2020 through 2023 when annual stock gains exceeded 125% and GDP growth exceeded 2.5%.
Find both data series, align the years, and return every year that meets both thresholds.
HLE
Specialist knowledge
In hummingbirds, how many tendon pairs does the sesamoid bone within the m. depressor caudae insertion support?
Interpret specialist anatomy terms and establish the precise count from relevant evidence.
WISER
Business research
For Salesforce’s FY2024, calculate the share of subscription and support revenue from Tableau’s reporting segment. Compare it with Salesforce’s 2023 CRM market share.
Identify the reporting segment, calculate its revenue share, and compare it with a separate market estimate.
Each model answers the same 100 questions per benchmark, with and without Parallel Search Fast mode and Extract. Tool budgets are fixed; code execution is disabled. Once the tool budget is exhausted, we request a final answer with tools disabled. Reasoning settings and token limits may vary by model to reflect provider recommendations, when provided.
DSQA answers are extracted without access to the reference answer, then graded. HLE and WISER use correct-or-incorrect grading.
Search Intelligence Score: DSQA F1, HLE accuracy, and WISER accuracy, weighted equally on a 0–100 scale. F1 allows partial credit for incomplete answers and penalizes incorrect items. Scores use all 100 questions per benchmark; failed and pending tasks count as zero. Lift: the difference between a model's score with search and without. Search Efficiency Leaderboard: models at or above the median score (61.9), ranked by cost per 1,000 tasks, cheapest first. Costs include inference and estimated Search and Extract usage, excluding grading. Cost records are incomplete and may omit recovery attempts. Time per task: average seconds from question to final answer across available records, with equal weight for each benchmark. Pareto frontier: no other model is both cheaper and higher scoring.