Claude Opus 5
Claude Opus 5 is at the frontier in agentic tasks, reasoning, and coding; strong in factuality; and capable in math, long context, and multimodal tasks. Too few results yet to rate safety, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Instruction Following.
Price
$5.00input$25.00outputper million tokens
From Anthropic's own price page · 4 providers tracked · All prices
Evidence
174results on78benchmarks
- 22 independently verified
- 79 aggregator
- 57 vendor-reported
- 16 cross-referenced
From 31 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 5 benchmark results
174 results on 78 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
Leads the field7 of 7 ranked benchmarks measured
- 90.80Oct 7, 2026
- 85.80Jul 24, 2026
- 83.40Sep 22, 2026
- τ-Bench V3 · Banking42.06Oct 8, 2026tauBanking
- Terminal-Bench 4.048.99Oct 8, 2026
- 61.19Oct 8, 2026
Show 13 more agentic resultsHide 13 agentic results
- 49.60Oct 8, 2026
- 54.79Oct 8, 2026
- 59.62Oct 8, 2026
- 40.22Oct 8, 2026
- 85.80Oct 8, 2026
- Terminal-Bench 4.034.34Oct 8, 2026
- Terminal-Bench 4.045.96Oct 8, 2026
- Terminal-Bench 4.046.46Oct 8, 2026
- Terminal-Bench 4.026.26Oct 8, 2026
- τ-Bench V3 · Banking38.56Oct 8, 2026tauBanking
- τ-Bench V3 · Banking44.74Oct 8, 2026tauBanking
- τ-Bench V3 · Banking30.31Oct 8, 2026tauBanking
- τ-Bench V3 · Banking43.30Oct 8, 2026tauBanking
2.3% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning91.21Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond93.23Oct 8, 2026gpqa
- 90.42Sep 1, 2026
- 80.60Jul 25, 2026
- 29.14Oct 8, 2026
- Humanity's Last Exam54.87Oct 8, 2026aa_hle
Show 21 more reasoning resultsHide 21 reasoning results
- 90.42Sep 21, 2026
- 88.33Sep 21, 2026
- 90.40Jul 24, 2026
- 30.16Sep 21, 2026
- 30.20Oct 7, 2026
- 30.16Jul 24, 2026
- 26.86Oct 8, 2026
- 28.29Oct 8, 2026
- 23.14Oct 8, 2026
- 27.71Oct 8, 2026
- GPQA Diamond91.92Oct 8, 2026gpqa
- GPQA Diamond93.74Oct 8, 2026gpqa
- GPQA Diamond88.89Oct 8, 2026gpqa
- Humanity's Last Exam51.30Oct 8, 2026aa_hle
- Humanity's Last Exam52.83Oct 8, 2026aa_hle
- Humanity's Last Exam54.40Oct 8, 2026aa_hle
- Humanity's Last Exam43.42Oct 8, 2026aa_hle
- 64.70Oct 7, 2026
- Humanity's Last Exam63.60Sep 22, 2026Humanity's Last Exam with tools
- Humanity's Last Exam56.60Sep 1, 2026Humanity's Last Exam (No tools)
- Humanity's Last Exam56.30Jul 24, 2026Humanity's Last Exam (No tools)
5.9% behind the leader8 of 10 ranked benchmarks measured
- 96.00Jul 24, 2026
- 1691.77Sep 21, 2026
- Terminal-Bench 2.189.14Oct 8, 2026terminalbenchV21
- 89.50Sep 22, 2026
- LiveBench · Coding81.45Oct 8, 2026livebench_coding@2026-06-25
- 79.20Sep 22, 2026
- LiveBench · Agentic Coding65.20Oct 8, 2026livebench_agentic_coding@2026-06-25
- SciCode56.37Oct 8, 2026aa_scicode
Show 10 more coding resultsHide 10 coding results
- 1657.70Sep 21, 2026
- SciCode51.50Oct 8, 2026aa_scicode
- SciCode55.44Oct 8, 2026aa_scicode
- SciCode49.19Oct 8, 2026aa_scicode
- SciCode55.67Oct 8, 2026aa_scicode
- Terminal-Bench 2.186.14Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.187.64Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.176.40Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.188.01Oct 8, 2026terminalbenchV21
- 89.10Sep 22, 2026
8.7% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy60.87Oct 8, 2026omniscienceAccuracy
- 59.90Sep 21, 2026
- AA-Omniscience · Non-hallucination39.18Oct 8, 2026omniscienceNonHallucination
Show 8 more factuality resultsHide 8 factuality results
- AA-Omniscience · Accuracy57.07Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy58.88Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy59.50Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy55.95Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination39.32Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination38.79Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination40.45Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination37.80Oct 8, 2026omniscienceNonHallucination
12.3% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics95.73Oct 8, 2026livebench_math@2026-06-25
- 85.61Sep 21, 2026
- 73.17Sep 21, 2026
12.5% behind the leader1 of 3 ranked benchmarks measured
- 79.33Oct 8, 2026
18.6% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro84.74Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)83.70Sep 9, 2026CharXiv Reasoning (No tools)
Show 5 more multimodal resultsHide 5 multimodal results
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following63.77Oct 8, 2026livebench_instruction_following@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1489Oct 8, 2026
- livebench_data_analysis74.55Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language88.69Oct 8, 2026livebench_language@2026-06-25
- 57.00Oct 8, 2026
- AA Intelligence51.00Oct 3, 2026Artificial Analysis Intelligence Index
- 97.50Sep 21, 2026
Show 77 more resultsHide 77 results
- 46.98Sep 9, 2026
- 56.22Sep 9, 2026
- 52.69Sep 9, 2026
- 55.47Sep 9, 2026
- 37.47Sep 9, 2026
- AA Intelligence44.83Oct 8, 2026aa_intelligence_index
- AA Intelligence50.78Oct 8, 2026aa_intelligence_index
- AA Intelligence48.12Oct 8, 2026aa_intelligence_index
- AA Intelligence39.35Oct 8, 2026aa_intelligence_index
- AA Intelligence49.68Oct 8, 2026aa_intelligence_index
- 31.02Oct 8, 2026
- 37.07Oct 8, 2026
- 33.72Oct 8, 2026
- 28.55Oct 8, 2026
- 35.38Oct 8, 2026
- Agentic coding (FrontierCode v1.1, Main)53.40Sep 23, 2026
- 31.60Sep 22, 2026
- 97.50Sep 1, 2026
- Artificial Analysis Coding Index74.33Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index77.98Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index76.52Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index66.95Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index77.00Sep 9, 2026aa_coding_index
- 8.00Aug 26, 2026
- 68.80Oct 7, 2026
- DeepSWE 1.174.00Sep 22, 2026DeepSWE v1.1
- GDP (Surge AI)85.50Sep 1, 2026GDP.pdf (with tools)
- GDP.PDF (All pass rate)37.00Sep 9, 2026
- GDPval-AA (Elo)1708.00Sep 22, 2026GDPval-AA (max effort)
- GDPval-AA (Elo)1827.00Jul 24, 2026GDPval-AA ELO (xhigh effort)
- GDPval-AA 2.11708.00Sep 22, 2026GDPval-AA v2.1
- GDPval-AA 2.11708.00Sep 22, 2026
- 1824.00Sep 4, 2026
- 1861.00Jul 24, 2026
- GDPval-AA v2 Elo1824.00Sep 1, 2026GDPval-AA v2 ELO (max effort)
- GDPval-AA v2 Elo1824.00Sep 11, 2026GDPVal-AA v2 (Elo)
- 92.50Sep 22, 2026
- HealthBench (raw score)67.10Sep 22, 2026
- 67.10Sep 1, 2026
- HealthBench length-adjusted57.80Jul 24, 2026HealthBench (length-adjusted)
- 59.80Oct 7, 2026
- HealthBench Professional (raw score)73.40Sep 22, 2026
- 73.40Sep 1, 2026
- HealthBench Professional length-adjusted59.80Jul 24, 2026HealthBench Professional (length-adjusted)
- HLE (with tools)63.60Sep 17, 2026Humanity's Last Exam (with tools)
- HLE (with tools)64.70Jul 24, 2026Humanity's Last Exam (With tools)
- 54.40Sep 9, 2026
- 65.70Sep 22, 2026
- 11.70Oct 7, 2026
- 75.40Sep 9, 2026
- 92.10Sep 22, 2026
- 78.10Sep 22, 2026
- 66.90Sep 22, 2026
- 2.30Sep 18, 2026
- 70.60Oct 7, 2026
- 70.57Jul 24, 2026
- OSWorld 2.0 (partial)74.00Sep 22, 2026OSWorld 2.0 partial
- OSWorld 2.0 (partial)75.40Sep 17, 2026OSWorld 2.0 partial
- OSWorld 2.0 (partial)75.40Sep 9, 2026OSWorld-2.0 (Partial score)
- OSWorld 2.0 (strict)39.60Sep 23, 2026OSWorld 2.0 strict
- OSWorld 2.0 (strict)37.20Sep 22, 2026OSWorld 2.0 (strict pass rate)
- OSWorld 2.1 partial74.00Sep 29, 2026
- 85.40Sep 22, 2026
- 37.00Sep 22, 2026
- 47.70Sep 1, 2026
- 59.40Sep 22, 2026
- Terminal-Bench 4.051.80Oct 7, 2026
- Terminal-Bench 4.052.30Sep 28, 2026
- Terminal-Bench 4.042.00Sep 4, 2026
- Terminal-Bench 4.052.00Sep 1, 2026
- Terminal-Bench 4.049.00Sep 22, 2026
- Terminal-Bench 4.051.80Sep 11, 2026
- Terminal-Bench-Science 0.130.00Oct 7, 2026
- Terminal-Bench-Science 0.129.00Sep 28, 2026
- Toolathlon Verified87.00Sep 22, 2026Toolathlon Verified Pass@3
- Toolathlon Verified80.60Sep 22, 2026Toolathlon Verified Pass@1
- 80.60Sep 22, 2026
Claude Opus 5: common questions
Who makes Claude Opus 5?
Claude Opus 5 is made by Anthropic.
When was Claude Opus 5 released?
Claude Opus 5 was released on Jul 24, 2026, according to Anthropic's own announcement.
What is Claude Opus 5 good at?
Claude Opus 5 is at the frontier in agentic tasks, reasoning, and coding; strong in factuality; and capable in math, long context, and multimodal tasks. Too few results yet to rate safety, multilingual tasks, or instruction following.
How much does Claude Opus 5 cost?
Claude Opus 5 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 330 priced models we track.
How many benchmarks has Claude Opus 5 been tested on?
We track 174 results for Claude Opus 5 on 78 benchmarks from 31 sources, 22 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 5.
About this record
Where Claude Opus 5's numbers come from, and every name it appears under.
- Tracked since
- Jul 17, 2026
- Newest source mention
- Oct 5, 2026
Where the results come from
Verification: 174 scores · 22 independently verified · 79 aggregator-attributed · 16 vendor cross-reference · 57 vendor-reported. How these tiers are assigned
From 31 sources on 13 sites. Artificial Analysis supplies 80 of them; the 22 independently verified results come from 8 sites. Bars are coloured by trust tier.
- artificialanalysis.ai80
- www-cdn.anthropic.com37
- anthropic.com12
- huggingface.co9
- api.llm-stats.com8
- deepmind.google7
- livebench.ai7
- arcprize.org4
- datasets-server.huggingface.co3
- epoch.ai3
- labs.scale.com2
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name Claude Opus 5 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude opus 5 (low) | claude opus 5 (adaptive reasoning, low effort) | claude-opus-5-low |
| medium | claude opus 5 (medium) | claude opus 5 (adaptive reasoning, medium effort) | claude-opus-5-medium |
| high | claude opus 5 (high) | claude opus 5 (adaptive reasoning, high effort) | claude-opus-5-high |
| xhigh | — | claude opus 5 (adaptive reasoning, xhigh effort) | claude-opus-5-xhigh claude-opus-5 (xhigh) |
| max | claude opus 5 (max) | claude opus 5 (adaptive reasoning, max effort) | claude-opus-5-max |