Claude Opus 4
Claude Opus 4 is behind the leaders in long context, coding, instruction following, and reasoning. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Math, Multimodal, Multilingual or Factuality.
Price
$15.00input$75.00outputper million tokens
From Anthropic's own price page · 2 providers tracked · All prices
Evidence
121results on69benchmarks
- 47 independently verified
- 15 aggregator
- 15 vendor-reported
- 44 cross-referenced
From 27 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingReasoning
As listed by OpenRouter
Claude Opus 4 benchmark results
121 results on 69 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
27.4% behind the leader2 of 3 ranked benchmarks measured
- 55.60Jun 5, 2026
- 69.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 40.00Sep 4, 2026
34.0% behind the leader4 of 10 ranked benchmarks measured
- 72.50Oct 7, 2026
- SciCode39.81Sep 4, 2026aa_scicode
- LiveCodeBench v647.40Jun 15, 2026LiveCodeBench v6 (Aug 24 - May 25)
- Terminal-Bench Hard31.06Oct 8, 2026aa_terminalbench_hard
Show 6 more coding resultsHide 6 coding results
- SciCode40.86Sep 4, 2026aa_scicode
- 67.60May 1, 2026
- SWE-bench Verified79.40Jun 13, 2026SWE-bench (high compute)
- SWE-bench Verified79.40Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified53.00Jun 15, 2026SWE-bench Verified (Agentless Coding)
- 72.50Jun 5, 2026
35.3% behind the leader2 of 3 ranked benchmarks measured
- 58.62Oct 8, 2026
- IFBench53.74Oct 8, 2026aa_ifbench
Show 3 more instruction following resultsHide 3 instruction following results
- IFBench43.27Oct 8, 2026aa_ifbench
- 49.00Jun 15, 2026
- 45.80Jun 5, 2026
41.8% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond79.60Oct 8, 2026gpqa
- 58.80May 10, 2026
- Humanity's Last Exam12.33Oct 8, 2026aa_hle
- ARC-AGI-28.60Oct 7, 2026ARC-AGI v2
Show 12 more reasoning resultsHide 12 reasoning results
- 0.00Sep 22, 2026
- 4.52Sep 22, 2026
- 8.61Sep 22, 2026
- 1.27May 10, 2026
- GPQA Diamond70.10Oct 8, 2026gpqa
- GPQA Diamond79.60Oct 7, 2026GPQA
- 74.90Jun 13, 2026
- 74.90Jun 15, 2026
- 79.60Jun 5, 2026
- Humanity's Last Exam6.16Oct 8, 2026aa_hle
- Humanity's Last Exam7.10Jun 15, 2026Humanity's Last Exam (Text Only)
- Humanity's Last Exam10.70Jun 5, 2026HLE (no tools)
0 of 6 ranked benchmarks measured
- 1191.54Sep 22, 2026
- 1175.23Aug 25, 2026
- 1207May 6, 2026
Show 2 more multimodal resultsHide 2 multimodal results
- 1187May 1, 2026
- 73.70Jun 13, 2026
0 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better12.00Jun 20, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 27.00Sep 22, 2026
- 30.67Sep 22, 2026
- 35.67Sep 22, 2026
- 70.70Aug 29, 2026
- AA Intelligence25.00Jul 3, 2026Artificial Analysis Intelligence Index
- 123.20Jun 20, 2026
Show 75 more resultsHide 75 results
- AA Intelligence20.64Oct 8, 2026aa_intelligence_index
- AA Intelligence16.61Oct 8, 2026aa_intelligence_index
- 75.60Jun 15, 2026
- 70.70Jun 15, 2026
- 33.90Jun 13, 2026
- 48.20Jun 15, 2026
- 76.00Jun 5, 2026
- 70.00May 2, 2026
- 75.50Oct 7, 2026
- 33.90Jun 15, 2026
- 75.50Jun 5, 2026
- 85.74May 19, 2026
- 98.90Jun 18, 2026
- 98.91May 10, 2026
- 22.50May 10, 2026
- Artificial Analysis Coding Index33.98Jun 18, 2026aa_coding_index
- 86.10Jun 15, 2026
- 97.10Jun 18, 2026
- 99.30May 10, 2026
- 57.60Jun 15, 2026
- frontiermath_tier_4_v14.17May 20, 2026frontiermath_tier_4
- 70.30Jun 5, 2026
- 91.69Jun 18, 2026
- 89.13May 10, 2026
- 15.90Jun 15, 2026
- 60.00May 11, 2026
- 87.40Jun 15, 2026
- 74.60Jun 15, 2026
- 62.37May 29, 2026
- 70.43May 3, 2026
- 56.60Jun 5, 2026
- 96.27May 29, 2026
- 99.07May 3, 2026
- 23.71May 29, 2026
- 33.14May 3, 2026
- 69.19May 29, 2026
- 80.42May 3, 2026
- MATH-500 (EM)94.40Jun 15, 2026MATH-500
- MATH-500 (EM)98.20Jun 5, 2026MATH-500
- 61.66May 10, 2026
- 20.00May 10, 2026
- 99.75May 10, 2026
- 92.90Jun 15, 2026
- 86.60Jun 15, 2026
- 85.00Jun 5, 2026
- 94.20Jun 15, 2026
- 88.80Oct 7, 2026
- 87.40Jun 13, 2026
- 89.60Jun 15, 2026
- 19.60Jun 15, 2026
- 48.90Jun 5, 2026
- 49.80Jun 15, 2026
- 100.00Jun 18, 2026
- 99.50May 10, 2026
- 22.80Jun 15, 2026
- 56.50Jun 15, 2026
- 67.60May 1, 2026
- 39.50Jul 29, 2026
- TAU-bench (airline)59.60Oct 7, 2026TAU-bench Airline
- 59.60Jun 5, 2026
- TAU-bench (retail)81.40Oct 7, 2026TAU-bench Retail
- 81.40Jun 5, 2026
- 60.00Jun 15, 2026
- 39.20Sep 23, 2026
- 43.20Jun 13, 2026
- 43.20Jun 15, 2026
- 91.00Jun 20, 2026
- 88.00Jun 20, 2026
- 97.00Jun 18, 2026
- 97.00May 10, 2026
- 59.30Jun 15, 2026
- 95.10Jun 5, 2026
- τ²-Bench57.00Jun 15, 2026Tau2 telecom
- τ²-Bench (Retail)81.80Jun 15, 2026Tau2 retail
- τ²-Bench Telecom (AA run)73.39Oct 8, 2026aa_tau2
Claude Opus 4: common questions
Who makes Claude Opus 4?
Claude Opus 4 is made by Anthropic.
When was Claude Opus 4 released?
Claude Opus 4 was released on May 22, 2025, according to Artificial Analysis.
What is Claude Opus 4 good at?
Claude Opus 4 is behind the leaders in long context, coding, instruction following, and reasoning. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
How much does Claude Opus 4 cost?
Claude Opus 4 costs $15.00 per million input tokens and $75.00 per million output tokens, according to Anthropic's own price page. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 97% of the 330 priced models we track.
How many benchmarks has Claude Opus 4 been tested on?
We track 121 results for Claude Opus 4 on 69 benchmarks from 27 sources, 47 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4 support?
OpenRouter lists tool calling and reasoning for Claude Opus 4.
About this record
Where Claude Opus 4's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 2, 2026
Where the results come from
Verification: 121 scores · 47 independently verified · 15 aggregator-attributed · 44 vendor cross-reference · 15 vendor-reported. How these tiers are assigned
From 27 sources on 17 sites. Hugging Face supplies 44 of them; the 47 independently verified results come from 13 sites. Bars are coloured by trust tier.
- huggingface.co44
- artificialanalysis.ai16
- storage.googleapis.com11
- api.llm-stats.com8
- arcprize.org8
- livecodebench.github.io8
- raw.githubusercontent.com7
- anthropic.com6
- datasets-server.huggingface.co2
- lmarena.ai2
- matharena.ai2
- swebench.com2
- aider.chat1
- epoch.ai1
- labs.scale.com1
- simple-bench.com1
- www-cdn.anthropic.com1