Claude Opus 4.1
Claude Opus 4.1 is capable in long context; and behind the leaders in coding, multimodal tasks, instruction following, and math. Too few results yet to rate reasoning, agentic tasks, safety, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Multilingual or Factuality.
Price
$15.00input$75.00outputper million tokens
From Anthropic's own price page · 3 providers tracked · All prices
Evidence
45results on40benchmarks
- 17 independently verified
- 12 aggregator
- 16 vendor-reported
From 15 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 4.1 benchmark results
45 results on 40 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
17.8% behind the leader1 of 3 ranked benchmarks measured
- 76.00Oct 8, 2026
31.5% behind the leader4 of 10 ranked benchmarks measured
- 74.50Oct 7, 2026
- SciCode40.90Sep 4, 2026aa_scicode
- 1390.19May 22, 2026
- Terminal-Bench Hard34.34Oct 8, 2026aa_terminalbench_hard
33.4% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro67.92Oct 8, 2026aa_mmmu_pro
34.8% behind the leader2 of 3 ranked benchmarks measured
- 57.20Oct 8, 2026
- IFBench55.44Oct 8, 2026aa_ifbench
56.0% behind the leader2 of 5 ranked benchmarks measured
- 12.63Sep 9, 2026
- 2.44Sep 8, 2026
0 of 6 ranked benchmarks measured
- 60.00May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond80.90Oct 8, 2026gpqa
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond80.90Oct 7, 2026GPQA
- 81.00Jul 29, 2026
- Humanity's Last Exam12.47Oct 8, 2026aa_hle
0 of 7 ranked benchmarks measured
- OSWorld-Verified44.40Jul 29, 2026OSWorld
- 40.90Jul 29, 2026
0 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better11.80Jun 18, 2026
- 34.80May 20, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence19.00Sep 4, 2026Artificial Analysis Intelligence Index
- 1447Jul 23, 2026
- 1449Jul 23, 2026
- 129.10Jun 18, 2026
- 92.40Jun 18, 2026
- 88.20Jun 18, 2026
Show 19 more resultsHide 19 results
- AA Intelligence18.57Oct 8, 2026aa_intelligence_index
- AA Intelligence22.85Oct 8, 2026aa_intelligence_index
- 78.00Oct 7, 2026
- Artificial Analysis Coding Index36.52Jun 18, 2026aa_coding_index
- 18.20Jul 29, 2026
- frontiermath_tier_4_v14.17May 20, 2026frontiermath_tier_4
- 0.71Jul 29, 2026
- 60.61Jun 1, 2026
- 15.00Jun 1, 2026
- 100.00Jun 1, 2026
- 89.50Oct 7, 2026
- 43.80Jul 29, 2026
- TAU-bench (airline)56.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)82.40Oct 7, 2026TAU-bench Retail
- 43.30Sep 23, 2026
- 46.50Jul 29, 2026
- τ²-Bench71.50Jul 29, 2026τ²-Bench (Telecom)
- 86.80Jul 29, 2026
- τ²-Bench Telecom (AA run)71.37Oct 8, 2026aa_tau2
Claude Opus 4.1: common questions
Who makes Claude Opus 4.1?
Claude Opus 4.1 is made by Anthropic.
When was Claude Opus 4.1 released?
Claude Opus 4.1 was released on Aug 5, 2025, according to Artificial Analysis.
What is Claude Opus 4.1 good at?
Claude Opus 4.1 is capable in long context; and behind the leaders in coding, multimodal tasks, instruction following, and math. Too few results yet to rate reasoning, agentic tasks, safety, multilingual tasks, or factuality.
How much does Claude Opus 4.1 cost?
Claude Opus 4.1 costs $15.00 per million input tokens and $75.00 per million output tokens, according to Anthropic's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 97% of the 330 priced models we track.
How many benchmarks has Claude Opus 4.1 been tested on?
We track 45 results for Claude Opus 4.1 on 40 benchmarks from 15 sources, 17 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4.1 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.1.
About this record
Where Claude Opus 4.1's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 31, 2026
Where the results come from
Verification: 45 scores · 17 independently verified · 12 aggregator-attributed · 16 vendor-reported. How these tiers are assigned
From 15 sources on 9 sites. Artificial Analysis supplies 13 of them; the 17 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai13
- www-cdn.anthropic.com9
- api.llm-stats.com7
- raw.githubusercontent.com7
- epoch.ai4
- lmarena.ai2
- datasets-server.huggingface.co1
- labs.scale.com1
- simple-bench.com1