Claude Opus 4.5
Claude Opus 4.5 is strong in long context; capable in factuality and multimodal tasks; and behind the leaders in reasoning, coding, agentic tasks, instruction following, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$5.00input$25.00outputper million tokens
From Anthropic's own price page · 3 providers tracked · All prices
Evidence
205results on128benchmarks
- 44 independently verified
- 32 aggregator
- 31 vendor-reported
- 98 cross-referenced
From 26 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 4.5 benchmark results
205 results on 128 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.5% behind the leader2 of 3 ranked benchmarks measured
- 64.40Jun 15, 2026
- 77.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 70.67Oct 8, 2026
22.3% behind the leader4 of 4 ranked benchmarks measured
- 10.90Jun 20, 2026
- AA-Omniscience · Accuracy46.58Oct 8, 2026omniscienceAccuracy
- 45.70Sep 22, 2026
- AA-Omniscience · Non-hallucination39.00Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy40.92Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination23.75Oct 8, 2026omniscienceNonHallucination
22.7% behind the leader5 of 6 ranked benchmarks measured
- 80.72Jul 29, 2026
- 86.50Jun 15, 2026
- MathVista80.20Jun 15, 2026MathVista (mini)
- MMMU-Pro74.05Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)68.50Aug 24, 2026CharXiv RQ
27.6% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond86.57Oct 8, 2026gpqa
- LiveBench · Reasoning80.09Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam30.12Oct 8, 2026aa_hle
- ARC-AGI-237.60Oct 7, 2026ARC-AGI v2
- 4.57Oct 8, 2026
Show 15 more reasoning resultsHide 15 reasoning results
- 37.64Sep 22, 2026
- 30.56Sep 22, 2026
- 22.78Sep 22, 2026
- 13.89Sep 22, 2026
- 9.44Sep 22, 2026
- 7.78May 10, 2026
- 0.29Oct 8, 2026
- GPQA Diamond81.01Oct 8, 2026gpqa
- GPQA Diamond87.00Oct 7, 2026GPQA
- 86.95Jul 29, 2026
- GPQA Diamond87.00Aug 24, 2026GPQA-D
- Humanity's Last Exam13.16Oct 8, 2026aa_hle
- Humanity's Last Exam28.40Aug 24, 2026HLE w/o tools
- 30.80Jun 15, 2026
- 62.00May 10, 2026
31.4% behind the leader9 of 10 ranked benchmarks measured
- 84.80Jun 15, 2026
- LiveBench · Coding79.65Oct 8, 2026livebench_coding@2026-06-25
- 80.90Oct 7, 2026
- SciCode49.54Sep 4, 2026aa_scicode
- Terminal-Bench Hard46.97Oct 8, 2026aa_terminalbench_hard
- 1469.10Jun 20, 2026
- 52.00Jul 29, 2026
- LiveBench · Agentic Coding39.70Oct 8, 2026livebench_agentic_coding@2026-06-25
- 7.00Oct 3, 2026
Show 19 more coding resultsHide 19 coding results
- LiveCodeBench v682.20Jun 15, 2026LiveCodeBench (v6)
- 1494.53May 22, 2026
- SciCode46.99Sep 4, 2026aa_scicode
- 50.00Aug 24, 2026
- 49.50Jun 15, 2026
- 70.70Aug 10, 2026
- 76.20Jul 29, 2026
- 77.50Jun 15, 2026
- 45.89Oct 8, 2026
- 57.10Jun 15, 2026
- 55.40Jun 15, 2026
- 74.40Sep 25, 2026
- 76.80Sep 25, 2026
- 81.50Jul 29, 2026
- SWE-bench Verified76.00Jul 29, 2026SWE-bench Verified (medium effort)
- 80.90Jun 15, 2026
- SWE-bench Verified74.40Jun 5, 2026SWE-bench Verified (mini-swe-agent)
- SWE-bench Verified75.20Jun 5, 2026SWE-bench Verified (Droid)
- Terminal-Bench Hard40.91Oct 8, 2026aa_terminalbench_hard
39.2% behind the leader4 of 7 ranked benchmarks measured
- OSWorld-Verified66.30Oct 7, 2026OSWorld
- 62.30Oct 7, 2026
- 47.29Jun 15, 2026
- 37.00Jun 15, 2026
Show 4 more agentic resultsHide 4 agentic results
- 45.97Jun 15, 2026
- 69.80Oct 8, 2026
- MCP Atlas65.20May 17, 2026MCP-Atlas (Public Set)
- OSWorld-Verified66.26Jul 29, 2026OSWorld
40.1% behind the leader3 of 3 ranked benchmarks measured
- LiveBench · Instruction Following62.55Oct 8, 2026livebench_instruction_following@2026-06-25
- 58.97Oct 8, 2026
- IFBench57.96Oct 8, 2026aa_ifbench
51.0% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics90.39Oct 8, 2026livebench_math@2026-06-25
- 34.39Sep 22, 2026
- 4.88Sep 22, 2026
Show 4 more math resultsHide 4 math results
- AIME 202695.10Jun 15, 2026AIME26
- HMMT Feb 202685.30Jun 15, 2026HMMT Feb 26
- 84.00Jun 15, 2026
- 78.50Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis74.44Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language81.26Oct 8, 2026livebench_language@2026-06-25
- 1470Oct 5, 2026
- 80.00Sep 22, 2026
- 75.83Sep 22, 2026
- 72.00Sep 22, 2026
Show 111 more resultsHide 111 results
- 59.64Jun 18, 2026
- 59.22Jun 18, 2026
- AA Intelligence35.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence23.68Oct 8, 2026aa_intelligence_index
- AA Intelligence29.10Oct 8, 2026aa_intelligence_index
- -4.13Oct 8, 2026
- 14.00Oct 8, 2026
- 92.80Jun 15, 2026
- AIME 202591.00Jun 5, 2026AIME25
- 93.30May 17, 2026
- AIME25 no tools91.00Aug 24, 2026AIME25
- 58.67Sep 22, 2026
- 35.17Sep 22, 2026
- 40.00May 10, 2026
- 82.00Jul 29, 2026
- 80.00Jul 29, 2026
- 37.60Jul 29, 2026
- 1473Aug 11, 2026
- 42.94Jun 18, 2026
- 47.83Jun 18, 2026
- BrowseComp (context management)59.20Aug 24, 2026BrowseComp (w/ctx manage)
- 57.80Jun 5, 2026
- browsecomp_with_context_manager67.80May 17, 2026BrowseComp (w/ Context Manage)
- 72.89Jul 29, 2026
- 67.59Jul 29, 2026
- 62.40May 17, 2026
- 92.20Jun 15, 2026
- 76.90Aug 24, 2026
- Claw Eval (pass@3)59.60Aug 24, 2026Claw-Eval Pass^3
- 76.60Jun 15, 2026
- 50.63Jul 29, 2026
- 50.60Jun 15, 2026
- DeepSearchQA (F1)76.10Jun 15, 2026DeepSearchQA
- 79.70Aug 24, 2026
- 46.80Aug 24, 2026
- 55.20Jul 29, 2026
- 66.20Jun 15, 2026
- frontiermath_tier_4_v14.17Aug 29, 2026frontiermath_tier_4
- HLE (with tools)43.20Jun 15, 2026HLE-Full (w/ tools)
- HLE (with tools)43.40May 17, 2026HLE (w/ Tools)
- HMMT Feb. 202592.90Jun 15, 2026HMMT Feb 25
- HMMT Nov. 202593.30Jun 15, 2026HMMT Nov 25
- 91.70May 17, 2026
- 76.90Jun 15, 2026
- 54.90Jul 29, 2026
- 69.20Jul 29, 2026
- 75.96Oct 7, 2026
- 87.00Jun 5, 2026
- 67.20Aug 24, 2026
- 77.10Jun 15, 2026
- 73.41May 10, 2026
- 40.00May 10, 2026
- 99.49May 10, 2026
- 81.70Aug 24, 2026
- 89.50Jun 15, 2026
- 89.30Jun 15, 2026
- 90.00Jun 5, 2026
- 95.60Jun 15, 2026
- 90.80Oct 7, 2026
- 90.77Jul 29, 2026
- 80.70Jul 29, 2026
- 77.30Jun 15, 2026
- 60.30Jun 15, 2026
- 50.00Jun 5, 2026
- 67.20Aug 24, 2026
- 43.20Jun 15, 2026
- 36.20Jun 5, 2026
- 54.60Aug 24, 2026
- 87.70Jun 15, 2026
- 72.90Jun 15, 2026
- 52.30Jun 15, 2026
- 1536.00Jun 15, 2026
- 77.00Aug 24, 2026
- 47.70Jun 15, 2026
- 65.70Aug 24, 2026
- 69.70Jun 15, 2026
- 45.30Jun 15, 2026
- 64.25Jul 29, 2026
- 70.60Jun 15, 2026
- 74.40Aug 29, 2026
- 76.80Aug 29, 2026
- 4.70Jun 5, 2026
- 80.20Jun 5, 2026
- 59.30Oct 7, 2026
- 59.30Jun 15, 2026
- 57.80Jun 5, 2026
- Terminal-Bench 2.057.90May 17, 2026Terminal-Bench 2.0 (Claude Code)
- 43.50May 17, 2026
- 43.50Jun 5, 2026
- vectara_answer_rate98.70Jun 20, 2026Answer Rate
- vectara_avg_summary_length114.50Jun 20, 2026Average Summary Length (Words)
- vectara_factual_consistency89.10Jun 20, 2026Factual Consistency Rate
- 29.00Jul 29, 2026
- 90.70Jun 5, 2026
- 92.20Jun 5, 2026
- 98.00Jun 5, 2026
- 90.00Jun 5, 2026
- 84.00Jun 5, 2026
- 89.10Jun 5, 2026
- 77.70Aug 24, 2026
- 84.40Aug 24, 2026
- 76.20Jun 15, 2026
- 36.80Jun 15, 2026
- 3.00Jun 15, 2026
- 9.00Jun 15, 2026
- τ²-Bench98.20Jul 29, 2026τ²-Bench (Telecom)
- 91.60Jun 13, 2026
- τ²-Bench90.00Jun 5, 2026𝜏²-Bench Telecom
- τ²-Bench (Retail)88.90Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)86.26Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)89.47Oct 8, 2026aa_tau2
Claude Opus 4.5: common questions
Who makes Claude Opus 4.5?
Claude Opus 4.5 is made by Anthropic.
When was Claude Opus 4.5 released?
Claude Opus 4.5 was released on Nov 24, 2025, according to Artificial Analysis.
What is Claude Opus 4.5 good at?
Claude Opus 4.5 is strong in long context; capable in factuality and multimodal tasks; and behind the leaders in reasoning, coding, agentic tasks, instruction following, and math. Too few results yet to rate safety or multilingual tasks.
How much does Claude Opus 4.5 cost?
Claude Opus 4.5 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 331 priced models we track.
How many benchmarks has Claude Opus 4.5 been tested on?
We track 205 results for Claude Opus 4.5 on 128 benchmarks from 26 sources, 44 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4.5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.5.
About this record
Where Claude Opus 4.5's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 7, 2026
Where the results come from
Verification: 205 scores · 44 independently verified · 32 aggregator-attributed · 98 vendor cross-reference · 31 vendor-reported. How these tiers are assigned
From 26 sources on 14 sites. Hugging Face supplies 98 of them; the 44 independently verified results come from 10 sites. Bars are coloured by trust tier.
- huggingface.co98
- artificialanalysis.ai33
- www-cdn.anthropic.com19
- arcprize.org12
- api.llm-stats.com9
- livebench.ai7
- raw.githubusercontent.com7
- swebench.com5
- epoch.ai4
- anthropic.com3
- labs.scale.com3
- datasets-server.huggingface.co2
- lmarena.ai2
- simple-bench.com1
Also known as
How our sources name Claude Opus 4.5 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| medium | claude 4.5 opus medium (20251101) claude 4.5 opus (20251101) (medium) | — |
| high | claude 4.5 opus (high reasoning) Claude 4.5 Opus (high) | claude-opus-4-5 (high) |