Claude Opus 4.6
Claude Opus 4.6 is capable in long context, reasoning, coding, and factuality; and behind the leaders in multimodal tasks, agentic tasks, and math. Too few results yet to rate safety, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Instruction Following.
Price
$5.00input$25.00outputper million tokens
From Anthropic's own price page · 3 providers tracked · All prices
Evidence
231results on134benchmarks
- 55 independently verified
- 33 aggregator
- 50 vendor-reported
- 93 cross-referenced
From 42 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 4.6 benchmark results
231 results on 134 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
10.5% behind the leader2 of 3 ranked benchmarks measured
- 78.00Oct 8, 2026
- 84.00May 1, 2026
Show 1 more long context resultHide 1 long context result
- 67.00Oct 8, 2026
13.7% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning88.67Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond89.60Oct 8, 2026gpqa
- 67.60May 10, 2026
- ARC-AGI-268.80Oct 7, 2026ARC-AGI v2
- Humanity's Last Exam39.94Oct 8, 2026aa_hle
- 12.57Oct 8, 2026
Show 16 more reasoning resultsHide 16 reasoning results
- 68.75Sep 21, 2026
- 69.17Sep 21, 2026
- 66.25Sep 21, 2026
- 64.58Sep 21, 2026
- 68.80May 1, 2026
- 0.51Sep 21, 2026
- 2.84Oct 8, 2026
- GPQA Diamond84.04Oct 8, 2026gpqa
- GPQA Diamond91.30Oct 7, 2026GPQA
- 91.30Aug 26, 2026
- GPQA Diamond90.00Aug 24, 2026GPQA-D
- Humanity's Last Exam19.09Oct 8, 2026aa_hle
- 53.10Oct 7, 2026
- Humanity's Last Exam30.70Aug 24, 2026HLE w/o tools
- Humanity's Last Exam40.00Jun 27, 2026HLE (Pass@1)
- 36.70May 18, 2026
21.1% behind the leader9 of 10 ranked benchmarks measured
- 88.80Aug 26, 2026
- LiveBench · Coding78.18Oct 8, 2026livebench_coding@2026-06-25
- 80.80Oct 7, 2026
- 77.83Oct 7, 2026
- 1546.39Sep 21, 2026
- SciCode51.85Sep 4, 2026aa_scicode
- Terminal-Bench Hard46.21Oct 8, 2026aa_terminalbench_hard
- LiveBench · Agentic Coding48.99Oct 8, 2026livebench_agentic_coding@2026-06-25
- 53.40Jun 21, 2026
Show 20 more coding resultsHide 20 coding results
- 1537.11May 22, 2026
- 1544.67May 22, 2026
- SciCode45.72Sep 4, 2026aa_scicode
- 52.00Aug 24, 2026
- 51.90May 3, 2026
- 72.00May 1, 2026
- 77.80Jun 21, 2026
- 77.50Aug 26, 2026
- 77.80May 3, 2026
- 51.90Oct 8, 2026
- 53.40Aug 26, 2026
- SWE-bench Pro57.30Jul 4, 2026SWE Pro (Resolved)
- 75.60May 1, 2026
- SWE-bench Verified81.42Jun 4, 2026SWE-bench (prompt modification)
- SWE-bench Verified80.80Jul 4, 2026SWE Verified (Resolved)
- SWE-bench Verified75.90Jun 5, 2026SWE-Bench Verified (OpenCode)
- SWE-bench Verified78.90Jun 5, 2026SWE-Bench Verified (Droid)
- Terminal-Bench 2.178.20Aug 14, 2026Terminal Bench 2.1 (Terminus)
- Terminal-Bench Hard48.48Oct 8, 2026aa_terminalbench_hard
- 65.40May 1, 2026
23.2% behind the leader4 of 4 ranked benchmarks measured
- 12.20May 2, 2026
- AA-Omniscience · Accuracy46.98Oct 8, 2026omniscienceAccuracy
- 47.00Sep 21, 2026
- AA-Omniscience · Non-hallucination37.16Oct 8, 2026omniscienceNonHallucination
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy45.78Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination19.92Oct 8, 2026omniscienceNonHallucination
- SimpleQA Verified46.20Jun 27, 2026SimpleQA-Verified (Pass@1)
30.0% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro75.43Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)77.40Oct 7, 2026CharXiv-R
Show 12 more multimodal resultsHide 12 multimodal results
- CharXiv (reasoning)69.10Jun 21, 2026CharXiv Reasoning (No tools)
- CharXiv (reasoning)66.00Aug 26, 2026CharXiv (RQ)
- CharXiv (reasoning)69.10May 3, 2026CharXiv (RQ)
- 1312.99Sep 22, 2026
- 1315.75Sep 21, 2026
- 1297Jun 17, 2026
- 1300May 25, 2026
- MMMU-Pro72.54Oct 8, 2026aa_mmmu_pro
- 77.30Oct 7, 2026
- 70.60Jun 3, 2026
- 73.90May 3, 2026
- 48.40May 22, 2026
31.5% behind the leader4 of 7 ranked benchmarks measured
- 84.00Oct 7, 2026
- OSWorld-Verified72.70Oct 7, 2026OSWorld
- 62.70Oct 7, 2026
- 55.94Jun 15, 2026
Show 13 more agentic resultsHide 13 agentic results
- 33.04Oct 8, 2026
- AA ApexAgents33.00May 3, 2026APEX-Agents
- 29.80May 1, 2026
- BrowseComp86.80Aug 13, 2026BrowseComp (multi-agent harness)
- 83.70Jun 21, 2026
- BrowseComp83.70Jul 4, 2026BrowseComp (Pass@1)
- 84.00May 1, 2026
- 54.33Jun 15, 2026
- 76.80Oct 8, 2026
- 75.80Jun 21, 2026
- MCP Atlas73.80May 18, 2026MCP-Atlas (Public Set)
- 59.50May 1, 2026
- 72.70Aug 24, 2026
36.4% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics89.32Oct 8, 2026livebench_math@2026-06-25
- 65.96Sep 21, 2026
- 26.83Sep 21, 2026
Show 11 more math resultsHide 11 math results
- 96.67Sep 2, 2026
- 96.67May 2, 2026
- 95.60May 18, 2026
- 96.70May 3, 2026
- 96.21Sep 2, 2026
- 96.21May 10, 2026
- HMMT Feb 202696.20Jul 4, 2026HMMT 2026 Feb (Pass@1)
- HMMT Feb 202684.30May 18, 2026HMMT Feb. 2026
- IMO-AnswerBench75.30Jul 4, 2026IMOAnswerBench (Pass@1)
- 47.02Sep 2, 2026
- 66.20May 3, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following63.31Oct 8, 2026livebench_instruction_following@2026-06-25
- 37.15Oct 8, 2026
- 56.02Oct 8, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1497Oct 8, 2026
- livebench_language83.27Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis69.89Oct 8, 2026livebench_data_analysis@2026-06-25
- 38.33Oct 8, 2026
- AA Intelligence26.00Sep 21, 2026Artificial Analysis Intelligence Index
- 93.00Sep 21, 2026
Show 112 more resultsHide 112 results
- 67.58Jun 18, 2026
- 64.22Jun 18, 2026
- AA Intelligence38.00Jun 2, 2026Artificial Analysis Intelligence Index
- AA Intelligence31.95Oct 8, 2026aa_intelligence_index
- AA Intelligence26.35Oct 8, 2026aa_intelligence_index
- 2.37Oct 8, 2026
- 13.67Oct 8, 2026
- 61.74Jun 24, 2026
- 69.90Jun 24, 2026
- 70.20Jun 24, 2026
- 57.80Jun 24, 2026
- 29.30Jun 24, 2026
- 64.55Jun 24, 2026
- 57.51Jun 24, 2026
- 51.42Jun 24, 2026
- 99.79Oct 7, 2026
- AIME25 no tools95.60Aug 24, 2026AIME25
- 62.00Aug 26, 2026
- 34.50Jul 4, 2026
- 85.90Jul 4, 2026
- 94.00Sep 21, 2026
- 92.00Sep 21, 2026
- 86.00Sep 21, 2026
- 93.00Jun 21, 2026
- Artificial Analysis Coding Index48.09Jun 18, 2026aa_coding_index
- 47.56Jun 18, 2026
- baby_vision_with_python38.40May 3, 2026BabyVision (w/ python)
- 12.60Aug 24, 2026
- 14.80May 3, 2026
- 90.20Jun 3, 2026
- 83.70Jun 6, 2026
- 86.57Jun 3, 2026
- 83.73Jun 3, 2026
- browsecomp_with_context_manager84.00May 18, 2026BrowseComp (w/ Context Manage)
- 84.70Jun 21, 2026
- charxiv_rq_with_python84.70May 3, 2026CharXiv (RQ) (w/ python)
- Chinese SimpleQA (C-SimpleQA)76.40Jul 4, 2026Chinese-SimpleQA (Pass@1)
- 82.40May 3, 2026
- 71.70Jul 4, 2026
- 73.80Oct 7, 2026
- 66.60May 18, 2026
- 80.60May 3, 2026
- DeepSearchQA (F1)91.30Oct 7, 2026DeepSearchQA
- DeepSearchQA (F1)91.30May 3, 2026DeepSearchQA (f1-score)
- 78.90Jun 12, 2026
- 40.80Aug 26, 2026
- 60.70Oct 7, 2026
- 60.10Jun 21, 2026
- Finance Agent76.7May 11, 2026Finance Agent evaluation (General Finance module)
- 39.60Aug 25, 2026
- frontiermath_tier_4_v122.90Aug 29, 2026frontiermath_tier_4
- frontiermath_tier_4_v120.83May 20, 2026frontiermath_tier_4
- 1619.00Jul 4, 2026
- 1606.00May 1, 2026
- HLE (with tools)53.30Jun 21, 2026Humanity's Last Exam (With tools)
- HLE (with tools)53.00Jun 3, 2026HLE with tools
- HLE (with tools)53.10Jul 4, 2026HLE w/ tools (Pass@1)
- 96.30May 18, 2026
- 36.60Aug 26, 2026
- 4.20Oct 7, 2026
- 76.33Oct 7, 2026
- LiveCodeBench88.80Jul 4, 2026LiveCodeBench (Pass@1)
- 63.00Aug 26, 2026
- 65.50Aug 26, 2026
- 71.20May 3, 2026
- 72.26Sep 2, 2026
- mathvision_with_python84.60May 3, 2026MathVision (w/ python)
- 73.80Jul 4, 2026
- 56.70May 3, 2026
- 76.99May 10, 2026
- 97.44May 10, 2026
- 98.48May 10, 2026
- 76.00Jun 3, 2026
- MMLU-Pro89.10Jul 4, 2026MMLU-Pro (EM)
- 91.10Oct 7, 2026
- MMMLU91.10May 1, 2026mmmlu_multilingual_qa
- mmmu_pro_with_python77.30May 3, 2026MMMU-Pro (w/ python)
- 76.00Jun 6, 2026
- 49.80May 18, 2026
- 59.80Sep 11, 2026
- 73.50Jun 21, 2026
- 57.10Jun 21, 2026
- 60.30May 3, 2026
- 86.60Aug 24, 2026
- OpenAI-MRCR (1M)92.90Jul 4, 2026MRCR 1M (MMR)
- 75.90Jun 12, 2026
- 73.90Aug 26, 2026
- 57.70Jun 21, 2026
- 83.10Jun 21, 2026
- 75.60May 1, 2026
- 27.10Jun 21, 2026
- 27.10Aug 24, 2026
- 65.40Oct 7, 2026
- Terminal-Bench 2.065.40May 18, 2026Terminal-Bench 2.0 (Terminus-2)
- 47.20May 18, 2026
- Toolathlon66.70Jun 30, 2026Toolathlon Pass@3
- Toolathlon56.80Jun 30, 2026Toolathlon Pass@1
- Toolathlon47.20Jul 4, 2026Toolathlon (Pass@1)
- 16.90Jun 30, 2026
- 47.20Jul 7, 2026
- 86.40May 3, 2026
- vectara_answer_rate99.80May 2, 2026Answer Rate
- vectara_avg_summary_length137.60May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency87.80May 2, 2026Factual Consistency Rate
- 801759.00Oct 7, 2026
- 54.50May 2, 2026
- 91.90May 6, 2026
- τ²-Bench (Retail)91.90Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)84.80Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)92.11Oct 8, 2026aa_tau2
- 99.30May 1, 2026
- 72.40May 18, 2026
Claude Opus 4.6: common questions
Who makes Claude Opus 4.6?
Claude Opus 4.6 is made by Anthropic.
When was Claude Opus 4.6 released?
Claude Opus 4.6 was released on Feb 5, 2026, according to Artificial Analysis.
What is Claude Opus 4.6 good at?
Claude Opus 4.6 is capable in long context, reasoning, coding, and factuality; and behind the leaders in multimodal tasks, agentic tasks, and math. Too few results yet to rate safety, multilingual tasks, or instruction following.
How much does Claude Opus 4.6 cost?
Claude Opus 4.6 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 331 priced models we track.
How many benchmarks has Claude Opus 4.6 been tested on?
We track 231 results for Claude Opus 4.6 on 134 benchmarks from 42 sources, 55 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4.6 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.6.
About this record
Where Claude Opus 4.6's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Sep 26, 2026
Where the results come from
Verification: 231 scores · 55 independently verified · 33 aggregator-attributed · 93 vendor cross-reference · 50 vendor-reported. How these tiers are assigned
From 42 sources on 19 sites. Hugging Face supplies 82 of them; the 55 independently verified results come from 12 sites. Bars are coloured by trust tier.
- huggingface.co82
- artificialanalysis.ai35
- www-cdn.anthropic.com22
- api.llm-stats.com20
- deepmind.google10
- arcprize.org9
- anthropic.com7
- livebench.ai7
- raw.githubusercontent.com7
- matharena.ai6
- datasets-server.huggingface.co5
- epoch.ai5
- labs.scale.com5
- lmarena.ai3
- swebench.com3
- 99franklin.github.io2
- cdn.sanity.io1
- mistral.ai1
- simple-bench.com1
Also known as
How our sources name Claude Opus 4.6 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude opus 4.6 (120k, low) | — | — |
| medium | claude opus 4.6 (120k, medium) opus4.6 medium | — | — |
| high | claude opus 4.6 (120k, high) claude opus 4.6 (non-reasoning, high) | Claude Opus 4.6 (Non-reasoning, High Effort) | claude-opus-4.6 (high) claude-opus-4-6-high |
| max | anthropic opus 4.6 (max) claude opus 4.6 (120k, max) Opus-4.6 Max Opus 4.6 Thinking (Max) claude opus 4.6 (max) opus4.6 max claude opus 4.6 max | claude opus 4.6 (adaptive reasoning, max effort) claude opus 4.6 (max effort) | claude-opus-4-6-thinking-max claude-opus-4-6 (max) |