Claude Sonnet 4.6
Claude Sonnet 4.6 is strong in long context; capable in coding; and behind the leaders in reasoning, agentic tasks, factuality, multimodal tasks, math, and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
152results on88benchmarks
- 24 independently verified
- 52 aggregator
- 45 vendor-reported
- 31 cross-referenced
From 20 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Sonnet 4.6 benchmark results
152 results on 88 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
5.7% behind the leader2 of 3 ranked benchmarks measured
- 80.00Oct 8, 2026
- MRCR v2 (8-needle, 128K)84.90Jul 6, 2026MRCR v2 (8-needle) Long context performance
22.5% behind the leader8 of 10 ranked benchmarks measured
- LiveBench · Coding79.27Oct 8, 2026livebench_coding@2026-06-25
- 79.60Oct 7, 2026
- Terminal-Bench Hard53.03Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.171.16Oct 8, 2026terminalbenchV21
- SciCode50.12Oct 8, 2026aa_scicode
- 1522.14May 22, 2026
- 58.10Aug 12, 2026
- LiveBench · Agentic Coding42.63Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 8 more coding resultsHide 8 coding results
- SciCode44.10Sep 4, 2026aa_scicode
- SciCode46.88Sep 4, 2026aa_scicode
- 47.00May 1, 2026
- 79.60May 1, 2026
- 67.00Aug 12, 2026
- Terminal-Bench Hard42.42Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard46.21Oct 8, 2026aa_terminalbench_hard
- 59.10May 1, 2026
25.2% behind the leader5 of 6 ranked benchmarks measured
- LiveBench · Reasoning84.77Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond87.47Oct 8, 2026gpqa
- ARC-AGI-258.30Oct 7, 2026ARC-AGI v2
- Humanity's Last Exam33.60Oct 8, 2026aa_hle
- 3.14Oct 8, 2026
Show 13 more reasoning resultsHide 13 reasoning results
- 58.33Sep 21, 2026
- 60.42May 10, 2026
- ARC-AGI-258.30Jul 6, 2026ARC-AGI-2 Abstract reasoning puzzles
- 0.86Oct 8, 2026
- GPQA Diamond79.70Oct 8, 2026gpqa
- GPQA Diamond79.90Oct 8, 2026gpqa
- GPQA Diamond89.90Oct 7, 2026GPQA
- Humanity's Last Exam11.21Oct 8, 2026aa_hle
- Humanity's Last Exam13.35Oct 8, 2026aa_hle
- 49.00Oct 7, 2026
- Humanity's Last Exam34.60Jul 7, 2026Humanity's Last Exam (No tools)
- Humanity's Last Exam33.20Jul 6, 2026Humanity’s Last Exam Academic reasoning (full set, text + MM)
- 49.00May 1, 2026
27.8% behind the leader7 of 7 ranked benchmarks measured
- OSWorld-Verified72.50Oct 7, 2026OSWorld
- 74.70Oct 7, 2026
- 61.30Oct 7, 2026
- τ-Bench V3 · Banking34.43Oct 8, 2026tauBanking
- 36.70Oct 8, 2026
- Terminal-Bench 4.03.03Oct 8, 2026
Show 11 more agentic resultsHide 11 agentic results
- 28.02Oct 8, 2026
- 39.80Oct 8, 2026
- 76.20Jun 30, 2026
- 74.70May 1, 2026
- 47.91Jun 15, 2026
- 54.76Jun 15, 2026
- 69.50Oct 8, 2026
- MCP Atlas69.50Jul 6, 2026MCP Atlas Multi-step workflows using MCP
- 61.30May 1, 2026
- 78.50Aug 12, 2026
- OSWorld-Verified72.50Jul 6, 2026OSWorld-Verified Agentic computer use
28.0% behind the leader4 of 4 ranked benchmarks measured
- 10.60May 2, 2026
- AA-Omniscience · Accuracy40.87Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination51.55Oct 8, 2026omniscienceNonHallucination
- 32.80Sep 21, 2026
Show 6 more factuality resultsHide 6 factuality results
- AA-Omniscience · Accuracy36.52Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy38.55Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination39.17Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination31.49Oct 8, 2026omniscienceNonHallucination
- 35.50Sep 21, 2026
- 29.00May 20, 2026
33.8% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro73.29Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)71.60Jul 7, 2026CharXiv Reasoning (without tools)
Show 7 more multimodal resultsHide 7 multimodal results
- CharXiv (reasoning)72.40Jul 6, 2026CharXiv Reasoning Information synthesis from complex charts
- 1282.60Aug 25, 2026
- 1278Jun 17, 2026
- MMMU-Pro69.19Oct 8, 2026aa_mmmu_pro
- MMMU-Pro70.58Oct 8, 2026aa_mmmu_pro
- 75.60Oct 7, 2026
- MMMU-Pro74.50Jul 6, 2026MMMU-Pro Multimodal understanding and reasoning
39.9% behind the leader1 of 5 ranked benchmarks measured
- LiveBench · Mathematics86.99Oct 8, 2026livebench_math@2026-06-25
Show 1 more math resultHide 1 math result
- 55.00Jul 7, 2026
42.1% behind the leader2 of 3 ranked benchmarks measured
- LiveBench · Instruction Following63.22Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench56.60Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language76.10Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis77.95Oct 8, 2026livebench_data_analysis@2026-06-25
- 1472Oct 5, 2026
- 86.00Sep 21, 2026
- frontiermath_tier_4_v18.30Jun 23, 2026frontiermath_tier_4
- 86.50May 10, 2026
Show 66 more resultsHide 66 results
- 33.12Sep 9, 2026
- 57.45Jun 18, 2026
- 61.62Jun 18, 2026
- AA Intelligence30.06Oct 8, 2026aa_intelligence_index
- AA Intelligence23.30Oct 8, 2026aa_intelligence_index
- AA Intelligence24.69Oct 8, 2026aa_intelligence_index
- 12.22Oct 8, 2026
- -2.10Oct 8, 2026
- -3.55Oct 8, 2026
- 58.03Jun 24, 2026
- 70.00Jun 24, 2026
- 63.17Jun 24, 2026
- 56.04Jun 24, 2026
- 28.79Jun 24, 2026
- 64.52Jun 24, 2026
- 56.98Jun 24, 2026
- 50.78Jun 24, 2026
- Artificial Analysis Coding Index63.03Sep 9, 2026aa_coding_index
- 42.98Jun 18, 2026
- 46.43Jun 18, 2026
- 0.33Jul 7, 2026
- 0.27Jul 7, 2026
- Blueprint-Bench 26.70Jul 6, 2026Blueprint-Bench 2 Agentic spatial reasoning
- 76.20Aug 12, 2026
- 80.90Jul 7, 2026
- 59.30Jul 7, 2026
- 85.30Jul 7, 2026
- 30.00Oct 7, 2026
- 63.30Oct 7, 2026
- 51.03Oct 7, 2026
- Finance Agent v251.00Jul 6, 2026Finance Agent v2 Financial analysis and decision-making
- 23.70Aug 25, 2026
- GDP (Surge AI)78.60Jul 7, 2026GDP.pdf (with tools)
- GDPval-AA (Elo)1676.00Jul 6, 2026GDPval-AA Economically valuable knowledge work
- 1633.00May 1, 2026
- 1395.00Aug 12, 2026
- 1381.00Jun 30, 2026
- 89.90May 1, 2026
- 44.20Aug 12, 2026
- HLE (with tools)46.80Aug 12, 2026Humanity's Last Exam (With tools)
- 33.20Jul 7, 2026
- 5.40Oct 7, 2026
- 5.40Jun 30, 2026
- 8.00Aug 12, 2026
- 88.48Aug 12, 2026
- Legal Agent Benchmark, Full Public Set8.00Aug 12, 2026Legal Agent Benchmark (Full Public Set)
- 75.47Oct 7, 2026
- 89.30Oct 7, 2026
- MMMLU89.30May 1, 2026mmmlu_multilingual_qa
- 68.70Jul 7, 2026
- 53.40Jul 7, 2026
- 59.10Oct 7, 2026
- Toolathlon41.00Aug 12, 2026Toolathlon Pass@1
- Toolathlon60.20Jun 30, 2026Toolathlon Pass@3
- Toolathlon49.40Jun 30, 2026Toolathlon Pass@1
- 16.50Jun 30, 2026
- 38.00Jul 7, 2026
- vectara_answer_rate99.90May 2, 2026Answer Rate
- vectara_avg_summary_length114.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency89.40May 2, 2026Factual Consistency Rate
- 91.70May 6, 2026
- τ²-Bench (Retail)91.70Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)78.95Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)75.73Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)79.53Oct 8, 2026aa_tau2
- 97.90May 1, 2026
Claude Sonnet 4.6: common questions
Who makes Claude Sonnet 4.6?
Claude Sonnet 4.6 is made by Anthropic.
When was Claude Sonnet 4.6 released?
Claude Sonnet 4.6 was released on Feb 17, 2026, according to Artificial Analysis.
What is Claude Sonnet 4.6 good at?
Claude Sonnet 4.6 is strong in long context; capable in coding; and behind the leaders in reasoning, agentic tasks, factuality, multimodal tasks, math, and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does Claude Sonnet 4.6 cost?
Claude Sonnet 4.6 costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 331 priced models we track.
How many benchmarks has Claude Sonnet 4.6 been tested on?
We track 152 results for Claude Sonnet 4.6 on 88 benchmarks from 20 sources, 24 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Sonnet 4.6 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Sonnet 4.6.
About this record
Where Claude Sonnet 4.6's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 152 scores · 24 independently verified · 52 aggregator-attributed · 31 vendor cross-reference · 45 vendor-reported. How these tiers are assigned
From 20 sources on 13 sites. Artificial Analysis supplies 52 of them; the 24 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai52
- www-cdn.anthropic.com29
- deepmind.google22
- api.llm-stats.com16
- huggingface.co8
- livebench.ai7
- arcprize.org4
- epoch.ai4
- raw.githubusercontent.com4
- datasets-server.huggingface.co2
- lmarena.ai2
- labs.scale.com1
- mistral.ai1
Also known as
How our sources name Claude Sonnet 4.6 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude sonnet 4.6 (non-reasoning, low) | claude sonnet 4.6 (non-reasoning, low effort) | claude-sonnet-4-6-non-reasoning-low-effort |
| high | claude sonnet 4.6 (high) claude sonnet 4.6 (non-reasoning, high) | Claude Sonnet 4.6 (Non-reasoning, High Effort) | — |
| max | Sonnet 4.6 Thinking (Max) claude sonnet 4.6 (max) | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | — |