Claude Sonnet 5
Claude Sonnet 5 is strong in long context; capable in agentic tasks, coding, reasoning, multimodal tasks, and factuality; and behind the leaders in math. Too few results yet to rate safety, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Instruction Following.
Price
$2.00input$10.00outputper million tokens
From Anthropic's own price page · 4 providers tracked · All prices
Evidence
153results on81benchmarks
- 20 independently verified
- 69 aggregator
- 50 vendor-reported
- 14 cross-referenced
From 23 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Sonnet 5 benchmark results
153 results on 81 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
5.6% behind the leader2 of 3 ranked benchmarks measured
- 82.00Oct 8, 2026
- MRCR v2 (8-needle, 128K)81.50Jul 22, 2026GDM-MRCR v2 (8-needle) (128k (average))
13.8% behind the leader6 of 7 ranked benchmarks measured
- 81.20Oct 7, 2026
- 84.70Oct 7, 2026
- τ-Bench V3 · Banking37.32Oct 8, 2026tauBanking
- 48.25Oct 8, 2026
- Terminal-Bench 4.014.14Oct 8, 2026
Show 10 more agentic resultsHide 10 agentic results
- 33.02Oct 8, 2026
- 35.69Oct 8, 2026
- 37.89Oct 8, 2026
- 43.10Oct 8, 2026
- 28.53Oct 8, 2026
- Terminal-Bench 4.02.02Oct 8, 2026
- Terminal-Bench 4.05.05Oct 8, 2026
- Terminal-Bench 4.07.07Oct 8, 2026
- Terminal-Bench 4.02.53Oct 8, 2026
- τ-Bench V3 · Banking15.67Oct 8, 2026tauBanking
14.3% behind the leader8 of 10 ranked benchmarks measured
- 85.20Oct 7, 2026
- LiveBench · Coding80.68Oct 8, 2026livebench_coding@2026-06-25
- Terminal-Bench 2.180.52Oct 8, 2026terminalbenchV21
- 78.30Oct 7, 2026
- SciCode54.28Oct 8, 2026aa_scicode
- 1540.86Sep 21, 2026
- LiveBench · Agentic Coding59.39Oct 8, 2026livebench_agentic_coding@2026-06-25
- 63.20Oct 7, 2026
Show 8 more coding resultsHide 8 coding results
- 1536.72Jul 3, 2026
- SciCode51.62Oct 8, 2026aa_scicode
- SciCode54.05Oct 8, 2026aa_scicode
- SciCode50.12Oct 8, 2026aa_scicode
- SciCode48.61Sep 4, 2026aa_scicode
- Terminal-Bench 2.175.28Oct 8, 2026terminalbenchV21
- 80.40Jul 7, 2026
- 80.40Sep 11, 2026
15.5% behind the leader5 of 6 ranked benchmarks measured
- LiveBench · Reasoning88.69Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond91.11Oct 8, 2026gpqa
- 60.60Jul 10, 2026
- Humanity's Last Exam41.29Oct 8, 2026aa_hle
- 16.86Oct 8, 2026
Show 13 more reasoning resultsHide 13 reasoning results
- 8.57Oct 8, 2026
- 1.14Oct 8, 2026
- 15.14Oct 8, 2026
- 15.43Oct 8, 2026
- 4.57Oct 8, 2026
- GPQA Diamond80.00Oct 8, 2026gpqa
- Humanity's Last Exam29.98Oct 8, 2026aa_hle
- Humanity's Last Exam19.05Oct 8, 2026aa_hle
- Humanity's Last Exam39.02Oct 8, 2026aa_hle
- Humanity's Last Exam35.73Oct 8, 2026aa_hle
- Humanity's Last Exam21.92Oct 8, 2026aa_hle
- 57.40Oct 7, 2026
- Humanity's Last Exam54.90Sep 28, 2026Humanity's Last Exam With tools
20.6% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)88.30Oct 7, 2026CharXiv-R
- MMMU-Pro77.28Oct 8, 2026aa_mmmu_pro
Show 4 more multimodal resultsHide 4 multimodal results
- CharXiv (reasoning)77.00Jul 7, 2026CharXiv Reasoning (without tools)
- CharXiv (reasoning)70.10Sep 11, 2026CharXiv Reasoning (No tools)
- 1275.15Sep 21, 2026
- MMMU-Pro71.91Oct 8, 2026aa_mmmu_pro
23.8% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination60.63Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy40.05Oct 8, 2026omniscienceAccuracy
- 33.70Sep 21, 2026
Show 11 more factuality resultsHide 11 factuality results
- AA-Omniscience · Accuracy37.10Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy33.80Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy38.97Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy37.43Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy37.35Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination30.10Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination47.96Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination34.28Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination41.37Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination27.24Oct 8, 2026omniscienceNonHallucination
- 32.90Sep 21, 2026
33.2% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics92.94Oct 8, 2026livebench_math@2026-06-25
- 65.61Sep 21, 2026
- 29.27Sep 21, 2026
Show 1 more math resultHide 1 math result
- 79.50Jul 7, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following63.86Oct 8, 2026livebench_instruction_following@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis71.74Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language74.97Oct 8, 2026livebench_language@2026-06-25
- AA Intelligence38.00Oct 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence23.00Sep 21, 2026Artificial Analysis Intelligence Index
- 1461Sep 13, 2026
- AA Intelligence43.00Jul 2, 2026Artificial Analysis Intelligence Index
Show 67 more resultsHide 67 results
- 44.28Sep 9, 2026
- 34.25Sep 4, 2026
- AA Intelligence53.00Jul 1, 2026Artificial Analysis Intelligence Index
- AA Intelligence23.20Oct 8, 2026aa_intelligence_index
- AA Intelligence28.05Oct 8, 2026aa_intelligence_index
- AA Intelligence31.66Oct 8, 2026aa_intelligence_index
- AA Intelligence34.38Oct 8, 2026aa_intelligence_index
- AA Intelligence24.26Oct 8, 2026aa_intelligence_index
- AA Intelligence38.16Oct 8, 2026aa_intelligence_index
- -6.87Oct 8, 2026
- -0.65Oct 8, 2026
- -3.68Oct 8, 2026
- 3.18Oct 8, 2026
- -8.23Oct 8, 2026
- 16.45Oct 8, 2026
- 33.30Aug 14, 2026
- Artificial Analysis Coding Index66.39Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index71.55Sep 9, 2026aa_coding_index
- 0.37Jul 7, 2026
- 0.27Jul 7, 2026
- 86.60Aug 12, 2026
- 86.70Jul 7, 2026
- 70.10Jul 7, 2026
- 88.30Jul 7, 2026
- 1541.00Aug 14, 2026
- 54.00Oct 7, 2026
- DeepSWE 1.153.80Jul 22, 2026DeepSWE v1.1
- GDP (Surge AI)81.60Oct 7, 2026GDP.pdf
- GDP.PDF (All pass rate)28.00Sep 9, 2026
- GDPval-AA 2.11449.00Sep 28, 2026GDPval-AA v2.1
- 1618.00Jul 7, 2026
- GDPval-AA v2 Elo1584.00Jul 22, 2026GDPVal-AA v2 (Elo)
- 89.20Sep 22, 2026
- 89.00Sep 1, 2026
- HealthBench (raw score)59.20Sep 22, 2026
- 59.20Sep 1, 2026
- 57.80Oct 7, 2026
- HealthBench Professional (raw score)62.40Sep 22, 2026
- 62.40Sep 1, 2026
- HLE (with tools)57.40Jul 7, 2026Humanity's Last Exam (With tools)
- 31.00Aug 14, 2026
- 5.80Oct 7, 2026
- Legal Agent Benchmark, Full Public Set8.90Jul 7, 2026Legal Agent Full Public Benchmark Set
- 68.50Aug 14, 2026
- 89.30Sep 22, 2026
- 66.90Jul 22, 2026
- 75.10Sep 28, 2026
- 73.30Jul 7, 2026
- 59.40Oct 7, 2026
- 62.10Sep 28, 2026
- OSWorld 2.0 (partial)42.60Sep 9, 2026OSWorld-2.0 (Partial score)
- OSWorld 2.157.00Sep 28, 2026
- OSWorld 2.1 partial57.00Sep 28, 2026
- OSWorld 2.1 partial score57.00Sep 28, 2026
- OSWorld 2.1 strict pass rate25.60Sep 28, 2026
- 77.30Sep 28, 2026
- 36.60Sep 1, 2026
- 28.10Oct 7, 2026
- 80.40Oct 7, 2026
- 14.60Aug 14, 2026
- Terminal-Bench 4.012.40Oct 7, 2026
- Terminal-Bench 4.012.40Sep 11, 2026
- 54.30Oct 7, 2026
- Toolathlon63.00Jul 7, 2026Toolathlon Pass@3
- 40.70Jul 7, 2026
- Toolathlon Verified84.30Sep 22, 2026Toolathlon Verified Pass@3
- Toolathlon Verified74.70Sep 22, 2026Toolathlon Verified Pass@1
Claude Sonnet 5: common questions
Who makes Claude Sonnet 5?
Claude Sonnet 5 is made by Anthropic.
When was Claude Sonnet 5 released?
Claude Sonnet 5 was released on Jun 30, 2026, according to Artificial Analysis.
What is Claude Sonnet 5 good at?
Claude Sonnet 5 is strong in long context; capable in agentic tasks, coding, reasoning, multimodal tasks, and factuality; and behind the leaders in math. Too few results yet to rate safety, multilingual tasks, or instruction following.
How much does Claude Sonnet 5 cost?
Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens, according to Anthropic's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 82% of the 330 priced models we track.
How many benchmarks has Claude Sonnet 5 been tested on?
We track 153 results for Claude Sonnet 5 on 81 benchmarks from 23 sources, 20 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Sonnet 5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Sonnet 5.
About this record
Where Claude Sonnet 5's numbers come from, and every name it appears under.
- Tracked since
- Jul 1, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 153 scores · 20 independently verified · 69 aggregator-attributed · 14 vendor cross-reference · 50 vendor-reported. How these tiers are assigned
From 23 sources on 10 sites. Artificial Analysis supplies 73 of them; the 20 independently verified results come from 6 sites. Bars are coloured by trust tier.
- artificialanalysis.ai73
- www-cdn.anthropic.com33
- api.llm-stats.com16
- deepmind.google14
- livebench.ai7
- epoch.ai4
- datasets-server.huggingface.co3
- anthropic.com1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name Claude Sonnet 5 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude sonnet 5 (low) | claude sonnet 5 (adaptive reasoning, low effort) | claude-sonnet-5-low |
| medium | claude sonnet 5 (medium) | claude sonnet 5 (adaptive reasoning, medium effort) | claude-sonnet-5-medium |
| high | claude sonnet 5 (high) | claude sonnet 5 (adaptive reasoning, high effort) claude sonnet 5 (non-reasoning, high effort) | claude-sonnet-5-high |
| xhigh | claude sonnet 5 (xhigh) | claude sonnet 5 (adaptive reasoning, xhigh effort) | claude-sonnet-5-xhigh |
| max | claude sonnet 5 (max) | claude sonnet 5 (adaptive reasoning, max effort) | — |