Claude Sonnet 4
Claude Sonnet 4 is capable in long context and factuality; and behind the leaders in multimodal tasks, instruction following, coding, and agentic tasks. Too few results yet to rate reasoning, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Safety, Math or Multilingual.
Price
$3.00input$15.00outputper million tokens
From anthropic · All prices
Evidence
146results on83benchmarks
- 45 independently verified
- 34 aggregator
- 14 vendor-reported
- 53 cross-referenced
From 27 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingReasoning
As listed by OpenRouter
Claude Sonnet 4 benchmark results
146 results on 83 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
24.9% behind the leader1 of 3 ranked benchmarks measured
- 70.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 44.00Oct 8, 2026
24.9% behind the leader3 of 4 ranked benchmarks measured
- 10.30Aug 31, 2026
- AA-Omniscience · Non-hallucination70.88Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy22.68Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy22.67Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination59.03Oct 8, 2026omniscienceNonHallucination
34.9% behind the leader2 of 6 ranked benchmarks measured
- 74.40Oct 7, 2026
- MMMU-Pro61.79Oct 8, 2026aa_mmmu_pro
Show 6 more multimodal resultsHide 6 multimodal results
- 1190.95Sep 22, 2026
- 1175.06Sep 1, 2026
- 1207May 23, 2026
- 1188May 1, 2026
- 72.60Jun 13, 2026
- MMMU-Pro62.37Oct 8, 2026aa_mmmu_pro
35.2% behind the leader2 of 3 ranked benchmarks measured
- 57.11Oct 8, 2026
- IFBench54.69Oct 8, 2026aa_ifbench
Show 3 more instruction following resultsHide 3 instruction following results
- IFBench45.37Oct 8, 2026aa_ifbench
- 55.00Jun 6, 2026
- 46.80Jun 15, 2026
40.1% behind the leader7 of 10 ranked benchmarks measured
- 72.70Oct 7, 2026
- SciCode40.05Sep 4, 2026aa_scicode
- 53.30May 15, 2026
- LiveCodeBench v648.50Jun 15, 2026LiveCodeBench v6 (Aug 24 - May 25)
- 42.70Oct 8, 2026
- Terminal-Bench Hard31.06Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.136.33Oct 8, 2026terminalbenchV21
Show 11 more coding resultsHide 11 coding results
- SciCode37.27Sep 4, 2026aa_scicode
- 40.00Jun 6, 2026
- SWE-bench Multilingual51.00Jun 15, 2026SWE-bench Multilingual (Agentic Coding)
- 56.90Jun 6, 2026
- 64.93Sep 1, 2026
- SWE-bench Verified80.20Jun 13, 2026SWE-bench (high compute)
- SWE-bench Verified80.20Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified50.20Jun 15, 2026SWE-bench Verified (Agentless Coding)
- SWE-bench Verified72.70Jun 15, 2026SWE-bench Verified (Agentic Coding)
- Terminal-Bench Hard27.27Oct 8, 2026aa_terminalbench_hard
- 30.00Jun 6, 2026
46.6% behind the leader4 of 7 ranked benchmarks measured
- OSWorld-Verified42.20Aug 13, 2026OSWorld
- τ-Bench V3 · Banking16.70Oct 8, 2026tauBanking
- 12.20Jun 6, 2026
- 9.73Oct 8, 2026
Show 1 more agentic resultHide 1 agentic result
- 31.18Jun 15, 2026
0 of 6 ranked benchmarks measured
- 5.93Sep 22, 2026
- 2.12Sep 22, 2026
- 0.85Sep 22, 2026
Show 14 more reasoning resultsHide 14 reasoning results
- 1.27May 10, 2026
- 1.14Oct 8, 2026
- 0.29Oct 8, 2026
- GPQA Diamond68.28Oct 8, 2026gpqa
- GPQA Diamond77.68Oct 8, 2026gpqa
- GPQA Diamond75.40Oct 7, 2026GPQA
- 70.00Jun 13, 2026
- 70.00Jun 15, 2026
- 78.00Jun 6, 2026
- Humanity's Last Exam4.26Oct 8, 2026aa_hle
- Humanity's Last Exam10.70Oct 8, 2026aa_hle
- Humanity's Last Exam5.80Jun 15, 2026Humanity's Last Exam (Text Only)
- Humanity's Last Exam9.60Jun 6, 2026HLE (w/o tools)
- 45.50May 10, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 40.00Sep 22, 2026
- 29.00Sep 22, 2026
- 28.00Sep 22, 2026
- 56.40Sep 1, 2026
- vectara_avg_summary_length145.80Aug 31, 2026Average Summary Length (Words)
- vectara_answer_rate98.60Aug 31, 2026Answer Rate
Show 80 more resultsHide 80 results
- 17.60Sep 4, 2026
- 39.21Jun 18, 2026
- AA Intelligence25.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence16.61Oct 8, 2026aa_intelligence_index
- AA Intelligence18.92Oct 8, 2026aa_intelligence_index
- -9.02Oct 8, 2026
- 0.17Oct 8, 2026
- 76.20Jun 15, 2026
- 70.70May 1, 2026
- 56.40Jun 15, 2026
- 33.10Jun 13, 2026
- 43.40Jun 15, 2026
- 70.50Oct 7, 2026
- 33.10Jun 15, 2026
- AIME 202574.00Jun 6, 2026AIME25
- 88.27May 19, 2026
- 99.15Jun 21, 2026
- 98.44May 10, 2026
- 23.83May 10, 2026
- 57.30Jun 6, 2026
- Artificial Analysis Coding Index37.57Sep 9, 2026aa_coding_index
- 30.60Jun 18, 2026
- 89.80Jun 15, 2026
- 96.60Jun 21, 2026
- 97.90May 10, 2026
- 29.10Jun 6, 2026
- 60.40Jun 15, 2026
- 42.00Jun 6, 2026
- frontiermath_tier_4_v10.00May 20, 2026frontiermath_tier_4
- 68.30Jun 6, 2026
- 98.00Jun 21, 2026
- 94.13May 10, 2026
- HLE (with tools)20.30Jun 6, 2026HLE (w/ tools)
- 15.90Jun 15, 2026
- 87.60Jun 15, 2026
- 74.80Jun 15, 2026
- 59.43May 15, 2026
- 68.53May 3, 2026
- LiveCodeBench66.00Jun 6, 2026LiveCodeBench (LCB)
- 96.27May 15, 2026
- 98.76May 3, 2026
- 20.57May 15, 2026
- 32.57May 3, 2026
- 63.97May 15, 2026
- 75.98May 3, 2026
- MATH-500 (EM)94.00Jun 15, 2026MATH-500
- 91.50Jun 15, 2026
- 83.70Jun 15, 2026
- 84.00Jun 6, 2026
- 93.60Jun 15, 2026
- 86.50Oct 7, 2026
- 85.40Jun 13, 2026
- 35.70Jun 6, 2026
- 88.60Jun 15, 2026
- 15.30Jun 15, 2026
- 52.80Jun 15, 2026
- 100.00Jun 21, 2026
- 99.50May 10, 2026
- 15.90Jun 15, 2026
- 55.70Jun 15, 2026
- 64.93May 1, 2026
- 58.33May 1, 2026
- 67.10May 15, 2026
- TAU-bench (airline)60.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)80.50Oct 7, 2026TAU-bench Retail
- 55.50Jun 15, 2026
- 35.50Sep 23, 2026
- 35.50Jun 15, 2026
- 36.40Jun 6, 2026
- theagentcompany37.00Jun 6, 2026AgentCompany
- vectara_factual_consistency89.70Aug 31, 2026Factual Consistency Rate
- 64.60Jun 6, 2026
- 96.61Jun 21, 2026
- 97.22May 10, 2026
- 73.70Jun 15, 2026
- τ²-Bench45.20Jun 15, 2026Tau2 telecom
- τ²-Bench65.00Jun 6, 2026τ²-Bench-Telecom
- τ²-Bench (Retail)75.00Jun 15, 2026Tau2 retail
- τ²-Bench Telecom (AA run)52.34Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)64.62Oct 8, 2026aa_tau2
Claude Sonnet 4: common questions
Who makes Claude Sonnet 4?
Claude Sonnet 4 is made by Anthropic.
When was Claude Sonnet 4 released?
Claude Sonnet 4 was released on May 22, 2025, according to Artificial Analysis.
What is Claude Sonnet 4 good at?
Claude Sonnet 4 is capable in long context and factuality; and behind the leaders in multimodal tasks, instruction following, coding, and agentic tasks. Too few results yet to rate reasoning, safety, math, or multilingual tasks.
How much does Claude Sonnet 4 cost?
Claude Sonnet 4 costs $3.00 per million input tokens and $15.00 per million output tokens, according to anthropic. At a mix of three input tokens to one output token, it costs more than 88% of the 331 priced models we track.
How many benchmarks has Claude Sonnet 4 been tested on?
We track 146 results for Claude Sonnet 4 on 83 benchmarks from 27 sources, 45 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Sonnet 4 support?
OpenRouter lists tool calling and reasoning for Claude Sonnet 4.
About this record
Where Claude Sonnet 4's numbers come from, and every name it appears under.
- Tracked since
- May 3, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 146 scores · 45 independently verified · 34 aggregator-attributed · 53 vendor cross-reference · 14 vendor-reported. How these tiers are assigned
From 27 sources on 15 sites. Hugging Face supplies 53 of them; the 45 independently verified results come from 12 sites. Bars are coloured by trust tier.
- huggingface.co53
- artificialanalysis.ai35
- storage.googleapis.com11
- api.llm-stats.com8
- arcprize.org8
- livecodebench.github.io8
- anthropic.com6
- raw.githubusercontent.com4
- swebench.com3
- aider.chat2
- datasets-server.huggingface.co2
- labs.scale.com2
- lmarena.ai2
- epoch.ai1
- simple-bench.com1