Claude Sonnet 4.5
Claude Sonnet 4.5 is capable in long context; and behind the leaders in multimodal tasks, coding, instruction following, factuality, reasoning, and math. Too few results yet to rate agentic tasks, safety, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety or Multilingual.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
197results on127benchmarks
- 44 independently verified
- 35 aggregator
- 26 vendor-reported
- 92 cross-referenced
From 40 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Sonnet 4.5 benchmark results
197 results on 127 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
16.9% behind the leader2 of 3 ranked benchmarks measured
- 61.80Jun 6, 2026
- 72.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 54.00Oct 8, 2026
30.9% behind the leader4 of 6 ranked benchmarks measured
- 79.60Aug 24, 2026
- MathVista79.80Aug 24, 2026Mathvista(mini)
- MMMU-Pro68.73Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)67.20Aug 24, 2026CharXiv(RQ)
31.9% behind the leader8 of 10 ranked benchmarks measured
- SWE-bench Verified77.20Oct 7, 2026SWE-bench Verified (Agentic Coding)
- 67.00Aug 10, 2026
- 64.00Jun 6, 2026
- SciCode45.72Oct 8, 2026aa_scicode
- Terminal-Bench 2.155.81Oct 8, 2026terminalbenchV21
- 1393.55Sep 21, 2026
- Terminal-Bench Hard35.61Oct 8, 2026aa_terminalbench_hard
- 43.60Oct 8, 2026
Show 15 more coding resultsHide 15 coding results
- 1384.77May 22, 2026
- 1392.36May 22, 2026
- SciCode42.82Sep 4, 2026aa_scicode
- 45.00Aug 24, 2026
- SciCode44.70Jun 12, 2026SciCode no tools
- SWE-bench Multilingual68.00Jun 12, 2026SWE-bench Multilingual w/ tools
- 71.40Sep 25, 2026
- 70.60May 1, 2026
- SWE-bench Verified82.00Jul 29, 2026SWE-bench Verified (high compute)
- SWE-bench Verified77.20Jun 12, 2026SWE-bench Verified w/ tools
- SWE-bench Verified70.60Jun 5, 2026SWE-bench Verified (mini-swe-agent)
- SWE-bench Verified72.30Jun 5, 2026SWE-bench Verified (Droid)
- Terminal-Bench Hard28.79Oct 8, 2026aa_terminalbench_hard
- 33.00Jun 6, 2026
- 33.30May 18, 2026
34.6% behind the leader2 of 3 ranked benchmarks measured
- IFBench57.28Oct 8, 2026aa_ifbench
- 55.32Oct 8, 2026
34.8% behind the leader4 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better12.00May 30, 2026
- AA-Omniscience · Non-hallucination50.92Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy32.87Oct 8, 2026omniscienceAccuracy
- 30.70Sep 22, 2026
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy28.38Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination47.29Oct 8, 2026omniscienceNonHallucination
- 23.70Sep 21, 2026
35.5% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond83.43Oct 8, 2026gpqa
- 54.30May 10, 2026
- Humanity's Last Exam17.79Oct 8, 2026aa_hle
- 13.61Sep 22, 2026
- 1.14Oct 8, 2026
Show 12 more reasoning resultsHide 12 reasoning results
- 6.94Sep 22, 2026
- 5.83Sep 22, 2026
- 3.75May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond72.73Oct 8, 2026gpqa
- GPQA Diamond83.40Oct 7, 2026GPQA
- GPQA Diamond83.00Aug 24, 2026GPQA-D
- 83.40Jun 6, 2026
- Humanity's Last Exam7.23Oct 8, 2026aa_hle
- Humanity's Last Exam17.30Aug 24, 2026HLE w/o tools
- Humanity's Last Exam19.80Jun 12, 2026HLE (Text-only) no tools
- Humanity's Last Exam13.70Jun 6, 2026HLE (no tools)
54.4% behind the leader2 of 5 ranked benchmarks measured
- 23.86Sep 22, 2026
- 2.44Sep 22, 2026
Show 2 more math resultsHide 2 math results
- IMO-AnswerBench65.90Jun 12, 2026IMO-AnswerBench no tools
- 65.80May 18, 2026
0 of 7 ranked benchmarks measured
- 59.50Oct 8, 2026
- 20.54Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
Show 6 more agentic resultsHide 6 agentic results
- 19.60Jun 6, 2026
- 24.10May 18, 2026
- 40.36Jun 15, 2026
- 43.80Jul 29, 2026
- OSWorld-Verified61.40Oct 7, 2026OSWorld
- τ-Bench V3 · Banking24.54Oct 8, 2026tauBanking
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 63.67Sep 22, 2026
- 48.33Sep 22, 2026
- 46.50Sep 22, 2026
- 31.00Sep 22, 2026
- 100.00Sep 7, 2026
- 96.39Sep 7, 2026
Show 118 more resultsHide 118 results
- 17.55Sep 9, 2026
- 50.64Jun 18, 2026
- AA Intelligence36.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence29.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence35.00Jun 28, 2026Artificial Analysis Intelligence Index
- AA Intelligence20.67Oct 8, 2026aa_intelligence_index
- AA Intelligence19.34Oct 8, 2026aa_intelligence_index
- -9.37Oct 8, 2026
- -0.08Oct 8, 2026
- AI2D87.00Aug 24, 2026AI2D_TEST
- 84.17May 2, 2026
- 87.00Oct 7, 2026
- 87.00Jun 6, 2026
- AIME 202588.00Jun 6, 2026AIME25
- AIME25 no tools88.00Aug 24, 2026AIME25
- 87.00Jun 12, 2026
- 100.00Jun 12, 2026
- 89.84May 19, 2026
- 98.21Sep 7, 2026
- 25.50May 10, 2026
- 13.60Jul 29, 2026
- 1455Jul 23, 2026
- 1456Jul 23, 2026
- 76.70Jun 6, 2026
- 63.30Jun 6, 2026
- 61.50Jun 6, 2026
- Artificial Analysis Coding Index52.10Sep 9, 2026aa_coding_index
- 33.47Jun 18, 2026
- 98.90Sep 7, 2026
- 26.10Jun 5, 2026
- 24.10Jun 12, 2026
- 67.23Jul 29, 2026
- 60.36Jul 29, 2026
- 40.80Jun 6, 2026
- 42.40May 18, 2026
- 42.40Jun 12, 2026
- 68.10Aug 24, 2026
- 60.80Jun 6, 2026
- 44.00May 15, 2026
- 44.00Jun 12, 2026
- 85.00May 15, 2026
- 85.00Jun 12, 2026
- frontiermath_tier_4_v14.17May 30, 2026frontiermath_tier_4
- 71.20Jun 6, 2026
- GPQA (unspecified)83.40Jun 12, 2026GPQA no tools
- 59.90Aug 24, 2026
- 92.00Sep 7, 2026
- 44.20May 15, 2026
- 44.20Jun 12, 2026
- HLE (with tools)24.50Jun 6, 2026HLE (w/ tools)
- HLE (with tools)32.00May 15, 2026HLE (Text-only) w/ tools
- HMMT 202574.60May 15, 2026HMMT25
- 67.50May 11, 2026
- 79.20Jun 6, 2026
- 81.70May 18, 2026
- 74.60Jun 12, 2026
- 88.80Jun 12, 2026
- 52.30Jul 29, 2026
- 63.70Jul 29, 2026
- LiveCodeBench71.00Jun 6, 2026LiveCodeBench (LCB)
- 64.00Jun 12, 2026
- Longform Writing eval (Kimi K2 Thinking system card)79.80Jun 12, 2026Longform Writing no tools
- 75.80Sep 2, 2026
- 18.50Jun 3, 2026
- 88.30Aug 24, 2026
- MMLU-Pro87.50Jun 12, 2026MMLU-Pro no tools
- 88.20Jun 6, 2026
- 88.00Jun 6, 2026
- MMLU-Redux95.60Jun 12, 2026MMLU-Redux no tools
- 89.10Oct 7, 2026
- MMMU (val) (Pass@1)77.80Aug 27, 2026MMMUval
- 55.40Jun 6, 2026
- Multi-SWE-Bench44.30Jun 12, 2026Multi-SWE-bench w/ tools
- 18.50Jun 6, 2026
- 22.80Jun 5, 2026
- 30.40May 15, 2026
- 30.40Jun 12, 2026
- 85.80Aug 24, 2026
- 70.30Aug 24, 2026
- 53.40May 15, 2026
- 53.40Jun 12, 2026
- 57.60Aug 24, 2026
- 71.40Aug 10, 2026
- 70.60May 1, 2026
- 45.30Jul 29, 2026
- 3.00Jun 5, 2026
- 69.50Jun 5, 2026
- TAU-bench (airline)70.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)86.20Oct 7, 2026TAU-bench Retail
- 50.00Sep 23, 2026
- 50.00Jun 6, 2026
- 51.00May 15, 2026
- 50.00Jul 29, 2026
- Terminal-Bench 2.042.80Jun 15, 2026Terminal Bench 2
- 50.00Jun 5, 2026
- 51.00Jun 12, 2026
- theagentcompany41.00Jun 6, 2026AgentCompany
- Toolathlon54.60Jun 30, 2026Toolathlon Pass@3
- Toolathlon41.00Jun 30, 2026Toolathlon Pass@1
- 38.90Jun 5, 2026
- 32.00Jun 30, 2026
- 28.70Jul 7, 2026
- 95.60May 30, 2026
- 127.80May 30, 2026
- 88.00May 30, 2026
- 85.20Jun 5, 2026
- 87.50Jun 5, 2026
- 90.80Jun 5, 2026
- 81.20Jun 5, 2026
- 79.10Jun 5, 2026
- 87.30Jun 5, 2026
- 66.00Jun 6, 2026
- τ²-Bench98.00Jul 29, 2026τ²-Bench (Telecom)
- 87.20Jun 13, 2026
- τ²-Bench78.00Jun 6, 2026τ²-Bench-Telecom
- 86.20Jul 29, 2026
- τ²-Bench Telecom (AA run)70.47Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)78.07Oct 8, 2026aa_tau2
Claude Sonnet 4.5: common questions
Who makes Claude Sonnet 4.5?
Claude Sonnet 4.5 is made by Anthropic.
When was Claude Sonnet 4.5 released?
Claude Sonnet 4.5 was released on Sep 29, 2025, according to Artificial Analysis.
What is Claude Sonnet 4.5 good at?
Claude Sonnet 4.5 is capable in long context; and behind the leaders in multimodal tasks, coding, instruction following, factuality, reasoning, and math. Too few results yet to rate agentic tasks, safety, or multilingual tasks.
How much does Claude Sonnet 4.5 cost?
Claude Sonnet 4.5 costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.
How many benchmarks has Claude Sonnet 4.5 been tested on?
We track 197 results for Claude Sonnet 4.5 on 127 benchmarks from 40 sources, 44 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Sonnet 4.5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Sonnet 4.5.
About this record
Where Claude Sonnet 4.5's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 27, 2026
Where the results come from
Verification: 197 scores · 44 independently verified · 35 aggregator-attributed · 92 vendor cross-reference · 26 vendor-reported. How these tiers are assigned
From 40 sources on 15 sites. Hugging Face supplies 92 of them; the 44 independently verified results come from 11 sites. Bars are coloured by trust tier.
- huggingface.co92
- artificialanalysis.ai38
- www-cdn.anthropic.com14
- api.llm-stats.com9
- arcprize.org9
- storage.googleapis.com6
- epoch.ai5
- swebench.com5
- raw.githubusercontent.com4
- anthropic.com3
- datasets-server.huggingface.co3
- labs.scale.com3
- matharena.ai3
- lmarena.ai2
- simple-bench.com1