Claude 3.7 Sonnet
Claude 3.7 Sonnet is capable in factuality; and behind the leaders in multimodal tasks, coding, long context, agentic tasks, and instruction following. Too few results yet to rate reasoning, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Safety, Math or Multilingual.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · All prices
Evidence
81results on50benchmarks
- 36 independently verified
- 30 aggregator
- 15 vendor-reported
From 25 sources · latest Oct 8, 2026 · How verification works
Claude 3.7 Sonnet benchmark results
81 results on 50 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
22.4% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination60.03Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy28.03Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy28.20Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination47.19Oct 8, 2026omniscienceNonHallucination
34.9% behind the leader2 of 6 ranked benchmarks measured
- 75.00Oct 7, 2026
- MMMU-Pro60.06Oct 8, 2026aa_mmmu_pro
Show 4 more multimodal resultsHide 4 multimodal results
- 1168.80Sep 22, 2026
- 1151.09Aug 25, 2026
- 1195May 6, 2026
- 1176May 1, 2026
35.0% behind the leader3 of 10 ranked benchmarks measured
- 70.30Oct 7, 2026
- SciCode40.28Sep 4, 2026aa_scicode
- Terminal-Bench Hard21.21Oct 8, 2026aa_terminalbench_hard
Show 4 more coding resultsHide 4 coding results
- LiveBench · Coding74.22Aug 23, 2026livebench_coding@2025-04-07
- SciCode37.62Sep 4, 2026aa_scicode
- 52.80Sep 1, 2026
- 63.70Jun 22, 2026
36.2% behind the leader1 of 3 ranked benchmarks measured
- 51.67Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 62.33Sep 4, 2026
37.9% behind the leader1 of 7 ranked benchmarks measured
- 27.38Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 27.37Jun 15, 2026
41.6% behind the leader2 of 3 ranked benchmarks measured
- 51.58Oct 8, 2026
- IFBench48.30Oct 8, 2026aa_ifbench
Show 2 more instruction following resultsHide 2 instruction following results
- IFBench44.01Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following85.68Aug 23, 2026livebench_instruction_following@2025-04-07
0 of 6 ranked benchmarks measured
- 0.90Sep 24, 2026
- 0.40Sep 24, 2026
- 0.70Sep 24, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 0.00May 10, 2026
- 0.00Oct 8, 2026
- 0.86Oct 8, 2026
- GPQA Diamond65.56Oct 8, 2026gpqa
- GPQA Diamond77.17Oct 8, 2026gpqa
- GPQA Diamond84.80Oct 7, 2026GPQA
- Humanity's Last Exam4.16Oct 8, 2026aa_hle
- Humanity's Last Exam9.67Oct 8, 2026aa_hle
- 44.90May 10, 2026
- 46.40May 10, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 21.20Sep 24, 2026
- 11.60Sep 24, 2026
- 28.60Sep 24, 2026
- 3.65Sep 2, 2026
- 64.90Aug 29, 2026
- livebench_language61.41Aug 23, 2026livebench_language@2025-04-07
Show 37 more resultsHide 37 results
- 35.73Jun 18, 2026
- 37.00Jun 18, 2026
- AA Intelligence24.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence15.27Oct 8, 2026aa_intelligence_index
- AA Intelligence17.71Oct 8, 2026aa_intelligence_index
- -9.72Oct 8, 2026
- -0.73Oct 8, 2026
- 60.40Jun 20, 2026
- 49.17May 2, 2026
- 54.80Oct 7, 2026
- 81.77May 19, 2026
- 99.66May 10, 2026
- 13.60May 10, 2026
- Artificial Analysis Coding Index36.37Sep 9, 2026aa_coding_index
- 26.68Jun 18, 2026
- 92.10May 10, 2026
- 98.80Jun 22, 2026
- 0.89Jun 22, 2026
- 84.00Jun 22, 2026
- -0.98Jun 22, 2026
- 84.25May 10, 2026
- 31.67May 11, 2026
- 93.20Aug 31, 2026
- MATH-500 (EM)96.20Oct 7, 2026MATH-500
- 54.23May 10, 2026
- 0.00May 10, 2026
- 98.74May 10, 2026
- 86.10Oct 7, 2026
- 100.00May 10, 2026
- 52.80Jun 21, 2026
- TAU-bench (airline)58.40Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)81.20Oct 7, 2026TAU-bench Retail
- 35.20Sep 23, 2026
- 0.00May 10, 2026
- 96.44May 10, 2026
- τ²-Bench Telecom (AA run)50.00Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)54.68Oct 8, 2026aa_tau2
Claude 3.7 Sonnet: common questions
Who makes Claude 3.7 Sonnet?
Claude 3.7 Sonnet is made by Anthropic.
When was Claude 3.7 Sonnet released?
Claude 3.7 Sonnet was released on Feb 24, 2025, according to Artificial Analysis.
What is Claude 3.7 Sonnet good at?
Claude 3.7 Sonnet is capable in factuality; and behind the leaders in multimodal tasks, coding, long context, agentic tasks, and instruction following. Too few results yet to rate reasoning, safety, math, or multilingual tasks.
How much does Claude 3.7 Sonnet cost?
Claude 3.7 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.
How many benchmarks has Claude 3.7 Sonnet been tested on?
We track 81 results for Claude 3.7 Sonnet on 50 benchmarks from 25 sources, 36 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Claude 3.7 Sonnet's numbers come from, and every name it appears under.
- Tracked since
- Jun 18, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 81 scores · 36 independently verified · 30 aggregator-attributed · 15 vendor-reported. How these tiers are assigned
From 25 sources on 15 sites. Artificial Analysis supplies 31 of them; the 36 independently verified results come from 12 sites. Bars are coloured by trust tier.
- artificialanalysis.ai31
- api.llm-stats.com10
- arcprize.org8
- storage.googleapis.com6
- matharena.ai4
- www-cdn.anthropic.com4
- huggingface.co3
- raw.githubusercontent.com3
- aider.chat2
- datasets-server.huggingface.co2
- lmarena.ai2
- simple-bench.com2
- swebench.com2
- anthropic.com1
- labs.scale.com1