Claude 3.5 Sonnet
Claude 3.5 Sonnet is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · All prices
Evidence
184results on133benchmarks
- 29 independently verified
- 10 aggregator
- 30 vendor-reported
- 115 cross-referenced
From 34 sources · latest Oct 8, 2026 · How verification works
Claude 3.5 Sonnet benchmark results
184 results on 133 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
50.4% behind the leader4 of 6 ranked benchmarks measured
- 790.00Jun 12, 2026
- 68.30Oct 7, 2026
- 67.70Oct 7, 2026
- 54.70Jun 12, 2026
Show 11 more multimodal resultsHide 11 multimodal results
0 of 6 ranked benchmarks measured
- 27.50May 10, 2026
- 41.40May 10, 2026
- GPQA Diamond55.96Oct 8, 2026gpqa
Show 5 more reasoning resultsHide 5 reasoning results
- GPQA Diamond59.90Oct 8, 2026gpqa
- GPQA Diamond67.20Oct 7, 2026GPQA
- GPQA Diamond65.00Jun 6, 2026GPQA (diamond)
- Humanity's Last Exam3.23Oct 8, 2026aa_hle
- Humanity's Last Exam3.69Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- LiveBench · Coding60.16Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding60.16Jun 17, 2026livebench_coding@2025-04-07
- SciCode31.60Sep 4, 2026aa_scicode
Show 4 more coding resultsHide 4 coding results
- LiveCodeBench v637.20Jun 5, 2026LiveCodeBench v6 (Pass@1)
- SciCode36.57Sep 4, 2026aa_scicode
- 49.00Oct 7, 2026
- SWE-bench Verified50.80Jun 4, 2026SWE Verified (Resolved)
0 of 3 ranked benchmarks measured
- 41.00May 3, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following71.12Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following70.63Jun 17, 2026livebench_instruction_following@2025-04-07
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language52.18Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence10.00Jul 3, 2026Artificial Analysis Intelligence Index
- 97.20Jul 2, 2026
- 94.90Jul 2, 2026
- livebench_language52.18Jun 17, 2026livebench_language@2025-04-07
- frontiermath_tier_4_v10.00May 20, 2026frontiermath_tier_4
Show 145 more resultsHide 145 results
- AA Intelligence7.21Oct 8, 2026aa_intelligence_index
- AA Intelligence7.88Oct 8, 2026aa_intelligence_index
- 94.70Oct 7, 2026
- AI2D68.90Aug 24, 2026AI2D (test)
- 82.00Jun 12, 2026
- 82.00Jun 5, 2026
- 84.20May 3, 2026
- Aider-Polyglot45.30Aug 31, 2026Aider-Polyglot (Acc.)
- AIME 202416.00Jun 5, 2026AIME 2024 (Pass@1)
- AIME 202426.70Jun 4, 2026AIME 2024 cons@64
- 3.33May 2, 2026
- AIME 20257.40Jun 5, 2026AIME 2025 (Pass@1)
- 85.90May 19, 2026
- AlpacaEval 2 LC52.40Sep 10, 2026
- 52.00Jun 4, 2026
- 99.78May 10, 2026
- 79.20Sep 10, 2026
- 87.60Jun 6, 2026
- Arena-Hard (GPT-4-1106 judge)85.20Jun 4, 2026ArenaHard (GPT-4-1106)
- Artificial Analysis Coding Index26.04Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index30.16Sep 9, 2026aa_coding_index
- BBH93.10Oct 7, 2026BIG-Bench Hard
- 94.90May 10, 2026
- 93.60Jun 22, 2026
- 0.87Jun 22, 2026
- 76.20Jun 22, 2026
- -3.70Jun 22, 2026
- BIG-Bench Hard 3-shot CoT93.10Sep 7, 2026
- C-Eval76.70Aug 31, 2026C-Eval (EM)
- 90.80Oct 7, 2026
- 90.80Jun 5, 2026
- 90.80Jun 12, 2026
- Chinese SimpleQA (C-SimpleQA)55.40Sep 11, 2026C-SimpleQA (Correct)
- Chinese SimpleQA (C-SimpleQA)56.80Jun 6, 2026C-SimpleQA
- 51.30May 3, 2026
- CLUEWSC85.40Aug 31, 2026CLUEWSC (EM)
- CNMO 202413.10Jul 3, 2026math_cnmo_2024_pass1
- 20.30May 3, 2026
- 717.00May 1, 2026
- Codeforces (Percentile)20.30Jul 3, 2026code_codeforces_percentile
- Codeforces (Rating)717.00Jul 3, 2026code_codeforces_rating
- 95.20Oct 7, 2026
- 94.20Jun 12, 2026
- 94.20Jun 5, 2026
- DocVQA (test, ANLS score)95.20Sep 7, 2026
- 95.20Aug 24, 2026
- 87.10Oct 7, 2026
- 88.30Jun 5, 2026
- 88.80Jun 6, 2026
- 72.50Jun 4, 2026
- frontiermath_tier_4_v10.00May 20, 2026frontiermath_tier_4
- GPQA (Diamond) 0-shot CoT59.40Sep 7, 2026
- GPQA (Diamond) Maj@32 5-shot CoT67.20Sep 7, 2026
- 96.40Oct 7, 2026
- 96.90Jun 6, 2026
- 49.90Aug 24, 2026
- 98.06May 10, 2026
- 1.67May 11, 2026
- 93.70Oct 7, 2026
- 93.70Jun 6, 2026
- HumanEval 0-shot92.00Sep 7, 2026
- 81.70May 3, 2026
- 90.10Jun 12, 2026
- IFEval86.50Jun 5, 2026IF-Eval (Prompt Strict)
- 90.10Jun 6, 2026
- 45.60Aug 24, 2026
- LiveCodeBench38.90Jun 4, 2026LiveCodeBench pass@1
- LiveCodeBench32.80May 3, 2026LiveCodeBench (Pass@1)
- 33.80Jun 4, 2026
- 36.30May 3, 2026
- LiveCodeBench (v5)38.90Jun 5, 2026LiveCodeBench v5 (Pass@1)
- 55.20Jun 6, 2026
- 46.90Jun 6, 2026
- 41.50Jun 6, 2026
- 37.30Jun 6, 2026
- 44.40Jun 6, 2026
- 37.00Jun 12, 2026
- 41.90Jun 6, 2026
- 38.60Jun 6, 2026
- 46.70Jun 6, 2026
- 41.00Jun 6, 2026
- 53.90Jun 6, 2026
- 46.10Jun 6, 2026
- M-LongDoc (multimodal long-document benchmark)31.40Jun 12, 2026M-LongDoc_acc
- 31.40Jun 5, 2026
- 81.28May 10, 2026
- 71.10Sep 7, 2026
- 74.10Jun 6, 2026
- 78.30May 1, 2026
- MATH-500 (EM)78.30Sep 11, 2026MATH-500 (Pass@1)
- 78.30May 16, 2026
- 75.10Jun 6, 2026
- 51.40Jun 12, 2026
- MEGA-Bench_macro51.40Jun 5, 2026MEGA-Bench
- 40.67May 10, 2026
- 5.00May 10, 2026
- 98.74May 10, 2026
- 91.60Oct 7, 2026
- MGSM 0-shot CoT91.60Sep 7, 2026
- 82.30Aug 24, 2026
- 80.70Aug 24, 2026
- 79.70Aug 24, 2026
- 78.50Aug 24, 2026
- 1920.00Aug 24, 2026
- 79.94May 10, 2026
- MMLU88.70Sep 7, 2026MMLU 5-shot
- 88.30Jun 6, 2026
- 88.30Sep 7, 2026
- 90.40Sep 7, 2026
- 77.60Oct 7, 2026
- MMLU-Pro78.00Aug 31, 2026MMLU-Pro (EM)
- MMLU-Redux88.90Aug 31, 2026MMLU-Redux (EM)
- MMMU (val) (Pass@1)68.30Aug 24, 2026MMMU_val
- MMMU (val) (Pass@1)52.67Aug 24, 2026MMMU (val)
- 68.30Sep 7, 2026
- MT-Bench8.81Sep 10, 2026MT-Bench (GPT-4-Turbo)
- 55.65Jun 6, 2026
- 53.62Jun 6, 2026
- 20.22Jun 6, 2026
- 62.30Jun 6, 2026
- 59.70Jun 6, 2026
- 31.42Jun 6, 2026
- 74.63May 10, 2026
- 50.16May 10, 2026
- OlympiadBench28.40Jun 12, 2026OlympiadBench_full
- 28.40Jun 5, 2026
- 76.60Aug 24, 2026
- 60.10Aug 24, 2026
- 96.50Jun 6, 2026
- 96.00Jun 6, 2026
- 95.70Jun 6, 2026
- 95.00Jun 6, 2026
- 95.20Jun 6, 2026
- 93.80Jun 6, 2026
- 73.80Aug 24, 2026
- 100.00May 10, 2026
- 28.10Jun 6, 2026
- SimpleQA28.40Jun 4, 2026SimpleQA (Correct)
- SuperGPQA48.20Jun 5, 2026SuperGPQA (Pass@1)
- TAU-bench (airline)46.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)69.20Oct 7, 2026TAU-bench Retail
- TextVQA-val70.50Aug 24, 2026TextVQA (val)
- 63.85Aug 24, 2026
- 55.90Aug 24, 2026
- 95.61May 10, 2026
Claude 3.5 Sonnet: common questions
Who makes Claude 3.5 Sonnet?
Claude 3.5 Sonnet is made by Anthropic.
When was Claude 3.5 Sonnet released?
Claude 3.5 Sonnet was released on Jun 21, 2024, according to Artificial Analysis.
What is Claude 3.5 Sonnet good at?
Claude 3.5 Sonnet is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How much does Claude 3.5 Sonnet cost?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.
How many benchmarks has Claude 3.5 Sonnet been tested on?
We track 184 results for Claude 3.5 Sonnet on 133 benchmarks from 34 sources, 29 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Claude 3.5 Sonnet's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 184 scores · 29 independently verified · 10 aggregator-attributed · 115 vendor cross-reference · 30 vendor-reported. How these tiers are assigned
From 34 sources on 11 sites. Hugging Face supplies 102 of them; the 29 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co102
- api.llm-stats.com15
- raw.githubusercontent.com15
- www-cdn.anthropic.com15
- storage.googleapis.com12
- artificialanalysis.ai11
- arxiv.org7
- epoch.ai2
- matharena.ai2
- simple-bench.com2
- lmarena.ai1