Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Sonnet 4

Score basis

Claude Sonnet 4 is capable in long context and factuality; and behind the leaders in multimodal tasks, instruction following, coding, and agentic tasks. Too few results yet to rate reasoning, safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable2 capabilities
  1. Long Context−24.9%1 of 3
  2. Factuality−24.9%3 of 4
Limited4 capabilities
  1. Multimodal−34.9%2 of 6
  2. Instruction Following−35.2%2 of 3
  3. Coding−40.1%7 of 10
  4. Agentic−46.6%4 of 7
Not rated4 capabilities

Too few results yet to rate Reasoning, Safety, Math or Multilingual.

Price

$3.00input$15.00outputper million tokens

From anthropic · All prices

Costs more than 88% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

146results on83benchmarks

  • 45 independently verified
  • 34 aggregator
  • 14 vendor-reported
  • 53 cross-referenced

From 27 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingReasoning

As listed by OpenRouter

Claude Sonnet 4 benchmark results

146 results on 83 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

CapableFull long context ranking

24.9% behind the leader1 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Factuality

CapableFull factuality ranking

24.9% behind the leader3 of 4 ranked benchmarks measured

Show 2 more factuality resultsHide 2 factuality results

Multimodal

LimitedFull multimodal ranking

34.9% behind the leader2 of 6 ranked benchmarks measured

Show 6 more multimodal resultsHide 6 multimodal results

Instruction Following

LimitedFull instruction following ranking

35.2% behind the leader2 of 3 ranked benchmarks measured

Show 3 more instruction following resultsHide 3 instruction following results

Coding

LimitedFull coding ranking

40.1% behind the leader7 of 10 ranked benchmarks measured

Show 11 more coding resultsHide 11 coding results

Agentic

LimitedFull agentic ranking

46.6% behind the leader4 of 7 ranked benchmarks measured

Show 1 more agentic resultHide 1 agentic result

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 14 more reasoning resultsHide 14 reasoning results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 80 more resultsHide 80 results

Claude Sonnet 4: common questions

Who makes Claude Sonnet 4?

Claude Sonnet 4 is made by Anthropic.

When was Claude Sonnet 4 released?

Claude Sonnet 4 was released on May 22, 2025, according to Artificial Analysis.

What is Claude Sonnet 4 good at?

Claude Sonnet 4 is capable in long context and factuality; and behind the leaders in multimodal tasks, instruction following, coding, and agentic tasks. Too few results yet to rate reasoning, safety, math, or multilingual tasks.

How much does Claude Sonnet 4 cost?

Claude Sonnet 4 costs $3.00 per million input tokens and $15.00 per million output tokens, according to anthropic. At a mix of three input tokens to one output token, it costs more than 88% of the 331 priced models we track.

How many benchmarks has Claude Sonnet 4 been tested on?

We track 146 results for Claude Sonnet 4 on 83 benchmarks from 27 sources, 45 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Claude Sonnet 4 support?

OpenRouter lists tool calling and reasoning for Claude Sonnet 4.

About this record

Where Claude Sonnet 4's numbers come from, and every name it appears under.

Tracked since
May 3, 2026
Newest source mention
Sep 1, 2026

Where the results come from

Verification: 146 scores · 45 independently verified · 34 aggregator-attributed · 53 vendor cross-reference · 14 vendor-reported. How these tiers are assigned

From 27 sources on 15 sites. Hugging Face supplies 53 of them; the 45 independently verified results come from 12 sites. Bars are coloured by trust tier.

  • huggingface.co53
  • artificialanalysis.ai35
  • storage.googleapis.com11
  • api.llm-stats.com8
  • arcprize.org8
  • livecodebench.github.io8
  • anthropic.com6
  • raw.githubusercontent.com4
  • swebench.com3
  • aider.chat2
  • datasets-server.huggingface.co2
  • labs.scale.com2
  • lmarena.ai2
  • epoch.ai1
  • simple-bench.com1

Also known as

anthropic claude sonnet 4claude 4 sonnetclaude 4 sonnet (20250514, extended thinking)claude 4 sonnet (20250514)claude 4 sonnet (non-reasoning)claude 4 sonnet (reasoning)claude 4 sonnet (thinking)Claude Sonnet 4 (Non-Reasoning)claude sonnet 4 (thinking 16k)claude sonnet 4 (thinking 1k)claude sonnet 4 (thinking 8k)claude sonnet 4 (thinking)Claude Sonnet 4-20250514 (no thinking)claude-4-sonnet-20250514claude-4-sonnet-thinkingclaude-sonnet-4 (reasoning)and 5 more