VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude 4 Opus or Claude Sonnet 4-20250514 (no thinking)?
Across 6 shared benchmarks, Claude 4 Opus scores higher on 2 and Claude Sonnet 4-20250514 (no thinking) on 4. The widest gap is aider_polyglot, where Claude 4 Opus scores 70.7 against 56.4.

Claude 4 Opus vs Claude Sonnet 4-20250514 (no thinking)

Across 6 shared benchmarks, Claude 4 Opus scores higher on 2 and Claude Sonnet 4-20250514 (no thinking) on 4. The widest gap is aider_polyglot, where Claude 4 Opus scores 70.7 against 56.4.

AnthropicvsAnthropic6 shared benchmarks24 head-to-head
BenchmarkClaude 4 OpusClaude Sonnet 4-20250514 (no thinking)
aider_polyglot70.756.4
arena_vision12071176
vectara_answer_rate9198.6
vectara_avg_summary_length123.2145.8
vectara_factual_consistency8889.7
vectara_hallucination_rate1210.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.