VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude Sonnet 4-20250514 (no thinking) or GPT-4o?
Across 6 shared benchmarks, Claude Sonnet 4-20250514 (no thinking) scores higher on 4 and GPT-4o on 2. The widest gap is aider_polyglot, where Claude Sonnet 4-20250514 (no thinking) scores 56.4 against 30.7.

Claude Sonnet 4-20250514 (no thinking) vs GPT-4o

Across 6 shared benchmarks, Claude Sonnet 4-20250514 (no thinking) scores higher on 4 and GPT-4o on 2. The widest gap is aider_polyglot, where Claude Sonnet 4-20250514 (no thinking) scores 56.4 against 30.7.

AnthropicvsOpenAI6 shared benchmarks42 head-to-head
BenchmarkClaude Sonnet 4-20250514 (no thinking)GPT-4o
aider_polyglot56.430.7
arena_vision11761162
vectara_answer_rate98.693.8
vectara_avg_summary_length145.886.6
vectara_factual_consistency89.790.4
vectara_hallucination_rate10.39.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.