VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude Sonnet 4-20250514 (no thinking) or o4-mini?
Across 6 shared benchmarks, Claude Sonnet 4-20250514 (no thinking) scores higher on 3 and o4-mini on 3. The widest gap is vectara_hallucination_rate, where Claude Sonnet 4-20250514 (no thinking) scores 10.3 against 18.6.

Claude Sonnet 4-20250514 (no thinking) vs o4-mini

Across 6 shared benchmarks, Claude Sonnet 4-20250514 (no thinking) scores higher on 3 and o4-mini on 3. The widest gap is vectara_hallucination_rate, where Claude Sonnet 4-20250514 (no thinking) scores 10.3 against 18.6.

AnthropicvsOpenAI6 shared benchmarks33 head-to-head
BenchmarkClaude Sonnet 4-20250514 (no thinking)o4-mini
aider_polyglot56.472
arena_vision11761201
vectara_answer_rate98.699.2
vectara_avg_summary_length145.8130.9
vectara_factual_consistency89.781.4
vectara_hallucination_rate10.318.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.