Across 6 shared benchmarks, Claude 4 Opus scores higher on 2 and Claude Sonnet 4-20250514 (no thinking) on 4. The widest gap is aider_polyglot, where Claude 4 Opus scores 70.7 against 56.4.
| Benchmark | Claude 4 Opus | Claude Sonnet 4-20250514 (no thinking) |
|---|---|---|
| aider_polyglot | 70.7 | 56.4 |
| arena_vision | 1207 | 1176 |
| vectara_answer_rate | 91 | 98.6 |
| vectara_avg_summary_length | 123.2 | 145.8 |
| vectara_factual_consistency | 88 | 89.7 |
| vectara_hallucination_rate ↓ | 12 | 10.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.