Across 6 shared benchmarks, Claude Sonnet 4-20250514 (no thinking) scores higher on 1 and Gemini 2.5 Flash on 5. The widest gap is vectara_avg_summary_length, where Claude Sonnet 4-20250514 (no thinking) scores 145.8 against 101.5.
| Benchmark | Claude Sonnet 4-20250514 (no thinking) | Gemini 2.5 Flash |
|---|---|---|
| aider_polyglot | 56.4 | 61.9 |
| arena_vision | 1176 | 1236 |
| vectara_answer_rate | 98.6 | 99 |
| vectara_avg_summary_length | 145.8 | 101.5 |
| vectara_factual_consistency | 89.7 | 92.2 |
| vectara_hallucination_rate ↓ | 10.3 | 7.8 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.