AIME 2025 — Kimi K2 Thinking 'heavy' = high-compute scaffold, distinct tools-mode from no-tools/python
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Grok 4 | SpaceXAI | 100 | 1 | 2026-05-15 |
| 2 | GPT-5 | OpenAI | 100 | 1 | 2026-05-15 |
| 3 | Kimi K2 (Reasoning) | Moonshot | 100 | 1 | 2026-05-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.