AIME 2025 — Explicit python-tool mode, distinct axis from no-tools/heavy siblings
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.5 | Anthropic | 100 | 2 | 2026-06-12 |
| 2 | GPT-5 | OpenAI | 99.6 | 1 | 2026-05-15 |
| 3 | Kimi K2 (Reasoning) | Moonshot | 99.1 | 2 | 2026-05-15 |
| 4 | Grok 4 | SpaceXAI | 98.8 | 1 | 2026-05-15 |
| 5 | DeepSeek-V3.2 | DeepSeek | 58.1 | 1 | 2026-05-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.