AIME 2025 — Explicit no-tools mode, distinct axis from 'w/ python' and 'heavy' siblings
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | GPT-5.2 | OpenAI | 98 | 1 | 2026-08-24 |
| 2 | Gemini 3 Pro | 96 | 1 | 2026-08-24 | |
| 3 | Claude Opus 4.6 | Anthropic | 95.6 | 1 | 2026-08-24 |
| 4 | GPT-5 | OpenAI | 94.6 | 1 | 2026-06-12 |
| 5 | Kimi K2 (Reasoning) | Moonshot | 94.5 | 2 | 2026-06-12 |
| 6 | Nemotron 3 Super 120B A12B | NVIDIA | 92.2 | 1 | 2026-07-07 |
| 7 | Grok 4 | SpaceXAI | 91.7 | 1 | 2026-06-12 |
| 8 | Claude Opus 4.5 | Anthropic | 91 | 1 | 2026-08-24 |
| 9 | NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 | NVIDIA | 89.7 | 1 | 2026-07-07 |
| 10 | DeepSeek-V3.2 | DeepSeek | 89.3 | 1 | 2026-06-12 |
| 11 | Claude Sonnet 4.5 | Anthropic | 88 | 2 | 2026-08-24 |
| 12 | MiniMax M2.5 | MiniMax | 86.3 | 1 | 2026-08-24 |
| 13 | MiniMax M2.1 | MiniMax | 83 | 1 | 2026-08-24 |
| 14 | Qwen3 30B A3B | Alibaba | 71.7 | 1 | 2026-08-09 |
| 15 | Qwen3.5 9B | Alibaba | 56.1 | 2 | 2026-08-10 |
| 16 | Qwen3.5 4B | Alibaba | 54.3 | 3 | 2026-08-10 |
| 17 | LFM2.5-2.6B | Liquid AI | 51.9 | 2 | 2026-08-10 |
| 18 | LFM2.5-8B-A1B | Liquid AI | 42.5 | 1 | 2026-08-09 |
| 19 | Qwen3 1.7B | Alibaba | 36.3 | 1 | 2026-08-09 |
| 20 | Gemma 4 E4B | 34.3 | 3 | 2026-08-10 | |
| 21 | lfm-2.5-1.2b-thinking:free | Liquid AI | 31.7 | 1 | 2026-08-09 |
| 22 | Gemma 4 E2B | 26.3 | 3 | 2026-08-10 | |
| 23 | Qwen3 1.7B (instruct) | Alibaba | 9.3 | 2 | 2026-08-09 |
| 24 | Granite 4.0 H Tiny | IBM | 4.9 | 1 | 2026-08-09 |
| 25 | Granite 4.0 1B | IBM | 3.3 | 2 | 2026-08-09 |
| 26 | Granite 4.0 H 1B | IBM | 1 | 1 | 2026-08-09 |
| 27 | Gemma 3 1B IT | 1 | 2 | 2026-08-09 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.