Across 19 models scored on t2-bench, Gemini 3.1 Pro leads at 99.3, ahead of Gemini 3 Flash Preview at 90.2. The median tracked score is 80.3, and the field spans 11.6 to 99.3.
Data as of August 25, 2026| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 99.3 | 2 | 2026-08-25 | |
| 2 | Gemini 3 Flash Preview | 90.2 | 1 | 2026-08-25 | |
| 3 | GLM-5 | Z.ai | 89.7 | 1 | 2026-08-25 |
| 4 | Qwen3.5 397B A17B | Alibaba | 86.7 | 1 | 2026-08-25 |
| 5 | Gemma 4 31B | 86.4 | 1 | 2026-08-25 | |
| 6 | Gemma 4 26B A4B | 85.5 | 1 | 2026-08-25 | |
| 7 | Gemini 3 Pro | 85.4 | 1 | 2026-08-25 | |
| 8 | Qwen3.5 35B A3B | Alibaba | 81.2 | 1 | 2026-08-25 |
| 9 | DeepSeek-V3.2-Speciale | DeepSeek | 80.3 | 1 | 2026-08-25 |
| 10 | DeepSeek-V3.2 | DeepSeek | 80.3 | 1 | 2026-08-25 |
| 11 | Qwen3.5 4B | Alibaba | 79.9 | 1 | 2026-08-25 |
| 12 | Qwen3.5 122B A10B | Alibaba | 79.5 | 1 | 2026-08-25 |
| 13 | Qwen3.5 9B | Alibaba | 79.1 | 1 | 2026-08-25 |
| 14 | Qwen3.5 27B | Alibaba | 79 | 1 | 2026-08-25 |
| 15 | Qwen3 Max (Reasoning) | Alibaba | 74.8 | 1 | 2026-08-25 |
| 16 | Gemma 4 E4B | 57.5 | 1 | 2026-08-25 | |
| 17 | Qwen3.5 2B | Alibaba | 48.8 | 1 | 2026-08-25 |
| 18 | Gemma 4 E2B | 29.4 | 1 | 2026-08-25 | |
| 19 | Qwen3.5 0.8B | Alibaba | 11.6 | 1 | 2026-08-25 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.