Across 8 models scored on DynaMath, Qwen3.6 Plus leads at 88, ahead of Qwen3.5 27B at 87.7. The median tracked score is 85.6, and the field spans 79.5 to 88.
Data as of August 25, 2026| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3.6 Plus | Alibaba | 88 | 1 | 2026-08-25 |
| 2 | Qwen3.5 27B | Alibaba | 87.7 | 2 | 2026-08-25 |
| 3 | Qwen3.5 397B A17B | Alibaba | 86.3 | 1 | 2026-08-24 |
| 4 | Qwen3.5 122B A10B | Alibaba | 85.9 | 1 | 2026-08-25 |
| 5 | Qwen3.6 27B | Alibaba | 85.6 | 2 | 2026-08-25 |
| 6 | Qwen3.5 35B A3B | Alibaba | 85 | 1 | 2026-08-25 |
| 7 | Claude Opus 4.5 | Anthropic | 79.7 | 1 | 2026-08-24 |
| 8 | Gemma 4 31B | 79.5 | 1 | 2026-08-24 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.