PostTrainBench — GLM-5.2 card eval; distinct from Program Bench (different score range/focus).
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 41.4 | 1 | 2026-07-27 |
| 2 | GLM-5.3 | Z.ai | 39.8 | 1 | 2026-08-23 |
| 3 | Claude Opus 4.8 | Anthropic | 37.2 | 3 | 2026-07-27 |
| 4 | MiniMax M3 | MiniMax | 37.1 | 1 | 2026-08-23 |
| 5 | Kimi K3 | Moonshot | 36.6 | 2 | 2026-08-23 |
| 6 | GPT-5.6 Sol | OpenAI | 34.6 | 1 | 2026-07-27 |
| 7 | GLM-5.2 Full Open Source | Z.ai | 34.3 | 6 | 2026-08-23 |
| 8 | GPT-5.5 | OpenAI | 28.4 | 3 | 2026-07-27 |
| 9 | Gemini 3.1 Pro | 21.6 | 2 | 2026-07-13 | |
| 10 | GLM-5.1 | Z.ai | 20.1 | 2 | 2026-07-13 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.