PaperBench — Research-paper-replication agentic benchmark; Kimi source family.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3.8 Max Preview | Alibaba | 93 | 2 | 2026-08-23 |
| 2 | GPT-5.6 Sol | OpenAI | 90.5 | 1 | 2026-08-12 |
| 3 | Claude Fable 5 | Anthropic | 88.8 | 1 | 2026-08-12 |
| 4 | Claude Opus 4.8 | Anthropic | 80.3 | 1 | 2026-08-12 |
| 5 | Claude Opus 4.5 | Anthropic | 72.9 | 1 | 2026-06-15 |
| 6 | Qwen3.7 Max | Alibaba | 64.8 | 1 | 2026-08-12 |
| 7 | GPT-5.2 | OpenAI | 63.7 | 1 | 2026-06-15 |
| 8 | Kimi K2.5 | Moonshot | 63.5 | 2 | 2026-08-23 |
| 9 | MiniMax M3 | MiniMax | 52.6 | 1 | 2026-08-23 |
| 10 | DeepSeek-V3.2 | DeepSeek | 47.1 | 1 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.