NL2Repo (natural-language-to-repository code-gen benchmark) — Repository-scale code generation from NL spec; multi-vendor aggregate.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 | Anthropic | 69.7 | 4 | 2026-08-13 |
| 2 | GLM-5.3 | Z.ai | 58 | 1 | 2026-08-23 |
| 3 | DeepSeek-V4-Flash-Vision-Exp | DeepSeek | 57.7 | 1 | 2026-08-23 |
| 4 | Qwen3.8 Max Preview | Alibaba | 55.9 | 1 | 2026-08-23 |
| 5 | DeepSeek-V4-Flash | DeepSeek | 54.2 | 6 | 2026-08-23 |
| 6 | GPT-5.5 | OpenAI | 50.7 | 2 | 2026-07-13 |
| 7 | Claude Opus 4.6 | Anthropic | 49.8 | 2 | 2026-05-18 |
| 8 | GLM-5.2 Full Open Source | Z.ai | 48.9 | 6 | 2026-08-23 |
| 9 | Qwen3.7 Max | Alibaba | 47.2 | 3 | 2026-08-23 |
| 10 | Hy3-preview | Tencent | 45.6 | 1 | 2026-08-23 |
| 11 | Claude Opus 4.5 | Anthropic | 43.2 | 1 | 2026-06-15 |
| 12 | GLM-5.1 | Z.ai | 42.7 | 5 | 2026-08-23 |
| 13 | Qwen3.8 27B | Alibaba | 42.3 | 1 | 2026-08-23 |
| 14 | MiniMax M3 | MiniMax | 42.1 | 3 | 2026-08-23 |
| 15 | GPT-5.4 | OpenAI | 41.3 | 2 | 2026-05-18 |
| 16 | Qwen3.7 Plus Preview | Alibaba | 41.1 | 1 | 2026-08-23 |
| 17 | MiniMax M2.7 | MiniMax | 39.8 | 3 | 2026-08-23 |
| 18 | DeepSeek-V4-Pro | DeepSeek | 38.5 | 4 | 2026-08-13 |
| 19 | Qwen3.6 Plus | Alibaba | 37.9 | 3 | 2026-08-23 |
| 20 | Qwen3.6 27B | Alibaba | 36.2 | 2 | 2026-08-23 |
| 21 | GLM-5 | Z.ai | 35.9 | 2 | 2026-05-18 |
| 22 | Gemini 3.1 Pro | 33.4 | 4 | 2026-07-13 | |
| 23 | Qwen3.5 397B A17B | Alibaba | 32.2 | 1 | 2026-06-15 |
| 24 | Kimi K2.5 | Moonshot | 32 | 2 | 2026-05-18 |
| 25 | Qwen3.6 35B A3B | Alibaba | 29.4 | 3 | 2026-08-23 |
| 26 | Qwen3.5 27B | Alibaba | 27.3 | 2 | 2026-06-15 |
| 27 | Qwen3.5 35B A3B | Alibaba | 20.5 | 1 | 2026-06-06 |
| 28 | Gemma 4 31B | 15.5 | 2 | 2026-06-15 | |
| 29 | Gemma 4 26B A4B | 11.6 | 1 | 2026-06-06 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.