SWE-QA — Rollup mixing 3 codebase-domain subsets (71.6-81.3 spans all three sub-ranges); domain axis unresolved
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | GPT-5.4 | OpenAI | 81.3 | 2 | 2026-06-15 |
| 2 | GLM-5.1 | Z.ai | 72.7 | 2 | 2026-06-15 |
| 3 | Kimi K2.6 | Moonshot | 71.6 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.