LiveCodeBench v6 — Explicit no-tools config (Kimi Thinking); distinct from general v6 pooled score (mixed configs).
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | GPT-5 | OpenAI | 87 | 1 | 2026-06-12 |
| 2 | Kimi K2 (Reasoning) | Moonshot | 83.1 | 2 | 2026-06-12 |
| 3 | DeepSeek-V3.2 | DeepSeek | 74.1 | 1 | 2026-06-12 |
| 4 | Qwen3.5 9B | Alibaba | 69.9 | 2 | 2026-08-10 |
| 5 | Claude Sonnet 4.5 | Anthropic | 64 | 1 | 2026-06-12 |
| 6 | Gemma 4 E4B | 63.8 | 2 | 2026-08-10 | |
| 7 | Qwen3.5 4B | Alibaba | 60.9 | 2 | 2026-08-10 |
| 8 | LFM2.5-2.6B | Liquid AI | 59.4 | 2 | 2026-08-10 |
| 9 | Gemma 4 E2B | 54.9 | 2 | 2026-08-10 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.