OJBench (Online Judge Bench) — Unscoped/language-unspecified reading (Kimi K2); much lower band than cpp/python.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2.6 | Moonshot | 60.6 | 1 | 2026-08-23 |
| 2 | Qwen3.5 27B | Alibaba | 40.1 | 1 | 2026-08-23 |
| 3 | Qwen3.5 122B A10B | Alibaba | 39.5 | 1 | 2026-08-23 |
| 4 | Qwen3.5 35B A3B | Alibaba | 36 | 1 | 2026-08-23 |
| 5 | Qwen3 Next 80B A3B | Alibaba | 29.7 | 1 | 2026-08-23 |
| 6 | Kimi K2 Instruct | Moonshot | 27.1 | 4 | 2026-08-23 |
| 7 | DeepSeek-V3 | DeepSeek | 24 | 2 | 2026-06-15 |
| 8 | Claude 4 Opus | Anthropic | 19.6 | 2 | 2026-06-15 |
| 9 | GPT-4.1 | OpenAI | 19.5 | 2 | 2026-06-15 |
| 10 | Claude Sonnet 4 | Anthropic | 15.3 | 2 | 2026-06-15 |
| 11 | Qwen3 235B A22B | Alibaba | 11.3 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.