PIQA — Physical Interaction QA, standard commonsense-reasoning benchmark.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Phi 3.5 MoE Instruct | Microsoft | 88.6 | 1 | 2026-08-23 |
| 2 | Hunyuan Large | Tencent | 88.3 | 1 | 2026-05-31 |
| 3 | Llama 3.1 Instruct 405B | Meta | 85.9 | 1 | 2026-05-01 |
| 4 | Kimi K2 Base | Moonshot | 85.5 | 3 | 2026-06-05 |
| 5 | DeepSeek-V3.2-Exp-Base | DeepSeek | 85.1 | 1 | 2026-06-05 |
| 6 | GLM-4.5-Base | Z.ai | 85.1 | 3 | 2026-06-05 |
| 7 | Mistral Large 3 675B Base 2512 | Mistral | 84.8 | 1 | 2026-06-05 |
| 8 | DeepSeek-V3 | DeepSeek | 84.7 | 1 | 2026-05-01 |
| 9 | Hy3 preview-Base | Tencent | 84.4 | 2 | 2026-06-05 |
| 10 | DeepSeek-V3-Base | DeepSeek | 84.2 | 2 | 2026-06-05 |
| 11 | DeepSeek-V2 | DeepSeek | 83.9 | 2 | 2026-05-31 |
| 12 | Nemotron 3 Ultra 550B A55B Base | NVIDIA | 83.8 | 1 | 2026-06-05 |
| 13 | NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16 | NVIDIA | 83.8 | 1 | 2026-08-24 |
| 14 | Gemma 2 27B IT | 83.2 | 1 | 2026-08-23 | |
| 15 | Qwen2.5 Instruct 72B | Alibaba | 82.6 | 1 | 2026-05-01 |
| 16 | Gemma 2 Instruct (9B) | 81.7 | 1 | 2026-08-23 | |
| 17 | Phi 3.5 Mini Instruct | Microsoft | 81 | 1 | 2026-08-23 |
| 18 | Phi 4 Mini Instruct | Microsoft | 77.6 | 1 | 2026-08-23 |
| 19 | ERNIE 4.5 | Baidu | 55.2 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.