TruthfulQA — Distinct named benchmark, IBM Granite 3.3 card
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | MAI-Thinking-1 | Microsoft | 88 | 1 | 2026-08-23 |
| 2 | Qwen3.5 2B | Alibaba | 88 | 1 | 2026-07-05 |
| 3 | Phi 3.5 MoE Instruct | Microsoft | 77.5 | 1 | 2026-08-23 |
| 4 | Granite 3.2 8B Instruct | IBM | 66.9 | 1 | 2026-06-11 |
| 5 | Granite 3.3 8B Instruct | IBM | 66.9 | 2 | 2026-08-23 |
| 6 | Phi 4 Mini Instruct | Microsoft | 66.4 | 1 | 2026-08-23 |
| 7 | Granite 3.1 8B Instruct | IBM | 65.8 | 1 | 2026-06-11 |
| 8 | Phi 3.5 Mini Instruct | Microsoft | 64 | 1 | 2026-08-23 |
| 9 | Granite 3.2 2B Instruct | IBM | 59.8 | 1 | 2026-06-11 |
| 10 | Granite 3.1 2B Instruct | IBM | 59.8 | 1 | 2026-06-11 |
| 11 | Granite 3.3 2B Instruct | IBM | 59 | 1 | 2026-06-11 |
| 12 | Llama 3.1 Nemotron Instruct 70B | NVIDIA | 58.6 | 1 | 2026-08-23 |
| 13 | Jamba 1.5 Large | AI21 | 58.3 | 1 | 2026-08-23 |
| 14 | Command R+ (Apr '24) | Cohere | 56.3 | 1 | 2026-08-23 |
| 15 | Qwen2 Instruct (72B) | Alibaba | 54.8 | 1 | 2026-08-23 |
| 16 | Qwen 2.5 Coder 32B Instruct | Alibaba | 54.2 | 1 | 2026-08-23 |
| 17 | Jamba 1.5 Mini | AI21 | 54.1 | 1 | 2026-08-23 |
| 18 | Llama 3.1 8B Instruct | Meta | 52.8 | 1 | 2026-06-11 |
| 19 | DeepSeek-R1-Distill-Llama-8B | DeepSeek | 47.4 | 1 | 2026-06-11 |
| 20 | DeepSeek-R1-Distill-Qwen-7B | DeepSeek | 47.1 | 1 | 2026-06-11 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.