DROP — Discrete Reasoning Over Paragraphs QA benchmark, unlabeled/default metric
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | DeepSeek-V3 | DeepSeek | 91.6 | 1 | 2026-08-23 |
| 2 | Llama 3.3 70B Instruct | Meta | 90.2 | 1 | 2026-05-30 |
| 3 | Hunyuan Large | Tencent | 88.9 | 1 | 2026-05-31 |
| 4 | DeepSeek-V4-Pro-Base | DeepSeek | 88.7 | 3 | 2026-06-05 |
| 5 | DeepSeek-V4-Flash-Base | DeepSeek | 88.6 | 3 | 2026-06-05 |
| 6 | Claude 3.5 Sonnet | Anthropic | 87.1 | 1 | 2026-08-23 |
| 7 | DeepSeek-V3.2-Exp-Base | DeepSeek | 86.6 | 2 | 2026-06-13 |
| 8 | DeepSeek-V3-Base | DeepSeek | 86.5 | 2 | 2026-06-05 |
| 9 | Kimi K2 Base | Moonshot | 86.4 | 7 | 2026-06-13 |
| 10 | MiMo V2.5 Pro Base | Xiaomi | 86.3 | 3 | 2026-06-05 |
| 11 | DeepSeek-V3.1-Base | DeepSeek | 86.3 | 2 | 2026-06-13 |
| 12 | MiMo V2.5 Pro | Xiaomi | 86.3 | 1 | 2026-08-23 |
| 13 | GPT-4 Turbo | OpenAI | 86 | 1 | 2026-08-23 |
| 14 | Hy3 preview-Base | Tencent | 85.5 | 2 | 2026-06-05 |
| 15 | Qwen 2.5 14B | Alibaba | 85.5 | 1 | 2026-05-30 |
| 16 | Nova Pro (Non-Reasoning) | Amazon | 85.4 | 1 | 2026-08-23 |
| 17 | Llama 3.1 Instruct 405B | Meta | 84.8 | 2 | 2026-08-23 |
| 18 | MiMo V2 Flash Base | Xiaomi | 84.7 | 2 | 2026-06-13 |
| 19 | MiMo V2.5 Base | Xiaomi | 83.7 | 3 | 2026-06-05 |
| 20 | GPT-4o | OpenAI | 83.4 | 2 | 2026-08-23 |
| 21 | Claude 3.5 Haiku | Anthropic | 83.1 | 1 | 2026-08-23 |
| 22 | Claude 3 Opus | Anthropic | 83.1 | 1 | 2026-08-23 |
| 23 | GLM-4.5-Base | Z.ai | 82.9 | 2 | 2026-06-05 |
| 24 | GPT4 | OpenAI | 80.9 | 1 | 2026-08-23 |
| 25 | Nova Lite (Non-Reasoning) | Amazon | 80.2 | 1 | 2026-08-23 |
| 26 | DeepSeek-V2 | DeepSeek | 80.1 | 1 | 2026-05-31 |
| 27 | GPT-4o mini | OpenAI | 79.7 | 2 | 2026-08-23 |
| 28 | Llama 3.1 70B Instruct | Meta | 79.6 | 2 | 2026-08-23 |
| 29 | Nova Micro (Non-Reasoning) | Amazon | 79.3 | 1 | 2026-08-23 |
| 30 | Claude 3 Sonnet | Anthropic | 78.9 | 1 | 2026-08-23 |
| 31 | Granite 4.1 30B Base | IBM | 78.6 | 2 | 2026-06-15 |
| 32 | Claude 3 Haiku | Anthropic | 78.4 | 1 | 2026-08-23 |
| 33 | Qwen2.5 Instruct 72B | Alibaba | 76.7 | 1 | 2026-05-30 |
| 34 | Phi 4 | Microsoft | 75.5 | 2 | 2026-08-23 |
| 35 | Gemini 1.5 Pro | 74.9 | 1 | 2026-08-23 | |
| 36 | Granite 4.1 8B Base | IBM | 72.4 | 2 | 2026-06-15 |
| 37 | Llama 3.1 8B Instruct | Meta | 71.2 | 2 | 2026-08-23 |
| 38 | Phi 3 Medium 4K Instruct | Microsoft | 68.3 | 1 | 2026-05-30 |
| 39 | Granite 4.1 3B Base | IBM | 66 | 1 | 2026-06-15 |
| 40 | Granite 3.3 8B Instruct | IBM | 59.4 | 2 | 2026-08-23 |
| 41 | Granite 3.1 8B Instruct | IBM | 58.6 | 1 | 2026-06-11 |
| 42 | Granite 3.2 8B Instruct | IBM | 58.3 | 1 | 2026-06-11 |
| 43 | DeepSeek-R1-Distill-Qwen-7B | DeepSeek | 51.8 | 1 | 2026-06-11 |
| 44 | DeepSeek-R1-Distill-Llama-8B | DeepSeek | 49.7 | 1 | 2026-06-11 |
| 45 | Granite 3.3 2B Instruct | IBM | 44.3 | 1 | 2026-06-11 |
| 46 | ERNIE 4.5 | Baidu | 28.6 | 1 | 2026-08-23 |
| 47 | Granite 3.2 2B Instruct | IBM | 23.8 | 1 | 2026-06-11 |
| 48 | Granite 3.1 2B Instruct | IBM | 21 | 1 | 2026-06-11 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.