NarrativeQA — HELM Lite reading-comprehension-over-long-narratives benchmark.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | GPT-4o | OpenAI | 80.4 | 2 | 2026-05-10 |
| 2 | Llama 3 (70B) | Meta | 79.8 | 1 | 2026-05-10 |
| 3 | DeepSeek-V3 | DeepSeek | 79.6 | 1 | 2026-05-10 |
| 4 | Llama 3.3 Instruct Turbo (70B) | Meta | 79.1 | 1 | 2026-05-10 |
| 5 | Nova Pro (Non-Reasoning) | Amazon | 79.1 | 2 | 2026-06-19 |
| 6 | Gemma 2 Instruct (27B) | 79 | 1 | 2026-05-10 | |
| 7 | Gemini 2.0 Flash (experimental) | 78.3 | 1 | 2026-06-24 | |
| 8 | Gemini 2.0 Flash Exp | 78.3 | 1 | 2026-05-10 | |
| 9 | Gemini 1.5 Flash (001) | 78.3 | 2 | 2026-05-10 | |
| 10 | Gemini 1.5 Pro | 78.3 | 2 | 2026-05-10 | |
| 11 | Mistral Large 2 | Mistral | 77.9 | 1 | 2026-05-10 |
| 12 | Llama 3.2 Vision Instruct Turbo (90B) | Meta | 77.7 | 1 | 2026-05-10 |
| 13 | Palmyra-X-004 | Writer | 77.3 | 1 | 2026-05-10 |
| 14 | Llama 3.1 Instruct Turbo (70B) | Meta | 77.2 | 1 | 2026-05-10 |
| 15 | Amazon Nova Lite | Amazon | 76.8 | 1 | 2026-06-23 |
| 16 | GPT-4o mini | OpenAI | 76.8 | 1 | 2026-05-10 |
| 17 | GPT4 | OpenAI | 76.8 | 2 | 2026-05-10 |
| 18 | Gemma 2 Instruct (9B) | 76.8 | 1 | 2026-05-10 | |
| 19 | Claude 3.5 Haiku | Anthropic | 76.3 | 2 | 2026-06-19 |
| 20 | GPT-4 Turbo | OpenAI | 76.1 | 2 | 2026-06-18 |
| 21 | Llama 3.1 Instruct Turbo (8B) | Meta | 75.6 | 1 | 2026-05-10 |
| 22 | Llama 3.2 Vision Instruct Turbo (11B) | Meta | 75.6 | 1 | 2026-05-10 |
| 23 | Llama (65B) | Meta | 75.5 | 1 | 2026-05-10 |
| 24 | Phi 3 Small 8K Instruct | Microsoft | 75.4 | 1 | 2026-05-10 |
| 25 | Llama 3 Instruct 8B | Meta | 75.4 | 1 | 2026-05-10 |
| 26 | Solar Pro | Upstage | 75.3 | 1 | 2026-05-10 |
| 27 | Palmyra X V2 (33B) | Writer | 75.3 | 1 | 2026-05-10 |
| 28 | Gemma (7B) | 75.2 | 1 | 2026-05-10 | |
| 29 | Llama 3.1 Instruct 405B | Meta | 74.9 | 1 | 2026-05-10 |
| 30 | Command | Cohere | 74.9 | 1 | 2026-05-10 |
| 31 | Claude 3.5 Sonnet | Anthropic | 74.6 | 1 | 2026-05-10 |
| 32 | Jamba 1.5 Mini | AI21 | 74.6 | 1 | 2026-05-10 |
| 33 | Qwen2.5 Instruct 72B | Alibaba | 74.5 | 1 | 2026-05-10 |
| 34 | Jurassic 2 Grande (17B) | AI21 | 74.5 | 1 | 2026-05-10 |
| 35 | Nova Micro (Non-Reasoning) | Amazon | 74.4 | 2 | 2026-06-19 |
| 36 | Luminous Supreme (70B) | Aleph Alpha | 74.3 | 1 | 2026-05-10 |
| 37 | Qwen2.5 Instruct Turbo (7B) | Alibaba | 74.2 | 1 | 2026-05-10 |
| 38 | Command R | Cohere | 74.2 | 1 | 2026-05-10 |
| 39 | Command R+ (Apr '24) | Cohere | 73.5 | 1 | 2026-05-10 |
| 40 | Mistral Nemo | Mistral | 73.1 | 1 | 2026-05-10 |
| 41 | Jurassic 2 Jumbo (178B) | AI21 | 72.8 | 1 | 2026-05-10 |
| 42 | Qwen2 Instruct (72B) | Alibaba | 72.7 | 1 | 2026-05-10 |
| 43 | Phi 3 Medium 4K Instruct | Microsoft | 72.4 | 1 | 2026-05-10 |
| 44 | Mistral V0.1 (7B) | Mistral | 71.6 | 1 | 2026-05-10 |
| 45 | Mistral Instruct V0.3 (7B) | Mistral | 71.6 | 1 | 2026-05-10 |
| 46 | Palmyra X V3 (72B) | Writer | 70.6 | 1 | 2026-05-10 |
| 47 | Phi 2 | Microsoft | 70.3 | 1 | 2026-05-10 |
| 48 | Luminous Extended (30B) | Aleph Alpha | 68.4 | 1 | 2026-05-10 |
| 49 | Jamba 1.5 Large | AI21 | 66.4 | 1 | 2026-05-10 |
| 50 | Jamba Instruct | AI21 | 65.8 | 1 | 2026-05-10 |
| 51 | Arctic Instruct | Snowflake | 65.3 | 1 | 2026-05-10 |
| 52 | Luminous Base (13B) | Aleph Alpha | 63.3 | 1 | 2026-05-10 |
| 53 | Command Light | Cohere | 63 | 1 | 2026-05-10 |
| 54 | OLMo (7B) | AllenAI | 59.7 | 1 | 2026-05-10 |
| 55 | DeepSeek-Chat | DeepSeek | 58.1 | 1 | 2026-05-10 |
| 56 | Mistral Small (2402) | Mistral | 51.9 | 1 | 2026-05-10 |
| 57 | DBRX Instruct | Databricks | 48.8 | 1 | 2026-05-10 |
| 58 | Mistral Large | Mistral | 45.4 | 1 | 2026-05-10 |
| 59 | Mistral Medium (Non-Reasoning) | Mistral | 44.9 | 1 | 2026-06-23 |
| 60 | Yi Large (Preview) | 01.AI | 37.3 | 1 | 2026-05-10 |
| 61 | Claude 3 Opus | Anthropic | 35.1 | 1 | 2026-05-10 |
| 62 | Claude 3 Haiku | Anthropic | 24.4 | 1 | 2026-05-10 |
| 63 | Claude 3 Sonnet | Anthropic | 11.1 | 1 | 2026-05-10 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.