VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

narrativeqa

63 models tracked

NarrativeQA — HELM Lite reading-comprehension-over-long-narratives benchmark.

#ModelVendorBest scoreRunsLast seen
1GPT-4oOpenAI80.422026-05-10
2Llama 3 (70B)Meta79.812026-05-10
3DeepSeek-V3DeepSeek79.612026-05-10
4Llama 3.3 Instruct Turbo (70B)Meta79.112026-05-10
5Nova Pro (Non-Reasoning)Amazon79.122026-06-19
6Gemma 2 Instruct (27B)Google7912026-05-10
7Gemini 2.0 Flash (experimental)Google78.312026-06-24
8Gemini 2.0 Flash ExpGoogle78.312026-05-10
9Gemini 1.5 Flash (001)Google78.322026-05-10
10Gemini 1.5 ProGoogle78.322026-05-10
11Mistral Large 2Mistral77.912026-05-10
12Llama 3.2 Vision Instruct Turbo (90B)Meta77.712026-05-10
13Palmyra-X-004Writer77.312026-05-10
14Llama 3.1 Instruct Turbo (70B)Meta77.212026-05-10
15Amazon Nova LiteAmazon76.812026-06-23
16GPT-4o miniOpenAI76.812026-05-10
17GPT4OpenAI76.822026-05-10
18Gemma 2 Instruct (9B)Google76.812026-05-10
19Claude 3.5 HaikuAnthropic76.322026-06-19
20GPT-4 TurboOpenAI76.122026-06-18
21Llama 3.1 Instruct Turbo (8B)Meta75.612026-05-10
22Llama 3.2 Vision Instruct Turbo (11B)Meta75.612026-05-10
23Llama (65B)Meta75.512026-05-10
24Phi 3 Small 8K InstructMicrosoft75.412026-05-10
25Llama 3 Instruct 8BMeta75.412026-05-10
26Solar ProUpstage75.312026-05-10
27Palmyra X V2 (33B)Writer75.312026-05-10
28Gemma (7B)Google75.212026-05-10
29Llama 3.1 Instruct 405BMeta74.912026-05-10
30CommandCohere74.912026-05-10
31Claude 3.5 SonnetAnthropic74.612026-05-10
32Jamba 1.5 MiniAI2174.612026-05-10
33Qwen2.5 Instruct 72BAlibaba74.512026-05-10
34Jurassic 2 Grande (17B)AI2174.512026-05-10
35Nova Micro (Non-Reasoning)Amazon74.422026-06-19
36Luminous Supreme (70B)Aleph Alpha74.312026-05-10
37Qwen2.5 Instruct Turbo (7B)Alibaba74.212026-05-10
38Command RCohere74.212026-05-10
39Command R+ (Apr '24)Cohere73.512026-05-10
40Mistral NemoMistral73.112026-05-10
41Jurassic 2 Jumbo (178B)AI2172.812026-05-10
42Qwen2 Instruct (72B)Alibaba72.712026-05-10
43Phi 3 Medium 4K InstructMicrosoft72.412026-05-10
44Mistral V0.1 (7B)Mistral71.612026-05-10
45Mistral Instruct V0.3 (7B)Mistral71.612026-05-10
46Palmyra X V3 (72B)Writer70.612026-05-10
47Phi 2Microsoft70.312026-05-10
48Luminous Extended (30B)Aleph Alpha68.412026-05-10
49Jamba 1.5 LargeAI2166.412026-05-10
50Jamba InstructAI2165.812026-05-10
51Arctic InstructSnowflake65.312026-05-10
52Luminous Base (13B)Aleph Alpha63.312026-05-10
53Command LightCohere6312026-05-10
54OLMo (7B)AllenAI59.712026-05-10
55DeepSeek-ChatDeepSeek58.112026-05-10
56Mistral Small (2402)Mistral51.912026-05-10
57DBRX InstructDatabricks48.812026-05-10
58Mistral LargeMistral45.412026-05-10
59Mistral Medium (Non-Reasoning)Mistral44.912026-06-23
60Yi Large (Preview)01.AI37.312026-05-10
61Claude 3 OpusAnthropic35.112026-05-10
62Claude 3 HaikuAnthropic24.412026-05-10
63Claude 3 SonnetAnthropic11.112026-05-10

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.