Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 62 measured
11.1479.82

62 models measured, most on the third-party eval harness. Llama 3 (70B) tops the board at 79.82.

62 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 11.14–79.82 · ◆ solid = first-party or better · outlined = arm's-length
01Llama 3 (70B)Meta79.82———
02DeepSeek-V3DeepSeek79.63$0.24/M$0.90/M$0.007
03GPT-4o+1 altOpenAI79.54$5.0/M$15.0/M$0.126
04Llama 3.3 70B InstructMeta79.13$0.71/M$0.72/M$0.009
05Nova ProAmazon79.13$0.80/M$3.2/M$0.025
06Gemma 2 27B ITGoogle78.97$0.65/M$0.65/M$0.008
07Gemini 2.0 Flash ExpGoogle78.3———
08Mistral Large 2Mistral77.87$2.0/M$6.0/M$0.051
09Llama 3.2 Instruct 90B (Vision)Meta77.69———
10Palmyra-X-004Writer77.26———
11Llama 3.1 70B InstructMeta77.24$0.56/M$0.56/M$0.007
12Nova LiteAmazon76.81$0.06/M$0.24/M$0.002
13GPT-4o miniOpenAI76.79$0.15/M$0.60/M$0.005
14GPT-4OpenAI76.78$30.0/M$60.0/M$0.586
15Gemma 2 Instruct (9B)Google76.77———
16Claude 3.5 HaikuAnthropic76.29$0.80/M$4.0/M$0.031
17GPT-4 Turbo+1 altOpenAI76.12$10.0/M$30.0/M$0.263
18Llama 3.1 8B InstructMeta75.59$0.02/M$0.04/M$0.000
19Gemini 1.5 Pro+1 altGoogle75.57———
20Llama 3.2 Instruct 11B (Vision)Meta75.57$0.34/M$0.34/M$0.005
21Llama (65B)Meta75.52———
22Phi 3 Small 8K InstructMicrosoft75.4———
23Llama 3 Instruct 8BMeta75.39$0.04/M$0.14/M$0.001
24Solar ProUpstage75.31———
25Palmyra X V2 (33B)Writer75.25———
26Gemma (7B)Google75.16———
27Llama 3.1 Instruct 405BMeta74.94———
28CommandCohere74.88———
29Claude 3.5 SonnetAnthropic74.63$3.0/M$15.0/M$0.121
30Jamba 1.5 MiniAI2174.62$0.20/M$0.40/M$0.004
31Gemini 1.5 Flash (001)+1 altGoogle74.62———
32Qwen2.5 Instruct 72BAlibaba74.47$0.47/M$0.49/M$0.007
33Jurassic 2 Grande (17B)AI2174.45———
34Nova MicroAmazon74.38$0.04/M$0.14/M$0.001
35Luminous Supreme (70B)Aleph Alpha74.32———
36Qwen2.5 Instruct Turbo (7B)Alibaba74.24———
37Command RCohere74.17$0.15/M$0.60/M$0.005
38Command R+ (Apr '24)Cohere73.52$2.5/M$10.0/M$0.085
39Mistral NemoMistral73.1$0.02/M$0.03/M$0.000
40Jurassic 2 Jumbo (178B)AI2172.82———
41Qwen2 Instruct (72B)Alibaba72.72———
42Phi 3 Medium 4K InstructMicrosoft72.41———
43Mistral V0.1 (7B)Mistral71.64———
44Mistral Instruct V0.3 (7B)Mistral71.63———
45Palmyra X V3 (72B)Writer70.59———
46Phi 2Microsoft70.26———
47Luminous Extended (30B)Aleph Alpha68.39———
48Jamba 1.5 LargeAI2166.36$2.0/M$8.0/M$0.075
49Jamba InstructAI2165.76———
50Arctic InstructSnowflake65.35———
51Luminous Base (13B)Aleph Alpha63.33———
52Command LightCohere62.95———
53OLMo (7B)AllenAI59.71———
54DeepSeek-LLM-67B-Chat (V1)DeepSeek58.1———
55Mistral Small (2402)Mistral51.92———
56DBRX InstructDatabricks48.84———
57Mistral LargeMistral45.35$4.0/M$12.0/M$0.176
58Mistral Medium (Non-Reasoning)Mistral44.94———
59Yi Large (Preview)01.AI37.28———
60Claude 3 OpusAnthropic35.14$15.0/M$75.0/M$1.281
61Claude 3 HaikuAnthropic24.41$0.25/M$1.3/M$0.031
62Claude 3 SonnetAnthropic11.14———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 62 models scored · 62 independently verified · 0 with source disagreement. How these tiers are assigned