Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 62 measured
2.8350.16

62 models measured, most on the third-party eval harness. Claude 3.5 Sonnet tops the board at 50.16.

62 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 2.83–50.16 · ◆ solid = first-party or better · outlined = arm's-length
01Claude 3.5 SonnetAnthropic50.16$3.0/M$15.0/M$0.179
02GPT-4o+1 altOpenAI49.58$5.0/M$15.0/M$0.202
03GPT-4 Turbo+1 altOpenAI48.21$10.0/M$30.0/M$0.415
04Llama 3 (70B)Meta47.53———
05DeepSeek-V3DeepSeek46.7$0.24/M$0.90/M$0.012
06GPT-4OpenAI45.69$30.0/M$60.0/M$0.985
07Llama 3.2 Instruct 90B (Vision)Meta45.68———
08Palmyra-X-004Writer45.66———
09Llama 3.1 Instruct 405BMeta45.6———
10Gemini 1.5 Pro+1 altGoogle45.55———
11Mistral Large 2Mistral45.27$2.0/M$6.0/M$0.088
12Llama 3.1 70B InstructMeta45.17$0.56/M$0.56/M$0.012
13Gemini 2.0 Flash ExpGoogle44.31———
14Claude 3 OpusAnthropic44.05$15.0/M$75.0/M$1.022
15Llama (65B)Meta43.26———
16Llama 3.3 70B InstructMeta43.07$0.71/M$0.72/M$0.017
17Palmyra X V2 (33B)Writer42.81———
18Yi Large (Preview)01.AI42.77———
19DeepSeek-LLM-67B-Chat (V1)DeepSeek41.21———
20Palmyra X V3 (72B)Writer40.72———
21Nova ProAmazon40.52$0.80/M$3.2/M$0.049
22Jamba 1.5 LargeAI2139.37$2.0/M$8.0/M$0.127
23CommandCohere39.11———
24Arctic InstructSnowflake38.98———
25Qwen2 Instruct (72B)Alibaba38.97———
26Jamba 1.5 MiniAI2138.79$0.20/M$0.40/M$0.008
27GPT-4o miniOpenAI38.55$0.15/M$0.60/M$0.010
28Jurassic 2 Jumbo (178B)AI2138.53———
29Jamba InstructAI2138.36———
30Llama 3 Instruct 8BMeta37.8$0.04/M$0.14/M$0.003
31Mistral V0.1 (7B)Mistral36.7———
32Qwen2.5 Instruct 72BAlibaba35.94$0.47/M$0.49/M$0.013
33Gemma 2 27B ITGoogle35.29$0.65/M$0.65/M$0.018
34Nova LiteAmazon35.24$0.06/M$0.24/M$0.004
35Command RCohere35.22$0.15/M$0.60/M$0.011
36Jurassic 2 Grande (17B)AI2135———
37Claude 3.5 HaikuAnthropic34.39$0.80/M$4.0/M$0.070
38Command R+ (Apr '24)Cohere34.32$2.5/M$10.0/M$0.182
39Gemma (7B)Google33.56———
40Gemma 2 Instruct (9B)Google32.78———
41Phi 3 Small 8K InstructMicrosoft32.37———
42Gemini 1.5 Flash (001)+1 altGoogle32.28———
43Mistral LargeMistral31.11$4.0/M$12.0/M$0.257
44Mistral Small (2402)Mistral30.42———
45Luminous Supreme (70B)Aleph Alpha29.92———
46Solar ProUpstage29.73———
47Mistral Medium (Non-Reasoning)Mistral28.97———
48Nova MicroAmazon28.47$0.04/M$0.14/M$0.003
49DBRX InstructDatabricks28.39———
50Phi 3 Medium 4K InstructMicrosoft27.84———
51Mistral NemoMistral26.48$0.02/M$0.03/M$0.001
52OLMo (7B)AllenAI25.86———
53Mistral Instruct V0.3 (7B)Mistral25.33———
54Luminous Extended (30B)Aleph Alpha25.29———
55Llama 3.2 Instruct 11B (Vision)Meta23.42$0.34/M$0.34/M$0.015
56Llama 3.1 8B InstructMeta20.9$0.02/M$0.04/M$0.001
57Qwen2.5 Instruct Turbo (7B)Alibaba20.51———
58Luminous Base (13B)Aleph Alpha19.72———
59Command LightCohere19.51———
60Phi 2Microsoft15.52———
61Claude 3 HaikuAnthropic14.37$0.25/M$1.3/M$0.052
62Claude 3 SonnetAnthropic2.83———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 62 models scored · 62 independently verified · 0 with source disagreement. How these tiers are assigned