Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 77 measured
22.297.2

77 models measured, most on the third-party eval harness. Claude 3.5 Sonnet tops the board at 97.2.

77 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 22.2–97.2 · ◆ solid = first-party or better · outlined = arm's-length
01Claude 3.5 SonnetAnthropic97.2$3.0/M$15.0/M$0.093
02GPT-4 Turbo+1 altOpenAI97$10.0/M$30.0/M$0.206
03GPT-4o+1 altOpenAI96.8$5.0/M$15.0/M$0.103
04Qwen2.5 Instruct 72BAlibaba96.2$0.47/M$0.49/M$0.005
05GPT-4OpenAI96$30.0/M$60.0/M$0.469
06Nova ProAmazon96$0.80/M$3.2/M$0.021
07Claude 3 OpusAnthropic95.6$15.0/M$75.0/M$0.471
08Qwen2 Instruct (72B)Alibaba95.4———
09DeepSeek-V3DeepSeek95.4$0.24/M$0.90/M$0.006
10Gemini 1.5 Pro+1 altGoogle95.2———
11Jamba 1.5 LargeAI2194.8$2.0/M$8.0/M$0.053
12Yi Large (Preview)01.AI94.6———
13Gemini 2.0 Flash ExpGoogle94.6———
14Llama 3.2 Instruct 90B (Vision)Meta94.2———
15Llama 3.1 Instruct 405BMeta94———
16Palmyra X V3 (72B)Writer93.8———
17Llama 3.1 70B InstructMeta93.8$0.56/M$0.56/M$0.006
18Llama 3 (70B)Meta93.4———
19Mistral Large 2Mistral93.2$2.0/M$6.0/M$0.043
20Llama 3.3 70B InstructMeta92.8$0.71/M$0.72/M$0.008
21Nova LiteAmazon92.8$0.06/M$0.24/M$0.002
22Palmyra-X-004Writer92.6———
23Solar ProUpstage92.2———
24GPT-4o miniOpenAI92$0.15/M$0.60/M$0.004
25Claude 3 SonnetAnthropic91.8———
26Gemma 2 27B ITGoogle91.8$0.65/M$0.65/M$0.007
27Phi 3 Medium 4K InstructMicrosoft91.6———
28Gemini 1.5 Flash (001)+1 altGoogle91.4———
29Phi 3 Small 8K InstructMicrosoft91.2———
30Gemma 2 Instruct (9B)Google91———
31DBRX InstructDatabricks91———
32Phi 3.5 MoE InstructMicrosoft89.6VENDOR———
33Mistral LargeMistral89.4$4.0/M$12.0/M$0.089
34Jamba 1.5 MiniAI2189$0.20/M$0.40/M$0.003
35Nova MicroAmazon88.8$0.04/M$0.14/M$0.001
36DeepSeek-LLM-67B-Chat (V1)DeepSeek88———
37Palmyra X V2 (33B)Writer87.8———
38Mistral Small (2402)Mistral86.2———
39Qwen2.5 Instruct Turbo (7B)Alibaba86.2———
40Claude 3.5 HaikuAnthropic85.4$0.80/M$4.0/M$0.028
41Claude 3 HaikuAnthropic83.8$0.25/M$1.3/M$0.009
42Mistral Medium (Non-Reasoning)Mistral83———
43Command R+ (Apr '24)Cohere82.8$2.5/M$10.0/M$0.075
44Arctic InstructSnowflake82.8———
45Phi 3 Mini 4K InstructMicrosoft82.2———
46Mistral NemoMistral82.2$0.02/M$0.03/M$0.000
47Gemma (7B)+1 altGoogle80.8———
48Phi 2Microsoft79.8———
49Mistral 7BMistral79.8———
50Jamba InstructAI2179.6———
51Phi 4 Mini InstructMicrosoft79.2VENDOR———
52Phi 3.5 Mini InstructMicrosoft79.2VENDOR———
53Mistral Instruct V0.3 (7B)Mistral79———
54Command RCohere78.2$0.15/M$0.60/M$0.005
55Mistral V0.1 (7B)Mistral77.6———
56CommandCohere77.4———
57Llama 3 Instruct 8B+1 altMeta76.6$0.04/M$0.14/M$0.001
58Llama (65B)Meta75.4———
59Llama 3.1 8B InstructMeta74$0.02/M$0.04/M$0.000
60Llama 3.2 Instruct 11B (Vision)Meta72.4$0.34/M$0.34/M$0.005
61Jurassic 2 Jumbo (178B)AI2168.8———
62Jurassic 2 Grande (17B)AI2161.4———
63Mistral Large 3 675B Base 2512Mistral51.4———
64Kimi K2 BaseMoonshot50.8———
65GLM-4.5-BaseZ.ai49.6———
66Gemma 4 26B A4BGoogle48.8$0.07/M$0.34/M$0.004
67Nemotron 3 Ultra 550B A55B BaseNVIDIA48.6———
68Nemotron 3 Super 120B A12BNVIDIA48.6$0.30/M$0.90/M$0.012
69DeepSeek-V3.2-Exp-BaseDeepSeek48.2———
70Nemotron 3.5 LightningNVIDIA47.6$0.06/M$0.20/M$0.003
71NVIDIA Nemotron 3 Nano 30B A3BNVIDIA46.8$0.05/M$0.20/M$0.003
72Qwen3.5 35B A3BAlibaba44.2$0.25/M$2.0/M$0.025
73Command LightCohere39.8———
74Luminous Base (13B)Aleph Alpha28.6———
75Luminous Supreme (70B)Aleph Alpha28.4———
76Luminous Extended (30B)Aleph Alpha27.2———
77OLMo (7B)AllenAI22.2———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 77 models scored · 62 independently verified · 7 vendor cross-reference · 8 vendor-reported · 0 with source disagreement. How these tiers are assigned