Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 101 measured
54.7420.2

101 models measured, most on the third-party eval harness. Phi 4 Mini Instruct tops the board at 420.2.

101 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 54.7–420.2 · ◆ solid = first-party or better · outlined = arm's-length
01Phi 4 Mini Instruct+1 altMicrosoft420.2———
02GPT-5.1+3 altsOpenAI254.4$1.3/M$10.0/M$0.022
03Ministral 8B 2410Mistral196VENDOR_EVA———
04GPT-5.2+2 altsOpenAI186.3$1.8/M$14.0/M$0.042
05Grok 4 Fast+2 altsSpaceXAI173.9$0.20/M$0.50/M$0.002
06Finix_s1_32bantgroup172.4———
07GPT-5 miniOpenAI169.7$0.25/M$2.0/M$0.007
08Ministral 3B 2410Mistral167.9———
09GPT-5+3 altsOpenAI162.7$1.3/M$10.0/M$0.035
10DeepSeek-V4-ProDeepSeek153.8$0.43/M$0.87/M$0.004
11Mercury 2Inception149.1$0.25/M$0.75/M$0.003
12Claude Opus 4.7+1 altAnthropic149.1$5.0/M$25.0/M$0.101
13GPT-6 AstraOpenAI148.5VENDOR_EVA$10.0/M$50.0/M$0.202
14GPT-5.4 ProOpenAI148.5VENDOR_EVA$30.0/M$180.0/M$0.707
15Claude Sonnet 4+1 altAnthropic145.8$3.0/M$15.0/M$0.062
16GPT-5.4 nanoOpenAI144.4$0.20/M$1.3/M$0.005
17Mistral Medium 3.1 (Non-Reasoning)+1 altMistral142.9$0.40/M$2.0/M$0.008
18Ministral 3 8BMistral139.4$0.15/M$0.15/M$0.001
19Claude Opus 4.6Anthropic137.6$5.0/M$25.0/M$0.109
20Llama 4 ScoutMeta137.3$0.19/M$0.68/M$0.003
21MiniMax M2.5MiniMax137.2$0.27/M$1.1/M$0.005
22Jamba 1.7 MiniAI21136.4———
23Ministral 3 14BMistral135.8$0.20/M$0.20/M$0.001
24GPT Oss 120bOpenAI135.2VENDOR_EVA$0.15/M$0.59/M$0.003
25MiniMax M2.7MiniMax131.9$0.21/M$0.84/M$0.004
26Granite 3.3 8B InstructIBM131.4$0.03/M$0.25/M$0.001
27GPT-5.6 SolOpenAI131.1VENDOR_EVA$4.0/M$20.0/M$0.092
28GPT-5.5OpenAI129.6$5.0/M$30.0/M$0.135
29Claude Opus 4.1Anthropic129.1$15.0/M$75.0/M$0.349
30Claude Sonnet 4.5+1 altAnthropic127.8$3.0/M$15.0/M$0.070
31O4 Mini+2 altsOpenAI127.7$1.1/M$4.4/M$0.022
32O3 Pro+1 altOpenAI127.4$20.0/M$80.0/M$0.392
33Jamba 1.7 LargeAI21124.8———
34Claude Opus 4+1 altAnthropic123.2$15.0/M$75.0/M$0.365
35Phi 4Microsoft120.9$0.13/M$0.50/M$0.003
36Trinity Large ThinkingArcee AI117.3$0.25/M$0.80/M$0.004
37Kimi K2.6Moonshot116.7$0.95/M$4.0/M$0.021
38Qwen3 32BAlibaba115.8$0.16/M$0.64/M$0.003
39Claude Haiku 4.5Anthropic115.1$1.0/M$5.0/M$0.026
40Claude Sonnet 4.6Anthropic114.7$3.0/M$15.0/M$0.078
41Claude Opus 4.5+1 altAnthropic114.5$5.0/M$25.0/M$0.131
42Aya Expanse 32BCohere112.7———
43Mistral Large 3Mistral112.7$0.50/M$1.5/M$0.009
44Kimi K2.5Moonshot112$0.45/M$2.3/M$0.012
45Qwen3 14BAlibaba111.1$0.35/M$1.4/M$0.008
46Jamba Mini 2AI21109.4———
47Gemini 3.1 ProGoogle107.7$2.0/M$12.0/M$0.065
48Granite 4.0 H Small+1 altIBM107.4$0.06/M$0.25/M$0.001
49MiniMax M2.1MiniMax106.9$0.30/M$1.2/M$0.007
50Gemini 2.5 ProGoogle106.4$1.3/M$10.0/M$0.053
51Llama 4 Maverick 17B 128E Instruct FP8Meta106$0.27/M$0.85/M$0.005
52GPT-5 nanoOpenAI105.7$0.05/M$0.40/M$0.002
53Qwen3 235B A22BAlibaba105.6$0.70/M$2.8/M$0.017
54Qwen3 4BAlibaba104.7———
55NVIDIA Nemotron 3 Nano 30B A3B+1 altNVIDIA104.2$0.05/M$0.20/M$0.001
56Gemini 3 ProGoogle101.9$2.0/M$12.0/M$0.069
57Command ACohere101.7$2.5/M$10.0/M$0.061
58Gemini 2.5 FlashGoogle101.5$0.30/M$2.5/M$0.014
59Nova MicroAmazon100$0.04/M$0.14/M$0.001
60Grok 4.1 Fast+2 altsSpaceXAI99.5$0.20/M$0.50/M$0.004
61Mistral Small 3 (Non-Reasoning)Mistral98.8VENDOR_EVA$0.05/M$0.08/M$0.001
62Gemma 3 27BGoogle96.4$0.08/M$0.16/M$0.001
63Grok 3+1 altSpaceXAI95.9$4.0/M$20.0/M$0.125
64Gemini 2.5 Flash LiteGoogle95.7$0.10/M$0.40/M$0.003
65Qwen3.5 FlashAlibaba95$0.10/M$0.40/M$0.003
66Qwen3.5 35B A3B+1 altAlibaba94.9$0.25/M$2.0/M$0.012
67Qwen3.5 27BAlibaba94.4$0.30/M$2.4/M$0.014
68Nova 2.0 LiteAmazon94.1$0.30/M$2.5/M$0.015
69DeepSeek-R1+1 altDeepSeek93.5$2.0/M$4.0/M$0.032
70Qwen3.5 Plus 02.15Alibaba92.1$0.26/M$1.6/M$0.010
71Nova LiteAmazon91.8$0.06/M$0.24/M$0.002
72GPT-4.1OpenAI91.7$2.0/M$8.0/M$0.055
73Command R+ (Apr '24)Cohere91.5$2.5/M$10.0/M$0.068
74Gemini 3 Flash PreviewGoogle90.2$0.50/M$3.0/M$0.019
75Gemma 3 12BGoogle89.7$0.05/M$0.15/M$0.001
76Aya Expanse 8BCohere88.2———
77GPT-4oOpenAI86.6$5.0/M$15.0/M$0.115
78Qwen3.5 122B A10B+1 altAlibaba86.4$0.40/M$3.2/M$0.021
79Mistral Large 2.1Mistral85$2.0/M$6.0/M$0.047
80Qwen3 8BAlibaba83.6$0.18/M$0.70/M$0.005
81DeepSeek-V3DeepSeek81.7$0.24/M$0.90/M$0.007
82GPT-5.4OpenAI81.7$2.5/M$15.0/M$0.107
83Arctic InstructSnowflake81.4———
84Gemma 3 4BGoogle77.4$0.05/M$0.10/M$0.001
85GLM-4.6Z.ai77.2$0.57/M$2.2/M$0.018
86Gemma 4 31BGoogle75.8$0.17/M$0.40/M$0.004
87GLM-5Z.ai74.4$1.0/M$3.2/M$0.028
88GLM-4.7-FlashZ.ai71.8$0.06/M$0.40/M$0.003
89GPT-6 SolOpenAI71.4VENDOR_EVA$2.0/M$10.0/M$0.084
90Qwen3 Next 80B A3B+1 altAlibaba70.9$0.15/M$1.2/M$0.010
91GLM-4.5-AirZ.ai70.6$0.17/M$0.98/M$0.008
92GLM-4.7Z.ai70.6$0.60/M$2.2/M$0.020
93Gemma 4 26B A4BGoogle67.1$0.07/M$0.34/M$0.003
94Nova ProAmazon66.2$0.80/M$3.2/M$0.030
95DeepSeek-V3.2-ExpDeepSeek64.6$0.28/M$0.42/M$0.005
96Llama 3.3 70B InstructMeta64.6$0.71/M$0.72/M$0.011
97DeepSeek-V3.1DeepSeek63.7$0.56/M$1.7/M$0.018
98Gemini 3.1 Flash Lite PreviewGoogle62.6$0.25/M$1.5/M$0.014
99DeepSeek-V3.2DeepSeek62$0.28/M$0.42/M$0.006
100Kimi K2 InstructMoonshot59.2$0.57/M$2.3/M$0.024
101GPT-5.4 miniOpenAI54.7$0.75/M$4.5/M$0.048
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 101 models scored · 101 independently verified · 0 with source disagreement. How these tiers are assigned