Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 101 measured
1.823.5

101 models measured, most on the third-party eval harness. Phi 4 Mini Instruct tops the board at 23.5.

101 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 1.8–23.5 · ◆ solid = first-party or better · outlined = arm's-length
01Phi 4 Mini Instruct+1 altMicrosoft23.5———
02O3 ProOpenAI23.3VENDOR_EVA$20.0/M$80.0/M$2.146
03Mistral Medium 3.1 (Non-Reasoning)+1 altMistral22.7$0.40/M$2.0/M$0.053
04Ministral 3 8BMistral21.7$0.15/M$0.15/M$0.007
05Grok 4 Fast+2 altsSpaceXAI20.2$0.20/M$0.50/M$0.017
06Ministral 3 14BMistral19.4$0.20/M$0.20/M$0.010
07Grok 4.1 Fast+2 altsSpaceXAI19.2$0.20/M$0.50/M$0.018
08O4 Mini+2 altsOpenAI18.6$1.1/M$4.4/M$0.148
09Kimi K2 InstructMoonshot17.9$0.57/M$2.3/M$0.080
10GPT-5+3 altsOpenAI15.1$1.3/M$10.0/M$0.373
11Jamba 1.7 MiniAI2114.7———
12Mistral Large 3Mistral14.5$0.50/M$1.5/M$0.069
13GPT Oss 120bOpenAI14.2$0.15/M$0.59/M$0.026
14Kimi K2.5Moonshot14.2$0.45/M$2.3/M$0.095
15Gemini 3 ProGoogle13.6$2.0/M$12.0/M$0.515
16Gemini 3 Flash PreviewGoogle13.5$0.50/M$3.0/M$0.130
17MiniMax M2.7MiniMax12.9$0.21/M$0.84/M$0.041
18GPT-5 miniOpenAI12.9$0.25/M$2.0/M$0.087
19GPT-5.6 SolOpenAI12.4VENDOR_EVA$4.0/M$20.0/M$0.968
20Mercury 2Inception12.3$0.25/M$0.75/M$0.041
21Claude Opus 4.6Anthropic12.2$5.0/M$25.0/M$1.230
22Qwen3.5 27BAlibaba12.1$0.30/M$2.4/M$0.112
23GPT-5.1+3 altsOpenAI12.1$1.3/M$10.0/M$0.465
24Claude Opus 4+1 altAnthropic12$15.0/M$75.0/M$3.750
25Claude Sonnet 4.5+1 altAnthropic12$3.0/M$15.0/M$0.750
26Claude Opus 4.7+1 altAnthropic12$5.0/M$25.0/M$1.250
27MiniMax M2.1MiniMax11.8$0.30/M$1.2/M$0.064
28Claude Opus 4.1Anthropic11.8VENDOR_EVA$15.0/M$75.0/M$3.814
29GLM-4.7Z.ai11.7$0.60/M$2.2/M$0.120
30DeepSeek-R1+1 altDeepSeek11.3$2.0/M$4.0/M$0.265
31Qwen3.5 122B A10B+1 altAlibaba11.2$0.40/M$3.2/M$0.161
32Aya Expanse 32BCohere10.9———
33Claude Opus 4.5+1 altAnthropic10.9$5.0/M$25.0/M$1.376
34Kimi K2.6Moonshot10.8$0.95/M$4.0/M$0.229
35GPT-5.2+2 altsOpenAI10.8$1.8/M$14.0/M$0.729
36Qwen3.5 Plus 02.15Alibaba10.7$0.26/M$1.6/M$0.085
37Claude Sonnet 4.6Anthropic10.6$3.0/M$15.0/M$0.849
38Granite 3.3 8B InstructIBM10.6$0.03/M$0.25/M$0.013
39Qwen3.5 35B A3B+1 altAlibaba10.5$0.25/M$2.0/M$0.107
40GPT-5 nanoOpenAI10.5VENDOR_EVA$0.05/M$0.40/M$0.021
41Qwen3.5 FlashAlibaba10.5$0.10/M$0.40/M$0.024
42Gemini 3.1 ProGoogle10.4$2.0/M$12.0/M$0.673
43Claude Sonnet 4+1 altAnthropic10.3$3.0/M$15.0/M$0.874
44GLM-5Z.ai10.1$1.0/M$3.2/M$0.208
45Claude Haiku 4.5Anthropic9.8$1.0/M$5.0/M$0.306
46Jamba 1.7 LargeAI219.7———
47GPT-4oOpenAI9.6$5.0/M$15.0/M$1.042
48NVIDIA Nemotron 3 Nano 30B A3B+1 altNVIDIA9.6$0.05/M$0.20/M$0.013
49GLM-4.6Z.ai9.5$0.57/M$2.2/M$0.146
50Aya Expanse 8BCohere9.5———
51Qwen3 235B A22BAlibaba9.3$0.70/M$2.8/M$0.188
52Command ACohere9.3$2.5/M$10.0/M$0.672
53Qwen3 Next 80B A3B+1 altAlibaba9.3$0.15/M$1.2/M$0.073
54GLM-4.7-FlashZ.ai9.3$0.06/M$0.40/M$0.025
55GLM-4.5-AirZ.ai9.3$0.17/M$0.98/M$0.062
56GPT-5.5OpenAI9.3$5.0/M$30.0/M$1.882
57MiniMax M2.5MiniMax9.1$0.27/M$1.1/M$0.074
58GPT-6 AstraOpenAI8.7VENDOR_EVA$10.0/M$50.0/M$3.448
59DeepSeek-V4-ProDeepSeek8.6$0.43/M$0.87/M$0.076
60GPT-5.4 ProOpenAI8.3$30.0/M$180.0/M$12.651
61Llama 4 Maverick 17B 128E Instruct FP8Meta8.2VENDOR_EVA$0.27/M$0.85/M$0.068
62Gemini 3.1 Flash Lite PreviewGoogle8.2$0.25/M$1.5/M$0.107
63Gemini 2.5 FlashGoogle7.8$0.30/M$2.5/M$0.179
64Llama 4 ScoutMeta7.7$0.19/M$0.68/M$0.056
65Gemma 4 31BGoogle7.4$0.17/M$0.40/M$0.038
66Ministral 8B 2410Mistral7.4———
67Gemma 3 27BGoogle7.4$0.08/M$0.16/M$0.016
68Ministral 3B 2410Mistral7.3———
69Gemini 2.5 ProGoogle7$1.3/M$10.0/M$0.804
70GPT-5.4OpenAI7$2.5/M$15.0/M$1.250
71Trinity Large ThinkingArcee AI6.9$0.25/M$0.80/M$0.076
72Command R+ (Apr '24)Cohere6.9$2.5/M$10.0/M$0.906
73GPT-6 SolOpenAI6.5VENDOR_EVA$2.0/M$10.0/M$0.923
74Gemma 3 4BGoogle6.4$0.05/M$0.10/M$0.012
75DeepSeek-V3.2DeepSeek6.3VENDOR_EVA$0.28/M$0.42/M$0.056
76Nova LiteAmazon6.1$0.06/M$0.24/M$0.025
77DeepSeek-V3DeepSeek6.1$0.24/M$0.90/M$0.093
78Qwen3 32BAlibaba5.9$0.16/M$0.64/M$0.068
79Grok 3+1 altSpaceXAI5.8$4.0/M$20.0/M$2.069
80Qwen3 4BAlibaba5.7———
81GPT-4.1OpenAI5.6$2.0/M$8.0/M$0.893
82Nova MicroAmazon5.5$0.04/M$0.14/M$0.016
83DeepSeek-V3.1DeepSeek5.5$0.56/M$1.7/M$0.204
84GPT-5.4 miniOpenAI5.5VENDOR_EVA$0.75/M$4.5/M$0.477
85Qwen3 14BAlibaba5.4$0.35/M$1.4/M$0.162
86DeepSeek-V3.2-ExpDeepSeek5.3$0.28/M$0.42/M$0.066
87Jamba Mini 2AI215.3———
88Granite 4.0 H Small+1 altIBM5.2$0.06/M$0.25/M$0.030
89Gemma 4 26B A4BGoogle5.2$0.07/M$0.34/M$0.039
90Nova ProAmazon5.1$0.80/M$3.2/M$0.392
91Nova 2.0 LiteAmazon5.1VENDOR_EVA$0.30/M$2.5/M$0.275
92Mistral Small 3 (Non-Reasoning)Mistral5.1$0.05/M$0.08/M$0.013
93Qwen3 8BAlibaba4.8$0.18/M$0.70/M$0.092
94Mistral Large 2.1Mistral4.5$2.0/M$6.0/M$0.889
95Gemma 3 12BGoogle4.4$0.05/M$0.15/M$0.023
96Arctic InstructSnowflake4.3———
97Llama 3.3 70B InstructMeta4.1$0.71/M$0.72/M$0.174
98Phi 4Microsoft3.7$0.13/M$0.50/M$0.084
99Gemini 2.5 Flash LiteGoogle3.3$0.10/M$0.40/M$0.076
100GPT-5.4 nanoOpenAI3.1$0.20/M$1.3/M$0.234
101Finix_s1_32bantgroup1.8———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 101 models scored · 101 independently verified · 0 with source disagreement. How these tiers are assigned