Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 51 measured
1.667100

51 models measured, most on the third-party eval harness. GPT-5.2 tops the board at 100.

51 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 1.667–100 · ◆ solid = first-party or better · outlined = arm's-length
01GPT-5.2+1 altOpenAI100$1.8/M$14.0/M$0.079
02Step 3.5 Flash+2 altsStepFun98.4VENDOR$0.10/M$0.30/M$0.002
03Gemini 3 Flash PreviewGoogle97.5$0.50/M$3.0/M$0.018
04Gemini 3 Pro+4 altsGoogle97.5$2.0/M$12.0/M$0.072
05DeepSeek-V3.2-SpecialeDeepSeek97.5$0.29/M$0.43/M$0.004
06GLM-4.7+3 altsZ.ai97.1$0.60/M$2.2/M$0.014
07GPT Oss 120bOpenAI96.667$0.15/M$0.59/M$0.004
08GPT-5.1OpenAI96.3$1.3/M$10.0/M$0.058
09Kimi K2.5+4 altsMoonshot95.4$0.45/M$2.3/M$0.014
10Qwen3.5 397B A17BAlibaba94.8$0.60/M$3.6/M$0.022
11Qwen3.6 27BAlibaba93.8$0.60/M$3.6/M$0.022
12Kimi K2 Thinking+5 altsMoonshot93.333$0.60/M$2.5/M$0.017
13Claude Opus 4.5+1 altAnthropic92.9$5.0/M$25.0/M$0.161
14DeepSeek-V3.2+6 altsDeepSeek92.5$0.28/M$0.42/M$0.004
15Qwen3.5 27B+1 altAlibaba92$0.30/M$2.4/M$0.015
16Gemma 4 26B A4BGoogle91.7$0.07/M$0.34/M$0.002
17Qwen3.6 35B A3B+1 altAlibaba90.7$0.38/M$2.3/M$0.014
18DeepSeek-V3.2-ExpDeepSeek90$0.28/M$0.42/M$0.004
19GLM-4.6Z.ai89.2$0.57/M$2.2/M$0.016
20Qwen3.5 35B A3BAlibaba89$0.25/M$2.0/M$0.013
21Gemma 4 31B+1 altGoogle88.7$0.17/M$0.40/M$0.003
22GPT-5+2 altsOpenAI88.333$1.3/M$10.0/M$0.064
23DeepSeek-V3.1DeepSeek85.833$0.56/M$1.7/M$0.013
24MiMo V2 Flash+3 altsXiaomi84.4$0.10/M$0.30/M$0.002
25Gemini 2.5 Pro+1 altGoogle80.833$1.3/M$10.0/M$0.070
26O3OpenAI77.5$2.0/M$8.0/M$0.065
27GPT Oss 20bOpenAI76.667$0.07/M$0.18/M$0.002
28QED-NanoLM-Provers76.667———
29MiniMax M2.1+2 altsMiniMax71VENDOR$0.30/M$1.2/M$0.011
30Claude Sonnet 4.5+2 altsAnthropic67.5$3.0/M$15.0/M$0.133
31K2-ThinkMBZUAI65———
32Gemini 2.5 Flash+2 altsGoogle64.167$0.30/M$2.5/M$0.022
33Qwen3 235B A22B+1 altAlibaba62.5$0.70/M$2.8/M$0.028
34DeepSeek-R1 0528-Qwen3-8BDeepSeek61.5VENDOR———
35Claude Opus 4Anthropic60$15.0/M$75.0/M$0.750
36Qwen3 30B A3BAlibaba50.833$0.20/M$0.80/M$0.010
37Phi 4 Reasoning Plus+1 altMicrosoft46.667———
38DeepSeek-R1DeepSeek41.667$2.0/M$4.0/M$0.072
39Gemini 2.0 Flash (Non-Reasoning)Google35.833———
40DeepSeek-R1-Distill-Llama-70BDeepSeek33.333$0.70/M$1.1/M$0.027
41DeepSeek-R1-Distill-Qwen-32BDeepSeek33.333$0.29/M$0.29/M$0.009
42DeepSeek-R1-Distill-Qwen-14BDeepSeek31.667$0.20/M$0.20/M$0.006
43Claude 3.7 SonnetAnthropic31.667$3.0/M$15.0/M$0.284
44DeepSeek-V3+1 altDeepSeek29.167$0.24/M$0.90/M$0.020
45O3 Mini+1 altOpenAI28.333$1.1/M$4.4/M$0.097
46QwQ 32B PreviewAlibaba18.333$0.66/M$1.0/M$0.045
47Gemini 2.0 Flash (Reasoning)Google13.333———
48DeepSeek-R1-Distill-Qwen-1.5BDeepSeek11.667———
49Llama 4 MaverickMeta8.333$0.25/M$0.87/M$0.067
50Gemini 2.0 Pro Exp 02.05Google7.5———
51Claude 3.5 SonnetAnthropic1.667$3.0/M$15.0/M$5.400
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 51 models scored · 35 independently verified · 5 vendor cross-reference · 11 vendor-reported · 0 with source disagreement. How these tiers are assigned