Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 28 measured
22.686.2

28 models measured, most on the official card harness. Claude Sonnet 4.5 tops the board at 86.2.

28 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 22.6–86.2 · ◆ solid = first-party or better · outlined = arm's-length
01Claude Sonnet 4.5Anthropic86.2$3.0/M$15.0/M$0.104
02Claude Opus 4.1Anthropic82.4$15.0/M$75.0/M$0.546
03Claude Opus 4+2 altsAnthropic81.4$15.0/M$75.0/M$0.553
04Claude 3.7 SonnetAnthropic81.2$3.0/M$15.0/M$0.111
05Claude Sonnet 4Anthropic80.5$3.0/M$15.0/M$0.112
06GLM-4.5Z.ai79.7$0.60/M$2.2/M$0.018
07GLM-4.5-AirZ.ai77.9$0.17/M$0.98/M$0.007
08Qwen3 Coder 480B A35B InstructAlibaba77.5$1.5/M$7.5/M$0.058
09O3+1 altOpenAI73.93P$2.0/M$8.0/M$0.068
10O4 MiniOpenAI71.8$1.1/M$4.4/M$0.038
11O1OpenAI70.8$15.0/M$60.0/M$0.530
12Seed Oss 36B InstructByteDance70.43P$0.21/M$0.57/M$0.006
13Qwen3 Next 80B A3BAlibaba69.6$0.15/M$1.2/M$0.010
14Claude 3.5 SonnetAnthropic69.2$3.0/M$15.0/M$0.130
15GPT-4.5 PreviewOpenAI68.4———
16GPT-4.1OpenAI68$2.0/M$8.0/M$0.074
17MiniMax M1 40K+2 altsMiniMax67.8———
18GPT Oss 120bOpenAI67.8$0.15/M$0.59/M$0.005
19Gemini 2.5 Pro+1 altGoogle673P$1.3/M$10.0/M$0.084
20MiniMax M1 80K+2 altsMiniMax63.5$0.55/M$2.2/M$0.022
21Qwen3 Next 80B A3B InstructAlibaba60.9$0.15/M$1.2/M$0.011
22GPT-4oOpenAI60.3$5.0/M$15.0/M$0.166
23Qwen3 235B A22B+1 altAlibaba58.63P$0.70/M$2.8/M$0.030
24O3 MiniOpenAI57.6$1.1/M$4.4/M$0.048
25GPT-4.1 miniOpenAI55.8$0.40/M$1.6/M$0.018
26GPT Oss 20bOpenAI54.8$0.07/M$0.18/M$0.002
27Claude 3.5 HaikuAnthropic51$0.80/M$4.0/M$0.047
28GPT-4.1 nanoOpenAI22.6$0.10/M$0.40/M$0.011
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 28 models scored · 0 independently verified · 3 vendor cross-reference · 25 vendor-reported · 0 with source disagreement. How these tiers are assigned