Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 39 measured
2.9298.2

39 models measured, most on the third-party eval harness. Claude Opus 4.5 tops the board at 98.2.

39 measured·5 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 2.92–98.2 · ◆ solid = first-party or better · outlined = arm's-length
01Claude Opus 4.5+2 altsAnthropic98.2VENDOR$5.0/M$25.0/M$0.153
02Gemini 3 Pro+3 altsGoogle98VENDOR$2.0/M$12.0/M$0.071
03Claude Sonnet 4.5+3 altsAnthropic98VENDOR$3.0/M$15.0/M$0.092
04GLM-5Z.ai89.7$1.0/M$3.2/M$0.023
05Step 3.5 Flash+2 altsStepFun88.2VENDOR$0.10/M$0.30/M$0.002
06GLM-4.7+4 altsZ.ai87.4$0.60/M$2.2/M$0.016
07MiniMax M2.1+3 altsMiniMax87$0.30/M$1.2/M$0.009
08MiniMax M2+1 altMiniMax87$0.30/M$1.2/M$0.009
09GPT-5.2+1 altOpenAI85.5$1.8/M$14.0/M$0.092
10Kimi K2.5+3 altsMoonshot85.4VENDOR$0.45/M$2.3/M$0.016
11GPT-5.1OpenAI82.7$1.3/M$10.0/M$0.068
12GPT-5+1 altOpenAI82.4$1.3/M$10.0/M$0.068
13MiMo V2 Flash+2 altsXiaomi80.3VENDOR$0.10/M$0.30/M$0.002
14DeepSeek-V3.2+6 altsDeepSeek80.3VENDOR$0.28/M$0.42/M$0.004
15GLM-4.7-FlashZ.ai79.5$0.06/M$0.40/M$0.003
16Gemma 4 31BNEW+26 altsGoogle76.9VENDOR$0.17/M$0.40/M$0.004
17GLM-4.6+1 altZ.ai75.2$0.57/M$2.2/M$0.018
18Kimi K2 Thinking+3 altsMoonshot74.3VENDOR$0.60/M$2.5/M$0.021
19Claude Opus 4.1Anthropic71.5VENDOR$15.0/M$75.0/M$0.629
20Gemma 4 12B+22 altsGoogle69VENDOR$0.10/M$0.30/M$0.003
21Gemma 4 26B A4BNEW+30 altsGoogle68.2VENDOR$0.07/M$0.34/M$0.003
22Kimi K2 Instruct+2 altsMoonshot65.8$0.57/M$2.3/M$0.022
23Claude Opus 4+1 altAnthropic57$15.0/M$75.0/M$0.789
24Gemini 2.5 ProGoogle54$1.3/M$10.0/M$0.104
25Qwen3 30B A3B 2507 Thinking+1 altAlibaba49$0.20/M$2.4/M$0.027
26GPT Oss 20bOpenAI47.7$0.07/M$0.18/M$0.003
27Claude Sonnet 4+2 altsAnthropic45.2$3.0/M$15.0/M$0.199
28Gemma 4 E4BNEW+30 altsGoogle42.2VENDOR$0.02/M$0.10/M$0.001
29DeepSeek-V3+1 altDeepSeek32.5$0.24/M$0.90/M$0.018
30Gemma 4 E2BNEW+35 altsGoogle24.5VENDOR———
31Qwen3 235B A22B+1 altAlibaba22.1$0.70/M$2.8/M$0.079
32LFM2.5-350M+1 altLiquid AI18.86———
33Gemma 3 27BNEW+27 altsGoogle16.2VENDOR$0.08/M$0.16/M$0.007
34Qwen3.5 0.8BAlibaba14.33———
35Qwen3.5 0.8B (Instruct)+1 altAlibaba12.57———
36LFM2-350M+1 altLiquid AI10.82———
37Gemma 3 1B Instruct+1 altGoogle9.36———
38LFM2.5-230MLiquid AI5.26———
39Granite 4.0 350M+1 altIBM2.92———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 39 models scored · 0 independently verified · 20 vendor cross-reference · 19 vendor-reported · 0 with source disagreement. How these tiers are assigned