Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 31 measured
31.292.4

31 models measured, most on the third-party eval harness. GPT-5.2 tops the board at 92.4.

31 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 31.2–92.4 · ◆ solid = first-party or better · outlined = arm's-length
01GPT-5.2OpenAI92.4VENDOR$1.8/M$14.0/M$0.085
02Kimi K2.6Moonshot91$0.95/M$4.0/M$0.027
03Claude Sonnet 4.6Anthropic89.9VENDOR$3.0/M$15.0/M$0.100
04DeepSeek-V4-FlashDeepSeek88.5$0.44/M$1.3/M$0.010
05DeepSeek-V4-ProDeepSeek87.8$0.43/M$0.87/M$0.007
06Qwen3.5 397B A17BAlibaba87.1$0.60/M$3.6/M$0.024
07MiniMax M2.7MiniMax86.6$0.21/M$0.84/M$0.006
08GLM-5.1Z.ai86.1$1.4/M$4.4/M$0.034
09Qwen3.6 35B A3BAlibaba86VENDOR$0.38/M$2.3/M$0.015
10GPT-5OpenAI85.7VENDOR$1.3/M$10.0/M$0.066
11Gemma 4 31BGoogle84.3$0.17/M$0.40/M$0.003
12Qwen3.5 35B A3BAlibaba84.2$0.25/M$2.0/M$0.013
13Claude Sonnet 4.5+1 altAnthropic83.4VENDOR$3.0/M$15.0/M$0.108
14Gemma 4 26B A4BGoogle82.3$0.07/M$0.34/M$0.002
15DeepSeek-V3.2DeepSeek79.9VENDOR$0.28/M$0.42/M$0.004
16GLM-4.7-FlashZ.ai75.2$0.06/M$0.40/M$0.003
17Kimi K2 InstructMoonshot74.2VENDOR$0.57/M$2.3/M$0.019
18Qwen3 30B A3B 2507 ThinkingAlibaba73.4$0.20/M$2.4/M$0.018
19GPT Oss 20bOpenAI71.5$0.07/M$0.18/M$0.002
20Nemotron 3 Ultra 550B A55B BaseNVIDIA50———
21Granite 4.1 30BIBM45.76———
22MiMo V2 Flash Base+1 altXiaomi43.5———
23Kimi K2 Base+2 altsMoonshot43.1———
24DeepSeek-V3.1-Base+1 altDeepSeek43.1———
25Qwen 2.5 14BAlibaba42.9———
26Step3.5 Flash Base+1 altStepFun41.7———
27DeepSeek-V3.2-Exp-Base+2 altsDeepSeek37.3———
28Mistral Large 3 675B Base 2512Mistral34.85———
29GLM-4.5-Base+2 altsZ.ai33.5———
30Granite 4.1 3BIBM31.7———
31Phi 3 Medium 4K InstructMicrosoft31.2———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 31 models scored · 0 independently verified · 22 vendor cross-reference · 9 vendor-reported · 0 with source disagreement. How these tiers are assigned