Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 14 measured
48.693.1

14 models measured, most on the official card harness. DeepSeek-R1-Zero tops the board at 93.1.

14 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 48.6–93.1 · ◆ solid = first-party or better · outlined = arm's-length
01DeepSeek-R1-ZeroDeepSeek93.1———
02DeepSeek-R1+5 altsDeepSeek92.8$2.0/M$4.0/M$0.032
03DeepSeek-V3+7 altsDeepSeek90.9$0.24/M$0.90/M$0.006
04DeepSeek-V2.5DeepSeek90.4———
05O1 Mini+4 altsOpenAI89.9———
06GPT-4o+5 altsOpenAI87.9$5.0/M$15.0/M$0.114
07Claude 3.5 Sonnet+5 altsAnthropic85.4$3.0/M$15.0/M$0.105
08DeepSeek-V4-Pro-Base+3 altsDeepSeek85.2———
09DeepSeek-V3.2-Base+2 altsDeepSeek83.5———
10Llama 3.1 Instruct 405B+1 altMeta83———
11Qwen2.5 Instruct 72B+1 altAlibaba82.5$0.47/M$0.49/M$0.006
12DeepSeek-V4-Flash-Base+2 altsDeepSeek82.2———
13DeepSeek-V2+1 altDeepSeek82———
14ERNIE 4.5Baidu48.6———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 14 models scored · 0 independently verified · 5 vendor cross-reference · 9 vendor-reported · 2 with source disagreement. How these tiers are assigned