Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 9 measured
2.0824.4

9 models measured, most on the third-party eval harness. Gemini 2.5 Pro tops the board at 24.4.

9 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 2.08–24.4 · ◆ solid = first-party or better · outlined = arm's-length
01Gemini 2.5 Pro+1 altGoogle24.4$1.3/M$10.0/M$0.231
02O4 Mini+1 altOpenAI19.05$1.1/M$4.4/M$0.144
03Grok 3+1 altSpaceXAI4.76$4.0/M$20.0/M$2.521
04DeepSeek-R1+1 altDeepSeek4.76$2.0/M$4.0/M$0.630
05Gemini 2.0 Flash (Reasoning)+1 altGoogle4.17———
06Claude 3.7 Sonnet+1 altAnthropic3.65$3.0/M$15.0/M$2.466
07QwQ 32B Preview+1 altAlibaba2.98$0.66/M$1.0/M$0.279
08O1 Pro+1 altOpenAI2.83$150.0/M$600.0/M$132.509
09O3 Mini+1 altOpenAI2.08$1.1/M$4.4/M$1.322
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 9 models scored · 9 independently verified · 0 with source disagreement. How these tiers are assigned