Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 16 measured
043.33

16 models measured, most on the third-party eval harness. GPT-5.1 Codex Max tops the board at 43.33.

16 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 0–43.33 · ◆ solid = first-party or better · outlined = arm's-length
01GPT-5.1 Codex MaxOpenAI43.33$1.3/M$10.0/M$0.130
02Claude Opus 4.5Anthropic40$5.0/M$25.0/M$0.375
03GPT-5OpenAI40$1.3/M$10.0/M$0.141
04GPT-5.2OpenAI35$1.8/M$14.0/M$0.225
05Gemini 3 ProGoogle35$2.0/M$12.0/M$0.200
06GPT-5.3 CodexOpenAI25$1.8/M$14.0/M$0.315
07Claude Opus 4Anthropic20$15.0/M$75.0/M$2.250
08Claude Opus 4.1+1 altAnthropic15$15.0/M$75.0/M$3.000
09O3OpenAI15$2.0/M$8.0/M$0.333
10Claude 3.5 SonnetAnthropic5$3.0/M$15.0/M$1.800
11GPT-4+1 altOpenAI0$30.0/M$60.0/M—
12O1OpenAI0$15.0/M$60.0/M—
13Claude 3 OpusAnthropic0$15.0/M$75.0/M—
14Claude 3.7 SonnetAnthropic0$3.0/M$15.0/M—
15GPT-4oOpenAI0$5.0/M$15.0/M—
16GPT-4 TurboOpenAI0$10.0/M$30.0/M—
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 16 models scored · 16 independently verified · 0 with source disagreement. How these tiers are assigned