Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 13 measured
37.797.3

13 models measured, most on the third-party eval harness. Qwen3 VL 235B A22B Reasoning tops the board at 97.3.

13 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 37.7–97.3 · ◆ solid = first-party or better · outlined = arm's-length
01Qwen3 VL 235B A22B ReasoningAlibaba97.3VENDOR$0.40/M$4.0/M$0.023
02O3+1 altOpenAI95.8$2.0/M$8.0/M$0.052
03Qwen3 235B A22B Instruct 2507Alibaba95VENDOR$0.23/M$0.92/M$0.006
04Gemini 2.5 Pro+1 altGoogle91.6$1.3/M$10.0/M$0.061
05Kimi K2 Instruct+3 altsMoonshot89VENDOR$0.57/M$2.3/M$0.016
06Qwen3 14BAlibaba88.5VENDOR$0.35/M$1.4/M$0.010
07MiniMax M1 80K+2 altsMiniMax86.8VENDOR$0.55/M$2.2/M$0.016
08DeepSeek-V3+1 altDeepSeek84$0.24/M$0.90/M$0.007
09MiniMax M1 40K+2 altsMiniMax80.1VENDOR———
10DeepSeek-R1+1 altDeepSeek78.7$2.0/M$4.0/M$0.038
11Claude Sonnet 4+1 altAnthropic73.7$3.0/M$15.0/M$0.122
12Claude Opus 4+3 altsAnthropic59.3$15.0/M$75.0/M$0.759
13Qwen3 235B A22B+3 altsAlibaba37.7$0.70/M$2.8/M$0.046
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 13 models scored · 0 independently verified · 7 vendor cross-reference · 6 vendor-reported · 0 with source disagreement. How these tiers are assigned