Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 34 measured
35.277.8

34 models measured, most on the official card harness. Qwen3.8 Max Preview tops the board at 77.8.

34 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 35.2–77.8 · ◆ solid = first-party or better · outlined = arm's-length
01Qwen3.8 Max PreviewAlibaba77.8$3.3/M$9.9/M$0.085
02Qwen3.8 Flash Next+1 altAlibaba72.3$0.15/M$0.47/M$0.004
03Seed 2.1 Pro Preview+1 altByteDance72———
04Seed2.1ByteDance71.3———
05Gemini 3.1 ProGoogle70.8$2.0/M$12.0/M$0.099
06Qwen3.7 Plus Preview+1 altAlibaba69.8$0.40/M$1.6/M$0.014
07Qwen3.6 PlusAlibaba65.7$0.50/M$3.0/M$0.027
08GPT-5OpenAI65.7$1.3/M$10.0/M$0.086
09Qwen3.8 27B+1 altAlibaba65.5$0.50/M$3.0/M$0.027
10Qwen3.5 35B A3BAlibaba64.8$0.25/M$2.0/M$0.017
11Muse SparkMeta64.7———
12GPT-5.5OpenAI64.5$5.0/M$30.0/M$0.271
13O3OpenAI64$2.0/M$8.0/M$0.078
14Qwen3.6 27BAlibaba62.5$0.60/M$3.6/M$0.034
15Qwen3.5 122B A10BAlibaba62$0.40/M$3.2/M$0.029
16Qwen3.5 27BAlibaba60.5$0.30/M$2.4/M$0.022
17Seed 1.8ByteDance58.8$0.25/M$2.0/M$0.019
18HY-Embodied 0.5 MoT-2BTencent54.5———
19Claude Opus 4.7Anthropic52.5$5.0/M$25.0/M$0.286
20Qwen3 VL 235B A22B ReasoningAlibaba52.5$0.40/M$4.0/M$0.042
21Qwen3 VL 32B ReasoningAlibaba52.3$0.16/M$0.64/M$0.008
22Qwen3 VL 235B A22B InstructAlibaba51.3$0.40/M$1.6/M$0.019
23Qwen3 VL 32B InstructAlibaba48.8$0.16/M$0.64/M$0.008
24Qwen3 VL 4B (Reasoning)+1 altAlibaba47.3———
25MiMo Embodied 7B+1 altXiaomi46.8———
26Qwen3 VL Thinking (8B)Alibaba46.8$0.18/M$2.1/M$0.024
27Qwen3.5 4BAlibaba46.33P$0.03/M$0.15/M$0.002
28Qwen3 VL 8B InstructAlibaba45.8$0.18/M$0.70/M$0.010
29Qwen3 VL 30B A3B ReasoningAlibaba45.3$0.20/M$2.4/M$0.029
30Qwen3 VL 30B A3B InstructAlibaba43$0.20/M$0.80/M$0.012
31Qwen3 VL 2BAlibaba41.8———
32Qwen3 VL 4B InstructAlibaba41.3———
33Claude Opus 4.6Anthropic40.8$5.0/M$25.0/M$0.368
34GPT-4oOpenAI35.2$5.0/M$15.0/M$0.284
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 34 models scored · 0 independently verified · 6 vendor cross-reference · 28 vendor-reported · 0 with source disagreement. How these tiers are assigned