Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 69 measured
43.295.6

69 models measured, most on the official card harness. Claude Sonnet 4.5 tops the board at 95.6.

69 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 43.2–95.6 · ◆ solid = first-party or better · outlined = arm's-length
01Claude Sonnet 4.5+1 altAnthropic95.6$3.0/M$15.0/M$0.094
02Claude Opus 4.5Anthropic95.63P$5.0/M$25.0/M$0.157
03GPT-5OpenAI95.3$1.3/M$10.0/M$0.059
04Qwen3.7 MaxAlibaba95$2.5/M$7.5/M$0.053
05Qwen3.5 397B A17B+1 altAlibaba94.9$0.60/M$3.6/M$0.022
06Qwen3.7 Plus PreviewAlibaba94.5$0.40/M$1.6/M$0.011
07Qwen3.6 PlusAlibaba94.5$0.50/M$3.0/M$0.019
08Kimi K2 ThinkingMoonshot94.4$0.60/M$2.5/M$0.016
09Claude Opus 4+1 altAnthropic94.23P$15.0/M$75.0/M$0.478
10Qwen3.5 122B A10BAlibaba94$0.40/M$3.2/M$0.019
11Qwen3 VL 235B A22B ReasoningAlibaba93.7$0.40/M$4.0/M$0.023
12Gemma 4 31B+1 altGoogle93.73P$0.17/M$0.40/M$0.003
13DeepSeek-V3.2DeepSeek93.7$0.28/M$0.42/M$0.004
14Claude Sonnet 4+1 altAnthropic93.63P$3.0/M$15.0/M$0.096
15Qwen3.6 27B+1 altAlibaba93.5$0.60/M$3.6/M$0.022
16Qwen3.6 35B A3B+2 altsAlibaba93.3$0.38/M$2.3/M$0.014
17Qwen3.5 35B A3B+1 altAlibaba93.3$0.25/M$2.0/M$0.012
18Qwen3.5 27B+2 altsAlibaba93.2$0.30/M$2.4/M$0.014
19Qwen3 235B A22B Instruct 2507Alibaba93.1$0.23/M$0.92/M$0.006
20DeepSeek-R1+5 altsDeepSeek92.9$2.0/M$4.0/M$0.032
21MiMo V2.5 ProXiaomi92.8$0.43/M$0.87/M$0.007
22Qwen3 Max (Reasoning)Alibaba92.8$1.2/M$6.0/M$0.039
23MiMo V2.5 Pro Base+2 altsXiaomi92.8———
24Gemma 4 26B A4BGoogle92.73P$0.07/M$0.34/M$0.002
25Kimi K2 Instruct+4 altsMoonshot92.7$0.57/M$2.3/M$0.015
26Qwen3 Next 80B A3BAlibaba92.5$0.15/M$1.2/M$0.007
27Qwen3 VL 235B A22B InstructAlibaba92.2$0.40/M$1.6/M$0.011
28Qwen3 VL 32B ReasoningAlibaba91.9$0.16/M$0.64/M$0.004
29DeepSeek-V3.1+2 altsDeepSeek91.8$0.56/M$1.7/M$0.012
30Qwen3.5 9BAlibaba91.1$0.14/M$0.20/M$0.002
31Qwen3 VL 30B A3B ReasoningAlibaba90.9$0.20/M$2.4/M$0.014
32Qwen3 Next 80B A3B InstructAlibaba90.9$0.15/M$1.2/M$0.007
33DeepSeek-V4-Pro-Base+6 altsDeepSeek90.8———
34MiMo V2 Flash Base+2 altsXiaomi90.63P———
35DeepSeek-V3.2-Exp-Base+2 altsDeepSeek90.43P———
36Kimi K2 Base+7 altsMoonshot90.2———
37DeepSeek-V3.1-Base+2 altsDeepSeek903P———
38MiMo V2.5 Base+2 altsXiaomi89.8———
39Qwen3 VL 32B InstructAlibaba89.8$0.16/M$0.64/M$0.004
40DeepSeek-V4-Flash-Base+5 altsDeepSeek89.4———
41Step3.5 Flash Base+1 altStepFun89.23P———
42DeepSeek-V3+9 altsDeepSeek89.1$0.24/M$0.90/M$0.006
43Claude 3.5 Sonnet+4 altsAnthropic88.9$3.0/M$15.0/M$0.101
44Qwen3.5 4BAlibaba88.8$0.03/M$0.15/M$0.001
45Qwen3 VL Thinking (8B)Alibaba88.8$0.18/M$2.1/M$0.013
46Qwen3 14BAlibaba88.6$0.35/M$1.4/M$0.010
47Qwen3 VL 30B A3B InstructAlibaba88.4$0.20/M$0.80/M$0.006
48GPT-4o+4 altsOpenAI88$5.0/M$15.0/M$0.114
49DeepSeek-V3.2-Base+2 altsDeepSeek87.5———
50Qwen3 235B A22B+2 altsAlibaba87.4$0.70/M$2.8/M$0.020
51Hy3-preview-Base+1 altTencent86.86———
52DeepSeek-V3-Base+1 altDeepSeek86.81———
53Qwen2.5 Instruct 72B+1 altAlibaba86.8$0.47/M$0.49/M$0.006
54O1 Mini+3 altsOpenAI86.7———
55GLM-4.5-Base+1 altZ.ai86.56———
56Llama 3.1 Instruct 405BMeta86.2———
57Qwen3 VL 4B (Reasoning)Alibaba86———
58DeepSeek-R1-ZeroDeepSeek85.6———
59Qwen3 VL 8B InstructAlibaba84.9$0.18/M$0.70/M$0.005
60Mistral Large 3Mistral82$0.50/M$1.5/M$0.012
61Qwen3 VL 4B InstructAlibaba81.5———
62DeepSeek-V2.5DeepSeek80.3———
63Qwen3.5 2BAlibaba79.6———
64DeepSeek-V2DeepSeek77.9———
65Qwen 2.5 Coder 32B InstructAlibaba77.5$0.06/M$0.20/M$0.002
66Qwen2.5 Omni 7BAlibaba71$0.10/M$6.8/M$0.048
67Qwen3 4B Base DenseAlibaba61.66———
68Qwen3.5 0.8BAlibaba59.5———
69ERNIE 4.5Baidu43.2———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 69 models scored · 0 independently verified · 17 vendor cross-reference · 52 vendor-reported · 0 with source disagreement. How these tiers are assigned