Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 31 measured
15.3992.7

31 models measured, most on the official card harness. GPT-6 Astra tops the board at 92.7.

31 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 15.39–92.7 · ◆ solid = first-party or better · outlined = arm's-length
01GPT-6 AstraOpenAI92.7$10.0/M$50.0/M$0.324
02Claude Opus 4.8+1 altAnthropic87.9$5.0/M$25.0/M$0.171
03GPT-5.2OpenAI86.3$1.8/M$14.0/M$0.091
04Qwen3.8 Max PreviewAlibaba84.5$3.3/M$9.9/M$0.078
05Muse SparkMeta84.1———
06Claude Opus 4.7+2 altsAnthropic79.5$5.0/M$25.0/M$0.189
07Qwen3.7 Plus PreviewAlibaba79$0.40/M$1.6/M$0.013
08Muse GlimmerMeta75.4$0.33/M$1.4/M$0.011
09Gemini 3 ProGoogle72.7$2.0/M$12.0/M$0.096
10Qwen3.5 122B A10BAlibaba70.4$0.40/M$3.2/M$0.026
11Qwen3.5 27BAlibaba70.3$0.30/M$2.4/M$0.019
12Gemini 3 Flash PreviewGoogle69.1$0.50/M$3.0/M$0.025
13Qwen3.5 35B A3BAlibaba68.6$0.25/M$2.0/M$0.016
14Qwen3.6 PlusAlibaba68.2$0.50/M$3.0/M$0.026
15Qwen3 VL 235B A22B InstructAlibaba62$0.40/M$1.6/M$0.016
16Qwen3 VL 235B A22B ReasoningAlibaba61.8$0.40/M$4.0/M$0.036
17Qwen3 VL 30B A3B InstructAlibaba60.5$0.20/M$0.80/M$0.008
18Qwen3 VL 4B InstructAlibaba59.5———
19Qwen3 VL 32B InstructAlibaba57.9$0.16/M$0.64/M$0.007
20Claude Opus 4.6+1 altAnthropic57.7$5.0/M$25.0/M$0.260
21Qwen3 VL 30B A3B ReasoningAlibaba57.3$0.20/M$2.4/M$0.023
22Qwen3 VL 32B ReasoningAlibaba57.1$0.16/M$0.64/M$0.007
23Qwen3 VL 8B InstructAlibaba54.6$0.18/M$0.70/M$0.008
24Step3 VL 10B+1 altStepFun51.55———
25Qwen3 VL 4B (Reasoning)Alibaba49.2———
26Qwen3 VL Thinking (8B)+2 altsAlibaba46.6$0.18/M$2.1/M$0.024
27GLM-4.6V-Flash (9B)+1 altZ.ai45.68$0.30/M$0.90/M$0.013
28Qwen2.5 VL 72BAlibaba43.6———
29Qwen2.5 VL 32B Instruct+1 altAlibaba39.4———
30MiMo VL RL 2508 (7B)+1 altXiaomi34.84———
31InternVL-3.5 (8B)+1 altOpenGVLab15.39———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 31 models scored · 0 independently verified · 3 vendor cross-reference · 28 vendor-reported · 1 with source disagreement. How these tiers are assigned