Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 44 measured
25.3191.7

44 models measured, most on the official card harness. Gemini 4 Argon tops the board at 91.7.

44 measured·1 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 25.31–91.7 · ◆ solid = first-party or better · outlined = arm's-length
01Gemini 4 ArgonNEW+2 altsGoogle91.7$2.0/M$10.0/M$0.065
02GPT-6 AstraOpenAI87.5$10.0/M$50.0/M$0.343
03Gemini 3.8 Flash+2 altsGoogle87.1$0.75/M$3.8/M$0.026
04Gemini 3.7 Flash+2 altsGoogle85.4$0.75/M$3.8/M$0.026
05Gemini 3.6 FlashGoogle84.2$0.75/M$3.8/M$0.027
06Claude Opus 5.5Anthropic83.7$4.0/M$20.0/M$0.143
07GPT-5.6 SolOpenAI82.1$4.0/M$20.0/M$0.146
08Qwen3.8 Max PreviewAlibaba81.8$3.3/M$9.9/M$0.081
09Claude Fable 5.1Anthropic79.7$10.0/M$50.0/M$0.376
10GPT-5.6 TerraOpenAI78.9$2.0/M$12.0/M$0.089
11Seed 2.1 Pro Preview+1 altByteDance78———
12Seed2.1ByteDance76.8———
13Qwen3.8 Flash Next+1 altAlibaba76.6$0.15/M$0.47/M$0.004
14Gemini 3.5 FlashGoogle76.3$1.5/M$9.0/M$0.069
15Qwen3.7 Plus Preview+1 altAlibaba76.2$0.40/M$1.6/M$0.013
16Kimi K2.5+1 altMoonshot75.9$0.45/M$2.3/M$0.018
17Claude Opus 5Anthropic75.4$5.0/M$25.0/M$0.199
18Qwen3.5 122B A10BAlibaba74.4$0.40/M$3.2/M$0.024
19Qwen3.5 27BAlibaba73.6$0.30/M$2.4/M$0.018
20Gemini 3 ProGoogle73.53P$2.0/M$12.0/M$0.095
21Seed 1.8ByteDance73$0.25/M$2.0/M$0.015
22Qwen3.8 27BAlibaba72.4$0.50/M$3.0/M$0.024
23Qwen3.5 35B A3BAlibaba71.4$0.25/M$2.0/M$0.016
24Qwen3.6 35B A3BAlibaba71.4$0.38/M$2.3/M$0.018
25Claude Sonnet 5Anthropic68.5$2.0/M$10.0/M$0.088
26Qwen3 VL 235B A22B InstructAlibaba67.7$0.40/M$1.6/M$0.015
27Gemini 3.1 ProGoogle66.2$2.0/M$12.0/M$0.106
28Qwen3 VL 32B InstructAlibaba63.8$0.16/M$0.64/M$0.006
29Qwen3 VL 235B A22B Reasoning+1 altAlibaba63.6$0.40/M$4.0/M$0.035
30Claude Opus 4.6Anthropic63$5.0/M$25.0/M$0.238
31Qwen3 VL 32B ReasoningAlibaba62.6$0.16/M$0.64/M$0.006
32Qwen3 VL 30B A3B InstructAlibaba62.5$0.20/M$0.80/M$0.008
33Qwen3 VL 30B A3B ReasoningAlibaba59.2$0.20/M$2.4/M$0.022
34Qwen3 VL 8B InstructAlibaba58$0.18/M$0.70/M$0.008
35Qwen3 VL 4B InstructAlibaba56.2———
36Qwen3 VL Thinking (8B)Alibaba55.8$0.18/M$2.1/M$0.020
37Qwen3 VL 4B (Reasoning)+1 altAlibaba53.5———
38Qwen2.5 VL 32B Instruct+1 altAlibaba49———
39Qwen2.5 VL 72BAlibaba47.3———
40Mage-VL-4BMicrosoft41.83P———
41Nova Pro+1 altAmazon41.6$0.80/M$3.2/M$0.048
42Nova Lite+1 altAmazon40.4$0.06/M$0.24/M$0.004
43Phi 4 R V 15BMicrosoft34.43P———
44Phi 4 Multimodal InstructMicrosoft25.313P———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 44 models scored · 0 independently verified · 11 vendor cross-reference · 33 vendor-reported · 0 with source disagreement. How these tiers are assigned