Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 29 measured
37.564.1

29 models measured, most on the third-party eval harness. Qwen3.5 9B tops the board at 64.1.

29 measured·4 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 37.5–64.1 · ◆ solid = first-party or better · outlined = arm's-length
01Qwen3.5 9BAlibaba64.1$0.14/M$0.20/M$0.003
02Gemini 3 ProGoogle63.8$2.0/M$12.0/M$0.110
03Qwen3 VL 235B A22B ReasoningAlibaba63.5VENDOR$0.40/M$4.0/M$0.035
04Gemini 2.5 ProGoogle62.2$1.3/M$10.0/M$0.090
05Qwen3 VL 32B ReasoningAlibaba62.1VENDOR$0.16/M$0.64/M$0.006
06Qwen3 VL 235B A22B InstructAlibaba61.8VENDOR$0.40/M$1.6/M$0.016
07Qwen3.5 35B A3BAlibaba61.7$0.25/M$2.0/M$0.018
08MiniCPM-o-4.5OpenBMB61.5———
09Qwen3 VL 8B InstructAlibaba61.2VENDOR$0.18/M$0.70/M$0.007
10Qwen3.6 35B A3BAlibaba60.8$0.38/M$2.3/M$0.022
11Seed1.6 VisionByteDance60.5———
12Qwen3 VL 30B A3B ReasoningAlibaba60.4VENDOR$0.20/M$2.4/M$0.022
13Qwen3 Omni 30B A3B InstructAlibaba60$0.25/M$0.97/M$0.010
14Claude Opus 4.6Anthropic59.8$5.0/M$25.0/M$0.251
15GLM-4.6V-Flash (9B)Z.ai59.5$0.30/M$0.90/M$0.010
16Qwen3 VL Thinking (8B)Alibaba59.2VENDOR$0.18/M$2.1/M$0.019
17Qwen3 VL 32B InstructAlibaba59.2VENDOR$0.16/M$0.64/M$0.007
18Qwen2.5 VL 32B InstructAlibaba59.1VENDOR———
19Nemotron 3 Nano OmniNVIDIA59$0.30/M$0.90/M$0.010
20MiniCPM-V-4.5-8BOpenBMB58.8———
21Qwen3.6 27BNEWAlibaba58.7$0.60/M$3.6/M$0.036
22Qwen3 VL 30B A3B InstructAlibaba57.8VENDOR$0.20/M$0.80/M$0.009
23Qwen3 VL 4B InstructAlibaba57.6VENDOR———
24Qwen3.8 27BNEWAlibaba56.9$0.50/M$3.0/M$0.031
25Muse GlimmerNEWMeta56.8$0.33/M$1.4/M$0.015
26Qwen3 VL 4B (Reasoning)Alibaba55.8VENDOR———
27MiniCPM-V 4.6OpenBMB52.3———
28LFM2.5-VL-3BNEWLiquid AI37.9———
29Gemma 4 12BGoogle37.5$0.10/M$0.30/M$0.005
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 29 models scored · 18 independently verified · 11 vendor-reported · 0 with source disagreement. How these tiers are assigned