Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 15 measured
43.0292.6

15 models measured, most on the third-party eval harness. Kimi K2.5 tops the board at 92.6.

15 measured·3 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 43.02–92.6 · ◆ solid = first-party or better · outlined = arm's-length
01Kimi K2.5Moonshot92.6$0.45/M$2.3/M$0.015
02Qwen3 VL 235B A22B ReasoningAlibaba89.5$0.40/M$4.0/M$0.025
03GPT-5.2OpenAI84$1.8/M$14.0/M$0.094
04Mage-VL-4BMicrosoft80.33———
05Qwen3 VL 4B (Reasoning)Alibaba79.5———
06Claude Opus 4.5Anthropic76.9$5.0/M$25.0/M$0.195
07Phi 4 Multimodal InstructMicrosoft71.84———
08InternVL3-2BNEWOpenGVLab66.1———
09LFM2.5-VL-1.6BLiquid AI62.71———
10InternVL3.5-1BOpenGVLab60.99———
11LFM2-VL-1.6BNEW+1 altLiquid AI58.35———
12Gemini 3 ProGoogle57.2$2.0/M$12.0/M$0.122
13Phi 4 R V 15BMicrosoft55.41———
14LFM2-VL-450MNEW+1 altLiquid AI44.56———
15LFM2.5-VL-450M-ExtractLiquid AI43.02———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 15 models scored · 0 independently verified · 8 vendor cross-reference · 7 vendor-reported · 0 with source disagreement. How these tiers are assigned