Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 25 measured
53.9778.8

25 models measured, most on the third-party eval harness. DeepSeek-V3 tops the board at 78.8.

25 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 53.97–78.8 · ◆ solid = first-party or better · outlined = arm's-length
01DeepSeek-V3DeepSeek78.8$0.24/M$0.90/M$0.007
02Hy3-preview-Base+1 altTencent78.71VENDOR———
03GLM-4.5-Base+1 altZ.ai78.05VENDOR———
04Qwen2.5 Instruct 72BAlibaba77$0.47/M$0.49/M$0.006
05GPT-4oOpenAI76.2$5.0/M$15.0/M$0.131
06Gemini 2.0 Flash ExpGoogle75.9———
07DeepSeek-V3-Base+1 altDeepSeek75.47VENDOR———
08Gemini 1.5 ProGoogle75.4———
09Claude 3.5 SonnetAnthropic75.1$3.0/M$15.0/M$0.120
10MiMo V2.5 Pro Base+2 altsXiaomi74.1VENDOR———
11Granite 4.1 8BIBM73.81$0.05/M$0.10/M$0.001
12Kimi K2 Base+5 altsMoonshot73.8VENDOR———
13Granite 4.1 30BIBM73.54———
14Llama 3.1 Instruct 405BMeta73———
15Seed Coder 8B InstructByteDance72.8———
16DeepSeek-V3.1-Base+2 altsDeepSeek72.2———
17MiniMax Text 01MiniMax71.7———
18MiMo V2 Flash Base+2 altsXiaomi71.4———
19MiMo V2.5 Base+2 altsXiaomi70.9VENDOR———
20Step3.5 Flash Base+1 altStepFun70.6———
21DeepSeek-V3.2-Exp-Base+2 altsDeepSeek69.8———
22Granite 4.1 30B Base+1 altIBM69.58———
23Granite 4.1 3B BaseIBM68.25———
24Granite 4.1 3BIBM62.17———
25Granite 4.1 8B Base+1 altIBM53.97———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 25 models scored · 0 independently verified · 12 vendor cross-reference · 13 vendor-reported · 0 with source disagreement. How these tiers are assigned