Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensofficial card harness
The field · 54 measured
1.9793.1

54 models measured, most on the official card harness. DeepSeek-V3.1 tops the board at 93.1.

54 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 1.97–93.1 · ◆ solid = first-party or better · outlined = arm's-length
01DeepSeek-V3.1+1 altDeepSeek93.1$0.56/M$1.7/M$0.012
02Gemini 2.5 Pro+1 altGoogle923P$1.3/M$10.0/M$0.061
03Seed Oss 36B InstructByteDance91.73P$0.21/M$0.57/M$0.004
04O3+1 altOpenAI91.63P$2.0/M$8.0/M$0.055
05Ministral 3 14B+3 altsMistral89.8$0.20/M$0.20/M$0.002
06DeepSeek-R1 0528-Qwen3-8BDeepSeek86———
07Qwen3 VL Thinking (8B)+3 altsAlibaba863P$0.18/M$2.1/M$0.013
08MiniMax M1 80K+1 altMiniMax863P$0.55/M$2.2/M$0.016
09Ministral 3 8B+2 altsMistral863P$0.15/M$0.15/M$0.002
10Qwen3 235B A22B+4 altsAlibaba85.7$0.70/M$2.8/M$0.020
11Qwen3 14B+3 altsAlibaba83.73P$0.35/M$1.4/M$0.010
12MiniMax M1 40K+1 altMiniMax83.33P———
13Gemini 2.5 FlashGoogle82.3$0.30/M$2.5/M$0.017
14Qwen3 32BAlibaba81.4$0.16/M$0.64/M$0.005
15Phi 4 Reasoning PlusMicrosoft81.3———
16DeepSeek-R1+8 altsDeepSeek79.8$2.0/M$4.0/M$0.038
17O3 MiniOpenAI79.6$1.1/M$4.4/M$0.035
18O1+4 altsOpenAI79.2$15.0/M$60.0/M$0.473
19DeepSeek-R1-ZeroDeepSeek77.9———
20Ministral 3 3B+3 altsMistral77.5$0.10/M$0.10/M$0.001
21Qwen3 8BAlibaba76$0.18/M$0.70/M$0.006
22Qwen3 VL 4B (Reasoning)+3 altsAlibaba72.93P———
23DeepSeek-R1-Distill-Qwen-32B+4 altsDeepSeek72.6$0.29/M$0.29/M$0.004
24Magistral Medium 1Mistral70.7———
25DeepSeek-R1-Distill-Llama-70B+3 altsDeepSeek70$0.70/M$1.1/M$0.013
26DeepSeek-R1-Distill-Qwen-14B+8 altsDeepSeek69.7$0.20/M$0.20/M$0.003
27Kimi K2 Instruct+1 altMoonshot69.63P$0.57/M$2.3/M$0.021
28MiMo 7B RL+8 altsXiaomi68.2———
29O1 Mini+9 altsOpenAI63.6———
30MiMo 7B SFT+4 altsXiaomi58.7———
31MiMo 7B RL Zero+4 altsXiaomi56.4———
32DeepSeek-R1-Distill-Qwen-7B+8 altsDeepSeek55.5$0.15/M$0.15/M$0.003
33DeepSeek-R1-Distill-Llama-8B+3 altsDeepSeek50.4$0.05/M$0.05/M$0.001
34QwQ 32B Preview+8 altsAlibaba50$0.66/M$1.0/M$0.017
35Claude Opus 4+3 altsAnthropic48.23P$15.0/M$75.0/M$0.934
36DeepSeek-R1-Zero-Qwen-32BDeepSeek473P———
37Claude Sonnet 4+1 altAnthropic43.43P$3.0/M$15.0/M$0.207
38DeepSeek-V3+9 altsDeepSeek39.2$0.24/M$0.90/M$0.015
39SynLogic Mix 3 32BMiniMax35.83P———
40MiMo 7B Base+4 altsXiaomi32.9———
41DeepSeek-R1-Distill-Qwen-1.5B+3 altsDeepSeek28.9———
42Llama 3.1 Instruct 405BMeta23.3———
43Qwen2.5 Instruct 72BAlibaba23.3$0.47/M$0.49/M$0.021
44Qwen3.5 2BAlibaba17———
45DeepSeek-V2.5DeepSeek16.7———
46Claude 3.5 Sonnet+10 altsAnthropic16$3.0/M$15.0/M$0.563
47SynLogic 7BMiniMax103P———
48GPT-4o+10 altsOpenAI9.3$5.0/M$15.0/M$1.075
49Granite 3.3 8B InstructIBM8.123P$0.03/M$0.25/M$0.017
50Qwen3.5 0.8BAlibaba5.67———
51DeepSeek-V2DeepSeek4.6———
52Granite 3.3 2B InstructIBM3.283P———
53Granite 3.2 8B InstructIBM2.433P———
54Granite 3.1 8B InstructIBM1.973P———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 54 models scored · 0 independently verified · 24 vendor cross-reference · 30 vendor-reported · 8 with source disagreement. How these tiers are assigned