Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 39 measured
23.1792.3

39 models measured, most on the third-party eval harness. Phi 4 Reasoning Plus tops the board at 92.3.

39 measured·10 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 23.17–92.3 · ◆ solid = first-party or better · outlined = arm's-length
01Phi 4 Reasoning PlusMicrosoft92.3VENDOR———
02Granite 3.3 8B Instruct+1 altIBM86.09VENDOR$0.03/M$0.25/M$0.002
03Granite 3.1 8B InstructIBM85.79———
04Granite 3.2 8B InstructIBM85.72———
05Granite 4.1 30BIBM85.37———
06Kimi K2 Base+3 altsMoonshot84.8VENDOR———
07Phi 4Microsoft82.8VENDOR$0.13/M$0.50/M$0.004
08Qwen3 4B 2507 InstructNEWAlibaba82.32$0.01/M$0.03/M$0.000
09Seed Coder 8B InstructByteDance82.3———
10Llama 3.1 8B InstructMeta80.15$0.02/M$0.04/M$0.000
11Granite 4.1 8BIBM79.88$0.05/M$0.10/M$0.001
12DeepSeek-R1-Distill-Qwen-7BDeepSeek78.43$0.15/M$0.15/M$0.002
13Granite 4.1 3BIBM76.83———
14Granite 3.3 2B InstructIBM75.68———
15MiMo V2.5 ProXiaomi75.6VENDOR$0.43/M$0.87/M$0.009
16MiMo V2.5 Pro Base+2 altsXiaomi75.6VENDOR———
17Granite 3.1 2B InstructIBM75.26———
18Granite 4.0 H TinyNEWIBM73.78———
19Granite 3.2 2B InstructIBM73.39———
20Step3.5 Flash Base+1 altStepFun72———
21Qwen3 4BNEWAlibaba71.95———
22MiMo V2.5 Base+2 altsXiaomi71.3VENDOR———
23MiMo V2 Flash Base+2 altsXiaomi70.7———
24LFM2-8B-A1B (Non-Reasoning)NEWLiquid AI69.51———
25DeepSeek-V3.2-Exp-Base+2 altsDeepSeek67.7———
26Qwen3.5 2BAlibaba66VENDOR———
27DeepSeek-V3.1-Base+2 altsDeepSeek64.6———
28DeepSeek-R1-Distill-Llama-8BDeepSeek62.91$0.05/M$0.05/M$0.001
29Granite 4.1 8B Base+1 altIBM62.8———
30Gemma 3 4BNEWGoogle62.8$0.05/M$0.10/M$0.001
31Granite 4.1 30B Base+1 altIBM62.2———
32Qwen3 1.7BNEWAlibaba60.98———
33LFM2-2.6BNEWLiquid AI57.93———
34Granite 4.1 3B BaseIBM54.27———
35Qwen3.5 0.8BAlibaba46.81VENDOR———
36Gemma 3 1B InstructNEWGoogle37.2———
37ERNIE 4.5Baidu25VENDOR———
38Llama 3.2 3B InstructNEWMeta24.06$0.05/M$0.33/M$0.008
39Llama 3.2 Instruct 1BNEWMeta23.17$0.03/M$0.20/M$0.005
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 39 models scored · 0 independently verified · 20 vendor cross-reference · 19 vendor-reported · 0 with source disagreement. How these tiers are assigned