Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 10 measured
3.6450.8

10 models measured, most on the third-party eval harness. Qwen3 4B 2507 Instruct tops the board at 50.8.

10 measured·10 new this week·Lifecycle Updated Oct 9, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 3.64–50.8 · ◆ solid = first-party or better · outlined = arm's-length
01Qwen3 4B 2507 InstructNEWAlibaba50.8$0.01/M$0.03/M$0.000
02Qwen3 4BNEWAlibaba38.64———
03Granite 4.0 H TinyNEWIBM27.27———
04Qwen3 1.7BNEWAlibaba26.48———
05LFM2-8B-A1B (Non-Reasoning)NEWLiquid AI21.36———
06Gemma 3 4BNEWGoogle19.09$0.05/M$0.10/M$0.004
07LFM2-2.6BNEWLiquid AI14.43———
08Llama 3.2 3B InstructNEWMeta11.48$0.05/M$0.33/M$0.017
09Gemma 3 1B InstructNEWGoogle4.43———
10Llama 3.2 Instruct 1BNEWMeta3.64$0.03/M$0.20/M$0.031
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 10 models scored · 0 independently verified · 9 vendor cross-reference · 1 vendor-reported · 0 with source disagreement. How these tiers are assigned