Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 8 measured
50.665.7

8 models measured, most on the third-party eval harness. Llama 3.1 Instruct 405B tops the board at 65.7.

8 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 50.6–65.7 · ◆ solid = first-party or better · outlined = arm's-length
01Llama 3.1 Instruct 405BMeta65.7VENDOR———
02Llama 3.1 70B InstructMeta62VENDOR$0.56/M$0.56/M$0.009
03Kimi K2 Base+2 altsMoonshot58.8———
04Step3.5 Flash Base+1 altStepFun58———
05MiMo V2 Flash Base+2 altsXiaomi56.7———
06DeepSeek-V3.1-Base+2 altsDeepSeek52.5———
07Llama 3.1 8B InstructMeta52.4VENDOR$0.02/M$0.04/M$0.001
08DeepSeek-V3.2-Exp-Base+2 altsDeepSeek50.6———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 8 models scored · 0 independently verified · 3 vendor cross-reference · 5 vendor-reported · 0 with source disagreement. How these tiers are assigned