Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 8 measured
45.775.2

8 models measured, most on the third-party eval harness. Llama 3.1 Instruct 405B tops the board at 75.2.

8 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 45.7–75.2 · ◆ solid = first-party or better · outlined = arm's-length
01Llama 3.1 Instruct 405BMeta75.2VENDOR———
02Step3.5 Flash Base+1 altStepFun67.7———
03Llama 3.1 70B InstructMeta65.5VENDOR$0.56/M$0.56/M$0.009
04Kimi K2 Base+2 altsMoonshot60.5———
05MiMo V2 Flash Base+2 altsXiaomi59.5———
06Llama 3.1 8B InstructMeta50.8VENDOR$0.02/M$0.04/M$0.001
07DeepSeek-V3.1-Base+2 altsDeepSeek45.9———
08DeepSeek-V3.2-Exp-Base+2 altsDeepSeek45.7———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 8 models scored · 0 independently verified · 3 vendor cross-reference · 5 vendor-reported · 0 with source disagreement. How these tiers are assigned