Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 17 measured
0.82619.937

17 models measured, most on the third-party eval harness. Llama 3 Instruct 8B tops the board at 19.937.

17 measured·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 0.826–19.937 · ◆ solid = first-party or better · outlined = arm's-length
01Llama 3 Instruct 8BMeta19.937$0.04/M$0.14/M$0.005
02Yi 1.5 34B 32K01.AI14.078———
03Yi 1.5 6B Chat01.AI14.031———
04Yi 1.5 34B Chat 16K01.AI13.737———
05Yi 1.5 6B01.AI13.309———
06Yi 1.5 34B Chat01.AI13.058———
07Yi 1.5 9B Chat01.AI12.838———
08Yi 9B 200K01.AI12.109———
09Yi 1.5 9B01.AI12.031———
10Yi 1.5 34B01.AI11.217———
11Yi 1.5 9B 32K01.AI10.827———
12Yi 1.5 9B Chat 16K01.AI10.038———
13Marco-o1AIDC-AI9.964———
14Yi 9B01.AI8.909———
15Yi Coder 9B Chat01.AI7.964———
16NuminaMath-7B-TIRAI-MO4.199———
17NuminaMath-7B-CoTAI-MO0.826———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 17 models scored · 17 independently verified · 0 with source disagreement. How these tiers are assigned