Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 16 measured
25.8437.92

16 models measured, most on the third-party eval harness. Yi 1.5 9B tops the board at 37.92.

16 measured·Lifecycle Updated Oct 7, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 25.84–37.92 · ◆ solid = first-party or better · outlined = arm's-length
01Yi 1.5 9B01.AI37.92———
02Yi 1.5 34B01.AI36.58———
03Yi 1.5 34B Chat01.AI36.49———
04Yi 1.5 34B 32K01.AI36.33———
05Yi 1.5 9B 32K01.AI35.91———
06Yi 1.5 34B Chat 16K01.AI33.81———
07Yi 1.5 9B Chat01.AI33.47———
08Yi 9B01.AI31.8———
09Yi 9B 200K01.AI31.54———
10Yi 1.5 6B01.AI31.38———
11Yi 1.5 9B Chat 16K01.AI30.87———
12Yi 1.5 6B Chat01.AI30.2———
13Llama 3 Instruct 8BMeta26.59$0.04/M$0.14/M$0.004
14NuminaMath-7B-CoTAI-MO26.59———
15Marco-o1AIDC-AI25.92———
16NuminaMath-7B-TIRAI-MO25.84———
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 16 models scored · 16 independently verified · 0 with source disagreement. How these tiers are assigned