Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 59 measured
53.2582.97

59 models measured, most on the third-party eval harness. GPT-6 Astra tops the board at 82.97.

59 measured·1 new this week·Lifecycle Updated Oct 7, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 53.25–82.97 · ◆ solid = first-party or better · outlined = arm's-length
01GPT-6 AstraOpenAI82.97$10.0/M$50.0/M$0.362
02GPT-6.1 Sol+1 altOpenAI82.66$2.0/M$10.0/M$0.073
03GPT-5.5OpenAI81.58$5.0/M$30.0/M$0.215
04GPT-6 SolOpenAI81.19$2.0/M$10.0/M$0.074
05Claude Fable 5Anthropic80.54$10.0/M$50.0/M$0.372
06Claude Opus 5.5+1 altAnthropic80.31$4.0/M$20.0/M$0.149
07Claude Fable 5.1Anthropic80.28$10.0/M$50.0/M$0.374
08GPT-5.6 SolOpenAI79.84$4.0/M$20.0/M$0.150
09Muse Spark 1.3Meta79.57$1.3/M$4.3/M$0.035
10DeepSeek-V4-Flash-Vision-ExpDeepSeek79.48$0.22/M$0.65/M$0.005
11DeepSeek-V4-Flash+1 altDeepSeek79.33$0.44/M$1.3/M$0.011
12GPT-5.4OpenAI79.31$2.5/M$15.0/M$0.110
13GPT-5.6 TerraOpenAI79.31$2.0/M$12.0/M$0.088
14DeepSeek-V4.1-FlashDeepSeek79.25$0.05/M$1.2/M$0.008
15Kimi K3Moonshot78.73$0.62/M$15.0/M$0.099
16Gemini 3.1 ProGoogle78.54$2.0/M$12.0/M$0.089
17Qwen3.8 Max PreviewAlibaba78.41$3.3/M$9.9/M$0.084
18Claude Opus 4.7Anthropic78.26$5.0/M$25.0/M$0.192
19GPT-5.2 CodexOpenAI78.2$1.8/M$14.0/M$0.101
20GPT-5.2OpenAI78.16$1.8/M$14.0/M$0.101
21GPT-5.6 LunaOpenAI78.03$0.20/M$1.2/M$0.009
22Claude Sonnet 4.6Anthropic77.95$3.0/M$15.0/M$0.115
23Grok 4.7SpaceXAI76.88$2.0/M$6.0/M$0.052
24Qwen3.8 27BAlibaba76.59$0.50/M$3.0/M$0.023
25Muse Spark 1.2Meta76.46$1.3/M$4.3/M$0.036
26GLM 5.3 FlashZ.ai76.4$0.15/M$0.50/M$0.004
27MiniMax M3MiniMax76.17$0.28/M$1.1/M$0.009
28Ox Alpha MaxStealth75.77———
29Mistral Large 4NEWMistral74.91$0.68/M$2.1/M$0.018
30Union AlphaStealth74.6———
31Claude Opus 5Anthropic74.55$5.0/M$25.0/M$0.201
32DeepSeek-V4-ProDeepSeek74.54$0.43/M$0.87/M$0.009
33Claude Opus 4.5Anthropic74.44$5.0/M$25.0/M$0.202
34Qwen3.8 Flash NextAlibaba74.24$0.15/M$0.47/M$0.004
35Grok 4.6SpaceXAI73.86$2.0/M$6.0/M$0.054
36GLM-5.2Z.ai73.74$1.4/M$4.4/M$0.039
37GPT-6 LunaOpenAI73.37$0.10/M$0.50/M$0.004
38Grok4.5SpaceXAI73.04$2.0/M$6.0/M$0.055
39InklingThinking Machines72.78$0.95/M$4.0/M$0.034
40Muse Spark 1.1Meta72.55$1.3/M$4.3/M$0.038
41Qwen3.7 MaxAlibaba71.79$2.5/M$7.5/M$0.070
42Claude Sonnet 5Anthropic71.74$2.0/M$10.0/M$0.084
43GPT-5.4 miniOpenAI70.79$0.75/M$4.5/M$0.037
44Qwen3.6 27BAlibaba70.43$0.60/M$3.6/M$0.030
45GLM-5.3Z.ai70.24$1.4/M$4.4/M$0.041
46Qwen3.6 PlusAlibaba69.91$0.50/M$3.0/M$0.025
47Claude Opus 4.6Anthropic69.89$5.0/M$25.0/M$0.215
48Gemini 3.7 FlashGoogle67.96$0.75/M$3.8/M$0.033
49GPT-5.4 nanoOpenAI67.64$0.20/M$1.3/M$0.011
50Claude Opus 4.8Anthropic66.03$5.0/M$25.0/M$0.227
51Kimi K2.6Moonshot65.13$0.95/M$4.0/M$0.038
52Gemini 3.5 FlashGoogle64.86$1.5/M$9.0/M$0.081
53Gemini 3.6 FlashGoogle63$0.75/M$3.8/M$0.036
54Kimi K2.7 CodeMoonshot62.66$1.9/M$8.1/M$0.080
55Sonnet 5.5+1 altAnthropic59.49$2.0/M$10.0/M$0.101
56Grok 4.3SpaceXAI55.77$1.3/M$2.5/M$0.034
57Nemotron 3 Ultra 550B A55BNVIDIA54.46$0.60/M$2.5/M$0.028
58Gemini 3.8 FlashGoogle54.01$0.75/M$3.8/M$0.042
59Gemini 3.5 Flash LiteGoogle53.25$0.30/M$2.5/M$0.026
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 59 models scored · 59 independently verified · 0 with source disagreement. How these tiers are assigned