Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Pro Mode — raw scores, one benchmark under the lensthird-party eval harness
The field · 84 measured
17.488.4

84 models measured, most on the third-party eval harness. Claude Opus 5.5 tops the board at 88.4.

84 measured·2 new this week·Lifecycle Updated Oct 8, 2026Score basis
#◆Vendor$/M in$/M out
Leader — ignites whiteField — fill cools with scoreBars show 17.4–88.4 · ◆ solid = first-party or better · outlined = arm's-length
01Claude Opus 5.5Anthropic88.4$4.0/M$20.0/M$0.136
02Claude Fable 5.1Anthropic86.6$10.0/M$50.0/M$0.346
03GPT-6 AstraOpenAI83.6$10.0/M$50.0/M$0.359
04GPT-6.1 SolNEWOpenAI82.9$2.0/M$10.0/M$0.072
05Gemini 3.8 FlashGoogle82.4$0.75/M$3.8/M$0.027
06Claude Fable 5Anthropic81.9$10.0/M$50.0/M$0.366
07Muse Spark 1.3Meta81.8$1.3/M$4.3/M$0.034
08Claude Opus 5Anthropic80.6$5.0/M$25.0/M$0.186
09Gemini 3.1 ProGoogle79.6$2.0/M$12.0/M$0.088
10GPT-5.5 ProOpenAI76.9$30.0/M$180.0/M$1.365
11Gemini 3.5 FlashGoogle76.7$1.5/M$9.0/M$0.068
12Gemini 3 ProGoogle76.4$2.0/M$12.0/M$0.092
13Sonnet 5.5NEWAnthropic75.9$2.0/M$10.0/M$0.079
14Grok 4.6SpaceXAI75.9$2.0/M$6.0/M$0.053
15Muse Spark 1.2Meta74.5$1.3/M$4.3/M$0.037
16GPT-6 SolOpenAI73.1$2.0/M$10.0/M$0.082
17GPT-5.6 Sol Pro+1 altOpenAI71.7$2.0/M$10.0/M$0.084
18Qwen3.7 MaxAlibaba70.4$2.5/M$7.5/M$0.071
19Grok4.5SpaceXAI70$2.0/M$6.0/M$0.057
20GPT-5.5OpenAI69$5.0/M$30.0/M$0.254
21Claude Opus 4.6Anthropic67.6$5.0/M$25.0/M$0.222
22DeepSeek-V4.1-FlashDeepSeek66.7$0.20/M$0.60/M$0.006
23GLM-5.3Z.ai66.2$1.4/M$4.4/M$0.044
24GPT-5.6 Sol+1 altOpenAI64.8$4.0/M$20.0/M$0.185
25Claude Opus 4.8Anthropic64.8$5.0/M$25.0/M$0.231
26Qwen3.6 Max PreviewAlibaba63$1.3/M$7.8/M$0.072
27Qwen3.8 2.4T A95BAlibaba62.5$2.0/M$6.0/M$0.064
28Claude Opus 4.7Anthropic61.7$5.0/M$25.0/M$0.243
29GPT-5 ProOpenAI61.6$15.0/M$120.0/M$1.096
30Gemini 3 Flash PreviewGoogle61.1$0.50/M$3.0/M$0.029
31Kimi K3Moonshot60.7$0.58/M$12.3/M$0.106
32Claude Sonnet 5Anthropic60.6$2.0/M$10.0/M$0.099
33Qwen3.8 27BAlibaba60.2$0.50/M$3.0/M$0.029
34GLM-5.2Z.ai58.8$1.4/M$4.4/M$0.049
35Claude Opus 4Anthropic58.8$15.0/M$75.0/M$0.765
36Kimi K2.7 CodeMoonshot57.9$1.9/M$8.1/M$0.086
37GPT-5.2 ProOpenAI57.4$21.0/M$168.0/M$1.646
38GPT-5OpenAI56.7$1.3/M$10.0/M$0.099
39Grok 4.1 FastSpaceXAI56$0.20/M$0.50/M$0.006
40GLM-5.1Z.ai55.1$1.4/M$4.4/M$0.052
41Claude Sonnet 4.5Anthropic54.3$3.0/M$15.0/M$0.166
42GLM-5Z.ai53.2$1.0/M$3.2/M$0.039
43GPT-5.1OpenAI53.2$1.3/M$10.0/M$0.106
44O3OpenAI53.1$2.0/M$8.0/M$0.094
45DeepSeek-V3.2-SpecialeDeepSeek52.6$0.29/M$0.43/M$0.007
46Gemini 2.5 ProGoogle51.6$1.3/M$10.0/M$0.109
47DeepSeek-V4-ProDeepSeek50.9$0.43/M$0.87/M$0.013
48InklingThinking Machines50$0.95/M$4.0/M$0.050
49GPT-5.6 Terra+1 altOpenAI48.9$2.0/M$12.0/M$0.143
50GLM-4.7Z.ai47.7$0.60/M$2.2/M$0.029
51GPT-5.6 Luna+1 altOpenAI46.8$0.20/M$1.2/M$0.015
52Claude 3.7 Sonnet+1 altAnthropic46.4$3.0/M$15.0/M$0.194
53DeepSeek-V4-FlashDeepSeek46.3$0.44/M$1.3/M$0.019
54MiniMax M3MiniMax45.8$0.28/M$1.1/M$0.015
55GPT-5.2OpenAI45.8$1.8/M$14.0/M$0.172
56Claude Sonnet 4Anthropic45.5$3.0/M$15.0/M$0.198
57Nemotron 3 Ultra 550B A55BNVIDIA41.7$0.60/M$2.5/M$0.037
58Gemini 2.5 FlashGoogle41.2$0.30/M$2.5/M$0.034
59O1+1 altOpenAI40.1$15.0/M$60.0/M$0.935
60DeepSeek-V3.1DeepSeek40$0.56/M$1.7/M$0.028
61Grok 3SpaceXAI36.1$4.0/M$20.0/M$0.332
62Qwen 3.6 FlashAlibaba35.2$0.25/M$1.5/M$0.025
63MiniMax M2.1MiniMax34.7$0.30/M$1.2/M$0.022
64GPT-4.5 PreviewOpenAI34.5———
65Gemini Exp 1206Google31.1———
66Gemini 2.0 Flash (Reasoning)Google30.7———
67Llama 4 MaverickMeta27.7$0.25/M$0.87/M$0.020
68Claude 3.5 Sonnet+1 altAnthropic27.5$3.0/M$15.0/M$0.327
69Gemini 1.5 ProGoogle27.1———
70GPT-4.1OpenAI27$2.0/M$8.0/M$0.185
71Kimi K2 ThinkingMoonshot26.3$0.60/M$2.5/M$0.059
72GPT-4 TurboOpenAI25.1$10.0/M$30.0/M$0.797
73Claude 3 OpusAnthropic23.5$15.0/M$75.0/M$1.915
74Llama 3.1 Instruct 405BMeta23———
75O3 MiniOpenAI22.8$1.1/M$4.4/M$0.121
76Grok 2SpaceXAI22.7———
77Mistral Large 2Mistral22.5$2.0/M$6.0/M$0.178
78GPT Oss 120bOpenAI22.1$0.15/M$0.59/M$0.017
79Mistral Large 3Mistral20.4$0.50/M$1.5/M$0.049
80Llama 3.3 70B InstructMeta19.9$0.71/M$0.72/M$0.036
81Gemini 2.0 Flash ExpGoogle18.9———
82DeepSeek-V3+1 altDeepSeek18.9$0.24/M$0.90/M$0.030
83O1 MiniOpenAI18.1———
84Command R+ (Apr '24)Cohere17.4$2.5/M$10.0/M$0.359
◆ solid = independently verified · aggregator attested · vendor attributed — outlined = cross-reference · source-attributed. Dimmed ◆ = confidence inferred. “+N alts” = alternate disclosures in the drawer with why-demoted. Prices per million tokens; $/point = input price ÷ score. Every row opens its model dossier.

Verification: 84 models scored · 84 independently verified · 0 with source disagreement. How these tiers are assigned