VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.8 vs GPT-5.5

AnthropicvsOpenAI85 shared benchmarks4536 head-to-head
BenchmarkClaude Opus 4.8GPT-5.5
AA Agentic Index49.447.4
AA Intelligence57.356.3
AA-LCR7379
AA-Omniscience28.820.5
Agents' Last Exam2726.6
aime_2026100100
AIRS-Bench8486
apex_agents39.438.5
arc_agi_192.595
arc_agi_31.50.4
ARC-AGI-272.185
arena_elo14821482
arena_text_factuality14641482
arena_vision12801283
Artificial Analysis Coding Index74.374.9
browsecomp88.584.4
BrowseComp (single-agent)88.584.4
coding_arena_elo15391457
critpt20.927.1
cybergym83.181.8
DeepSWE5970
DeepSWE 1.15967
Finance Agent v253.951.8
finance_agent53.960
Fortress18.216.3
frontiermath_tier_431.335.4
GDP (Surge AI)84.816.7
gdpval54.249.5
GDPval-AA (Elo)18901769
GDPval-AA v215931509
GDPval-AA v2 (Elo)15931491
GPQA Diamond93.693.6
GraphWalks — Parents task, 256K-token context subset99.390.1
GraphWalks BFS 1M subset68.145.4
GraphWalks BFS 256K85.973.7
GraphWalks Parents 1M subset83.358.5
HealthBench59.356.5
healthbench_professional57.451.8
HiL-Bench35.339.7
HLE57.952.2
HLE (with tools)57.952.2
HLE Calibration26.544.2
hmmt_feb_202696.798.5
hmmt_nov_202596.596.5
IFBench62.275.9
JobBench48.438.3
Legal Agent Benchmark102.1
livebench77.280.7
livebench_agentic_coding50.554
livebench_coding81.882.2
livebench_data_analysis6681.6
livebench_instruction_following7270.7
livebench_language79.787.4
livebench_math94.395.9
livebench_reasoning89.289.7
matharena_visual_math_overall81.694.9
MCP Atlas83.682.8
MCP Mark Verified76.492.9
MLS Bench Lite42.835.5
MMMU-Pro78.983.2
nl2repo69.750.7
officeqa_pro66.260.9
OmniScience Accuracy48.858
OmniScience Non-Hallucination60.712
OR-Bench (FRR)3.36.7
OSWorld-Verified83.478.7
PostTrainBench37.228.4
Program Bench71.970.8
ResearchRubrics73.564
scicode53.556.1
simplebench64.869
simpleqa_verified39.563.1
SpreadsheetBench 231.629.1
SWE-bench Pro69.258.6
SWE-bench Verified88.680.6
SWE-Marathon4014
TauBench V3 - Banking34.239
Terminal-Bench 2.074.682.7
Terminal-Bench 2.18584.3
Terminal-Bench Hard58.360.6
tool_decathlon59.955.6
toolathlon59.955.6
usamo_202696.798.2
τ²-Bench Telecom (AA run)94.493.9
τ³-Bench27.631.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.