VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.8 vs GPT-5.4

AnthropicvsOpenAI72 shared benchmarks4329 head-to-head
BenchmarkClaude Opus 4.8GPT-5.4
AA Agentic Index49.458.2
AA Intelligence57.353.1
AA-LCR7377.7
AA-Omniscience28.85.8
AgentWorldBench - Android61.560
AgentWorldBench - MCP54.970.1
AgentWorldBench - OS66.668.6
AgentWorldBench - Overall56.658.3
AgentWorldBench - Search35.137.3
AgentWorldBench - SWE64.166.3
AgentWorldBench - Terminal59.253.7
AgentWorldBench - Web54.751.8
aime_202610099.2
arc_agi_192.593.7
arc_agi_31.50.2
ARC-AGI-272.174
arena_elo14821477
arena_text_factuality14641476
arena_vision12801277
Artificial Analysis Coding Index74.371.1
baby_vision_with_python81.280.2
browsecomp88.582.7
charxiv_rq80.582.8
coding_arena_elo15391457
critpt20.923.4
cybergym83.166.3
deepsearchqa_f193.178.6
DeepSWE 1.15952
finance_agent53.957.2
frontiermath_tier_431.327.1
gdpval54.250.1
GDPval-AA (Elo)18901674
GPQA Diamond93.693
HiL-Bench35.39.7
HLE57.943.7
HLE (with tools)57.952.1
hmmt_feb_202696.797.7
hmmt_nov_202596.595.8
IFBench62.273.9
imo_answer_bench83.591.4
Legal Agent Benchmark100.4
livebench77.280.3
livebench_agentic_coding50.553.8
livebench_coding81.877.5
livebench_data_analysis6679.3
livebench_instruction_following7270.2
livebench_language79.782.6
livebench_math94.394.2
livebench_reasoning89.288.1
matharena_visual_math_overall81.692.5
mathvision86.792
MCP Atlas83.670.6
MMMU-Pro78.981.2
nl2repo69.741.3
officeqa77.668.1
officeqa_pro66.251.1
OmniScience Accuracy48.850.9
OmniScience Non-Hallucination60.717.4
OSWorld-Verified83.475
scicode53.556.6
simpleqa_verified39.545.3
SWE-bench Multilingual84.471.7
SWE-bench Pro69.259.1
TauBench V3 - Banking34.239.6
Terminal-Bench 2.074.675.1
Terminal-Bench 2.18578.3
Terminal-Bench Hard58.357.6
tool_decathlon59.954.6
toolathlon59.954.6
usamo_202696.795.2
τ²-Bench Telecom (AA run)94.487.1
τ³-Bench27.672.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.