VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs GPT-5.4

AnthropicvsOpenAI92 shared benchmarks4052 head-to-head
BenchmarkClaude Opus 4.6GPT-5.4
AA Agentic Index67.658.2
AA Intelligence44.953.1
AA-LCR74.377.7
AA-Omniscience13.75.8
AgentWorldBench - Android61.760
AgentWorldBench - MCP69.970.1
AgentWorldBench - OS70.268.6
AgentWorldBench - Overall57.858.3
AgentWorldBench - Search29.337.3
AgentWorldBench - SWE64.566.3
AgentWorldBench - Terminal57.553.7
AgentWorldBench - Web51.451.8
aime_202696.799.2
Apex (Pass@1)34.554.1
Apex Shortlist (Pass@1)85.978.1
apexAgents3333.3
arc_agi_19393.7
arc_agi_30.50.2
ARC-AGI-268.874
arena_elo14971477
arena_text_factuality14871476
arena_vision13001277
Artificial Analysis Coding Index48.171.1
baby_vision14.849.7
baby_vision_with_python38.480.2
browsecomp86.882.7
BrowseComp (Agent Swarm)83.782.7
browsecomp_with_context_manager8482.7
charxiv_rq69.182.8
charxiv_rq_with_python84.790
Chinese SimpleQA (C-SimpleQA)76.476.8
coding_arena_elo15451457
critpt12.623.4
cybergym73.866.3
deepsearchqa_accuracy80.663.7
deepsearchqa_f191.378.6
finance_agent76.757.2
frontiermath_tier_422.927.1
gdpval55.950.1
GDPval-AA (Elo)16191674
GPQA Diamond91.393
HiL-Bench38.39.7
HLE53.143.7
HLE (with tools)62.752.1
hmmt_feb_202696.297.7
hmmt_nov_202596.395.8
IFBench62.573.9
imo_answer_bench75.391.4
Legal Agent Benchmark4.20.4
livebench76.380.3
livebench_agentic_coding4953.8
livebench_coding78.277.5
livebench_data_analysis69.979.3
livebench_instruction_following63.370.2
livebench_language83.382.6
livebench_math89.394.2
livebench_reasoning88.788.1
matharena_visual_math_overall72.392.5
mathvision71.292
mathvision_with_python84.696.1
MCP Atlas76.870.6
MCPAtlas Public (Pass@1)73.867.2
mcpmark56.762.5
MMLU-Pro89.187.5
mmmu_pro_with_python77.382.1
MMMU-Pro77.381.2
MultiNRC57.158.3
nl2repo49.841.3
officeqa73.568.1
officeqa_pro57.151.1
OmniDocBench 1.586.689.1
OmniScience Accuracy4750.9
OmniScience Non-Hallucination37.217.4
OSWorld-Verified72.775
scicode5256.6
SEAL VISTA46.150.9
simpleqa_verified46.545.3
SWE-bench Multilingual77.871.7
SWE-bench Pro57.359.1
Terminal-Bench 2.065.475.1
Terminal-Bench Hard65.457.6
tool_decathlon47.254.6
toolathlon56.854.6
usamo_202666.295.2
V* (w/ python)86.498.4
vectara_answer_rate99.899.9
vectara_avg_summary_length137.681.7
vectara_factual_consistency87.893
vectara_hallucination_rate12.27
VisualToolBench27.529.2
τ²-Bench Telecom (AA run)99.387.1
τ³-Bench72.472.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.