VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Opus 5

AnthropicvsAnthropic41 shared benchmarks437 head-to-head
BenchmarkClaude Opus 4.6Opus 5
AA Agentic Index67.659.2
AA Intelligence44.963.1
AA-LCR74.378.7
AA-Omniscience13.737.1
arc_agi_19397.5
arc_agi_30.530.2
ARC-AGI-268.890.4
arena_elo14971487
arena_text_factuality14871481
Artificial Analysis Coding Index48.178
browsecomp86.890.8
coding_arena_elo15451662
critpt12.629.1
gdpval55.967.2
GDPval-AA (Elo)16191827
GPQA Diamond91.393.7
HiL-Bench38.357
HLE53.164.7
HLE (with tools)62.764.7
Legal Agent Benchmark4.223
livebench_agentic_coding4965.2
livebench_coding78.281.5
livebench_data_analysis69.974.5
livebench_instruction_following63.363.8
livebench_language83.388.7
livebench_math89.395.7
livebench_reasoning88.791.2
MCP Atlas76.885.8
MMMU-Pro77.384.7
officeqa73.578.1
officeqa_pro57.166.9
OmniScience Accuracy4760.9
OmniScience Non-Hallucination37.240.5
OSWorld-Verified72.74
scicode5255.7
simplebench67.680.6
simpleqa_verified46.556.7
swe_bench_multimodal27.159.4
SWE-bench Multilingual77.889.5
SWE-bench Pro57.379.2
SWE-bench Verified80.896

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.