VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Claude Opus 4.8

AnthropicvsAnthropic80 shared benchmarks2555 head-to-head
BenchmarkClaude Opus 4.6Claude Opus 4.8
AA Agentic Index67.649.4
AA Intelligence44.957.3
AA-LCR74.373
AA-Omniscience13.728.8
AgentWorldBench - Android61.761.5
AgentWorldBench - MCP69.954.9
AgentWorldBench - OS70.266.6
AgentWorldBench - Overall57.856.6
AgentWorldBench - Search29.335.1
AgentWorldBench - SWE64.564.1
AgentWorldBench - Terminal57.559.2
AgentWorldBench - Web51.454.7
aime_202696.7100
arc_agi_19392.5
arc_agi_30.51.5
ARC-AGI-268.872.1
arena_elo14971482
arena_text_factuality14871464
arena_vision13001280
Artificial Analysis Coding Index48.174.3
baby_vision_with_python38.481.2
browsecomp86.888.5
BrowseComp (multi-agent)86.688.5
BrowseComp (single-agent)83.788.5
charxiv_reasoning_with_tools84.789.9
charxiv_rq69.180.5
coding_arena_elo15451539
critpt12.620.9
cybergym73.883.1
deepsearchqa_f191.393.1
finance_agent76.753.9
Fortress20.518.2
frontiermath_tier_422.931.3
gdpval55.954.2
GDPval-AA (Elo)16191890
GPQA Diamond91.393.6
HiL-Bench38.335.3
HLE53.157.9
HLE (with tools)62.757.9
hmmt_feb_202696.296.7
hmmt_nov_202596.396.5
IFBench62.562.2
imo_answer_bench75.383.5
Legal Agent Benchmark4.210
livebench76.377.2
livebench_agentic_coding4950.5
livebench_coding78.281.8
livebench_data_analysis69.966
livebench_instruction_following63.372
livebench_language83.379.7
livebench_math89.394.3
livebench_reasoning88.789.2
matharena_visual_math_overall72.381.6
mathvision71.286.7
MCP Atlas76.883.6
MMMU-Pro77.378.9
nl2repo49.869.7
officeqa73.577.6
officeqa_pro57.166.2
OmniScience Accuracy4748.8
OmniScience Non-Hallucination37.260.7
OSWorld-Verified72.783.4
scicode5253.5
screenspot_pro_no_tools57.787.9
screenspot_pro_with_tools83.187.9
simplebench67.664.8
simpleqa_verified46.539.5
swe_bench_multimodal27.138.4
SWE-bench Multilingual77.884.4
SWE-bench Pro57.369.2
SWE-bench Verified80.888.6
Terminal-Bench 2.065.474.6
Terminal-Bench Hard65.458.3
tool_decathlon47.259.9
toolathlon56.859.9
Toolathlon Avg turns16.924.5
Toolathlon Pass@∞47.248.1
usamo_202666.296.7
τ²-Bench Telecom (AA run)99.394.4
τ³-Bench72.427.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.