VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.7 vs Claude Opus 4.8

AnthropicvsAnthropic68 shared benchmarks2543 head-to-head
BenchmarkClaude Opus 4.7Claude Opus 4.8
AA Agentic Index64.649.4
AA Intelligence5557.3
AA-LCR75.373
AA-Omniscience27.328.8
aime_202695.8100
arc_agi_19292.5
arc_agi_30.21.5
ARC-AGI-275.872.1
arena_elo14941482
arena_text_factuality14741464
arena_vision13041280
Artificial Analysis Coding Index73.674.3
browsecomp79.888.5
ChartQAPro (No tools)67.669.4
ChartQAPro (With tools)69.872.3
charxiv_reasoning_with_tools9189.9
charxiv_rq82.180.5
coding_arena_elo15571539
critpt1220.9
cybergym73.883.1
Finance Agent v251.553.9
finance_agent71.553.9
frontiermath_tier_422.931.3
gdpval58.654.2
GDPval-AA (Elo)17531890
GPQA Diamond94.293.6
GraphWalks — Parents task, 256K-token context subset93.699.3
GraphWalks BFS 256K76.985.9
HiL-Bench41.735.3
HLE54.757.9
HLE (with tools)54.757.9
hmmt_feb_202693.996.7
IFBench58.662.2
Legal Agent Benchmark7.110
livebench76.977.2
livebench_agentic_coding50.750.5
livebench_coding82.181.8
livebench_data_analysis78.366
livebench_instruction_following66.772
livebench_language77.979.7
livebench_math92.894.3
livebench_reasoning87.289.2
MCP Atlas79.183.6
MMMU-Pro78.878.9
officeqa86.377.6
officeqa_pro80.666.2
OmniScience Accuracy48.948.8
OmniScience Non-Hallucination57.760.7
OSWorld-Verified82.883.4
scicode54.553.5
screenspot_pro_no_tools79.587.9
screenspot_pro_with_tools87.687.9
simplebench61.764.8
simpleqa_verified50.639.5
swe_bench_multimodal34.538.4
SWE-bench Multilingual80.584.4
SWE-bench Pro64.369.2
SWE-bench Verified87.688.6
TauBench V3 - Banking34.634.2
Terminal-Bench 2.069.474.6
Terminal-Bench 2.183.185
Terminal-Bench Hard54.558.3
toolathlon59.359.9
Toolathlon Avg turns25.924.5
Toolathlon Pass@∞52.848.1
usamo_202669.396.7
τ²-Bench Telecom (AA run)88.694.4
τ³-Bench28.927.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.