VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.7 vs Claude Sonnet 5

AnthropicvsAnthropic44 shared benchmarks3113 head-to-head
BenchmarkClaude Opus 4.7Claude Sonnet 5
AA Agentic Index64.649.7
AA Intelligence5555.3
AA-LCR75.377
AA-Omniscience27.316.4
arena_elo14941463
arena_text_factuality14741453
Artificial Analysis Coding Index73.671.5
browsecomp79.884.7
charxiv_reasoning_with_tools9188.3
charxiv_rq82.177
coding_arena_elo15571540
critpt1216.9
gdpval58.654.8
GPQA Diamond94.291.1
HLE54.757.4
Legal Agent Benchmark7.15.8
livebench_agentic_coding50.759.4
livebench_coding82.180.7
livebench_data_analysis78.371.7
livebench_instruction_following66.763.9
livebench_language77.975
livebench_math92.892.9
livebench_reasoning87.288.7
MMMU-Pro78.877.3
mrcr_v2_8needle_128k_average59.381.5
officeqa86.373.3
officeqa_pro80.659.4
OmniScience Accuracy48.940
OmniScience Non-Hallucination57.760.6
OSWorld-Verified82.881.2
scicode54.553.6
simplebench61.760.6
simpleqa_verified50.625
swe_bench_multimodal34.528.1
SWE-bench Multilingual80.578.3
SWE-bench Pro64.363.2
SWE-bench Verified87.685.2
TauBench V3 - Banking34.637.3
Terminal-Bench 2.069.480.4
Terminal-Bench 2.183.180.5
toolathlon59.354.3
Toolathlon Pass@∞52.840.7
usamo_202669.379.5
τ³-Bench28.928.2

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.