VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Claude Opus 4.7

AnthropicvsAnthropic69 shared benchmarks2048 head-to-head
BenchmarkClaude Opus 4.6Claude Opus 4.7
AA Agentic Index67.664.6
AA Intelligence44.955
AA-LCR74.375.3
AA-Omniscience13.727.3
aime_202696.795.8
arc_agi_19392
arc_agi_30.50.2
ARC-AGI-268.875.8
arena_elo14971494
arena_text_factuality14871474
arena_vision13001304
Artificial Analysis Coding Index48.173.6
BigLaw Bench90.290.9
browsecomp86.879.8
charxiv_reasoning_with_tools84.791
charxiv_rq69.182.1
coding_arena_elo15451557
critpt12.612
cybergym73.873.8
finance_agent76.771.5
frontiermath_tier_422.922.9
gdpval55.958.6
GDPval-AA (Elo)16191753
GPQA Diamond91.394.2
HiL-Bench38.341.7
HLE53.154.7
HLE (with tools)62.754.7
hmmt_feb_202696.293.9
IFBench62.558.6
Legal Agent Benchmark4.27.1
livebench76.376.9
livebench_agentic_coding4950.7
livebench_coding78.282.1
livebench_data_analysis69.978.3
livebench_instruction_following63.366.7
livebench_language83.377.9
livebench_math89.392.8
livebench_reasoning88.787.2
MCP Atlas76.879.1
mmmlu91.191.5
MMMU-Pro77.378.8
mrcr_v2_8needle_128k_average8459.3
officeqa73.586.3
officeqa_pro57.180.6
OmniScience Accuracy4748.9
OmniScience Non-Hallucination37.257.7
OSWorld-Verified72.782.8
scicode5254.5
screenspot_pro_no_tools57.779.5
screenspot_pro_with_tools83.187.6
simplebench67.661.7
simpleqa_verified46.550.6
swe_bench_multimodal27.134.5
SWE-bench Multilingual77.880.5
SWE-bench Pro57.364.3
SWE-bench Verified80.887.6
Terminal-Bench 2.065.469.4
Terminal-Bench Hard65.454.5
toolathlon56.859.3
Toolathlon Avg turns16.925.9
Toolathlon Pass@∞47.252.8
usamo_202666.269.3
vectara_answer_rate99.898
vectara_avg_summary_length137.6149.1
vectara_factual_consistency87.888
vectara_hallucination_rate12.212
XBOW visual-acuity benchmark54.598.5
τ²-Bench Telecom (AA run)99.388.6
τ³-Bench72.428.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.