VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Claude Sonnet 4.6

AnthropicvsAnthropic67 shared benchmarks5116 head-to-head
BenchmarkClaude Opus 4.6Claude Sonnet 4.6
AA Agentic Index67.661.6
AA Intelligence44.948.4
AA-LCR74.374
AA-Omniscience13.712.2
AgentWorldBench - Android61.758
AgentWorldBench - MCP69.970
AgentWorldBench - OS70.263.2
AgentWorldBench - Overall57.856
AgentWorldBench - Search29.328.8
AgentWorldBench - SWE64.564.5
AgentWorldBench - Terminal57.557
AgentWorldBench - Web51.450.8
apexAgents3328
arc_agi_19386.5
ARC-AGI-268.860.4
arena_elo14971472
arena_text_factuality14871460
arena_vision13001278
Artificial Analysis Coding Index48.163
browsecomp86.876.2
charxiv_reasoning_with_tools84.785.3
charxiv_rq69.172.4
coding_arena_elo15451523
critpt12.63.1
finance_agent76.763.3
frontiermath_tier_422.98.3
gdpval55.954.8
GDPval-AA (Elo)16191676
GPQA Diamond91.389.9
HLE53.149
HLE (with tools)62.746.8
IFBench62.556.6
Legal Agent Benchmark4.25.4
livebench76.375.5
livebench_agentic_coding4942.6
livebench_coding78.279.3
livebench_data_analysis69.978
livebench_instruction_following63.363.2
livebench_language83.376.1
livebench_math89.387
livebench_reasoning88.784.8
MCP Atlas76.869.5
mmmlu91.189.3
MMMU-Pro77.375.6
mrcr_v2_8needle_128k_average8484.9
officeqa73.568.7
officeqa_pro57.153.4
OmniScience Accuracy4740.9
OmniScience Non-Hallucination37.251.6
OSWorld-Verified72.778.5
scicode5247
simpleqa_verified46.529
SWE-bench Pro57.358.1
SWE-bench Verified80.879.6
Tau2 retail91.991.7
Terminal-Bench 2.065.459.1
Terminal-Bench Hard65.459.1
toolathlon56.849.4
Toolathlon Avg turns16.916.5
Toolathlon Pass@∞47.238
usamo_202666.255
vectara_answer_rate99.899.9
vectara_avg_summary_length137.6114.7
vectara_factual_consistency87.889.4
vectara_hallucination_rate12.210.6
τ²-Bench Telecom (AA run)99.397.9
τ³-Bench72.430.5

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.