VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4.6 vs DeepSeek-V4-Pro

AnthropicvsDeepSeek54 shared benchmarks3023 head-to-head
BenchmarkClaude Sonnet 4.6DeepSeek-V4-Pro
AA Agentic Index61.663.3
AA Intelligence48.453
AA-LCR7470
AA-Omniscience12.2-10.6
AgentWorldBench - Android5855.2
AgentWorldBench - MCP7063.3
AgentWorldBench - OS63.263.7
AgentWorldBench - Overall5653
AgentWorldBench - Search28.827.6
AgentWorldBench - SWE64.559.4
AgentWorldBench - Terminal5751.3
AgentWorldBench - Web50.850.3
apexAgents2824.3
arena_elo14721457
arena_text_factuality14601450
Artificial Analysis Coding Index6359.4
browsecomp76.283.4
coding_arena_elo15231582
critpt3.113
gdpval54.849
GDPval-AA (Elo)16761554
GDPval-AA v213811307
GPQA Diamond89.990.5
HLE4948.2
HLE (with tools)46.848.2
IFBench56.676.5
itbenchSre39.838.3
livebench75.573.6
livebench_agentic_coding42.642.6
livebench_coding79.370
livebench_data_analysis7874.5
livebench_instruction_following63.262.4
livebench_language76.178.1
livebench_math8790.7
livebench_reasoning84.882.7
MCP Atlas69.574.2
OmniScience Accuracy40.943
OmniScience Non-Hallucination51.612.2
scicode4750
simpleqa_verified2957.9
SWE-bench Pro58.155.4
SWE-bench Verified79.680.6
TauBench V3 - Banking34.430.1
Terminal-Bench 2.059.167.9
Terminal-Bench 2.171.272.1
Terminal-Bench Hard59.146.2
toolathlon49.451.8
usamo_20265560.7
vectara_answer_rate99.997.2
vectara_avg_summary_length114.7153.8
vectara_factual_consistency89.491.4
vectara_hallucination_rate10.68.6
τ²-Bench Telecom (AA run)97.996.2
τ³-Bench30.525.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.