VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Pro vs GPT-5.6 Sol

DeepSeekvsOpenAI49 shared benchmarks643 head-to-head
BenchmarkDeepSeek-V4-ProGPT-5.6 Sol
AA Agentic Index63.357.8
AA Intelligence5361
AA-LCR7077.7
AA-Omniscience-10.622
Agents' Last Exam16.552.7
aime_202696.799.9
arena_elo14571482
arena_text_factuality14501474
Artificial Analysis Coding Index59.478.3
browsecomp83.490.4
BrowseComp (w/ Ctx)83.489.4
coding_arena_elo15821619
critpt1332.3
cybergym52.783.6
DeepSWE12.873
FORTRESS (Adversarial)3682.4
FORTRESS (Benign)98.598.1
gdpval4961.1
GDPval-AA v213071748
Global-MMLU-Lite89.391.8
GPQA Diamond90.594.6
HLE48.249.5
HLE (with tools)48.258
IFBench76.572.7
itbenchSre38.356.2
livebench_agentic_coding42.656.2
livebench_coding7083.9
livebench_data_analysis74.579.8
livebench_instruction_following62.471.8
livebench_language78.187.7
livebench_math90.796.2
livebench_reasoning82.791.7
MCP Atlas74.283.6
OmniScience Accuracy4359.4
OmniScience Non-Hallucination12.210.6
Program Bench47.877.6
scicode5056.9
simplebench50.964.8
simpleqa_verified57.971.6
strongreject98.698.5
SWE-bench Pro55.464.6
SWEBench Pro (Public)55.464.6
Tau 3 Banking2633
TauBench V3 - Banking30.144.3
Terminal Bench 2.1 (Best Harness)6489.5
Terminal-Bench 2.172.189.5
Terminal-Bench Hard46.265.9
toolathlon51.858
τ²-Bench Telecom (AA run)96.285.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.