VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Pro vs Gemini 3.1 Pro

DeepSeekvsGoogle82 shared benchmarks3150 head-to-head
BenchmarkDeepSeek-V4-ProGemini 3.1 Pro
AA Agentic Index63.323
AA Intelligence5347.7
AA-LCR7079
AA-Omniscience-10.632.9
AgentWorldBench - Android55.261.4
AgentWorldBench - MCP63.359.1
AgentWorldBench - OS63.766.9
AgentWorldBench - Overall5354.6
AgentWorldBench - Search27.630.2
AgentWorldBench - SWE59.459.1
AgentWorldBench - Terminal51.352.5
AgentWorldBench - Web50.352.8
aime_202696.798.3
Apex (Pass@1)38.360.9
Apex Shortlist (Pass@1)90.289.1
apexAgents24.333.5
arena_elo14571486
arena_text_factuality14501471
Artificial Analysis Coding Index59.468.8
browsecomp83.485.9
BrowseComp (w/ Ctx)83.485.9
Chinese SimpleQA (C-SimpleQA)77.785.9
Codeforces32063052
coding_arena_elo15821447
CorpusQA 1M (ACC)6253.8
critpt1317.7
cybergym52.738.8
DeepSWE12.810
FORTRESS (Adversarial)3665.2
FORTRESS (Benign)98.598
gdpval4923.3
GDPval-AA (Elo)15541317
GDPval-AA v21307962
Global-MMLU-Lite89.392.7
GPQA Diamond90.594.3
HLE48.251.4
HLE (with tools)48.251.6
hmmt_feb_202695.294.7
hmmt_nov_202594.494.8
IFBench76.577.1
imo_answer_bench89.891
itbenchSre38.330.3
livebench73.679.9
livebench_agentic_coding42.644.1
livebench_coding7076.5
livebench_data_analysis74.578.5
livebench_instruction_following62.479.1
livebench_language78.185.4
livebench_math90.791
livebench_reasoning82.784
LiveCodeBench93.591.7
MCP Atlas74.278.2
MCPAtlas Public (Pass@1)73.669.2
MMLU-Pro87.591
MRCR 1M (MMR)83.576.3
nl2repo38.533.4
OmniScience Accuracy4355.3
OmniScience Non-Hallucination12.250.1
Program Bench47.839.5
scicode5059
simplebench50.979.6
simpleqa_verified57.977.3
strongreject98.698
SWE-bench Multilingual76.276.9
SWE-bench Pro55.454.2
SWE-bench Verified80.680.6
SWEBench Pro (Public)55.454.2
Tau 3 Banking2616.5
TauBench V3 - Banking30.121.4
Terminal Bench 2.1 (Best Harness)6473.8
Terminal-Bench 2.067.968.5
Terminal-Bench 2.172.174
Terminal-Bench Hard46.268.5
tool_decathlon52.848.8
toolathlon51.848.8
usamo_202660.774.4
vectara_answer_rate97.299.4
vectara_avg_summary_length153.8107.7
vectara_factual_consistency91.489.6
vectara_hallucination_rate8.610.4
τ²-Bench Telecom (AA run)96.299.3
τ³-Bench25.867.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.