VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs Gemini 3.1 Pro

DeepSeekvsGoogle60 shared benchmarks2138 head-to-head
BenchmarkDeepSeek-V4-FlashGemini 3.1 Pro
AA Agentic Index61.323
AA Intelligence51.847.7
AA-LCR74.379
AA-Omniscience-14.332.9
aime_202695.898.3
Apex (Pass@1)3360.9
Apex Shortlist (Pass@1)85.789.1
arc_agi_18498
ARC-AGI-24677.1
arena_elo14391486
arena_text_factuality14341471
Artificial Analysis Coding Index69.168.8
browsecomp73.285.9
Chinese SimpleQA (C-SimpleQA)78.985.9
Codeforces30523052
coding_arena_elo14311447
CorpusQA 1M (ACC)60.553.8
critpt16.617.7
cybergym76.738.8
DeepSWE54.410
gdpval52.923.3
GDPval-AA (Elo)13951317
GDPval-AA v21189962
Global-MMLU-Lite88.492.7
GPQA Diamond90.894.3
HLE38.651.4
HLE (with tools)45.151.6
hmmt_feb_202694.894.7
IFBench79.277.1
imo_answer_bench88.491
itbenchSre31.530.3
livebench_agentic_coding46.844.1
livebench_coding7576.5
livebench_data_analysis79.378.5
livebench_instruction_following65.579.1
livebench_language79.285.4
livebench_math86.891
livebench_reasoning86.684
LiveCodeBench91.691.7
MCP Atlas6978.2
MMLU-Pro86.491
MRCR 1M (MMR)78.776.3
nl2repo54.233.4
OmniScience Accuracy40.455.3
OmniScience Non-Hallucination15.650.1
scicode49.959
simplebench46.379.6
simpleqa_verified4577.3
strongreject97.498
SWE-bench Multilingual73.376.9
SWE-bench Pro52.654.2
SWE-bench Verified7980.6
Tau 3 Banking22.916.5
TauBench V3 - Banking39.421.4
Terminal-Bench 2.056.968.5
Terminal-Bench 2.182.774
Terminal-Bench Hard38.668.5
toolathlon70.348.8
τ²-Bench Telecom (AA run)95.699.3
τ³-Bench22.967.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.