VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs DeepSeek-V4-Pro

DeepSeekvsDeepSeek64 shared benchmarks2638 head-to-head
BenchmarkDeepSeek-V4-FlashDeepSeek-V4-Pro
AA Agentic Index61.363.3
AA Intelligence51.853
AA-LCR74.370
AA-Omniscience-14.3-10.6
Agents' Last Exam25.216.5
aime_202695.896.7
Apex (Pass@1)3338.3
Apex Shortlist (Pass@1)85.790.2
arena_elo14391457
arena_text_factuality14341450
Artificial Analysis Coding Index69.159.4
AutomationBench (Public)25.112.8
AutomationBench Public25.112.8
browsecomp73.283.4
Chinese SimpleQA (C-SimpleQA)78.977.7
Codeforces30523206
coding_arena_elo14311582
CorpusQA 1M (ACC)60.562
critpt16.613
cybergym76.752.7
DeepSWE54.412.8
DSBench-FullStack68.741.8
DSBench-Hard59.631.1
gdpval52.949
GDPval-AA (Elo)13951554
GDPval-AA v211891307
Global-MMLU-Lite88.489.3
GPQA Diamond90.890.5
HLE38.648.2
HLE (with tools)45.148.2
HLE (wo / w tools)51.548.2
hmmt_feb_202694.895.2
IFBench79.276.5
imo_answer_bench88.489.8
itbenchSre31.538.3
livebench_agentic_coding46.842.6
livebench_coding7570
livebench_data_analysis79.374.5
livebench_instruction_following65.562.4
livebench_language79.278.1
livebench_math86.890.7
livebench_reasoning86.682.7
LiveCodeBench91.693.5
MCP Atlas6974.2
MMLU-Pro86.487.5
MRCR 1M (MMR)78.783.5
nl2repo54.238.5
OmniScience Accuracy40.443
OmniScience Non-Hallucination15.612.2
scicode49.950
simplebench46.350.9
simpleqa_verified4557.9
strongreject97.498.6
SWE-bench Multilingual73.376.2
SWE-bench Pro52.655.4
SWE-bench Verified7980.6
Tau 3 Banking22.926
TauBench V3 - Banking39.430.1
Terminal-Bench 2.056.967.9
Terminal-Bench 2.182.772.1
Terminal-Bench Hard38.646.2
toolathlon70.351.8
τ²-Bench Telecom (AA run)95.696.2
τ³-Bench22.925.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.