DeepSeek-V4-Flash
DeepSeek-V4-Flash is capable in long context, reasoning, agentic tasks, and instruction following; and behind the leaders in coding and math. Too few results yet to rate safety, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multimodal, Multilingual or Factuality.
Price
$0.44input$1.32outputper million tokens
From Artificial Analysis · 7 providers tracked · All prices
Evidence
258results on85benchmarks
- 33 independently verified
- 110 aggregator
- 66 vendor-reported
- 49 cross-referenced
From 24 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
DeepSeek-V4-Flash benchmark results
258 results on 85 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
11.7% behind the leader1 of 3 ranked benchmarks measured
- 79.67Oct 8, 2026
17.3% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond90.81Oct 8, 2026gpqa
- LiveBench · Reasoning86.63Oct 8, 2026livebench_reasoning@2026-06-25
- 61.39Sep 21, 2026
- Humanity's Last Exam38.55Oct 8, 2026aa_hle
- 46.30Jun 3, 2026
- 16.57Oct 8, 2026
Show 29 more reasoning resultsHide 29 reasoning results
- 2.08Sep 22, 2026
- 55.97Sep 21, 2026
- 7.14Oct 8, 2026
- 3.43Oct 8, 2026
- 0.29Oct 8, 2026
- 3.43Sep 9, 2026
- 7.14Sep 9, 2026
- 7.10Jul 30, 2026
- GPQA Diamond89.39Oct 8, 2026gpqa
- GPQA Diamond71.62Oct 8, 2026gpqa
- GPQA Diamond86.67Oct 8, 2026gpqa
- GPQA Diamond86.67Sep 9, 2026gpqa
- GPQA Diamond89.39Sep 9, 2026gpqa
- GPQA Diamond88.10Jun 27, 2026GPQA Diamond (Pass@1)
- GPQA Diamond71.20Jun 23, 2026GPQA Diamond (Pass@1)
- 90.80Aug 26, 2026
- 89.40Jul 30, 2026
- Humanity's Last Exam34.85Oct 8, 2026aa_hle
- Humanity's Last Exam30.26Oct 8, 2026aa_hle
- Humanity's Last Exam7.78Oct 8, 2026aa_hle
- Humanity's Last Exam30.26Sep 9, 2026aa_hle
- Humanity's Last Exam34.85Sep 9, 2026aa_hle
- Humanity's Last Exam36.84Jul 31, 2026aa_hle
- Humanity's Last Exam34.80Jun 27, 2026HLE (Pass@1)
- Humanity's Last Exam8.10Jun 23, 2026HLE (Pass@1)
- 33.80Aug 26, 2026
- Humanity's Last Exam32.10Jul 30, 2026HLE text only
- Humanity's Last Exam32.20Jun 6, 2026HLE (no tools)
- LiveBench · Reasoning70.58Oct 8, 2026livebench_reasoning@2026-06-25
22.1% behind the leader6 of 7 ranked benchmarks measured
- BrowseComp73.20Jun 27, 2026BrowseComp (Pass@1)
- MCP Atlas69.00Jul 4, 2026MCPAtlas (Pass@1)
- τ-Bench V3 · Banking39.38Oct 8, 2026tauBanking
- 46.91Oct 8, 2026
- Terminal-Bench 4.012.12Oct 8, 2026
Show 21 more agentic resultsHide 21 agentic results
- 31.52Oct 8, 2026
- 31.52Sep 9, 2026
- 46.90Jun 6, 2026
- 26.95Oct 8, 2026
- 25.40Oct 8, 2026
- 29.13Sep 9, 2026
- 30.78Sep 9, 2026
- 52.93Jul 31, 2026
- 44.77Jun 15, 2026
- GDPval (win rate)50.20Jun 6, 2026GDPVal
- MCP Atlas64.00Jun 23, 2026MCPAtlas (Pass@1)
- Terminal-Bench 4.02.53Oct 8, 2026
- Terminal-Bench 4.03.03Oct 8, 2026
- Terminal-Bench 4.03.03Sep 9, 2026
- Terminal-Bench 4.02.53Sep 9, 2026
- τ-Bench V3 · Banking30.93Oct 8, 2026tauBanking
- τ-Bench V3 · Banking26.19Oct 8, 2026tauBanking
- τ-Bench V3 · Banking26.19Sep 9, 2026tauBanking
- τ-Bench V3 · Banking30.93Sep 9, 2026tauBanking
- τ-Bench V3 · Banking31.13Jul 31, 2026tauBanking
- 26.70Jun 6, 2026
23.6% behind the leader2 of 3 ranked benchmarks measured
- IFBench79.18Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following65.52Oct 8, 2026livebench_instruction_following@2026-06-25
Show 7 more instruction following resultsHide 7 instruction following results
- IFBench73.47Oct 8, 2026aa_ifbench
- IFBench47.21Oct 8, 2026aa_ifbench
- IFBench79.18Sep 9, 2026aa_ifbench
- IFBench73.47Sep 9, 2026aa_ifbench
- 79.20Aug 26, 2026
- 79.20Aug 24, 2026
- LiveBench · Instruction Following63.14Oct 8, 2026livebench_instruction_following@2026-06-25
25.9% behind the leader10 of 10 ranked benchmarks measured
- 90.60Aug 26, 2026
- Terminal-Bench 2.178.65Oct 8, 2026terminalbenchV21
- SWE-bench Verified79.00Jun 27, 2026SWE Verified (Resolved)
- LiveBench · Coding74.98Oct 8, 2026livebench_coding@2026-06-25
- SWE-bench Multilingual73.30Jun 27, 2026SWE Multilingual (Resolved)
- SciCode50.35Oct 8, 2026aa_scicode
- 1430.29Sep 21, 2026
- LiveBench · Agentic Coding46.77Oct 8, 2026livebench_agentic_coding@2026-06-25
- SWE-bench Pro52.60Jun 27, 2026SWE Pro (Resolved)
- Terminal-Bench Hard35.61Oct 8, 2026aa_terminalbench_hard
Show 29 more coding resultsHide 29 coding results
- LiveBench · Agentic Coding37.63Oct 8, 2026livebench_agentic_coding@2026-06-25
- LiveBench · Coding69.23Oct 8, 2026livebench_coding@2026-06-25
- LiveCodeBench v690.90Jun 6, 2026LiveCodeBench (v6)
- SciCode45.25Oct 8, 2026aa_scicode
- SciCode40.16Oct 8, 2026aa_scicode
- SciCode40.16Sep 9, 2026aa_scicode
- SciCode45.25Sep 9, 2026aa_scicode
- SciCode37.27Sep 4, 2026aa_scicode
- SciCode49.88Jul 31, 2026aa_scicode
- 44.90Jul 30, 2026
- SWE-bench Multilingual69.70Jun 23, 2026SWE Multilingual (Resolved)
- 72.10Jun 6, 2026
- SWE-bench Pro49.10Jun 23, 2026SWE Pro (Resolved)
- 56.00Aug 26, 2026
- SWE-bench Verified73.70Jun 23, 2026SWE Verified (Resolved)
- 79.00Jul 30, 2026
- 72.40Jun 6, 2026
- Terminal-Bench 2.161.80Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.156.93Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.156.93Sep 9, 2026terminalbenchV21
- Terminal-Bench 2.161.80Sep 9, 2026terminalbenchV21
- 82.70Oct 7, 2026
- 61.80Aug 13, 2026
- 61.80Jul 30, 2026
- 54.20Jun 6, 2026
- Terminal-Bench Hard34.09Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard38.64Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard35.61Sep 9, 2026aa_terminalbench_hard
- Terminal-Bench Hard38.64Sep 9, 2026aa_terminalbench_hard
48.5% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics86.79Oct 8, 2026livebench_math@2026-06-25
- 57.54Sep 21, 2026
- 24.39Sep 21, 2026
Show 11 more math resultsHide 11 math results
- 95.83Sep 2, 2026
- 95.83Jun 25, 2026
- 95.80Jul 30, 2026
- 93.94Sep 2, 2026
- 93.94Aug 22, 2026
- HMMT Feb 202694.80Jun 27, 2026HMMT 2026 Feb (Pass@1)
- HMMT Feb 202640.80Jun 23, 2026HMMT 2026 Feb (Pass@1)
- 93.90Jul 30, 2026
- IMO-AnswerBench88.40Jun 27, 2026IMOAnswerBench (Pass@1)
- IMO-AnswerBench41.90Jun 23, 2026IMOAnswerBench (Pass@1)
- LiveBench · Mathematics79.65Oct 8, 2026livebench_math@2026-06-25
0 of 4 ranked benchmarks measured
- 33.63Sep 21, 2026
- AA-Omniscience · Accuracy40.38Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy36.82Oct 8, 2026omniscienceAccuracy
Show 17 more factuality resultsHide 17 factuality results
- AA-Omniscience · Accuracy34.87Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy26.00Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy34.87Sep 9, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy36.82Sep 9, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy37.23Jul 31, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.30Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination3.90Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination5.45Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination11.03Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination11.03Sep 9, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination3.90Sep 9, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination15.64Jul 31, 2026omniscienceNonHallucination
- SimpleQA Verified34.10Jun 27, 2026SimpleQA-Verified (Pass@1)
- SimpleQA Verified23.10Jun 23, 2026SimpleQA-Verified (Pass@1)
- SimpleQA Verified28.90Jun 6, 2026SimpleQA-Verified (Pass@1)
- SimpleQA Verified45.00Jun 6, 2026SimpleQA-Verified (Pass@1)
- 34.10Aug 24, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language79.18Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis79.33Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language70.12Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis68.02Oct 8, 2026livebench_data_analysis@2026-06-25
- AA Intelligence24.00Oct 1, 2026Artificial Analysis Intelligence Index
- 11.83Sep 22, 2026
Show 102 more resultsHide 102 results
- 41.68Sep 9, 2026
- 27.87Sep 9, 2026
- 23.74Sep 9, 2026
- 45.67Jul 31, 2026
- 61.33Jun 18, 2026
- AA Intelligence40.00Jul 9, 2026Artificial Analysis Intelligence Index
- AA Intelligence24.17Oct 8, 2026aa_intelligence_index
- AA Intelligence34.33Oct 8, 2026aa_intelligence_index
- AA Intelligence24.38Oct 8, 2026aa_intelligence_index
- AA Intelligence18.88Oct 8, 2026aa_intelligence_index
- AA Intelligence24.84Sep 9, 2026aa_intelligence_index
- AA Intelligence24.63Sep 9, 2026aa_intelligence_index
- AA Intelligence49.93Jul 31, 2026aa_intelligence_index
- -23.90Oct 8, 2026
- -14.28Oct 8, 2026
- -23.08Oct 8, 2026
- -43.97Oct 8, 2026
- -23.08Sep 9, 2026
- -23.90Sep 9, 2026
- -15.72Jul 31, 2026
- 25.20Oct 7, 2026
- 15.80Aug 13, 2026
- 25.20Aug 26, 2026
- 33.00Jun 27, 2026
- 1.00Jun 23, 2026
- 85.70Jun 27, 2026
- 9.30Jun 23, 2026
- 82.40Jun 6, 2026
- 82.00Jun 6, 2026
- 87.00Sep 21, 2026
- 89.00Sep 21, 2026
- 84.00Aug 9, 2026
- 1439Jul 23, 2026
- Artificial Analysis Coding Index69.06Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index51.96Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index56.17Sep 9, 2026aa_coding_index
- 35.15Jun 18, 2026
- AutomationBench Public25.10Aug 31, 2026AutomationBench (Public)
- AutomationBench Public10.80Aug 13, 2026AutomationBench (Public)
- BrowseComp (context management)73.20Aug 24, 2026BrowseComp (with context management)
- Chinese SimpleQA (C-SimpleQA)78.90Jul 4, 2026Chinese-SimpleQA (Pass@1)
- Chinese SimpleQA (C-SimpleQA)71.50Jun 23, 2026Chinese-SimpleQA (Pass@1)
- Chinese SimpleQA (C-SimpleQA)73.20Jun 6, 2026Chinese-SimpleQA (Pass@1)
- Chinese SimpleQA (C-SimpleQA)75.80Jun 6, 2026Chinese-SimpleQA (Pass@1)
- 3052.00Jun 27, 2026
- 60.50Jun 27, 2026
- 15.50Jun 23, 2026
- 10.60Jun 6, 2026
- 76.70Oct 7, 2026
- 38.70Aug 13, 2026
- 54.40Oct 7, 2026
- 7.30Aug 13, 2026
- 54.40Aug 26, 2026
- 68.70Oct 7, 2026
- 37.00Aug 13, 2026
- 59.60Oct 7, 2026
- 25.80Aug 13, 2026
- 1395.00Jun 27, 2026
- 1189.00Jul 30, 2026
- 88.40Aug 24, 2026
- GPQA (unspecified)88.50Jun 6, 2026GPQA (no tools)
- HLE (with tools)45.10Jun 27, 2026HLE w/ tools (Pass@1)
- HLE (with tools)45.10Jul 30, 2026HLE with tools
- 45.10Aug 13, 2026
- 51.50Aug 13, 2026
- 89.60Jun 6, 2026
- 41.30Aug 26, 2026
- LiveCodeBench91.60Jun 27, 2026LiveCodeBench (Pass@1)
- LiveCodeBench55.20Jun 23, 2026LiveCodeBench (Pass@1)
- MMLU-Pro86.20Jul 4, 2026MMLU-Pro (EM)
- MMLU-Pro83.00Jun 23, 2026MMLU-Pro (EM)
- MMLU-Pro86.40Jun 6, 2026MMLU-Pro (EM)
- 86.40Jun 6, 2026
- 54.20Oct 7, 2026
- 39.40Aug 13, 2026
- OpenAI-MRCR (1M)78.70Jun 27, 2026MRCR 1M (MMR)
- OpenAI-MRCR (1M)37.50Jun 23, 2026MRCR 1M (MMR)
- 91.30Jun 6, 2026
- 57.00Jun 6, 2026
- 48.20Jun 6, 2026
- 97.40Aug 24, 2026
- 52.60Jul 30, 2026
- 80.80Jun 6, 2026
- 73.70Jun 6, 2026
- 89.10Jun 6, 2026
- Terminal-Bench 2.056.90Jun 27, 2026Terminal Bench 2.0 (Acc)
- Terminal-Bench 2.049.10Jun 23, 2026Terminal Bench 2.0 (Acc)
- 70.30Oct 7, 2026
- Toolathlon47.80Jun 27, 2026Toolathlon (Pass@1)
- Toolathlon40.70Jun 23, 2026Toolathlon (Pass@1)
- 70.30Aug 31, 2026
- 49.70Aug 13, 2026
- Toolathlon Verified70.30Aug 26, 2026Toolathlon Verified (Pass@1)
- 50.90Jul 30, 2026
- 60.10Jun 6, 2026
- 58.40Jun 6, 2026
- τ²-Bench Telecom (AA run)95.03Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.61Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)94.44Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.03Sep 9, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.61Sep 9, 2026aa_tau2
- τ³-Bench Banking22.90Jul 30, 2026τ³-Banking
DeepSeek-V4-Flash: common questions
Who makes DeepSeek-V4-Flash?
DeepSeek-V4-Flash is made by DeepSeek.
When was DeepSeek-V4-Flash released?
DeepSeek-V4-Flash was released on Apr 24, 2026, according to Artificial Analysis.
What is DeepSeek-V4-Flash good at?
DeepSeek-V4-Flash is capable in long context, reasoning, agentic tasks, and instruction following; and behind the leaders in coding and math. Too few results yet to rate safety, multimodal tasks, multilingual tasks, or factuality.
How much does DeepSeek-V4-Flash cost?
DeepSeek-V4-Flash costs $0.44 per million input tokens and $1.32 per million output tokens, according to Artificial Analysis. We track its price at 7 providers. At a mix of three input tokens to one output token, it is cheaper than 52% of the 330 priced models we track.
How many benchmarks has DeepSeek-V4-Flash been tested on?
We track 258 results for DeepSeek-V4-Flash on 85 benchmarks from 24 sources, 33 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-V4-Flash support?
OpenRouter lists tool calling, structured outputs, and reasoning for DeepSeek-V4-Flash.
About this record
Where DeepSeek-V4-Flash's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Oct 8, 2026
Where the results come from
Verification: 258 scores · 33 independently verified · 110 aggregator-attributed · 49 vendor cross-reference · 66 vendor-reported. How these tiers are assigned
From 24 sources on 11 sites. Artificial Analysis supplies 112 of them; the 33 independently verified results come from 8 sites. Bars are coloured by trust tier.
- artificialanalysis.ai112
- huggingface.co94
- livebench.ai14
- thinkingmachines.ai13
- api.llm-stats.com8
- arcprize.org7
- matharena.ai4
- epoch.ai3
- datasets-server.huggingface.co1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name DeepSeek-V4-Flash at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | DeepSeek V4 Flash 0731 (Low) | — | — |
| high | deepseek v4 flash 0731 (high) deepseek-v4-flash high deepseek v4 flash (high) | deepseek v4 flash (reasoning, high effort) deepseek v4 flash 0420 (reasoning, high effort) | deepseek-v4-flash-0420-high deepseek-v4-flash-high-preview |
| max | deepseek v4 flash 0731 (max) deepseek v4 flash max v4-flash max deepseek v4 flash vision (max) | deepseek v4 flash (reasoning, max effort) deepseek v4 flash 0731 (reasoning, max effort) deepseek v4 flash vision (reasoning, max effort) | deepseek-v4-flash (max) |