DeepSeek-V4-Pro
DeepSeek-V4-Pro is capable in long context, reasoning, and agentic tasks; and behind the leaders in coding, instruction following, factuality, and math. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multimodal or Multilingual.
Price
$0.43input$0.87outputper million tokens
From Artificial Analysis · 5 providers tracked · All prices
Evidence
256results on111benchmarks
- 27 independently verified
- 84 aggregator
- 86 vendor-reported
- 59 cross-referenced
From 23 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
DeepSeek-V4-Pro benchmark results
256 results on 111 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
19.1% behind the leader1 of 3 ranked benchmarks measured
- 74.67Oct 8, 2026
22.1% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond88.79Oct 8, 2026gpqa
- LiveBench · Reasoning82.69Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam37.53Oct 8, 2026aa_hle
- 50.90Jun 28, 2026
- 12.86Oct 8, 2026
Show 23 more reasoning resultsHide 23 reasoning results
- 10.00Oct 8, 2026
- 0.86Oct 8, 2026
- 12.86Sep 9, 2026
- 10.00Sep 9, 2026
- 12.90Sep 2, 2026
- 13.00Jul 30, 2026
- GPQA Diamond71.72Oct 8, 2026gpqa
- GPQA Diamond90.51Oct 8, 2026gpqa
- GPQA Diamond88.79Sep 9, 2026gpqa
- GPQA Diamond90.51Sep 9, 2026gpqa
- GPQA Diamond90.10Oct 7, 2026GPQA
- GPQA Diamond92.40Sep 10, 2026GPQA Diamond (Pass@1)
- 90.10Sep 2, 2026
- Humanity's Last Exam8.25Oct 8, 2026aa_hle
- Humanity's Last Exam35.22Oct 8, 2026aa_hle
- Humanity's Last Exam37.53Sep 9, 2026aa_hle
- Humanity's Last Exam35.22Sep 9, 2026aa_hle
- 48.20Oct 7, 2026
- Humanity's Last Exam42.70Sep 10, 2026HLE (Pass@1)
- Humanity's Last Exam37.70Jun 27, 2026HLE (Pass@1)
- Humanity's Last Exam7.70Jun 27, 2026HLE (Pass@1)
- Humanity's Last Exam34.50Jun 27, 2026HLE (Pass@1)
- 37.70Sep 2, 2026
24.2% behind the leader6 of 7 ranked benchmarks measured
- 83.40Oct 7, 2026
- 73.60Oct 7, 2026
- τ-Bench V3 · Banking30.10Oct 8, 2026tauBanking
- 32.95Oct 8, 2026
- Terminal-Bench 4.014.65Oct 8, 2026
Show 20 more agentic resultsHide 20 agentic results
- 24.26Oct 8, 2026
- 24.26Sep 9, 2026
- 38.32Oct 8, 2026
- 38.32Sep 9, 2026
- BrowseComp80.40Jun 27, 2026BrowseComp (Pass@1)
- 59.40Jun 6, 2026
- 32.58Oct 8, 2026
- 36.17Sep 9, 2026
- 35.94Sep 9, 2026
- 48.98Jun 15, 2026
- GDPval (win rate)54.60Jun 6, 2026GDPVal
- MCP Atlas69.40Jun 27, 2026MCPAtlas (Pass@1)
- MCP Atlas74.20Jun 27, 2026MCPAtlas (Pass@1)
- MCP Atlas73.60Sep 2, 2026MCP-Atlas (Public Set)
- 73.20Jul 16, 2026
- Terminal-Bench 4.014.65Sep 9, 2026
- τ-Bench V3 · Banking26.19Oct 8, 2026tauBanking
- τ-Bench V3 · Banking30.10Sep 9, 2026tauBanking
- τ-Bench V3 · Banking26.19Sep 9, 2026tauBanking
- 25.90Jun 6, 2026
25.5% behind the leader10 of 10 ranked benchmarks measured
- LiveCodeBench v692.50Jun 6, 2026LiveCodeBench (v6)
- 80.60Oct 7, 2026
- 76.20Oct 7, 2026
- LiveBench · Coding69.99Oct 8, 2026livebench_coding@2026-06-25
- SciCode50.81Oct 8, 2026aa_scicode
- Terminal-Bench Hard46.21Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.164.04Oct 8, 2026terminalbenchV21
- 1463.70Sep 21, 2026
- 55.40Oct 7, 2026
- LiveBench · Agentic Coding42.63Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 24 more coding resultsHide 24 coding results
- 1581.92Sep 21, 2026
- 1446.40Jun 28, 2026
- 1458.25May 22, 2026
- SciCode50.81Sep 9, 2026aa_scicode
- SciCode46.41Sep 4, 2026aa_scicode
- SciCode42.36Sep 4, 2026aa_scicode
- SWE-bench Multilingual69.80Jun 27, 2026SWE Multilingual (Resolved)
- SWE-bench Multilingual74.10Jun 27, 2026SWE Multilingual (Resolved)
- SWE-bench Pro52.10Jun 27, 2026SWE Pro (Resolved)
- SWE-bench Pro54.40Jun 27, 2026SWE Pro (Resolved)
- 55.40Sep 2, 2026
- SWE-bench Verified73.60Jun 27, 2026SWE Verified (Resolved)
- SWE-bench Verified79.40Jun 27, 2026SWE Verified (Resolved)
- 80.60Jul 16, 2026
- Terminal-Bench 2.164.79Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.164.04Sep 9, 2026terminalbenchV21
- Terminal-Bench 2.164.79Sep 9, 2026terminalbenchV21
- Terminal-Bench 2.187.90Sep 10, 2026Terminal-Bench 2.1 (Pass@1)
- 72.10Aug 13, 2026
- Terminal-Bench 2.164.00Sep 2, 2026Terminal Bench 2.1 (Terminus-2)
- Terminal-Bench Hard36.36Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard41.67Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard46.21Sep 9, 2026aa_terminalbench_hard
- Terminal-Bench Hard41.67Sep 9, 2026aa_terminalbench_hard
30.5% behind the leader2 of 3 ranked benchmarks measured
- IFBench76.46Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following62.35Oct 8, 2026livebench_instruction_following@2026-06-25
33.2% behind the leader4 of 4 ranked benchmarks measured
- 8.60Jun 29, 2026
- AA-Omniscience · Accuracy42.95Oct 8, 2026omniscienceAccuracy
- SimpleQA Verified46.20Jun 27, 2026SimpleQA-Verified (Pass@1)
- AA-Omniscience · Non-hallucination5.90Oct 8, 2026omniscienceNonHallucination
Show 12 more factuality resultsHide 12 factuality results
- AA-Omniscience · Accuracy41.40Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy30.87Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy42.95Sep 9, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy41.40Sep 9, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination12.15Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination11.32Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination5.90Sep 9, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination11.32Sep 9, 2026omniscienceNonHallucination
- 46.99Sep 21, 2026
- SimpleQA Verified57.90Jun 27, 2026SimpleQA-Verified (Pass@1)
- SimpleQA Verified45.00Jun 27, 2026SimpleQA-Verified (Pass@1)
- 57.00Jul 16, 2026
51.2% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics90.68Oct 8, 2026livebench_math@2026-06-25
- 45.26Sep 21, 2026
- 2.44Sep 21, 2026
Show 13 more math resultsHide 13 math results
- 96.67Sep 2, 2026
- 96.67Jul 1, 2026
- 94.60Sep 2, 2026
- 96.70Jul 16, 2026
- 93.94Sep 2, 2026
- HMMT Feb 202695.20Oct 7, 2026HMMT Feb 26
- HMMT Feb 202631.70Jun 27, 2026HMMT 2026 Feb (Pass@1)
- HMMT Feb 202695.20Sep 2, 2026HMMT Feb. 2026
- 89.80Oct 7, 2026
- IMO-AnswerBench35.30Jun 27, 2026IMOAnswerBench (Pass@1)
- IMO-AnswerBench88.00Jun 27, 2026IMOAnswerBench (Pass@1)
- 89.80Sep 2, 2026
- 60.71Sep 2, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language78.13Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis74.54Oct 8, 2026livebench_data_analysis@2026-06-25
- 1462Sep 13, 2026
- AA Intelligence53.00Aug 13, 2026Artificial Analysis Intelligence Index
- 1457Aug 11, 2026
- vectara_avg_summary_length153.80Jun 29, 2026Average Summary Length (Words)
Show 118 more resultsHide 118 results
- 27.69Sep 9, 2026
- 35.26Sep 4, 2026
- 63.27Jun 18, 2026
- AA Intelligence44.00Jun 29, 2026Artificial Analysis Intelligence Index
- AA Intelligence30.45Oct 8, 2026aa_intelligence_index
- AA Intelligence30.11Oct 8, 2026aa_intelligence_index
- AA Intelligence20.83Oct 8, 2026aa_intelligence_index
- AA Intelligence30.87Sep 9, 2026aa_intelligence_index
- AA Intelligence30.11Sep 9, 2026aa_intelligence_index
- -10.73Oct 8, 2026
- -10.57Oct 8, 2026
- -29.87Oct 8, 2026
- -10.73Sep 9, 2026
- -10.57Sep 9, 2026
- Agents' Last Exam25.70Sep 10, 2026Agent's Last Exam (Pass@1)
- 16.50Aug 13, 2026
- 55.17Jun 24, 2026
- 63.27Jun 24, 2026
- 63.70Jun 24, 2026
- 52.97Jun 24, 2026
- 27.61Jun 24, 2026
- 59.44Jun 24, 2026
- 51.26Jun 24, 2026
- 50.32Jun 24, 2026
- 38.30Jun 27, 2026
- 0.40Jun 27, 2026
- 27.40Jun 27, 2026
- 90.20Jun 27, 2026
- 9.20Jun 27, 2026
- 85.50Jun 27, 2026
- 85.80Jun 6, 2026
- 86.50Jun 6, 2026
- Artificial Analysis Coding Index59.36Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index58.67Sep 9, 2026aa_coding_index
- 38.36Jun 18, 2026
- AutomationBench Public12.80Aug 13, 2026AutomationBench (Public)
- 83.40Jul 16, 2026
- Chinese SimpleQA (C-SimpleQA)77.70Jul 4, 2026Chinese-SimpleQA (Pass@1)
- Chinese SimpleQA (C-SimpleQA)84.40Jun 27, 2026Chinese-SimpleQA (Pass@1)
- Chinese SimpleQA (C-SimpleQA)75.80Jun 27, 2026Chinese-SimpleQA (Pass@1)
- 3348.00Sep 10, 2026
- 3206.00Jun 27, 2026
- 2919.00Jun 27, 2026
- 62.00Jun 27, 2026
- 35.60Jun 27, 2026
- 56.50Jun 27, 2026
- 14.00Jun 6, 2026
- CyberGym83.30Sep 10, 2026CyberGym (Pass@1)
- 52.70Aug 13, 2026
- 12.80Aug 13, 2026
- 8.00Sep 2, 2026
- DeepSWE v1.1 (Resolved)62.70Sep 10, 2026
- 41.80Aug 13, 2026
- 31.10Aug 13, 2026
- 1554.00Jun 27, 2026
- 1306.00Jul 30, 2026
- 1307.00Jul 16, 2026
- 89.30Jul 16, 2026
- GPQA (unspecified)87.80Jun 6, 2026GPQA (no tools)
- HLE (with tools)60.00Sep 10, 2026HLE w/ tools (Pass@1)
- HLE (with tools)48.20Jun 27, 2026HLE w/ tools (Pass@1)
- HLE (with tools)44.70Jun 27, 2026HLE w/ tools (Pass@1)
- HLE (with tools)48.20Sep 2, 2026HLE (w/ Tools)
- 48.20Aug 13, 2026
- 94.40Sep 2, 2026
- 79.10Aug 24, 2026
- 85.40Jun 6, 2026
- 580.10Jun 6, 2026
- 73.58Oct 7, 2026
- 93.50Aug 23, 2026
- LiveCodeBench56.80Jun 27, 2026LiveCodeBench (Pass@1)
- LiveCodeBench89.80Jun 27, 2026LiveCodeBench (Pass@1)
- MathArena Apex (Pass@1)90.20Oct 7, 2026MathArena Apex
- 65.30Sep 10, 2026
- 73.60Jul 4, 2026
- 87.50Oct 7, 2026
- MMLU-Pro87.10Jul 4, 2026MMLU-Pro (EM)
- MMLU-Pro82.90Jun 27, 2026MMLU-Pro (EM)
- 87.50Jun 6, 2026
- 38.50Aug 13, 2026
- 35.50Sep 2, 2026
- NL2Repo-Bench (Score)61.50Sep 10, 2026
- OpenAI-MRCR (1M)83.50Jun 27, 2026MRCR 1M (MMR)
- OpenAI-MRCR (1M)44.70Jun 27, 2026MRCR 1M (MMR)
- OpenAI-MRCR (1M)83.30Jun 27, 2026MRCR 1M (MMR)
- 88.60Jun 6, 2026
- 59.90Jun 6, 2026
- 47.80Sep 2, 2026
- ProgramBench (Almost@1)15.50Sep 10, 2026
- 50.50Jun 6, 2026
- 57.90Oct 7, 2026
- 98.60Jul 16, 2026
- SWEBench Pro Public55.40Jul 16, 2026SWEBench Pro (Public)
- 80.80Jun 6, 2026
- 73.20Jun 6, 2026
- 88.90Jun 6, 2026
- TauBench V3 - Telecom96.30Aug 24, 2026Telecom
- 67.90Oct 7, 2026
- Terminal-Bench 2.059.10Jun 27, 2026Terminal Bench 2.0 (Acc)
- Terminal-Bench 2.063.30Jun 27, 2026Terminal Bench 2.0 (Acc)
- Terminal-bench 3.011.80Sep 10, 2026Terminal-Bench 3.0 (Pass@1)
- Terminal-Bench 4.012.40Sep 10, 2026Terminal-Bench 4.0 (Pass@1)
- 52.80Sep 2, 2026
- 51.80Oct 7, 2026
- Toolathlon46.30Jun 27, 2026Toolathlon (Pass@1)
- Toolathlon49.00Jun 27, 2026Toolathlon (Pass@1)
- 55.90Aug 13, 2026
- 62.30Jun 6, 2026
- 58.90Jun 6, 2026
- vectara_answer_rate97.20Jun 29, 2026Answer Rate
- vectara_factual_consistency91.40Jun 29, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)96.20Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)94.15Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)91.23Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)96.20Sep 9, 2026aa_tau2
- τ²-Bench Telecom (AA run)94.15Sep 9, 2026aa_tau2
- τ³-Bench Banking26.00Jul 30, 2026τ³-Banking
- τ³-Bench Banking25.80Jul 16, 2026Tau 3 Banking
DeepSeek-V4-Pro: common questions
Who makes DeepSeek-V4-Pro?
DeepSeek-V4-Pro is made by DeepSeek.
When was DeepSeek-V4-Pro released?
DeepSeek-V4-Pro was released on Apr 24, 2026, according to Artificial Analysis.
What is DeepSeek-V4-Pro good at?
DeepSeek-V4-Pro is capable in long context, reasoning, and agentic tasks; and behind the leaders in coding, instruction following, factuality, and math. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
How much does DeepSeek-V4-Pro cost?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens, according to Artificial Analysis. We track its price at 5 providers. At a mix of three input tokens to one output token, it is cheaper than 55% of the 330 priced models we track.
How many benchmarks has DeepSeek-V4-Pro been tested on?
We track 256 results for DeepSeek-V4-Pro on 111 benchmarks from 23 sources, 27 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-V4-Pro support?
OpenRouter lists tool calling, structured outputs, and reasoning for DeepSeek-V4-Pro.
About this record
Where DeepSeek-V4-Pro's numbers come from, and every name it appears under.
- Tracked since
- Jun 27, 2026
- Newest source mention
- Sep 30, 2026
Where the results come from
Verification: 256 scores · 27 independently verified · 84 aggregator-attributed · 59 vendor cross-reference · 86 vendor-reported. How these tiers are assigned
From 23 sources on 11 sites. Hugging Face supplies 126 of them; the 27 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co126
- artificialanalysis.ai86
- api.llm-stats.com16
- livebench.ai7
- datasets-server.huggingface.co4
- matharena.ai4
- raw.githubusercontent.com4
- epoch.ai3
- thinkingmachines.ai3
- lmarena.ai2
- simple-bench.com1
Also known as
How our sources name DeepSeek-V4-Pro at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| high | deepseek v4 pro (high) | deepseek v4 pro (reasoning, high effort) | deepseek-v4-pro-high deepseek-v4-pro-high-preview deepseek-v4-pro-high-20260813 |
| max | deepseek v4 pro (max) deepseek v4 pro max ds-v4-pro max | deepseek v4 pro (reasoning, max effort) | — |