DeepSeek-V3.1
DeepSeek-V3.1 is behind the leaders in factuality, long context, reasoning, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.56input$1.68outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
67results on39benchmarks
- 8 independently verified
- 30 aggregator
- 24 vendor-reported
- 5 cross-referenced
From 9 sources · latest Oct 8, 2026 · How verification works
DeepSeek-V3.1 benchmark results
67 results on 39 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
26.2% behind the leader3 of 4 ranked benchmarks measured
- 5.50May 2, 2026
- AA-Omniscience · Accuracy29.00Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination17.49Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy23.13Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination14.29Oct 8, 2026omniscienceNonHallucination
33.1% behind the leader1 of 3 ranked benchmarks measured
- 56.67Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 47.00Oct 8, 2026
37.3% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond77.88Oct 8, 2026gpqa
- 40.00May 15, 2026
- Humanity's Last Exam14.27Oct 8, 2026aa_hle
- 2.00Oct 8, 2026
Show 6 more reasoning resultsHide 6 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond73.54Oct 8, 2026gpqa
- GPQA Diamond74.90Oct 7, 2026GPQA
- GPQA Diamond80.10Jun 4, 2026GPQA-Diamond (Pass@1)
- Humanity's Last Exam6.67Oct 8, 2026aa_hle
- 15.90Oct 7, 2026
38.0% behind the leader4 of 10 ranked benchmarks measured
- 66.00Oct 7, 2026
- SciCode39.12Sep 4, 2026aa_scicode
- 54.50Oct 7, 2026
- Terminal-Bench Hard25.00Oct 8, 2026aa_terminalbench_hard
Show 4 more coding resultsHide 4 coding results
- SciCode36.69Sep 4, 2026aa_scicode
- 54.50May 15, 2026
- 66.00May 15, 2026
- Terminal-Bench Hard24.24Oct 8, 2026aa_terminalbench_hard
45.8% behind the leader2 of 7 ranked benchmarks measured
- 30.00Oct 7, 2026
- 5.62Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 28.73Jun 15, 2026
0 of 3 ranked benchmarks measured
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 85.83May 11, 2026
- 90.83May 2, 2026
- vectara_avg_summary_length63.70May 2, 2026Average Summary Length (Words)
- vectara_answer_rate94.50May 2, 2026Answer Rate
- vectara_factual_consistency94.50May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)34.80Oct 8, 2026aa_tau2
Show 30 more resultsHide 30 results
- 31.94Jun 18, 2026
- 18.85Jun 18, 2026
- AA Intelligence13.71Oct 8, 2026aa_intelligence_index
- AA Intelligence13.48Oct 8, 2026aa_intelligence_index
- -42.75Oct 8, 2026
- -29.58Oct 8, 2026
- 68.40Oct 7, 2026
- Aider-Polyglot76.30Jun 4, 2026Aider-Polyglot (Acc.)
- AIME 202493.10Jun 4, 2026AIME 2024 (Pass@1)
- 49.80Oct 7, 2026
- AIME 202588.40Jun 4, 2026AIME 2025 (Pass@1)
- 28.39Jun 18, 2026
- Artificial Analysis Coding Index29.71Jun 18, 2026aa_coding_index
- 49.20Oct 7, 2026
- 2091.00Jun 4, 2026
- HLE (with tools)29.80Jun 4, 2026Humanity's Last Exam (Python + Search)
- 33.50Oct 7, 2026
- HMMT 202584.20Jun 4, 2026HMMT 2025 (Pass@1)
- 56.40Aug 23, 2026
- LiveCodeBench74.80Jun 4, 2026LiveCodeBench (2408-2505) (Pass@1)
- 83.70Oct 7, 2026
- MMLU-Pro84.80Jun 4, 2026MMLU-Pro (EM)
- 91.80Oct 7, 2026
- MMLU-Redux93.70Jun 4, 2026MMLU-Redux (EM)
- 29.00May 15, 2026
- 93.40Oct 7, 2026
- 53.30May 15, 2026
- 31.30Sep 23, 2026
- 31.30May 15, 2026
- τ²-Bench Telecom (AA run)37.43Oct 8, 2026aa_tau2
DeepSeek-V3.1: common questions
Who makes DeepSeek-V3.1?
DeepSeek-V3.1 is made by DeepSeek.
When was DeepSeek-V3.1 released?
DeepSeek-V3.1 was released on Aug 21, 2025, according to Artificial Analysis.
What is DeepSeek-V3.1 good at?
DeepSeek-V3.1 is behind the leaders in factuality, long context, reasoning, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, multilingual tasks, or instruction following.
How much does DeepSeek-V3.1 cost?
DeepSeek-V3.1 costs $0.56 per million input tokens and $1.68 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 54% of the 331 priced models we track.
How many benchmarks has DeepSeek-V3.1 been tested on?
We track 67 results for DeepSeek-V3.1 on 39 benchmarks from 9 sources, 8 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where DeepSeek-V3.1's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 67 scores · 8 independently verified · 30 aggregator-attributed · 5 vendor cross-reference · 24 vendor-reported. How these tiers are assigned
From 9 sources on 7 sites. Artificial Analysis supplies 30 of them; the 8 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai30
- huggingface.co15
- api.llm-stats.com14
- raw.githubusercontent.com4
- matharena.ai2
- labs.scale.com1
- simple-bench.com1