DeepSeek-V3.1-Terminus
DeepSeek-V3.1-Terminus is behind the leaders in long context, factuality, coding, reasoning, instruction following, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal or Multilingual.
Price
$0.27input$1.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
37results on21benchmarks
- 1 independently verified
- 32 aggregator
- 4 cross-referenced
From 3 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
DeepSeek-V3.1-Terminus benchmark results
37 results on 21 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
26.2% behind the leader1 of 3 ranked benchmarks measured
- 69.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 45.33Oct 8, 2026
29.0% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy27.70Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination25.17Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy23.52Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination12.42Oct 8, 2026omniscienceNonHallucination
35.1% behind the leader3 of 10 ranked benchmarks measured
- SciCode37.96Oct 8, 2026aa_scicode
- Terminal-Bench 2.144.94Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard30.30Oct 8, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- SciCode32.06Sep 4, 2026aa_scicode
- Terminal-Bench Hard31.82Oct 8, 2026aa_terminalbench_hard
35.2% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond79.19Oct 8, 2026gpqa
- Humanity's Last Exam16.36Oct 8, 2026aa_hle
- 1.71Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond75.05Oct 8, 2026gpqa
- Humanity's Last Exam8.67Oct 8, 2026aa_hle
- Humanity's Last Exam19.30Jun 25, 2026HLE_text
36.0% behind the leader2 of 3 ranked benchmarks measured
- IFBench57.01Oct 8, 2026aa_ifbench
- 54.40Jun 25, 2026
Show 1 more instruction following resultHide 1 instruction following result
- IFBench41.22Oct 8, 2026aa_ifbench
42.3% behind the leader3 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking21.03Oct 8, 2026tauBanking
- 10.61Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
Show 1 more agentic resultHide 1 agentic result
- 23.74Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence14.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence13.93Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)37.13Oct 8, 2026aa_tau2
- -43.47Oct 8, 2026
- -26.40Oct 8, 2026
- AA Intelligence14.76Oct 8, 2026aa_intelligence_index
Show 6 more resultsHide 6 results
- 8.94Sep 9, 2026
- 28.56Jun 18, 2026
- Artificial Analysis Coding Index43.50Sep 9, 2026aa_coding_index
- 31.90Jun 18, 2026
- 86.10Jun 25, 2026
- 74.90Jun 25, 2026
DeepSeek-V3.1-Terminus: common questions
Who makes DeepSeek-V3.1-Terminus?
DeepSeek-V3.1-Terminus is made by DeepSeek.
When was DeepSeek-V3.1-Terminus released?
DeepSeek-V3.1-Terminus was released on Sep 22, 2025, according to Artificial Analysis.
What is DeepSeek-V3.1-Terminus good at?
DeepSeek-V3.1-Terminus is behind the leaders in long context, factuality, coding, reasoning, instruction following, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
How much does DeepSeek-V3.1-Terminus cost?
DeepSeek-V3.1-Terminus costs $0.27 per million input tokens and $1.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 60% of the 330 priced models we track.
How many benchmarks has DeepSeek-V3.1-Terminus been tested on?
We track 37 results for DeepSeek-V3.1-Terminus on 21 benchmarks from 3 sources, 1 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-V3.1-Terminus support?
OpenRouter lists tool calling, structured outputs, and reasoning for DeepSeek-V3.1-Terminus.
About this record
Where DeepSeek-V3.1-Terminus's numbers come from, and every name it appears under.
- Tracked since
- Jun 18, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 37 scores · 1 independently verified · 32 aggregator-attributed · 4 vendor cross-reference. How these tiers are assigned
From 3 sources on 2 sites. Artificial Analysis supplies 33 of them; the 1 independently verified result comes from 1 site. Bars are coloured by trust tier.
- artificialanalysis.ai33
- huggingface.co4