DeepSeek-V3.2-Exp
DeepSeek-V3.2-Exp is capable in long context and factuality; and behind the leaders in reasoning, instruction following, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal or Multilingual.
Price
$0.28input$0.42outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
50results on33benchmarks
- 9 independently verified
- 28 aggregator
- 13 vendor-reported
From 8 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
DeepSeek-V3.2-Exp benchmark results
50 results on 33 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
22.5% behind the leader1 of 3 ranked benchmarks measured
- 72.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 44.67Oct 8, 2026
24.7% behind the leader3 of 4 ranked benchmarks measured
- 5.30May 2, 2026
- AA-Omniscience · Accuracy27.55Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination19.23Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy22.83Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.95Oct 8, 2026omniscienceNonHallucination
36.0% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond79.70Oct 8, 2026gpqa
- Humanity's Last Exam14.87Oct 8, 2026aa_hle
- 1.43Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond73.84Oct 8, 2026gpqa
- GPQA Diamond79.90Oct 7, 2026GPQA
- Humanity's Last Exam9.04Oct 8, 2026aa_hle
- 19.80Oct 7, 2026
37.0% behind the leader1 of 3 ranked benchmarks measured
- IFBench54.15Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench43.13Oct 8, 2026aa_ifbench
39.8% behind the leader5 of 10 ranked benchmarks measured
- 67.80Oct 7, 2026
- 57.90Oct 7, 2026
- SciCode37.73Sep 4, 2026aa_scicode
- Terminal-Bench Hard31.06Oct 8, 2026aa_terminalbench_hard
- 1271.59May 22, 2026
Show 2 more coding resultsHide 2 coding results
- SciCode39.93Sep 4, 2026aa_scicode
- Terminal-Bench Hard25.00Oct 8, 2026aa_terminalbench_hard
41.1% behind the leader2 of 7 ranked benchmarks measured
- 40.10Oct 7, 2026
- 24.95Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 28.32Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 70.20Sep 11, 2026
- 84.17May 30, 2026
- 90.00May 30, 2026
- 91.67May 30, 2026
- vectara_avg_summary_length64.60May 2, 2026Average Summary Length (Words)
- vectara_answer_rate96.60May 2, 2026Answer Rate
Show 18 more resultsHide 18 results
- 28.73Jun 18, 2026
- 31.03Jun 18, 2026
- AA Intelligence16.58Oct 8, 2026aa_intelligence_index
- AA Intelligence13.88Oct 8, 2026aa_intelligence_index
- -30.97Oct 8, 2026
- -48.20Oct 8, 2026
- 74.50Oct 7, 2026
- 89.30Oct 7, 2026
- Artificial Analysis Coding Index33.28Jun 18, 2026aa_coding_index
- Artificial Analysis Coding Index29.98Jun 18, 2026aa_coding_index
- 47.90Oct 7, 2026
- 83.60Oct 7, 2026
- 74.10Aug 23, 2026
- 85.00Oct 7, 2026
- 97.10Oct 7, 2026
- 37.70Sep 23, 2026
- vectara_factual_consistency94.70May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)33.92Oct 8, 2026aa_tau2
DeepSeek-V3.2-Exp: common questions
Who makes DeepSeek-V3.2-Exp?
DeepSeek-V3.2-Exp is made by DeepSeek.
When was DeepSeek-V3.2-Exp released?
DeepSeek-V3.2-Exp was released on Sep 29, 2025, according to Artificial Analysis.
What is DeepSeek-V3.2-Exp good at?
DeepSeek-V3.2-Exp is capable in long context and factuality; and behind the leaders in reasoning, instruction following, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
How much does DeepSeek-V3.2-Exp cost?
DeepSeek-V3.2-Exp costs $0.28 per million input tokens and $0.42 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 70% of the 330 priced models we track.
How many benchmarks has DeepSeek-V3.2-Exp been tested on?
We track 50 results for DeepSeek-V3.2-Exp on 33 benchmarks from 8 sources, 9 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-V3.2-Exp support?
OpenRouter lists tool calling, structured outputs, and reasoning for DeepSeek-V3.2-Exp.
About this record
Where DeepSeek-V3.2-Exp's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 50 scores · 9 independently verified · 28 aggregator-attributed · 13 vendor-reported. How these tiers are assigned
From 8 sources on 6 sites. Artificial Analysis supplies 28 of them; the 9 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai28
- api.llm-stats.com13
- raw.githubusercontent.com4
- matharena.ai3
- aider.chat1
- datasets-server.huggingface.co1