DeepSeek-V3.2
DeepSeek-V3.2 is capable in long context; and behind the leaders in reasoning, instruction following, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal, Multilingual or Factuality.
Price
$0.28input$0.42outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
182results on102benchmarks
- 26 independently verified
- 33 aggregator
- 17 vendor-reported
- 106 cross-referenced
From 26 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
DeepSeek-V3.2 benchmark results
182 results on 102 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
21.5% behind the leader2 of 3 ranked benchmarks measured
- 58.40Jun 13, 2026
- 73.33Oct 7, 2026
Show 2 more long context resultsHide 2 long context results
- 45.67Oct 7, 2026
- 59.80Jun 15, 2026
32.5% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond84.04Oct 7, 2026gpqa
- Humanity's Last Exam24.56Oct 7, 2026aa_hle
- 2.86Oct 7, 2026
- 4.03May 10, 2026
Show 13 more reasoning resultsHide 13 reasoning results
- 0.86Oct 7, 2026
- GPQA Diamond75.05Oct 7, 2026gpqa
- GPQA Diamond82.40Oct 6, 2026GPQA
- 82.40Jun 15, 2026
- 80.00Jun 6, 2026
- GPQA Diamond84.00Jun 5, 2026GPQA-D
- Humanity's Last Exam11.21Oct 7, 2026aa_hle
- 40.80Oct 6, 2026
- 25.10Oct 6, 2026
- Humanity's Last Exam25.10Jun 15, 2026HLE-Full
- Humanity's Last Exam19.80Jun 12, 2026HLE (Text-only) no tools
- Humanity's Last Exam13.80Jun 6, 2026HLE (w/o tools)
- Humanity's Last Exam22.20Jun 5, 2026HLE w/o tools
33.7% behind the leader1 of 3 ranked benchmarks measured
- IFBench60.68Oct 7, 2026aa_ifbench
35.7% behind the leader8 of 10 ranked benchmarks measured
- LiveCodeBench v683.30Jun 15, 2026LiveCodeBench (v6)
- 73.10Oct 6, 2026
- 70.20Oct 6, 2026
- SciCode38.89Sep 4, 2026aa_scicode
- Terminal-Bench Hard35.61Oct 7, 2026aa_terminalbench_hard
- 1361.56May 22, 2026
- Terminal-Bench 2.146.82Oct 7, 2026terminalbenchV21
- 15.56Oct 7, 2026
Show 19 more coding resultsHide 19 coding results
- 74.10May 15, 2026
- 1360.53May 22, 2026
- SciCode38.66Sep 4, 2026aa_scicode
- 38.90Jun 15, 2026
- 38.00Jun 6, 2026
- 39.00Jun 5, 2026
- SciCode37.70May 15, 2026SciCode no tools
- 59.00May 1, 2026
- 70.20Jun 15, 2026
- SWE-bench Multilingual57.90Jun 12, 2026SWE-bench Multilingual w/ tools
- 70.00Sep 25, 2026
- 60.00Sep 11, 2026
- 73.10Jun 15, 2026
- SWE-bench Verified67.80Jun 12, 2026SWE-bench Verified w/ tools
- SWE-bench Verified60.00Jun 5, 2026SWE-bench Verified (mini-swe-agent)
- SWE-bench Verified67.00Jun 5, 2026SWE-bench Verified (Droid)
- Terminal-Bench Hard32.58Oct 7, 2026aa_terminalbench_hard
- 35.40Jun 13, 2026
- 29.00Jun 6, 2026
42.6% behind the leader4 of 7 ranked benchmarks measured
- MCP Atlas62.20May 18, 2026MCP-Atlas (Public Set)
- 51.40Oct 6, 2026
- τ-Bench V3 · Banking18.76Aug 10, 2026tauBanking
- 9.75Oct 7, 2026
Show 4 more agentic resultsHide 4 agentic results
- 14.53Oct 7, 2026
- 51.40Jun 15, 2026
- 40.10Jun 6, 2026
- 18.78Jun 15, 2026
0 of 5 ranked benchmarks measured
- 84.09Sep 2, 2026
- 94.17Sep 2, 2026
- 80.81May 10, 2026
Show 6 more math resultsHide 6 math results
- 94.17May 2, 2026
- 95.10May 18, 2026
- HMMT Feb 202679.90May 18, 2026HMMT Feb. 2026
- 78.30Oct 6, 2026
- 78.30Jun 15, 2026
- IMO-AnswerBench76.00May 15, 2026IMO-AnswerBench no tools
0 of 4 ranked benchmarks measured
- 27.50May 20, 2026
- 6.30May 2, 2026
- AA-Omniscience · Accuracy32.97Oct 7, 2026omniscienceAccuracy
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy24.00Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination17.28Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination6.73Oct 7, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- frontiermath_tier_4_v12.10Aug 29, 2026frontiermath_tier_4
- 60.00Aug 29, 2026
- 76.74Jun 9, 2026
- 73.58May 22, 2026
- 61.22May 22, 2026
- 90.32May 22, 2026
Show 101 more resultsHide 101 results
- 18.32Aug 10, 2026
- 39.82Jun 18, 2026
- AA Intelligence21.49Oct 7, 2026aa_intelligence_index
- AA Intelligence16.04Oct 7, 2026aa_intelligence_index
- -22.48Oct 7, 2026
- -46.88Oct 7, 2026
- 91.25May 2, 2026
- 93.10Oct 6, 2026
- 93.10Jun 15, 2026
- AIME 202588.00Jun 6, 2026AIME25
- AIME 202592.00Jun 5, 2026AIME25
- AIME 202589.30May 15, 2026AIME25
- 92.70May 17, 2026
- 89.30Jun 12, 2026
- 58.10May 15, 2026
- 57.00May 10, 2026
- 88.80Jun 13, 2026
- 53.40Jun 13, 2026
- 55.80Jun 6, 2026
- Artificial Analysis Coding Index44.17Sep 9, 2026aa_coding_index
- 34.60Jun 18, 2026
- 67.60Jun 5, 2026
- 40.10Jun 12, 2026
- browsecomp_with_context_manager67.60Jun 13, 2026BrowseComp (w/ Context Manage)
- 65.00Oct 6, 2026
- 47.90Jun 6, 2026
- 65.00May 30, 2026
- 47.90Jun 12, 2026
- 17.30Jun 15, 2026
- DeepSearchQA (F1)60.90Jun 15, 2026DeepSearchQA
- 26.20Jun 6, 2026
- 27.00May 15, 2026
- 27.00Jun 12, 2026
- 59.10Jun 15, 2026
- 80.20May 15, 2026
- 80.20Jun 12, 2026
- 75.10May 30, 2026
- 63.50Jun 6, 2026
- GPQA (unspecified)79.90May 15, 2026GPQA no tools
- 46.90May 15, 2026
- 46.90Jun 12, 2026
- HLE (with tools)40.80Jun 15, 2026HLE-Full (w/ tools)
- HLE (with tools)27.20Jun 6, 2026HLE (w/ tools)
- HLE (with tools)20.30May 15, 2026HLE (Text-only) w/ tools
- 90.20Oct 6, 2026
- HMMT 202583.60May 15, 2026HMMT25
- 92.50May 11, 2026
- HMMT Feb. 202592.50Jun 15, 2026HMMT 2025 (Feb)
- 90.00May 10, 2026
- HMMT Nov. 202590.20May 30, 2026HMMT 2025 (Nov.)
- 83.60Jun 12, 2026
- 49.50May 15, 2026
- 83.30Aug 23, 2026
- LiveCodeBench79.00Jun 6, 2026LiveCodeBench (LCB)
- 86.00Jun 5, 2026
- 74.10Jun 12, 2026
- Longform Writing eval (Kimi K2 Thinking system card)72.50Jun 12, 2026Longform Writing no tools
- 38.00Oct 6, 2026
- 85.00Oct 6, 2026
- 85.00Jun 15, 2026
- 86.00Jun 5, 2026
- MMLU-Redux93.70May 15, 2026MMLU-Redux no tools
- 55.50Jun 13, 2026
- Multi-SWE-Bench30.60Jun 12, 2026Multi-SWE-bench w/ tools
- 37.40Jun 5, 2026
- 26.00Jun 5, 2026
- 54.70Jun 15, 2026
- 38.20May 15, 2026
- 38.20Jun 12, 2026
- 47.10Jun 15, 2026
- 55.80May 30, 2026
- 49.50Jun 15, 2026
- 38.50May 15, 2026
- 38.50Jun 12, 2026
- 0.90Jun 5, 2026
- 62.00Jun 5, 2026
- 80.30Oct 6, 2026
- 80.20Oct 6, 2026
- 37.70Jun 6, 2026
- 46.40Oct 6, 2026
- Terminal-Bench 2.039.30May 18, 2026Terminal-Bench 2.0 (Terminus-2)
- Terminal-Bench 2.046.40May 17, 2026Terminal-Bench 2.0 (Claude Code)
- 37.70Jun 12, 2026
- theagentcompany34.00Jun 6, 2026AgentCompany
- 35.20May 18, 2026
- 35.20Oct 6, 2026
- 35.20Jun 5, 2026
- vectara_answer_rate92.60May 2, 2026Answer Rate
- vectara_avg_summary_length62.00May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency93.70May 2, 2026Factual Consistency Rate
- 32.50Jun 15, 2026
- 71.00Jun 6, 2026
- xbench-DeepSearch55.70May 30, 2026xbench-DeepSearch (2025.10)
- 55.70May 3, 2026
- 85.30Jun 13, 2026
- 80.30Jun 13, 2026
- τ²-Bench34.00Jun 6, 2026τ²-Bench-Telecom
- τ²-Bench91.00Jun 5, 2026𝜏²-Bench Telecom
- τ²-Bench Telecom (AA run)90.64Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)78.95Oct 7, 2026aa_tau2
- 69.20May 18, 2026
DeepSeek-V3.2: common questions
Who makes DeepSeek-V3.2?
DeepSeek-V3.2 is made by DeepSeek.
When was DeepSeek-V3.2 released?
DeepSeek-V3.2 was released on Dec 1, 2025, according to Artificial Analysis.
What is DeepSeek-V3.2 good at?
DeepSeek-V3.2 is capable in long context; and behind the leaders in reasoning, instruction following, coding, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, multilingual tasks, or factuality.
How much does DeepSeek-V3.2 cost?
DeepSeek-V3.2 costs $0.28 per million input tokens and $0.42 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 71% of the 328 priced models we track.
How many benchmarks has DeepSeek-V3.2 been tested on?
We track 182 results for DeepSeek-V3.2 on 102 benchmarks from 26 sources, 26 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does DeepSeek-V3.2 support?
OpenRouter lists tool calling, structured outputs, and reasoning for DeepSeek-V3.2.
About this record
Where DeepSeek-V3.2's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 182 scores · 26 independently verified · 33 aggregator-attributed · 106 vendor cross-reference · 17 vendor-reported. How these tiers are assigned
From 26 sources on 10 sites. Hugging Face supplies 110 of them; the 26 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co110
- artificialanalysis.ai33
- api.llm-stats.com17
- matharena.ai7
- raw.githubusercontent.com4
- swebench.com4
- arcprize.org2
- datasets-server.huggingface.co2
- epoch.ai2
- labs.scale.com1