DeepSeek-V2.5
DeepSeek-V2.5 is behind the leaders in coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Long Context, Math, Multimodal, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
45results on32benchmarks
- 2 independently verified
- 3 aggregator
- 31 vendor-reported
- 9 cross-referenced
From 6 sources · latest Oct 7, 2026 · How verification works
DeepSeek-V2.5 benchmark results
45 results on 32 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
45.1% behind the leader1 of 10 ranked benchmarks measured
- 16.80Oct 6, 2026
Show 1 more coding resultHide 1 coding result
- 22.60May 3, 2026
0 of 6 ranked benchmarks measured
- GPQA Diamond42.32Oct 7, 2026gpqa
- GPQA Diamond41.30May 3, 2026GPQA-Diamond (Pass@1)
0 of 3 ranked benchmarks measured
- 35.40May 3, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence7.00Sep 30, 2026Artificial Analysis Intelligence Index
- 17.80May 1, 2026
- AA Intelligence6.62Oct 7, 2026aa_intelligence_index
- AA Intelligence6.55Oct 7, 2026aa_intelligence_index
- 9.02Oct 6, 2026
- HumanEval-Mul (Pass@1)73.80Oct 6, 2026HumanEval-Mul
Show 34 more resultsHide 34 results
- 71.60May 3, 2026
- 18.20May 3, 2026
- 16.70May 3, 2026
- 80.40Oct 6, 2026
- 8.00May 31, 2026
- 50.50Oct 6, 2026
- 50.50May 31, 2026
- 76.20Sep 8, 2026
- 76.20May 31, 2026
- 84.30Oct 6, 2026
- 84.30May 31, 2026
- 79.50May 3, 2026
- 54.10May 3, 2026
- 90.40May 3, 2026
- 10.80May 3, 2026
- 35.60May 3, 2026
- 87.80May 3, 2026
- 65.40May 3, 2026
- 95.10Oct 6, 2026
- 90.30May 31, 2026
- 89.00Oct 6, 2026
- 89.00May 31, 2026
- 77.40May 3, 2026
- 80.60May 3, 2026
- LiveCodeBench41.80Aug 23, 2026LiveCodeBench(01-09)
- LiveCodeBench28.40May 3, 2026LiveCodeBench (Pass@1)
- 29.20May 3, 2026
- 74.70May 31, 2026
- 74.70May 3, 2026
- 80.40May 31, 2026
- 66.20May 3, 2026
- 80.30May 3, 2026
- 9.00May 31, 2026
- 10.20May 3, 2026
DeepSeek-V2.5: common questions
Who makes DeepSeek-V2.5?
DeepSeek-V2.5 is made by DeepSeek.
When was DeepSeek-V2.5 released?
DeepSeek-V2.5 was released on Sep 6, 2024, according to Artificial Analysis.
What is DeepSeek-V2.5 good at?
DeepSeek-V2.5 is behind the leaders in coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
How many benchmarks has DeepSeek-V2.5 been tested on?
We track 45 results for DeepSeek-V2.5 on 32 benchmarks from 6 sources, 2 of them independently verified. The latest was recorded on Oct 7, 2026.
About this record
Where DeepSeek-V2.5's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Sep 2, 2026
Where the results come from
Verification: 45 scores · 2 independently verified · 3 aggregator-attributed · 9 vendor cross-reference · 31 vendor-reported. How these tiers are assigned
From 6 sources on 5 sites. arxiv.org supplies 21 of them; the 2 independently verified results come from 2 sites. Bars are coloured by trust tier.
- arxiv.org21
- api.llm-stats.com10
- huggingface.co9
- artificialanalysis.ai4
- aider.chat1