DeepSeek-V2
DeepSeek-V2 has too few ranked results yet to rate it on any capability.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability.
No capability has enough ranked results to rate yet.
Price
No current price is tracked for this model. See the rate card
Evidence
89results on69benchmarks
- 65 vendor-reported
- 24 cross-referenced
From 4 sources · latest Sep 10, 2026 · How verification works
Research
7 papers reference DeepSeek-V2DeepSeek-V2 benchmark results
89 results on 69 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
0 of 6 ranked benchmarks measured
- GPQA Diamond35.30May 3, 2026GPQA-Diamond (Pass@1)
0 of 3 ranked benchmarks measured
- 31.60May 3, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- CCPM93.00Aug 31, 2026CCPM (Acc.)
- CMRC77.40Aug 31, 2026CMRC (EM)
- C-Eval81.40Aug 31, 2026C-Eval (Acc.)
- CLUEWSC82.00Aug 31, 2026CLUEWSC (EM)
- AGIEval57.50Aug 31, 2026AGIEval (Acc.)
- RACE-High52.60Aug 31, 2026RACE-High (Acc.)
Show 81 more resultsHide 81 results
- 60.30May 3, 2026
- 4.60May 3, 2026
- 92.20May 1, 2026
- ARC-Challenge70.00Sep 10, 2026ARC-C
- ARC-Challenge92.40May 31, 2026ARC-C
- 92.20Jul 6, 2026
- 97.60May 1, 2026
- 78.80May 1, 2026
- 78.90Sep 10, 2026
- 78.80Jul 6, 2026
- 78.60May 3, 2026
- 81.70Sep 10, 2026
- 77.40May 1, 2026
- 77.40May 31, 2026
- 77.40Jul 6, 2026
- 48.50May 3, 2026
- 89.90May 3, 2026
- 78.70May 31, 2026
- 78.70May 1, 2026
- CMMLU84.00May 1, 2026chinese_cmmlu_acc
- 84.00Sep 10, 2026
- 84.00Jul 6, 2026
- 2.80May 3, 2026
- 17.50May 3, 2026
- CRUXEval-I (input prediction)52.50May 1, 2026CRUXEval-I (Acc.)
- CRUXEval-O (output prediction)49.80May 1, 2026CRUXEval-O (Acc.)
- 80.10May 31, 2026
- 83.00May 3, 2026
- 80.40May 1, 2026
- 55.00Sep 10, 2026
- 66.90May 3, 2026
- 79.20Sep 10, 2026
- 81.60May 1, 2026
- 87.10May 1, 2026
- 87.80Sep 10, 2026
- 87.10Jul 6, 2026
- HumanEval43.30May 1, 2026HumanEval (Pass@1)
- 45.70Sep 10, 2026
- 48.80May 31, 2026
- 69.30May 3, 2026
- 57.70May 3, 2026
- LiveCodeBench20.30May 3, 2026LiveCodeBench (Pass@1)
- LiveCodeBench11.60May 1, 2026code_livecodebenchbase_pass1
- 18.80May 3, 2026
- 11.60Jul 6, 2026
- 43.40May 1, 2026
- 43.60Sep 10, 2026
- 43.40Jul 6, 2026
- 56.30May 3, 2026
- MBPP65.00May 1, 2026MBPP (Pass@1)
- 73.90Sep 10, 2026
- 66.60May 31, 2026
- 63.60May 1, 2026
- 63.60Jul 6, 2026
- 75.60May 1, 2026
- 78.50Sep 10, 2026
- 78.40Jul 6, 2026
- MMLU-Pro51.40Jul 3, 2026english_mmlupro_acc
- 58.50May 3, 2026
- 51.40Jul 6, 2026
- 77.90May 3, 2026
- 75.60Jul 6, 2026
- MMMLU64.00Jul 6, 2026Multilingual (MMMLU-non-English (Acc.))
- 64.00Jul 7, 2026
- 58.80Sep 10, 2026
- 44.40Sep 10, 2026
- 38.60May 1, 2026
- 38.70May 31, 2026
- 38.60Jul 6, 2026
- Pile-test (BPB)lower is better0.61Jul 6, 2026
- 83.90May 1, 2026
- 83.70May 31, 2026
- 83.90Jul 6, 2026
- RACE-Middle73.10Aug 31, 2026RACE-Middle (Acc.)
- The Pile (Test, BPB)lower is better0.61May 1, 2026
- 79.90May 31, 2026
- 80.00May 1, 2026
- 42.20Sep 10, 2026
- 86.30May 1, 2026
- 84.90May 31, 2026
- 86.30Jul 6, 2026
DeepSeek-V2: common questions
Who makes DeepSeek-V2?
DeepSeek-V2 is made by DeepSeek.
When was DeepSeek-V2 released?
DeepSeek-V2's weights were first published on Hugging Face on Apr 22, 2024.
How many benchmarks has DeepSeek-V2 been tested on?
We track 89 results for DeepSeek-V2 on 69 benchmarks from 4 sources. The latest was recorded on Sep 10, 2026.
About this record
Where DeepSeek-V2's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 26, 2026
Where the results come from
Verification: 89 scores · 0 independently verified · 24 vendor cross-reference · 65 vendor-reported. How these tiers are assigned
From 4 sources on 3 sites. raw.githubusercontent.com supplies 47 of them. Bars are coloured by trust tier.
- raw.githubusercontent.com47
- huggingface.co24
- arxiv.org18