DeepSeek-V3
DeepSeek-V3 is behind the leaders in factuality and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.24input$0.90outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
205results on131benchmarks
- 35 independently verified
- 36 aggregator
- 91 vendor-reported
- 43 cross-referenced
From 30 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
72 papers reference DeepSeek-V3DeepSeek-V3 benchmark results
205 results on 131 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
31.3% behind the leader3 of 4 ranked benchmarks measured
- 6.10May 2, 2026
- AA-Omniscience · Accuracy25.45Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination14.14Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy24.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.99Oct 8, 2026omniscienceNonHallucination
44.6% behind the leader5 of 10 ranked benchmarks measured
- SciCode39.00Oct 8, 2026aa_scicode
- LiveCodeBench v646.90Jun 15, 2026LiveCodeBench v6 (Aug 24 - May 25)
- 42.00Oct 7, 2026
- Terminal-Bench Hard15.15Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.116.85Oct 8, 2026terminalbenchV21
Show 10 more coding resultsHide 10 coding results
- LiveBench · Coding71.09Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding71.09Jun 17, 2026livebench_coding@2025-04-07
- SciCode35.76Oct 8, 2026aa_scicode
- SWE-bench Multilingual29.30Jun 4, 2026SWE-bench Multilingual (Agent mode)
- SWE-bench Multilingual25.80Jun 15, 2026SWE-bench Multilingual (Agentic Coding)
- SWE-bench Verified45.40Jun 4, 2026SWE Verified (Agent mode)
- SWE-bench Verified38.80Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified36.60Jun 15, 2026SWE-bench Verified (Agentless Coding)
- Terminal-Bench 2.113.86Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard6.82Oct 8, 2026aa_terminalbench_hard
0 of 6 ranked benchmarks measured
- 18.90May 10, 2026
- 27.20May 10, 2026
- 0.00Oct 8, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond55.66Oct 8, 2026gpqa
- GPQA Diamond65.45Oct 8, 2026gpqa
- GPQA Diamond68.40Oct 7, 2026GPQA
- GPQA Diamond59.10Oct 7, 2026GPQA
- 68.40Jun 15, 2026
- GPQA Diamond59.10Jun 6, 2026GPQA (diamond)
- Humanity's Last Exam2.90Oct 8, 2026aa_hle
- Humanity's Last Exam4.74Oct 8, 2026aa_hle
- Humanity's Last Exam5.20Jun 15, 2026Humanity's Last Exam (Text Only)
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
Show 3 more agentic resultsHide 3 agentic results
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking4.74Oct 8, 2026tauBanking
- τ-Bench V3 · Banking4.74Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 29.33Oct 8, 2026
- 40.67Oct 8, 2026
- 48.70Oct 7, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following86.78Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following86.78Jun 17, 2026livebench_instruction_following@2025-04-07
- IFBench41.02Oct 8, 2026aa_ifbench
Show 2 more instruction following resultsHide 2 instruction following results
- IFBench34.76Oct 8, 2026aa_ifbench
- 31.40Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language48.69Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence14.00Jul 10, 2026Artificial Analysis Intelligence Index
- 95.40Jul 2, 2026
- 94.00Jul 2, 2026
- livebench_language48.69Jun 17, 2026livebench_language@2025-04-07
- 40.79May 19, 2026
Show 152 more resultsHide 152 results
- 0.79Sep 9, 2026
- 0.79Sep 9, 2026
- AA Intelligence8.49Oct 8, 2026aa_intelligence_index
- AA Intelligence9.72Oct 8, 2026aa_intelligence_index
- -41.65Oct 8, 2026
- -40.67Oct 8, 2026
- 72.70Jun 15, 2026
- AGIEval79.60Aug 31, 2026AGIEval (Acc.)
- 79.70May 3, 2026
- 55.10May 1, 2026
- 49.60Oct 7, 2026
- Aider-Polyglot55.10Jun 4, 2026Aider-Polyglot (Acc.)
- 55.10Jun 15, 2026
- AIME 202459.40Jun 4, 2026AIME 2024 (Pass@1)
- AIME 202439.20Jun 4, 2026AIME 2024 (Pass@1)
- 59.40Jun 15, 2026
- 50.00May 2, 2026
- 25.00May 2, 2026
- AIME 202551.30Jun 4, 2026AIME 2025 (Pass@1)
- 46.70Jun 15, 2026
- 70.00Jun 4, 2026
- 97.14May 10, 2026
- 95.30May 1, 2026
- 95.30Jul 6, 2026
- 98.90May 1, 2026
- 91.40Jun 6, 2026
- Arena-Hard (GPT-4-1106 judge)85.50Jun 4, 2026ArenaHard (GPT-4-1106)
- Artificial Analysis Coding Index23.04Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index21.16Sep 9, 2026aa_coding_index
- 88.90Jun 15, 2026
- 87.50May 1, 2026
- 87.50Jul 6, 2026
- 96.70May 10, 2026
- 86.50Oct 7, 2026
- C-Eval90.10Aug 31, 2026C-Eval (Acc.)
- 78.60May 1, 2026
- 78.60Jul 6, 2026
- CCPM92.00Aug 31, 2026CCPM (Acc.)
- Chinese SimpleQA (C-SimpleQA)68.00Sep 11, 2026C-SimpleQA (Correct)
- 64.80May 3, 2026
- Chinese SimpleQA (C-SimpleQA)64.80Jun 6, 2026C-SimpleQA
- 90.90Oct 7, 2026
- CLUEWSC82.70Sep 11, 2026CLUEWSC (EM)
- 90.70May 1, 2026
- CMMLU88.80May 1, 2026chinese_cmmlu_acc
- 88.80Jul 6, 2026
- CMRC76.30Aug 31, 2026CMRC (EM)
- 43.20Oct 7, 2026
- 74.70Jun 15, 2026
- 51.60May 3, 2026
- 1134.00May 1, 2026
- Codeforces (Percentile)58.70Jul 3, 2026code_codeforces_percentile
- Codeforces (Rating)1134.00Jul 3, 2026code_codeforces_rating
- CRUXEval-I (input prediction)67.30May 1, 2026CRUXEval-I (Acc.)
- CRUXEval-O (output prediction)69.80May 1, 2026CRUXEval-O (Acc.)
- 91.60Oct 7, 2026
- 91.60Jun 4, 2026
- 89.00May 1, 2026
- 91.00Jun 6, 2026
- 73.30Oct 7, 2026
- 73.30Jun 4, 2026
- 96.70Jun 6, 2026
- 89.30May 1, 2026
- 49.69May 10, 2026
- 88.90May 1, 2026
- 88.90Jul 6, 2026
- HMMT 202529.20Jun 4, 2026HMMT 2025 (Pass@1)
- 27.50Jun 15, 2026
- 29.17May 11, 2026
- 13.33May 11, 2026
- HumanEval65.20May 1, 2026HumanEval (Pass@1)
- 92.10Jun 6, 2026
- HumanEval-Mul (Pass@1)82.60Oct 7, 2026HumanEval-Mul
- 86.10Aug 31, 2026
- 81.10Jun 15, 2026
- 87.30Jun 12, 2026
- 87.30Jun 6, 2026
- 72.40Jun 15, 2026
- 49.55May 3, 2026
- 49.20Aug 23, 2026
- LiveCodeBench43.00Jun 4, 2026LiveCodeBench (2408-2505) (Pass@1)
- LiveCodeBench37.60May 3, 2026LiveCodeBench (Pass@1)
- LiveCodeBench19.40May 1, 2026code_livecodebenchbase_pass1
- 40.50May 3, 2026
- 83.29May 3, 2026
- 14.20May 3, 2026
- 53.50May 3, 2026
- 19.40Jul 6, 2026
- 48.70Jun 6, 2026
- 91.21May 10, 2026
- 61.60May 1, 2026
- 90.20May 1, 2026
- 84.60Jun 6, 2026
- 61.60Jul 6, 2026
- MATH-500 (EM)94.00Oct 7, 2026MATH-500
- MATH-500 (EM)90.20Oct 7, 2026MATH-500
- MATH-500 (EM)94.00Jun 15, 2026MATH-500
- MBPP75.40May 1, 2026MBPP (Pass@1)
- 78.80Jun 6, 2026
- 79.80May 1, 2026
- 79.80Jul 6, 2026
- 80.31May 10, 2026
- MMLU88.50Jun 4, 2026MMLU (Pass@1)
- 86.20May 1, 2026
- 89.40Jun 15, 2026
- 88.50Jun 6, 2026
- 87.10Jul 6, 2026
- 81.20Oct 7, 2026
- MMLU-Pro75.90Aug 31, 2026MMLU-Pro (EM)
- 64.40May 1, 2026
- 81.20Jun 15, 2026
- 75.90Jun 6, 2026
- 64.40Jul 6, 2026
- 70.50May 26, 2025
- 89.10Oct 7, 2026
- MMLU-Redux90.50Jun 4, 2026MMLU-Redux (EM)
- 90.50Jun 15, 2026
- 86.20Jul 6, 2026
- MMMLU79.40Jul 6, 2026Multilingual (MMMLU-non-English (Acc.))
- 79.40Jul 7, 2026
- 83.10Jun 15, 2026
- 79.63May 10, 2026
- 40.00May 1, 2026
- 40.00Jul 6, 2026
- 46.70May 10, 2026
- 24.00Jun 15, 2026
- Pile-test (BPB)lower is better0.55Jul 6, 2026
- 84.70May 1, 2026
- 84.70Jul 6, 2026
- 59.50Jun 15, 2026
- RACE-High51.30Aug 31, 2026RACE-High (Acc.)
- RACE-Middle67.10Aug 31, 2026RACE-Middle (Acc.)
- 95.25May 10, 2026
- 24.90Oct 7, 2026
- 27.70Jun 15, 2026
- 24.90Jun 6, 2026
- 53.70Jun 15, 2026
- 39.00Jun 15, 2026
- Terminal-Bench13.30Jun 4, 2026Terminal-bench (Terminus 1 framework)
- The Pile (Test, BPB)lower is better0.55May 1, 2026
- 82.90May 1, 2026
- vectara_answer_rate97.50May 2, 2026Answer Rate
- vectara_avg_summary_length81.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency93.90May 2, 2026Factual Consistency Rate
- 84.90May 1, 2026
- 84.90Jul 6, 2026
- 97.11May 10, 2026
- 84.00Jun 15, 2026
- τ²-Bench32.50Jun 15, 2026Tau2 telecom
- τ²-Bench (Retail)69.10Jun 15, 2026Tau2 retail
- τ²-Bench Telecom (AA run)22.81Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)47.08Oct 8, 2026aa_tau2
DeepSeek-V3: common questions
Who makes DeepSeek-V3?
DeepSeek-V3 is made by DeepSeek.
When was DeepSeek-V3 released?
DeepSeek-V3 was released on Dec 26, 2024, according to Artificial Analysis.
What is DeepSeek-V3 good at?
DeepSeek-V3 is behind the leaders in factuality and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does DeepSeek-V3 cost?
DeepSeek-V3 costs $0.24 per million input tokens and $0.90 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 65% of the 331 priced models we track.
How many benchmarks has DeepSeek-V3 been tested on?
We track 205 results for DeepSeek-V3 on 131 benchmarks from 30 sources, 35 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-V3 support?
OpenRouter lists tool calling and structured outputs for DeepSeek-V3.
About this record
Where DeepSeek-V3's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Jul 19, 2026
Where the results come from
Verification: 205 scores · 35 independently verified · 36 aggregator-attributed · 43 vendor cross-reference · 91 vendor-reported. How these tiers are assigned
From 30 sources on 10 sites. Hugging Face supplies 64 of them; the 35 independently verified results come from 9 sites. Bars are coloured by trust tier.
- huggingface.co64
- raw.githubusercontent.com57
- artificialanalysis.ai37
- api.llm-stats.com18
- storage.googleapis.com12
- arxiv.org6
- livecodebench.github.io4
- matharena.ai4
- simple-bench.com2
- aider.chat1