DeepSeek-R1
DeepSeek-R1 is behind the leaders in long context. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Math, Multimodal, Multilingual, Instruction Following or Factuality.
Price
$2.00input$4.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
86results on69benchmarks
- 26 independently verified
- 18 aggregator
- 28 vendor-reported
- 14 cross-referenced
From 23 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingJSON modeReasoning
As listed by OpenRouter
Research
78 papers reference DeepSeek-R1DeepSeek-R1 benchmark results
86 results on 69 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
33.4% behind the leader2 of 3 ranked benchmarks measured
- 58.30Jun 5, 2026
- 57.67Oct 8, 2026
0 of 6 ranked benchmarks measured
- 30.90May 10, 2026
- 1.30May 10, 2026
- 0.57Oct 8, 2026
Show 6 more reasoning resultsHide 6 reasoning results
- GPQA Diamond70.81Oct 8, 2026gpqa
- GPQA Diamond71.50Jun 4, 2026GPQA-Diamond (Pass@1)
- 71.50Jun 5, 2026
- Humanity's Last Exam8.50Oct 8, 2026aa_hle
- Humanity's Last Exam8.50Jun 4, 2026Humanity's Last Exam (Pass@1)
- Humanity's Last Exam8.60Jun 5, 2026HLE (no tools)
0 of 10 ranked benchmarks measured
- LiveBench · Coding70.31Aug 23, 2026livebench_coding@2025-04-07
- SciCode38.31Oct 8, 2026aa_scicode
- Terminal-Bench 2.119.10Oct 8, 2026terminalbenchV21
Show 3 more coding resultsHide 3 coding results
- SWE-bench Verified49.20Jun 4, 2026SWE Verified (Resolved)
- 49.20Jun 5, 2026
- Terminal-Bench Hard6.06Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking6.39Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following80.60Aug 23, 2026livebench_instruction_following@2025-04-07
- IFBench38.98Oct 8, 2026aa_ifbench
- 40.70Jun 5, 2026
0 of 4 ranked benchmarks measured
- 11.30Aug 29, 2026
- AA-Omniscience · Accuracy30.52Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.71Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 94.44Sep 21, 2026
- 98.00Sep 21, 2026
- 47.88Sep 21, 2026
- 96.60Sep 21, 2026
- 97.24Sep 21, 2026
- 56.90Sep 11, 2026
Show 54 more resultsHide 54 results
- 1.07Sep 9, 2026
- AA Intelligence11.41Oct 8, 2026aa_intelligence_index
- -32.22Oct 8, 2026
- Aider-Polyglot53.30Aug 31, 2026Aider-Polyglot (Acc.)
- AIME 202479.80Jun 4, 2026AIME 2024 (Pass@1)
- 79.80Jun 5, 2026
- 70.00May 2, 2026
- AIME 202570.00Jun 4, 2026AIME 2025 (Pass@1)
- 70.00Jun 5, 2026
- 52.91May 19, 2026
- 87.60Jun 4, 2026
- 96.16May 10, 2026
- 15.80May 10, 2026
- Arena-Hard (GPT-4-1106 judge)92.30Jun 4, 2026ArenaHard (GPT-4-1106)
- Artificial Analysis Coding Index24.64Sep 9, 2026aa_coding_index
- C-Eval91.80Aug 31, 2026C-Eval (EM)
- Chinese SimpleQA (C-SimpleQA)63.70Sep 11, 2026C-SimpleQA (Correct)
- CLUEWSC92.80Aug 31, 2026CLUEWSC (EM)
- CNMO 202478.80Jul 3, 2026math_cnmo_2024_pass1
- 2029.00May 3, 2026
- Codeforces (Percentile)96.30Jul 3, 2026code_codeforces_percentile
- Codeforces (Rating)2029.00Jul 3, 2026code_codeforces_rating
- 1530.00Jun 4, 2026
- 92.20Jun 4, 2026
- 82.50Jun 4, 2026
- 70.10Jun 5, 2026
- 47.13May 10, 2026
- HMMT 202541.70Jun 4, 2026HMMT 2025 (Pass@1)
- 41.67May 11, 2026
- IFEval83.30Jun 4, 2026IF-Eval (Prompt Strict)
- livebench_language49.36Aug 23, 2026livebench_language@2025-04-07
- LiveCodeBench63.50Jun 4, 2026LiveCodeBench (2408-2505) (Pass@1)
- 55.90Jun 5, 2026
- 65.90Jun 4, 2026
- 97.30May 1, 2026
- MATH-500 (EM)97.30Sep 11, 2026MATH-500 (Pass@1)
- MATH-500 (EM)97.30Jun 5, 2026MATH-500
- MMLU90.80Jun 4, 2026MMLU (Pass@1)
- MMLU-Pro84.00Aug 31, 2026MMLU-Pro (EM)
- 84.00Jun 5, 2026
- 75.50May 26, 2025
- MMLU-Redux92.90Aug 31, 2026MMLU-Redux (EM)
- 35.80Jun 5, 2026
- 97.50May 10, 2026
- SimpleQA30.10Jun 4, 2026SimpleQA (Correct)
- 30.10Jun 5, 2026
- 4.76Sep 2, 2026
- 0.00May 10, 2026
- vectara_answer_rate97.00Aug 29, 2026Answer Rate
- vectara_avg_summary_length93.50Aug 29, 2026Average Summary Length (Words)
- vectara_factual_consistency88.70Aug 29, 2026Factual Consistency Rate
- 95.33May 10, 2026
- 78.70Jun 5, 2026
- τ²-Bench Telecom (AA run)11.40Oct 8, 2026aa_tau2
DeepSeek-R1: common questions
Who makes DeepSeek-R1?
DeepSeek-R1 is made by DeepSeek.
When was DeepSeek-R1 released?
DeepSeek-R1 was released on Jan 20, 2025, according to Artificial Analysis.
What is DeepSeek-R1 good at?
DeepSeek-R1 is behind the leaders in long context. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
How much does DeepSeek-R1 cost?
DeepSeek-R1 costs $2.00 per million input tokens and $4.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 74% of the 331 priced models we track.
How many benchmarks has DeepSeek-R1 been tested on?
We track 86 results for DeepSeek-R1 on 69 benchmarks from 23 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does DeepSeek-R1 support?
OpenRouter lists tool calling, json mode, and reasoning for DeepSeek-R1.
About this record
Where DeepSeek-R1's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Sep 14, 2026
Where the results come from
Verification: 86 scores · 26 independently verified · 18 aggregator-attributed · 14 vendor cross-reference · 28 vendor-reported. How these tiers are assigned
From 23 sources on 9 sites. Hugging Face supplies 33 of them; the 26 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co33
- artificialanalysis.ai18
- raw.githubusercontent.com15
- storage.googleapis.com10
- matharena.ai4
- arcprize.org2
- arxiv.org2
- aider.chat1
- simple-bench.com1