Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

DeepSeek-R1

Score basis

DeepSeek-R1 is behind the leaders in long context. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, instruction following, or factuality.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited1 capability
  1. Long Context−33.4%2 of 3
Not rated9 capabilities

Too few results yet to rate Reasoning, Coding, Agentic, Safety, Math, Multimodal, Multilingual, Instruction Following or Factuality.

Price

$2.00input$4.00outputper million tokens

From Artificial Analysis · 3 providers tracked · All prices

Costs more than 74% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

86results on69benchmarks

  • 26 independently verified
  • 18 aggregator
  • 28 vendor-reported
  • 14 cross-referenced

From 23 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingJSON modeReasoning

As listed by OpenRouter

DeepSeek-R1 benchmark results

86 results on 69 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

LimitedFull long context ranking

33.4% behind the leader2 of 3 ranked benchmarks measured

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 6 more reasoning resultsHide 6 reasoning results

Coding

Not enough dataFull coding ranking

0 of 10 ranked benchmarks measured

Show 3 more coding resultsHide 3 coding results

Agentic

Not enough dataFull agentic ranking

0 of 7 ranked benchmarks measured

Instruction Following

Not enough dataFull instruction following ranking

0 of 3 ranked benchmarks measured

Factuality

Not enough dataFull factuality ranking

0 of 4 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 54 more resultsHide 54 results

DeepSeek-R1: common questions

Who makes DeepSeek-R1?

DeepSeek-R1 is made by DeepSeek.

When was DeepSeek-R1 released?

DeepSeek-R1 was released on Jan 20, 2025, according to Artificial Analysis.

What is DeepSeek-R1 good at?

DeepSeek-R1 is behind the leaders in long context. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, instruction following, or factuality.

How much does DeepSeek-R1 cost?

DeepSeek-R1 costs $2.00 per million input tokens and $4.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 74% of the 331 priced models we track.

How many benchmarks has DeepSeek-R1 been tested on?

We track 86 results for DeepSeek-R1 on 69 benchmarks from 23 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does DeepSeek-R1 support?

OpenRouter lists tool calling, json mode, and reasoning for DeepSeek-R1.

About this record

Where DeepSeek-R1's numbers come from, and every name it appears under.

Tracked since
Apr 25, 2026
Newest source mention
Sep 14, 2026

Where the results come from

Verification: 86 scores · 26 independently verified · 18 aggregator-attributed · 14 vendor cross-reference · 28 vendor-reported. How these tiers are assigned

From 23 sources on 9 sites. Hugging Face supplies 33 of them; the 26 independently verified results come from 8 sites. Bars are coloured by trust tier.

  • huggingface.co33
  • artificialanalysis.ai18
  • raw.githubusercontent.com15
  • storage.googleapis.com10
  • matharena.ai4
  • arcprize.org2
  • arxiv.org2
  • aider.chat1
  • simple-bench.com1

Also known as

deepseek r1 (hide reasoning)R1DeepSeek R1 (Jan '25)deepseek's r1deep seek's r1deepseek r1 (jan 2025)deepseek r1 (jan)deep seek r1deepseek-ai/deepseek-r1