Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Kimi K2 Thinking

Score basis

Kimi K2 Thinking is behind the leaders in long context, factuality, instruction following, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited6 capabilities
  1. Long Context−28.0%2 of 3
  2. Factuality−28.0%2 of 4
  3. Instruction Following−29.1%2 of 3
  4. Coding−33.1%5 of 10
  5. Reasoning−33.2%4 of 6
  6. Agentic−41.2%2 of 7
Not rated4 capabilities

Too few results yet to rate Safety, Math, Multimodal or Multilingual.

Price

$0.60input$2.50outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 60% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

89results on67benchmarks

  • 9 independently verified
  • 15 aggregator
  • 38 vendor-reported
  • 27 cross-referenced

From 15 sources · latest Oct 8, 2026 · How verification works

API features

Tool calling

As listed by OpenRouter

Kimi K2 Thinking benchmark results

89 results on 67 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

LimitedFull long context ranking

28.0% behind the leader2 of 3 ranked benchmarks measured

Factuality

LimitedFull factuality ranking

28.0% behind the leader2 of 4 ranked benchmarks measured

Instruction Following

LimitedFull instruction following ranking

29.1% behind the leader2 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

Coding

LimitedFull coding ranking

33.1% behind the leader5 of 10 ranked benchmarks measured

Show 6 more coding resultsHide 6 coding results

Reasoning

LimitedFull reasoning ranking

33.2% behind the leader4 of 6 ranked benchmarks measured

Show 4 more reasoning resultsHide 4 reasoning results

Agentic

LimitedFull agentic ranking

41.2% behind the leader2 of 7 ranked benchmarks measured

Show 1 more agentic resultHide 1 agentic result

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 52 more resultsHide 52 results

Kimi K2 Thinking: common questions

Who makes Kimi K2 Thinking?

Kimi K2 Thinking is made by Moonshot.

When was Kimi K2 Thinking released?

Kimi K2 Thinking was released on Nov 6, 2025, according to Artificial Analysis.

What is Kimi K2 Thinking good at?

Kimi K2 Thinking is behind the leaders in long context, factuality, instruction following, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.

How much does Kimi K2 Thinking cost?

Kimi K2 Thinking costs $0.60 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 60% of the 331 priced models we track.

How many benchmarks has Kimi K2 Thinking been tested on?

We track 89 results for Kimi K2 Thinking on 67 benchmarks from 15 sources, 9 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Kimi K2 Thinking support?

OpenRouter lists tool calling for Kimi K2 Thinking.

About this record

Where Kimi K2 Thinking's numbers come from, and every name it appears under.

Tracked since
May 1, 2026
Newest source mention
Sep 1, 2026

Where the results come from

Verification: 89 scores · 9 independently verified · 15 aggregator-attributed · 27 vendor cross-reference · 38 vendor-reported. How these tiers are assigned

From 15 sources on 8 sites. Hugging Face supplies 65 of them; the 9 independently verified results come from 6 sites. Bars are coloured by trust tier.

  • huggingface.co65
  • artificialanalysis.ai15
  • matharena.ai3
  • swebench.com2
  • aider.chat1
  • epoch.ai1
  • labs.scale.com1
  • simple-bench.com1

Also known as

kimi k2 thinking (together)k2 thinkingkimi k2kimi-k2-thinking-20251106