Kimi K2 Thinking
Kimi K2 Thinking is behind the leaders in long context, factuality, instruction following, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal or Multilingual.
Price
$0.60input$2.50outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
89results on67benchmarks
- 9 independently verified
- 15 aggregator
- 38 vendor-reported
- 27 cross-referenced
From 15 sources · latest Oct 8, 2026 · How verification works
API features
Tool calling
As listed by OpenRouter
Kimi K2 Thinking benchmark results
89 results on 67 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
28.0% behind the leader2 of 3 ranked benchmarks measured
- 72.00Oct 8, 2026
- 45.10Jun 13, 2026
28.0% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy30.87Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination24.37Oct 8, 2026omniscienceNonHallucination
29.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench68.10Oct 8, 2026aa_ifbench
- 55.42Oct 8, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 66.40Jun 25, 2026
33.1% behind the leader5 of 10 ranked benchmarks measured
- 83.10Jun 13, 2026
- 63.40Sep 24, 2026
- 61.10Jun 15, 2026
- SciCode42.36Sep 4, 2026aa_scicode
- Terminal-Bench Hard31.06Oct 8, 2026aa_terminalbench_hard
Show 6 more coding resultsHide 6 coding results
- 83.10May 15, 2026
- SciCode44.80May 15, 2026SciCode no tools
- SWE-bench Multilingual61.10Jun 12, 2026SWE-bench Multilingual w/ tools
- SWE-bench Verified71.30Jun 12, 2026SWE-bench Verified w/ tools
- 71.30Jun 15, 2026
- 30.60Jun 13, 2026
33.2% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond83.84Oct 8, 2026gpqa
- Humanity's Last Exam23.82Oct 8, 2026aa_hle
- 26.30Sep 20, 2026
- 2.57Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond84.50May 15, 2026GPQA no tools
- 84.50Jun 13, 2026
- Humanity's Last Exam23.90Jun 12, 2026HLE (Text-only) no tools
- Humanity's Last Exam23.90Jun 25, 2026HLE_text
41.2% behind the leader2 of 7 ranked benchmarks measured
- 41.50May 30, 2026
- 24.51Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 60.20May 15, 2026
0 of 5 ranked benchmarks measured
- IMO-AnswerBench78.60May 15, 2026IMO-AnswerBench no tools
- 78.60May 30, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 59.10Oct 8, 2026
- frontiermath_tier_4_v10.00Sep 8, 2026frontiermath_tier_4
- 63.40Aug 31, 2026
- 92.50Jul 5, 2026
- 93.33May 11, 2026
- 89.17May 10, 2026
Show 52 more resultsHide 52 results
- 47.90Jun 18, 2026
- AA Intelligence22.02Oct 8, 2026aa_intelligence_index
- -21.42Oct 8, 2026
- AIME 202594.50May 15, 2026AIME25
- 94.50Jun 13, 2026
- 100.00May 15, 2026
- 94.50Jun 12, 2026
- 99.10May 15, 2026
- 80.10Jun 13, 2026
- 71.90Jun 13, 2026
- 34.83Jun 18, 2026
- 60.20Jun 12, 2026
- browsecomp_with_context_manager60.20Jul 9, 2026BrowseComp (w/ Context Manage)
- 62.30May 15, 2026
- 62.30May 30, 2026
- 62.30Jun 12, 2026
- 47.40May 15, 2026
- 47.40Jun 12, 2026
- 87.00May 15, 2026
- 87.00Jun 12, 2026
- 75.60May 30, 2026
- 58.00May 15, 2026
- 58.00Jun 12, 2026
- HLE (with tools)51.00May 15, 2026HLE (Text-only) heavy
- HLE (with tools)44.90May 15, 2026HLE (Text-only) w/ tools
- HLE (with tools)44.90May 18, 2026HLE (w/ Tools)
- HMMT 202589.40May 15, 2026HMMT25
- 86.50Jun 25, 2026
- 89.40Jun 13, 2026
- HMMT Nov. 202589.20May 30, 2026HMMT 2025 (Nov.)
- 97.50May 15, 2026
- 89.40Jun 12, 2026
- 95.10May 15, 2026
- 79.20Jun 25, 2026
- 83.10Jun 12, 2026
- Longform Writing eval (Kimi K2 Thinking system card)73.80Jun 12, 2026Longform Writing no tools
- MMLU-Pro84.60May 15, 2026MMLU-Pro no tools
- 84.60Jun 13, 2026
- MMLU-Redux94.40May 15, 2026MMLU-Redux no tools
- 44.20Jun 13, 2026
- Multi-SWE-Bench41.90Jun 12, 2026Multi-SWE-bench w/ tools
- 48.70May 15, 2026
- 48.70Jun 12, 2026
- 56.20May 30, 2026
- 56.30May 15, 2026
- 56.30Jun 12, 2026
- 47.10May 15, 2026
- Terminal-Bench 2.035.70Jun 15, 2026Terminal Bench 2
- 47.10Jun 12, 2026
- xbench-DeepSearch76.00May 30, 2026xbench-DeepSearch (2025.05)
- 74.30Jun 13, 2026
- τ²-Bench Telecom (AA run)92.98Oct 8, 2026aa_tau2
Kimi K2 Thinking: common questions
Who makes Kimi K2 Thinking?
Kimi K2 Thinking is made by Moonshot.
When was Kimi K2 Thinking released?
Kimi K2 Thinking was released on Nov 6, 2025, according to Artificial Analysis.
What is Kimi K2 Thinking good at?
Kimi K2 Thinking is behind the leaders in long context, factuality, instruction following, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
How much does Kimi K2 Thinking cost?
Kimi K2 Thinking costs $0.60 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 60% of the 331 priced models we track.
How many benchmarks has Kimi K2 Thinking been tested on?
We track 89 results for Kimi K2 Thinking on 67 benchmarks from 15 sources, 9 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Kimi K2 Thinking support?
OpenRouter lists tool calling for Kimi K2 Thinking.
About this record
Where Kimi K2 Thinking's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 89 scores · 9 independently verified · 15 aggregator-attributed · 27 vendor cross-reference · 38 vendor-reported. How these tiers are assigned
From 15 sources on 8 sites. Hugging Face supplies 65 of them; the 9 independently verified results come from 6 sites. Bars are coloured by trust tier.
- huggingface.co65
- artificialanalysis.ai15
- matharena.ai3
- swebench.com2
- aider.chat1
- epoch.ai1
- labs.scale.com1
- simple-bench.com1