Kimi K2 Instruct
Kimi K2 Instruct is behind the leaders in long context, factuality, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.60input$2.50outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
152results on89benchmarks
- 15 independently verified
- 15 aggregator
- 101 vendor-reported
- 21 cross-referenced
From 17 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Kimi K2 Instruct benchmark results
152 results on 89 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
35.0% behind the leader1 of 3 ranked benchmarks measured
- 53.00Oct 8, 2026
37.6% behind the leader3 of 4 ranked benchmarks measured
- 17.90May 2, 2026
- AA-Omniscience · Accuracy25.42Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination30.28Oct 8, 2026omniscienceNonHallucination
43.5% behind the leader2 of 7 ranked benchmarks measured
- 18.17Jun 15, 2026
- 14.10Jun 6, 2026
Show 1 more agentic resultHide 1 agentic result
- 7.40May 15, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond76.67Oct 8, 2026gpqa
- Humanity's Last Exam6.39Oct 8, 2026aa_hle
Show 7 more reasoning resultsHide 7 reasoning results
- GPQA Diamond75.80Oct 7, 2026GPQA
- GPQA Diamond75.10Oct 7, 2026GPQA
- 77.00Jun 6, 2026
- 4.70Oct 7, 2026
- 4.70Oct 7, 2026
- Humanity's Last Exam7.90Jun 12, 2026HLE (Text-only) no tools
- Humanity's Last Exam6.30Jun 6, 2026HLE (w/o tools)
0 of 10 ranked benchmarks measured
- 27.67Oct 8, 2026
- 43.80May 1, 2026
- Terminal-Bench Hard23.48Oct 8, 2026aa_terminalbench_hard
Show 18 more coding resultsHide 18 coding results
- 53.70Oct 7, 2026
- 56.10May 15, 2026
- SciCode30.67Sep 4, 2026aa_scicode
- SciCode30.70May 15, 2026SciCode no tools
- 31.00Jun 6, 2026
- 47.30Oct 7, 2026
- 47.30Oct 7, 2026
- SWE-bench Multilingual55.90Jun 12, 2026SWE-bench Multilingual w/ tools
- 47.30May 15, 2026
- 55.90Jun 6, 2026
- SWE-bench Verified65.80Oct 7, 2026SWE-bench Verified (Agentic Coding)
- 65.80Oct 7, 2026
- SWE-bench Verified71.60Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified51.80Jun 15, 2026SWE-bench Verified (Agentless Coding)
- SWE-bench Verified69.20Jun 12, 2026SWE-bench Verified w/ tools
- 65.80May 15, 2026
- 69.20Jun 6, 2026
- 23.00Jun 6, 2026
0 of 5 ranked benchmarks measured
- IMO-AnswerBench45.80May 15, 2026IMO-AnswerBench no tools
0 of 3 ranked benchmarks measured
- IFBench41.70Oct 8, 2026aa_ifbench
- 54.10Oct 7, 2026
- 54.10Oct 7, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 42.00Jun 6, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence13.00Jul 6, 2026Artificial Analysis Intelligence Index
- vectara_avg_summary_length59.20Jul 6, 2026Average Summary Length (Words)
- AA Intelligence24.00Jun 20, 2026Artificial Analysis Intelligence Index
- 74.11May 19, 2026
- 98.22May 10, 2026
- 100.00May 10, 2026
Show 103 more resultsHide 103 results
- 37.73Jun 18, 2026
- AA Intelligence15.29Oct 8, 2026aa_intelligence_index
- -26.58Oct 8, 2026
- 76.50Oct 7, 2026
- 76.50Oct 7, 2026
- 60.00Oct 7, 2026
- 60.00Oct 7, 2026
- 69.60Jun 15, 2026
- 49.50Oct 7, 2026
- 49.50Oct 7, 2026
- AIME 202551.00May 15, 2026AIME25
- AIME 202557.00Jun 6, 2026AIME25
- 51.00Jun 12, 2026
- 75.20May 15, 2026
- 99.25May 10, 2026
- 54.20Jun 6, 2026
- 25.88Jun 18, 2026
- 89.50Oct 7, 2026
- 89.50Oct 7, 2026
- 94.90May 10, 2026
- 7.40Jun 12, 2026
- 22.20May 15, 2026
- 28.80Jun 6, 2026
- 22.20Jun 12, 2026
- 74.30Oct 7, 2026
- 74.30Oct 7, 2026
- 29.50Jun 6, 2026
- 10.40May 15, 2026
- 10.40Jun 12, 2026
- 58.10May 15, 2026
- 58.10Jun 12, 2026
- 60.20Jun 6, 2026
- GPQA (unspecified)74.20May 15, 2026GPQA no tools
- 97.30Oct 7, 2026
- 97.38May 10, 2026
- 43.80May 15, 2026
- 43.80Jun 12, 2026
- HLE (with tools)21.70May 15, 2026HLE (Text-only) w/ tools
- HLE (with tools)26.90Jun 6, 2026HLE (w/ tools)
- 38.80Oct 7, 2026
- 38.80Oct 7, 2026
- 38.80Jun 12, 2026
- 70.40May 15, 2026
- 94.50Oct 7, 2026
- 93.30Oct 7, 2026
- 89.80Aug 31, 2026
- 89.80Aug 23, 2026
- 76.40Oct 7, 2026
- 76.40Oct 7, 2026
- 53.70Aug 23, 2026
- LiveCodeBench61.00Jun 6, 2026LiveCodeBench (LCB)
- 56.10Jun 12, 2026
- Longform Writing eval (Kimi K2 Thinking system card)62.80Jun 12, 2026Longform Writing no tools
- MATH-500 (EM)97.40Oct 7, 2026MATH-500
- MATH-500 (EM)97.40Oct 7, 2026MATH-500
- 89.50Jun 15, 2026
- 82.50Oct 7, 2026
- 81.10Oct 7, 2026
- MMLU-Pro81.90May 15, 2026MMLU-Pro no tools
- 82.00Jun 6, 2026
- 92.70Oct 7, 2026
- 92.70Oct 7, 2026
- Multi-SWE-Bench33.50Jun 12, 2026Multi-SWE-bench w/ tools
- 31.30May 15, 2026
- 33.50Jun 6, 2026
- 85.70Oct 7, 2026
- 85.70Oct 7, 2026
- 76.40Oct 7, 2026
- 25.50May 15, 2026
- 25.50Jun 12, 2026
- 27.10Oct 7, 2026
- 27.10Oct 7, 2026
- 65.10Oct 7, 2026
- 65.10Oct 7, 2026
- 25.20May 15, 2026
- 25.20Jun 12, 2026
- 31.00Oct 7, 2026
- 31.00Oct 7, 2026
- 57.20Oct 7, 2026
- 57.20Oct 7, 2026
- 43.80May 1, 2026
- 61.90May 15, 2026
- 66.60May 15, 2026
- 56.50Oct 7, 2026
- 56.50Oct 7, 2026
- 30.00Sep 23, 2026
- 25.00Sep 1, 2026
- 25.00Jun 15, 2026
- 44.50May 15, 2026
- 37.50May 15, 2026
- 44.50Jun 6, 2026
- 44.50Jun 12, 2026
- theagentcompany30.00Jun 6, 2026AgentCompany
- vectara_answer_rate98.60May 2, 2026Answer Rate
- vectara_factual_consistency82.10May 2, 2026Factual Consistency Rate
- 61.00Jun 6, 2026
- 89.00Oct 7, 2026
- 89.00Oct 7, 2026
- τ²-Bench65.80Jun 15, 2026Tau2 telecom
- τ²-Bench73.00Jun 6, 2026τ²-Bench-Telecom
- τ²-Bench (Retail)70.60Oct 7, 2026Tau2 Retail
- τ²-Bench (Retail)70.60Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)73.39Oct 8, 2026aa_tau2
Kimi K2 Instruct: common questions
Who makes Kimi K2 Instruct?
Kimi K2 Instruct is made by Moonshot.
When was Kimi K2 Instruct released?
Kimi K2 Instruct was released on Sep 5, 2025, according to Artificial Analysis.
What is Kimi K2 Instruct good at?
Kimi K2 Instruct is behind the leaders in long context, factuality, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, multilingual tasks, or instruction following.
How much does Kimi K2 Instruct cost?
Kimi K2 Instruct costs $0.60 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 60% of the 330 priced models we track.
How many benchmarks has Kimi K2 Instruct been tested on?
We track 152 results for Kimi K2 Instruct on 89 benchmarks from 17 sources, 15 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Kimi K2 Instruct support?
OpenRouter lists tool calling and structured outputs for Kimi K2 Instruct.
About this record
Where Kimi K2 Instruct's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 2, 2026
Where the results come from
Verification: 152 scores · 15 independently verified · 15 aggregator-attributed · 21 vendor cross-reference · 101 vendor-reported. How these tiers are assigned
From 17 sources on 7 sites. Hugging Face supplies 66 of them; the 15 independently verified results come from 5 sites. Bars are coloured by trust tier.
- huggingface.co66
- api.llm-stats.com56
- artificialanalysis.ai17
- storage.googleapis.com6
- raw.githubusercontent.com4
- swebench.com2
- labs.scale.com1