Kimi K2.6
Kimi K2.6 is strong in long context; capable in coding, multimodal tasks, instruction following, and reasoning; and behind the leaders in factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.95input$4.00outputper million tokens
From Artificial Analysis · 5 providers tracked · All prices
Evidence
181results on116benchmarks
- 26 independently verified
- 36 aggregator
- 44 vendor-reported
- 75 cross-referenced
From 24 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
17 papers reference Kimi K2.6Kimi K2.6 benchmark results
181 results on 116 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
9.6% behind the leader1 of 3 ranked benchmarks measured
- 81.00Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 69.67Oct 8, 2026
18.2% behind the leader10 of 10 ranked benchmarks measured
- 89.60Oct 7, 2026
- LiveBench · Coding78.57Oct 8, 2026livebench_coding@2026-06-25
- 80.20Oct 7, 2026
- 76.70Oct 7, 2026
- SciCode51.50Oct 8, 2026aa_scicode
- 1509.12May 22, 2026
- Terminal-Bench 2.165.92Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard43.94Oct 8, 2026aa_terminalbench_hard
- 58.60Oct 7, 2026
- LiveBench · Agentic Coding46.92Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 12 more coding resultsHide 12 coding results
- LiveCodeBench v690.20Jun 6, 2026LiveCodeBench (v6)
- SciCode39.47Sep 4, 2026aa_scicode
- 52.20Oct 7, 2026
- SWE-bench Multilingual76.70Jul 4, 2026SWE Multilingual (Resolved)
- 76.30Jun 15, 2026
- 77.10Jun 6, 2026
- SWE-bench Pro58.60Jul 4, 2026SWE Pro (Resolved)
- 31.00Jun 15, 2026
- 80.20Jul 16, 2026
- 75.70Jun 6, 2026
- 67.20Jun 6, 2026
- Terminal-Bench Hard37.88Oct 8, 2026aa_terminalbench_hard
19.1% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)86.70Oct 7, 2026CharXiv-R
- MMMU-Pro79.36Oct 8, 2026aa_mmmu_pro
Show 6 more multimodal resultsHide 6 multimodal results
- CharXiv (reasoning)80.40May 3, 2026CharXiv (RQ)
- 1282.39Aug 25, 2026
- 1264Jun 17, 2026
- 80.10Oct 7, 2026
- 79.40May 3, 2026
- MMMU-Pro79.00Jul 16, 2026MMMU Pro (Standard 10)
21.8% behind the leader3 of 3 ranked benchmarks measured
- IFBench75.99Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following64.36Oct 8, 2026livebench_instruction_following@2026-06-25
- 63.10Jun 10, 2026
24.1% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond91.11Oct 8, 2026gpqa
- LiveBench · Reasoning79.38Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam37.49Oct 8, 2026aa_hle
- 8.00Oct 8, 2026
Show 13 more reasoning resultsHide 13 reasoning results
- 1.43Oct 8, 2026
- 8.00Jul 30, 2026
- GPQA Diamond78.79Oct 8, 2026gpqa
- GPQA Diamond90.50Oct 7, 2026GPQA
- GPQA Diamond91.00Aug 24, 2026GPQA (no tools)
- 91.10Jul 16, 2026
- GPQA Diamond90.50Jul 4, 2026GPQA Diamond (Pass@1)
- Humanity's Last Exam19.56Oct 8, 2026aa_hle
- 36.40Oct 7, 2026
- Humanity's Last Exam34.70May 19, 2026HLE-Full
- Humanity's Last Exam35.90Jul 16, 2026HLE (text only)
- Humanity's Last Exam36.40Jun 27, 2026HLE (Pass@1)
- Humanity's Last Exam34.80Jun 6, 2026HLE (no tools)
26.4% behind the leader4 of 4 ranked benchmarks measured
- 10.80May 12, 2026
- AA-Omniscience · Non-hallucination59.50Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy32.60Oct 8, 2026omniscienceAccuracy
- 34.90May 20, 2026
Show 4 more factuality resultsHide 4 factuality results
- AA-Omniscience · Accuracy24.30Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination55.77Oct 8, 2026omniscienceNonHallucination
- 38.70Jul 16, 2026
- SimpleQA Verified36.90Jun 27, 2026SimpleQA-Verified (Pass@1)
29.5% behind the leader6 of 7 ranked benchmarks measured
- 86.30Oct 7, 2026
- 73.10Oct 7, 2026
- 68.10Jul 16, 2026
- τ-Bench V3 · Banking23.30Oct 8, 2026tauBanking
- 27.01Oct 8, 2026
- Terminal-Bench 4.00.51Oct 8, 2026
Show 9 more agentic resultsHide 9 agentic results
- 28.47Oct 8, 2026
- AA ApexAgents27.90Oct 7, 2026APEX-Agents
- 31.19Oct 8, 2026
- 83.20May 3, 2026
- BrowseComp83.20Jun 27, 2026BrowseComp (Pass@1)
- 61.30Jun 6, 2026
- 41.39Jun 15, 2026
- GDPval (win rate)50.40Jun 6, 2026GDPVal
- 23.10Jun 6, 2026
42.1% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics84.28Oct 8, 2026livebench_math@2026-06-25
- 57.19Sep 9, 2026
- 25.64Sep 8, 2026
Show 11 more math resultsHide 11 math results
- 95.83Sep 2, 2026
- 95.83May 2, 2026
- 96.40Oct 7, 2026
- 96.40Jul 16, 2026
- 94.70Sep 2, 2026
- 94.70May 10, 2026
- HMMT Feb 202692.70Oct 7, 2026HMMT Feb 26
- HMMT Feb 202692.70Jul 4, 2026HMMT 2026 Feb (Pass@1)
- 86.00Oct 7, 2026
- IMO-AnswerBench86.00Jul 4, 2026IMOAnswerBench (Pass@1)
- 51.19Sep 2, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis65.13Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language75.14Oct 8, 2026livebench_language@2026-06-25
- 18.67Oct 8, 2026
- AA Intelligence27.00Oct 1, 2026Artificial Analysis Intelligence Index
- 1461Sep 1, 2026
- frontiermath_tier_4_v114.58May 20, 2026frontiermath_tier_4
Show 84 more resultsHide 84 results
- 22.14Sep 9, 2026
- 58.73Jun 18, 2026
- AA Intelligence23.57Oct 8, 2026aa_intelligence_index
- AA Intelligence26.98Oct 8, 2026aa_intelligence_index
- -9.18Oct 8, 2026
- 5.30Oct 8, 2026
- 58.93Jun 24, 2026
- 65.23Jun 24, 2026
- 60.80Jun 24, 2026
- 53.42Jun 24, 2026
- 27.48Jun 24, 2026
- 58.77Jun 24, 2026
- 52.54Jun 24, 2026
- 50.20Jun 24, 2026
- 24.00Jul 4, 2026
- 75.50Jul 4, 2026
- 77.40Jun 6, 2026
- 73.20Jun 6, 2026
- Artificial Analysis Coding Index61.77Sep 9, 2026aa_coding_index
- 38.41Jun 18, 2026
- baby_vision_with_python68.50May 3, 2026BabyVision (w/ python)
- 68.50Oct 7, 2026
- 39.80May 3, 2026
- 86.30May 3, 2026
- 83.20Jul 16, 2026
- charxiv_rq_with_python86.70May 3, 2026CharXiv (RQ) (w/ python)
- charxiv_rq_with_python86.70Jul 16, 2026Charxiv RQ (with python)
- Chinese SimpleQA (C-SimpleQA)75.90Jul 4, 2026Chinese-SimpleQA (Pass@1)
- Claw Eval (pass@3)80.90Oct 7, 2026Claw-Eval
- 9.10Jun 6, 2026
- 83.00May 3, 2026
- DeepSearchQA (F1)83.00Oct 7, 2026DeepSearchQA
- DeepSearchQA (F1)92.50May 3, 2026DeepSearchQA (f1-score)
- 44.87Oct 7, 2026
- 1482.00Jun 27, 2026
- 1190.00Jul 30, 2026
- 88.40Jul 16, 2026
- GPQA (unspecified)91.00Jun 6, 2026GPQA (no tools)
- HLE (with tools)54.00May 3, 2026HLE-Full (w/ tools)
- 54.00Jul 16, 2026
- 73.70Jun 10, 2026
- 93.71Jun 6, 2026
- 585.00Jun 6, 2026
- 72.17Oct 7, 2026
- LiveCodeBench89.60Jul 4, 2026LiveCodeBench (Pass@1)
- 93.20Oct 7, 2026
- 87.40May 3, 2026
- mathvision_with_python93.20May 3, 2026MathVision (w/ python)
- 66.60Jun 27, 2026
- 55.90Oct 7, 2026
- MMLU-Pro87.10Jul 4, 2026MMLU-Pro (EM)
- 88.10Jun 6, 2026
- 85.00Jun 10, 2026
- 79.40Jun 3, 2026
- mmmu_pro_with_python80.10May 3, 2026MMMU-Pro (w/ python)
- 60.60Oct 7, 2026
- 60.60May 3, 2026
- 90.20Jun 6, 2026
- 56.00Jun 6, 2026
- 0.13Jul 30, 2026
- 52.00Jun 6, 2026
- 99.80Jul 16, 2026
- 71.60Jun 15, 2026
- SWEBench Pro Public58.60Jul 16, 2026SWEBench Pro (Public)
- 85.80Jun 6, 2026
- 72.40Jun 6, 2026
- 82.90Jun 6, 2026
- TauBench V3 - Telecom97.80Aug 24, 2026Telecom
- 66.70Oct 7, 2026
- Terminal-Bench 2.066.70Jul 4, 2026Terminal Bench 2.0 (Acc)
- 50.00Oct 7, 2026
- Toolathlon50.00Jun 6, 2026Toolathlon (Pass@1)
- 96.90May 3, 2026
- 58.80Jun 6, 2026
- 54.00Jun 6, 2026
- vectara_answer_rate99.70May 12, 2026Answer Rate
- vectara_avg_summary_length116.70May 12, 2026Average Summary Length (Words)
- vectara_factual_consistency89.20May 12, 2026Factual Consistency Rate
- 80.80Oct 7, 2026
- 80.80May 3, 2026
- 84.50Jun 10, 2026
- τ²-Bench Telecom (AA run)93.86Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.91Oct 8, 2026aa_tau2
- τ³-Bench Banking20.60Jul 30, 2026τ³-Banking
Kimi K2.6: common questions
Who makes Kimi K2.6?
Kimi K2.6 is made by Moonshot.
When was Kimi K2.6 released?
Kimi K2.6 was released on Apr 20, 2026, according to Artificial Analysis.
What is Kimi K2.6 good at?
Kimi K2.6 is strong in long context; capable in coding, multimodal tasks, instruction following, and reasoning; and behind the leaders in factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
How much does Kimi K2.6 cost?
Kimi K2.6 costs $0.95 per million input tokens and $4.00 per million output tokens, according to Artificial Analysis. We track its price at 5 providers. At a mix of three input tokens to one output token, it costs more than 68% of the 331 priced models we track.
How many benchmarks has Kimi K2.6 been tested on?
We track 181 results for Kimi K2.6 on 116 benchmarks from 24 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Kimi K2.6 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Kimi K2.6.
About this record
Where Kimi K2.6's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 181 scores · 26 independently verified · 36 aggregator-attributed · 75 vendor cross-reference · 44 vendor-reported. How these tiers are assigned
From 24 sources on 11 sites. Hugging Face supplies 89 of them; the 26 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co89
- artificialanalysis.ai37
- api.llm-stats.com26
- livebench.ai7
- matharena.ai5
- epoch.ai4
- raw.githubusercontent.com4
- thinkingmachines.ai4
- datasets-server.huggingface.co2
- lmarena.ai2
- labs.scale.com1