Kimi K2.5
Kimi K2.5 is at the frontier in multimodal tasks; capable in long context and instruction following; and behind the leaders in coding, reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$0.60input$3.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
172results on108benchmarks
- 26 independently verified
- 35 aggregator
- 66 vendor-reported
- 45 cross-referenced
From 27 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
18 papers reference Kimi K2.5Kimi K2.5 benchmark results
172 results on 108 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
6.5% behind the leader5 of 6 ranked benchmarks measured
- 92.30Oct 7, 2026
- MathVista90.10Oct 7, 2026MathVista-Mini
- 87.40Oct 7, 2026
- MMMU-Pro75.38Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)77.50Oct 7, 2026CharXiv-R
Show 5 more multimodal resultsHide 5 multimodal results
10.8% behind the leader2 of 3 ranked benchmarks measured
- 61.00Oct 7, 2026
- 78.00Oct 8, 2026
24.5% behind the leader2 of 3 ranked benchmarks measured
- IFBench70.20Oct 8, 2026aa_ifbench
- 61.39Oct 8, 2026
27.5% behind the leader8 of 10 ranked benchmarks measured
- 85.00Oct 7, 2026
- 76.80Oct 7, 2026
- 73.00Oct 7, 2026
- SciCode48.96Sep 4, 2026aa_scicode
- 1436.41May 22, 2026
- 50.70Oct 7, 2026
- Terminal-Bench Hard34.85Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.145.69Oct 8, 2026terminalbenchV21
Show 9 more coding resultsHide 9 coding results
- 85.00May 30, 2026
- SciCode39.58Sep 4, 2026aa_scicode
- 48.70Oct 7, 2026
- 67.30May 1, 2026
- 73.00May 17, 2026
- 53.80May 18, 2026
- 70.80Sep 25, 2026
- 76.80Jul 16, 2026
- Terminal-Bench Hard18.94Oct 8, 2026aa_terminalbench_hard
27.6% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond87.88Oct 8, 2026gpqa
- Humanity's Last Exam30.72Oct 8, 2026aa_hle
- 11.81May 10, 2026
- 3.14Oct 8, 2026
Show 12 more reasoning resultsHide 12 reasoning results
- 0.57Oct 8, 2026
- 3.10Jul 30, 2026
- GPQA Diamond78.89Oct 8, 2026gpqa
- GPQA Diamond87.60Oct 7, 2026GPQA
- 87.90Jul 16, 2026
- 87.60May 18, 2026
- Humanity's Last Exam13.16Oct 8, 2026aa_hle
- 50.20Oct 7, 2026
- Humanity's Last Exam30.10Jun 15, 2026HLE-Full
- Humanity's Last Exam29.40Jul 16, 2026HLE (text only)
- 31.50May 18, 2026
- 46.80May 10, 2026
34.9% behind the leader4 of 4 ranked benchmarks measured
- 14.20May 2, 2026
- AA-Omniscience · Accuracy35.22Oct 8, 2026omniscienceAccuracy
- 34.30May 20, 2026
- AA-Omniscience · Non-hallucination34.32Oct 8, 2026omniscienceNonHallucination
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy24.13Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination50.04Oct 8, 2026omniscienceNonHallucination
- 36.90Jul 16, 2026
40.3% behind the leader5 of 7 ranked benchmarks measured
- 74.90Oct 7, 2026
- 63.30May 3, 2026
- 64.40Oct 8, 2026
- τ-Bench V3 · Banking14.23Oct 8, 2026tauBanking
- 17.16Oct 8, 2026
Show 7 more agentic resultsHide 7 agentic results
- 11.50Oct 8, 2026
- AA ApexAgents11.50May 3, 2026APEX-Agents
- 60.60Jun 15, 2026
- 60.60May 30, 2026
- 38.28Jun 15, 2026
- 64.00Jul 16, 2026
- MCP Atlas63.80May 18, 2026MCP-Atlas (Public Set)
0 of 5 ranked benchmarks measured
- 87.12Sep 2, 2026
- 95.83Sep 2, 2026
- 87.12May 10, 2026
Show 8 more math resultsHide 8 math results
- 95.83May 2, 2026
- 95.80May 3, 2026
- 95.80Jul 16, 2026
- 94.50May 18, 2026
- HMMT Feb 202687.10May 3, 2026HMMT 2026 (Feb)
- HMMT Feb 202681.30May 18, 2026HMMT Feb. 2026
- 81.80Oct 7, 2026
- 81.80May 30, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 80.56Sep 2, 2026
- 70.80Aug 29, 2026
- 1450Jul 23, 2026
- 95.83Jul 5, 2026
- frontiermath_tier_4_v14.20May 20, 2026frontiermath_tier_4
- 93.33May 11, 2026
Show 85 more resultsHide 85 results
- 21.69Sep 4, 2026
- 52.84Jun 18, 2026
- AA Intelligence23.46Oct 8, 2026aa_intelligence_index
- AA Intelligence19.43Oct 8, 2026aa_intelligence_index
- -7.33Oct 8, 2026
- -13.77Oct 8, 2026
- 96.10Oct 7, 2026
- 96.10May 30, 2026
- 92.50May 17, 2026
- 65.33May 10, 2026
- Artificial Analysis Coding Index46.78Sep 9, 2026aa_coding_index
- 25.82Jun 18, 2026
- baby_vision_with_python40.50May 3, 2026BabyVision (w/ python)
- 36.50May 3, 2026
- 78.40Jun 15, 2026
- BrowseComp (context management)74.90Aug 24, 2026BrowseComp (w/ctx manage)
- 74.90Jul 16, 2026
- browsecomp_with_context_manager74.90May 30, 2026BrowseComp (w/ Context Manager)
- 62.30May 30, 2026
- charxiv_rq_with_python78.70May 3, 2026CharXiv (RQ) (w/ python)
- charxiv_rq_with_python78.70Jul 16, 2026Charxiv RQ (with python)
- 75.40May 3, 2026
- 41.30Oct 7, 2026
- 41.30May 18, 2026
- 77.10May 3, 2026
- DeepSearchQA (F1)77.10Oct 7, 2026DeepSearchQA
- DeepSearchQA (F1)89.00May 3, 2026DeepSearchQA (f1-score)
- 67.80Oct 7, 2026
- 75.90May 30, 2026
- 1009.00Jul 30, 2026
- 84.00Jul 16, 2026
- HLE (with tools)50.20Jun 15, 2026HLE-Full (w/ tools)
- 50.20Jul 16, 2026
- HLE (with tools)51.80May 18, 2026HLE (w/ Tools)
- 95.40Oct 7, 2026
- HMMT Feb. 202595.40Jun 15, 2026HMMT 2025 (Feb)
- HMMT Feb. 202595.40May 30, 2026HMMT 2025 (Feb.)
- 89.17May 10, 2026
- 91.10May 18, 2026
- 92.60Jun 15, 2026
- 92.60Oct 7, 2026
- 69.07Oct 7, 2026
- 79.80Oct 7, 2026
- 75.90Oct 7, 2026
- 84.20Oct 7, 2026
- mathvision_with_python85.00May 3, 2026MathVision (w/ python)
- 29.50May 3, 2026
- 87.10Oct 7, 2026
- mmmu_pro_with_python77.70May 3, 2026MMMU-Pro (w/ python)
- 80.40Oct 7, 2026
- 70.40Oct 7, 2026
- 32.00May 18, 2026
- 57.40Jun 15, 2026
- 54.70May 3, 2026
- 88.80Oct 7, 2026
- 63.50Oct 7, 2026
- 59.50May 30, 2026
- 57.40Oct 7, 2026
- 71.20Oct 7, 2026
- 99.50Jul 16, 2026
- SWEBench Pro Public50.70Jul 16, 2026SWEBench Pro (Public)
- 50.80Oct 7, 2026
- Terminal-Bench 2.050.80May 18, 2026Terminal-Bench 2.0 (Terminus-2)
- 27.80May 18, 2026
- 27.80May 3, 2026
- 86.90May 3, 2026
- vectara_answer_rate92.20May 2, 2026Answer Rate
- vectara_avg_summary_length112.00May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency85.80May 2, 2026Factual Consistency Rate
- 86.60Oct 7, 2026
- 79.00Oct 7, 2026
- 79.00Jun 15, 2026
- 72.70Jun 15, 2026
- 46.30Oct 7, 2026
- xbench-DeepSearch76.70May 30, 2026xbench-DeepSearch (2025.05)
- 11.00Oct 7, 2026
- 9.00Jun 15, 2026
- 11.00Jun 15, 2026
- 80.20Jun 13, 2026
- 85.40May 30, 2026
- τ²-Bench Telecom (AA run)81.29Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.91Oct 8, 2026aa_tau2
- 66.00May 18, 2026
- τ³-Bench Banking14.20Jul 30, 2026τ³-Banking
- τ³-Bench Banking13.20Jul 16, 2026Tau 3 Banking
Kimi K2.5: common questions
Who makes Kimi K2.5?
Kimi K2.5 is made by Moonshot.
When was Kimi K2.5 released?
Kimi K2.5 was released on Jan 27, 2026, according to Artificial Analysis.
What is Kimi K2.5 good at?
Kimi K2.5 is at the frontier in multimodal tasks; capable in long context and instruction following; and behind the leaders in coding, reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does Kimi K2.5 cost?
Kimi K2.5 costs $0.60 per million input tokens and $3.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 62% of the 331 priced models we track.
How many benchmarks has Kimi K2.5 been tested on?
We track 172 results for Kimi K2.5 on 108 benchmarks from 27 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Kimi K2.5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Kimi K2.5.
About this record
Where Kimi K2.5's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 172 scores · 26 independently verified · 35 aggregator-attributed · 45 vendor cross-reference · 66 vendor-reported. How these tiers are assigned
From 27 sources on 13 sites. Hugging Face supplies 70 of them; the 26 independently verified results come from 9 sites. Bars are coloured by trust tier.
- huggingface.co70
- api.llm-stats.com38
- artificialanalysis.ai35
- matharena.ai8
- raw.githubusercontent.com4
- swebench.com3
- thinkingmachines.ai3
- arcprize.org2
- datasets-server.huggingface.co2
- epoch.ai2
- labs.scale.com2
- lmarena.ai2
- simple-bench.com1