Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Kimi K2.5

Score basis

Kimi K2.5 is at the frontier in multimodal tasks; capable in long context and instruction following; and behind the leaders in coding, reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Frontier1 capability
  1. Multimodal−6.5%5 of 6
Capable2 capabilities
  1. Long Context−10.8%2 of 3
  2. Instruction Following−24.5%2 of 3
Limited4 capabilities
  1. Coding−27.5%8 of 10
  2. Reasoning−27.6%4 of 6
  3. Factuality−34.9%4 of 4
  4. Agentic−40.3%5 of 7
Not rated3 capabilities

Too few results yet to rate Safety, Math or Multilingual.

Price

$0.60input$3.00outputper million tokens

From Artificial Analysis · 3 providers tracked · All prices

Costs more than 62% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

172results on108benchmarks

  • 26 independently verified
  • 35 aggregator
  • 66 vendor-reported
  • 45 cross-referenced

From 27 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Kimi K2.5 benchmark results

172 results on 108 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Multimodal

FrontierFull multimodal ranking

6.5% behind the leader5 of 6 ranked benchmarks measured

Show 5 more multimodal resultsHide 5 multimodal results

Long Context

CapableFull long context ranking

10.8% behind the leader2 of 3 ranked benchmarks measured

Show 2 more long context resultsHide 2 long context results

Instruction Following

CapableFull instruction following ranking

24.5% behind the leader2 of 3 ranked benchmarks measured

Show 2 more instruction following resultsHide 2 instruction following results

Coding

LimitedFull coding ranking

27.5% behind the leader8 of 10 ranked benchmarks measured

Show 9 more coding resultsHide 9 coding results

Reasoning

LimitedFull reasoning ranking

27.6% behind the leader4 of 6 ranked benchmarks measured

Show 12 more reasoning resultsHide 12 reasoning results

Factuality

LimitedFull factuality ranking

34.9% behind the leader4 of 4 ranked benchmarks measured

Show 3 more factuality resultsHide 3 factuality results

Agentic

LimitedFull agentic ranking

40.3% behind the leader5 of 7 ranked benchmarks measured

Show 7 more agentic resultsHide 7 agentic results

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

Show 8 more math resultsHide 8 math results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 85 more resultsHide 85 results

Kimi K2.5: common questions

Who makes Kimi K2.5?

Kimi K2.5 is made by Moonshot.

When was Kimi K2.5 released?

Kimi K2.5 was released on Jan 27, 2026, according to Artificial Analysis.

What is Kimi K2.5 good at?

Kimi K2.5 is at the frontier in multimodal tasks; capable in long context and instruction following; and behind the leaders in coding, reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

How much does Kimi K2.5 cost?

Kimi K2.5 costs $0.60 per million input tokens and $3.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 62% of the 331 priced models we track.

How many benchmarks has Kimi K2.5 been tested on?

We track 172 results for Kimi K2.5 on 108 benchmarks from 27 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Kimi K2.5 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Kimi K2.5.

About this record

Where Kimi K2.5's numbers come from, and every name it appears under.

Tracked since
May 1, 2026
Newest source mention
Sep 1, 2026

Where the results come from

Verification: 172 scores · 26 independently verified · 35 aggregator-attributed · 45 vendor cross-reference · 66 vendor-reported. How these tiers are assigned

From 27 sources on 13 sites. Hugging Face supplies 70 of them; the 26 independently verified results come from 9 sites. Bars are coloured by trust tier.

  • huggingface.co70
  • api.llm-stats.com38
  • artificialanalysis.ai35
  • matharena.ai8
  • raw.githubusercontent.com4
  • swebench.com3
  • thinkingmachines.ai3
  • arcprize.org2
  • datasets-server.huggingface.co2
  • epoch.ai2
  • labs.scale.com2
  • lmarena.ai2
  • simple-bench.com1

Also known as

kimi k2.5 (fireworks)kimi k2.5 (high reasoning)kimi k2.5 (high)kimi k2.5 (non-reasoning)kimi k2.5 (reasoning)kimi k2.5 (think)kimi-k2-5kimi-k2-5-non-reasoningkimi-k2.5-thinkingkimi-k2p5kimik 2.5kimmy k 2.5kimmy k2.5