Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Kimi K3

Score basis

Kimi K3 is strong in long context, agentic tasks, and reasoning; capable in factuality, multimodal tasks, and coding; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong3 capabilities
  1. Long Context−3.9%1 of 3
  2. Agentic−5.6%7 of 7
  3. Reasoning−9.3%6 of 6
Capable3 capabilities
  1. Factuality−12.2%3 of 4
  2. Multimodal−14.8%2 of 6
  3. Coding−19.1%5 of 10
Limited2 capabilities
  1. Math−30.6%5 of 5
  2. Instruction Following−37.2%1 of 3
Not rated2 capabilities

Too few results yet to rate Safety or Multilingual.

Price

$3.00input$15.00outputper million tokens

From Artificial Analysis · 6 providers tracked · All prices

Costs more than 88% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

131results on78benchmarks

  • 26 independently verified
  • 43 aggregator
  • 45 vendor-reported
  • 17 cross-referenced

From 18 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Kimi K3 benchmark results

131 results on 78 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

StrongFull long context ranking

3.9% behind the leader1 of 3 ranked benchmarks measured

Show 2 more long context resultsHide 2 long context results

5.6% behind the leader7 of 7 ranked benchmarks measured

Show 9 more agentic resultsHide 9 agentic results

Reasoning

StrongFull reasoning ranking

9.3% behind the leader6 of 6 ranked benchmarks measured

Show 10 more reasoning resultsHide 10 reasoning results

Factuality

CapableFull factuality ranking

12.2% behind the leader3 of 4 ranked benchmarks measured

Show 4 more factuality resultsHide 4 factuality results

Multimodal

CapableFull multimodal ranking

14.8% behind the leader2 of 6 ranked benchmarks measured

Show 4 more multimodal resultsHide 4 multimodal results

Coding

CapableFull coding ranking

19.1% behind the leader5 of 10 ranked benchmarks measured

Show 7 more coding resultsHide 7 coding results

30.6% behind the leader5 of 5 ranked benchmarks measured

Show 2 more math resultsHide 2 math results

Instruction Following

LimitedFull instruction following ranking

37.2% behind the leader1 of 3 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 60 more resultsHide 60 results

Kimi K3: common questions

Who makes Kimi K3?

Kimi K3 is made by Moonshot.

When was Kimi K3 released?

Kimi K3 was released on Jul 16, 2026, according to Artificial Analysis.

What is Kimi K3 good at?

Kimi K3 is strong in long context, agentic tasks, and reasoning; capable in factuality, multimodal tasks, and coding; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.

How much does Kimi K3 cost?

Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 6 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.

How many benchmarks has Kimi K3 been tested on?

We track 131 results for Kimi K3 on 78 benchmarks from 18 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Kimi K3 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Kimi K3.

About this record

Where Kimi K3's numbers come from, and every name it appears under.

Tracked since
Jul 2, 2026
Newest source mention
Sep 23, 2026

Where the results come from

Verification: 131 scores · 26 independently verified · 43 aggregator-attributed · 17 vendor cross-reference · 45 vendor-reported. How these tiers are assigned

From 18 sources on 11 sites. Artificial Analysis supplies 45 of them; the 26 independently verified results come from 9 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai45
  • huggingface.co38
  • api.llm-stats.com24
  • livebench.ai7
  • arcprize.org6
  • epoch.ai3
  • matharena.ai3
  • datasets-server.huggingface.co2
  • labs.scale.com1
  • lmarena.ai1
  • simple-bench.com1

Also known as

How our sources name Kimi K3 at each reasoning setting.

SettingShort formAPI id
lowkimi k3 (low)kimi-k3-low
highkimi k3 (high) kimi k3 high—
maxkimi k3 (max) kimmy k3 maxkimi-k3-max
Also listed askimmy k3kimi k-3kimi k threereteetzad/kimi-k3kim k3kimk3kim kim k3unsloth/kimi-k3moonshotai/kimi-k3unsloth/kimi-k3-ggufkimik3kimi k3 (think)kimi k.3moonshot ai kimi k3