Kimi K3
Kimi K3 is strong in long context, agentic tasks, and reasoning; capable in factuality, multimodal tasks, and coding; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · 6 providers tracked · All prices
Evidence
131results on78benchmarks
- 26 independently verified
- 43 aggregator
- 45 vendor-reported
- 17 cross-referenced
From 18 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
27 papers reference Kimi K3Kimi K3 benchmark results
131 results on 78 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
3.9% behind the leader1 of 3 ranked benchmarks measured
- 88.67Oct 8, 2026
5.6% behind the leader7 of 7 ranked benchmarks measured
- 91.20Oct 7, 2026
- 84.80Jul 27, 2026
- 84.20Oct 7, 2026
- τ-Bench V3 · Banking45.98Oct 8, 2026tauBanking
- 51.67Oct 8, 2026
- Terminal-Bench 4.012.63Oct 8, 2026
Show 9 more agentic resultsHide 9 agentic results
- 41.30Oct 8, 2026
- AA ApexAgents37.60Oct 7, 2026APEX-Agents
- AA ApexAgents41.00Jul 27, 2026APEX-Agents
- 47.69Oct 8, 2026
- 31.41Oct 8, 2026
- 59.37Jul 30, 2026
- 82.30Oct 8, 2026
- τ-Bench V3 · Banking41.65Oct 8, 2026tauBanking
- τ-Bench V3 · Banking33.40Jul 30, 2026tauBanking
9.3% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning90.67Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond93.54Oct 8, 2026gpqa
- 23.43Oct 8, 2026
- Humanity's Last Exam46.90Oct 8, 2026aa_hle
- 60.70Sep 21, 2026
- 60.42Sep 21, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 12.36Sep 21, 2026
- 55.00Sep 21, 2026
- 3.14Oct 8, 2026
- 23.40Jul 27, 2026
- GPQA Diamond84.24Oct 8, 2026gpqa
- GPQA Diamond93.50Oct 7, 2026GPQA
- Humanity's Last Exam24.98Oct 8, 2026aa_hle
- Humanity's Last Exam44.35Jul 30, 2026aa_hle
- 56.00Oct 7, 2026
- Humanity's Last Exam43.50Jul 27, 2026HLE-Full
12.2% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy47.58Oct 8, 2026omniscienceAccuracy
- 50.60Sep 21, 2026
- AA-Omniscience · Non-hallucination46.80Oct 8, 2026omniscienceNonHallucination
Show 4 more factuality resultsHide 4 factuality results
- AA-Omniscience · Accuracy45.75Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy45.95Jul 30, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination22.89Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination49.06Jul 30, 2026omniscienceNonHallucination
14.8% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)91.30Oct 7, 2026CharXiv-R
- MMMU-Pro80.52Oct 8, 2026aa_mmmu_pro
Show 4 more multimodal resultsHide 4 multimodal results
- CharXiv (reasoning)84.80Aug 24, 2026CharXiv (RQ)
- MMMU-Pro78.50Oct 8, 2026aa_mmmu_pro
- MMMU-Pro83.40Oct 7, 2026MMMU-Pro (with tools)
- 81.60Aug 24, 2026
19.1% behind the leader5 of 10 ranked benchmarks measured
- 1655.02Sep 21, 2026
- Terminal-Bench 2.185.02Oct 8, 2026terminalbenchV21
- LiveBench · Coding81.45Oct 8, 2026livebench_coding@2026-06-25
- SciCode59.49Oct 8, 2026aa_scicode
- LiveBench · Agentic Coding62.17Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 7 more coding resultsHide 7 coding results
- 1674.26Jul 17, 2026
- SciCode52.66Oct 8, 2026aa_scicode
- SciCode58.68Jul 30, 2026aa_scicode
- 58.70Jul 27, 2026
- Terminal-Bench 2.182.40Oct 8, 2026terminalbenchV21
- 88.30Oct 7, 2026
- 88.30Aug 28, 2026
30.6% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics84.44Oct 8, 2026livebench_math@2026-06-25
- 72.18Sep 21, 2026
- 39.02Sep 21, 2026
Show 2 more math resultsHide 2 math results
- 96.67Sep 2, 2026
- 96.97Sep 2, 2026
37.2% behind the leader1 of 3 ranked benchmarks measured
- LiveBench · Instruction Following71.36Oct 8, 2026livebench_instruction_following@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1488Oct 8, 2026
- livebench_language85.53Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis78.73Oct 8, 2026livebench_data_analysis@2026-06-25
- AA Intelligence44.00Sep 21, 2026Artificial Analysis Intelligence Index
- 65.67Sep 21, 2026
- 86.67Sep 21, 2026
Show 60 more resultsHide 60 results
- 50.65Sep 9, 2026
- 30.71Sep 9, 2026
- 50.07Jul 30, 2026
- AA Intelligence60.00Jul 18, 2026Artificial Analysis Intelligence Index
- AA Intelligence43.59Oct 8, 2026aa_intelligence_index
- AA Intelligence30.07Oct 8, 2026aa_intelligence_index
- AA Intelligence57.11Jul 30, 2026aa_intelligence_index
- 19.70Oct 8, 2026
- 3.92Oct 8, 2026
- 18.42Jul 30, 2026
- 28.30Jul 27, 2026
- 27.60Aug 13, 2026
- 27.60Aug 28, 2026
- 94.50Sep 21, 2026
- Artificial Analysis Coding Index76.24Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index71.98Sep 9, 2026aa_coding_index
- AutomationBench Public30.80Aug 13, 2026AutomationBench (Public)
- baby_vision_with_python85.70Aug 24, 2026BabyVision w/ python
- 85.70Oct 7, 2026
- 80.00Aug 28, 2026
- DeepSearchQA (F1)95.00Oct 7, 2026DeepSearchQA
- 67.50Oct 7, 2026
- DeepSWE67.50Aug 28, 2026DeepSWE (v1.1)
- 69.00Oct 7, 2026
- 73.70Aug 13, 2026
- 63.00Aug 13, 2026
- 54.40Aug 24, 2026
- 1682.00Aug 28, 2026
- GDPval-AA v2 Elo1686.00Jul 27, 2026GDPval-AA v2 (Elo)
- HLE (with tools)59.80Aug 28, 2026HLE w/ Tools
- 56.00Aug 13, 2026
- 52.90Oct 7, 2026
- 54.30Jul 27, 2026
- 97.80Oct 7, 2026
- 94.30Aug 24, 2026
- 91.25Sep 2, 2026
- 94.50Jul 27, 2026
- MLS Bench Litelower is better48.30Oct 7, 2026
- 82.10Aug 24, 2026
- 58.00Aug 28, 2026
- 63.30Oct 7, 2026
- 91.10Oct 7, 2026
- 58.30Aug 24, 2026
- 36.60Oct 7, 2026
- 32.00Aug 28, 2026
- 77.80Oct 7, 2026
- 17.50Aug 28, 2026
- 76.20Jul 27, 2026
- 34.80Oct 7, 2026
- 42.00Oct 7, 2026
- SWE-Marathon48.10Aug 28, 2026SWE-Marathon (v1.1)
- 17.40Aug 28, 2026
- 73.20Oct 7, 2026
- 76.50Jul 27, 2026
- 76.50Aug 28, 2026
- VideoMME (w sub.)90.00Aug 24, 2026Video-MME (w. sub)
- 51.00Oct 7, 2026
- 41.00Oct 7, 2026
- ZeroBench23.00Aug 24, 2026ZeroBench (pass@5)
- τ³-Bench Banking33.40Aug 24, 2026τ³-Banking
Kimi K3: common questions
Who makes Kimi K3?
Kimi K3 is made by Moonshot.
When was Kimi K3 released?
Kimi K3 was released on Jul 16, 2026, according to Artificial Analysis.
What is Kimi K3 good at?
Kimi K3 is strong in long context, agentic tasks, and reasoning; capable in factuality, multimodal tasks, and coding; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does Kimi K3 cost?
Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 6 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.
How many benchmarks has Kimi K3 been tested on?
We track 131 results for Kimi K3 on 78 benchmarks from 18 sources, 26 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Kimi K3 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Kimi K3.
About this record
Where Kimi K3's numbers come from, and every name it appears under.
- Tracked since
- Jul 2, 2026
- Newest source mention
- Sep 23, 2026
Where the results come from
Verification: 131 scores · 26 independently verified · 43 aggregator-attributed · 17 vendor cross-reference · 45 vendor-reported. How these tiers are assigned
From 18 sources on 11 sites. Artificial Analysis supplies 45 of them; the 26 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai45
- huggingface.co38
- api.llm-stats.com24
- livebench.ai7
- arcprize.org6
- epoch.ai3
- matharena.ai3
- datasets-server.huggingface.co2
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name Kimi K3 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | kimi k3 (low) | kimi-k3-low |
| high | kimi k3 (high) kimi k3 high | — |
| max | kimi k3 (max) kimmy k3 max | kimi-k3-max |