Kimi K2.5 vs Muse Spark
Wins 0 of 7 areas
—
Wins 6 of 7 areas
Coding · Agents · Reasoning · Facts · Images and charts · Following instructions
Muse Spark is the stronger all-rounder.
Scores updated · 32 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software05Muse Spark5 of 6 tests · 1 tie
- ReasoningHard problems that need careful thinking03Muse Spark3 of 4 tests · 1 tie
- Images and chartsUnderstanding pictures, charts and video03Muse Spark3 of 3 tests
- AgentsCarrying out multi-step tasks on its own02Muse Spark2 of 2 tests
- Following instructionsDoing exactly what it is asked02Muse Spark2 of 2 tests
- FactsGetting facts right instead of making them up12Muse Spark2 of 3 tests
- Long documentsFinding answers in very long texts00Even0 each · 1 tie
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Muse Spark pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+32points ahead
- Abstract visual puzzles that people can solveARC-AGI-2+30.7points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+16.5points ahead
Where Kimi K2.5 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+18.5points ahead
Every test, side by side
All 32 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingMuse Spark
- Terminal-Bench 2.1Muse Spark by 16.545.762.2+16.5
- Terminal-Bench HardMuse Spark by 10.734.845.5+10.7
- LMArena · WebDevMuse Spark by 104 rating points14361540+104 rating
- SciCodeMuse Spark by 2.54951.5+2.5
- SWE-bench ProMuse Spark by 1.750.752.4+1.7
- SWE-bench Verifiedtie76.877.4tie
AgentsMuse Spark
- GDPValMuse Spark by 7.917.225.1+7.9
- τ-Bench V3 · BankingMuse Spark by 5.414.219.6+5.4
ReasoningMuse Spark
- ARC-AGI-2Muse Spark by 30.711.842.5+30.7
- Humanity's Last ExamMuse Spark by 1030.740.7+10
- CritPtMuse Spark by 8.23.111.3+8.2
- GPQA Diamondtie87.988.4tie
FactsMuse Spark
- SimpleQA VerifiedMuse Spark by 3234.366.3+32
- AA-Omniscience · Non-hallucinationKimi K2.5 by 18.534.315.8+18.5
- AA-Omniscience · AccuracyMuse Spark by 14.435.249.6+14.4
Images and chartsMuse Spark
- CharXiv (RQ)Muse Spark by 8.977.586.4+8.9
- MMMU-ProMuse Spark by 5.175.480.5+5.1
- LMArena · VisionMuse Spark by 36 rating points12691305+36 rating
Following instructionsMuse Spark
- Multi-ChallengeMuse Spark by 14.161.475.5+14.1
- IFBenchMuse Spark by 5.770.275.9+5.7
Other results11 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ZeroBenchMuse Spark by 221133+22
- AA-OmniscienceMuse Spark by 14.5-7.37.2+14.5
- Artificial Analysis Coding IndexMuse Spark by 11.846.858.6+11.8
- frontiermath_tier_4_v1Muse Spark by 10.44.214.6+10.4
- Terminal-Bench 2.0Muse Spark by 8.250.859+8.2
- AA IntelligenceMuse Spark by 7.823.531.3+7.8
- AA Agentic IndexMuse Spark by 721.728.7+7
- τ²-Bench Telecom (AA run)Kimi K2.5 by 4.495.991.5+4.4
- DeepSearchQA (F1)Kimi K2.5 by 2.377.174.8+2.3
- CyberGymMuse Spark by 2.241.343.5+2.2
- SimpleVQAtie71.271.3tie
Questions people ask
Which is better, Kimi K2.5 or Muse Spark?
Muse Spark wins six of the seven areas where both have results: coding, agents, reasoning, facts, images and charts and following instructions. Kimi K2.5 wins none. They are level on long documents.
Which is better for coding?
Muse Spark. It wins 5 of the 6 coding tests both models report; Kimi K2.5 wins none, and 1 is a tie.
How do you compare the two?
We use the 32 benchmark tests both models have published scores on. The verdict counts the 21 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 11 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.