Claude Sonnet 5 vs Kimi K2.6
Wins 6 of 8 areas
Coding · Agents · Reasoning · Facts · Math · Long documents
Wins 0 of 8 areas
—
Claude Sonnet 5 is the stronger all-rounder.Kimi K2.6 is cheaper.
Scores updated · 39 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software80Claude Sonnet 58 of 8 tests
- MathCompetition and research-level math40Claude Sonnet 54 of 4 tests
- AgentsCarrying out multi-step tasks on its own41Claude Sonnet 54 of 5 tests
- ReasoningHard problems that need careful thinking30Claude Sonnet 53 of 4 tests · 1 tie
- FactsGetting facts right instead of making them up21Claude Sonnet 52 of 3 tests
- Long documentsFinding answers in very long texts10Claude Sonnet 51 of 1 test
- Images and chartsUnderstanding pictures, charts and video11Even1 each · 1 tie
- Following instructionsDoing exactly what it is asked00Even0 each · 1 tie
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Kimi K2.6 costs 59% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Claude Sonnet 5 pulls ahead
- Proof problems from the US Math OlympiadUSAMO 2026+28.3points ahead
- Real work tasks from 44 professionsGDPVal+21.3points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+14.6points ahead
Where Kimi K2.6 pulls ahead
- Harder college exam questions with imagesMMMU-Pro+2.1points ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+1.6points ahead
- Short factual questions, answered correctlySimpleQA Verified+1.2points ahead
Every test, side by side
All 39 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Sonnet 5
- Terminal-Bench 2.1Claude Sonnet 5 by 14.680.565.9+14.6
- LiveBench · Agentic CodingClaude Sonnet 5 by 12.559.446.9+12.5
- SWE-bench VerifiedClaude Sonnet 5 by 585.280.2+5
- SWE-bench ProClaude Sonnet 5 by 4.663.258.6+4.6
- LMArena · WebDevClaude Sonnet 5 by 32 rating points15411509+32 rating
- SciCodeClaude Sonnet 5 by 2.854.351.5+2.8
- LiveBench · CodingClaude Sonnet 5 by 2.180.778.6+2.1
- SWE-bench MultilingualClaude Sonnet 5 by 1.678.376.7+1.6
AgentsClaude Sonnet 5
- GDPValClaude Sonnet 5 by 21.348.327+21.3
- τ-Bench V3 · BankingClaude Sonnet 5 by 1437.323.3+14
- Terminal-Bench 4.0Claude Sonnet 5 by 13.614.10.5+13.6
- OSWorld-VerifiedClaude Sonnet 5 by 8.181.273.1+8.1
- BrowseCompKimi K2.6 by 1.684.786.3+1.6
ReasoningClaude Sonnet 5
- LiveBench · ReasoningClaude Sonnet 5 by 9.388.779.4+9.3
- CritPtClaude Sonnet 5 by 8.916.98+8.9
- Humanity's Last ExamClaude Sonnet 5 by 3.841.337.5+3.8
- GPQA Diamondtie91.191.1tie
FactsClaude Sonnet 5
- AA-Omniscience · AccuracyClaude Sonnet 5 by 7.54032.6+7.5
- SimpleQA VerifiedKimi K2.6 by 1.233.734.9+1.2
- AA-Omniscience · Non-hallucinationClaude Sonnet 5 by 1.160.659.5+1.1
Images and chartsEven
- MMMU-ProKimi K2.6 by 2.177.379.4+2.1
- CharXiv (RQ)Claude Sonnet 5 by 1.688.386.7+1.6
- LMArena · Visiontie12751282tie
MathClaude Sonnet 5
- USAMO 2026Claude Sonnet 5 by 28.379.551.2+28.3
- LiveBench · MathematicsClaude Sonnet 5 by 8.692.984.3+8.6
- FrontierMath Tiers 1-3 (v2)Claude Sonnet 5 by 8.465.657.2+8.4
- FrontierMath Tier 4Claude Sonnet 5 by 3.729.325.6+3.7
Following instructionsEven
- LiveBench · Instruction Followingtie63.964.4tie
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDPval-AA v2Claude Sonnet 5 by 428 rating points16181190+428 rating
- AA Agentic IndexClaude Sonnet 5 by 22.244.322.1+22.2
- Terminal-Bench 2.0Claude Sonnet 5 by 13.780.466.7+13.7
- AA IntelligenceClaude Sonnet 5 by 11.238.227+11.2
- AA-OmniscienceClaude Sonnet 5 by 11.216.45.3+11.2
- Artificial Analysis Coding IndexClaude Sonnet 5 by 9.771.561.8+9.7
- livebench_data_analysisClaude Sonnet 5 by 6.671.765.1+6.6
- ToolathlonClaude Sonnet 5 by 4.354.350+4.3
- HLE (with tools)Claude Sonnet 5 by 3.457.454+3.4
- livebench_languagetie7575.1tie
Questions people ask
Which is better, Claude Sonnet 5 or Kimi K2.6?
Claude Sonnet 5 wins six of the eight areas we test: coding, agents, reasoning, facts, math and long documents. Kimi K2.6 wins none, but costs 59% less. They are level on images and charts and following instructions.
Which is better for coding?
Claude Sonnet 5. It wins 8 of the 8 coding tests both models report; Kimi K2.6 wins none.
Which is cheaper?
Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens; Kimi K2.6 costs $0.95 and $4.00. That makes Kimi K2.6 about 59% cheaper for the same work.
How do you compare the two?
We use the 39 benchmark tests both models have published scores on. The verdict counts the 29 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.