Kimi K2 Instruct vs Qwen3 Next 80B A3B
Wins 2 of 6 areas
Agents · Facts
Wins 4 of 6 areas
Coding · Reasoning · Long documents · Following instructions
Qwen3 Next 80B A3B is the stronger all-rounder.Kimi K2 Instruct is better at facts.
Scores updated · 34 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software12Qwen3 Next 80B A3B2 of 3 tests
- ReasoningHard problems that need careful thinking01Qwen3 Next 80B A3B1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts01Qwen3 Next 80B A3B1 of 1 test
- Following instructionsDoing exactly what it is asked01Qwen3 Next 80B A3B1 of 1 test
- FactsGetting facts right instead of making them up21Kimi K2 Instruct2 of 3 tests
- AgentsCarrying out multi-step tasks on its own10Kimi K2 Instruct1 of 1 test
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Qwen3 Next 80B A3B costs 56% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Qwen3 Next 80B A3B pulls ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+19points ahead
- Recent programming contest problemsLiveCodeBench v6+15points ahead
- Reasons across sets of long documentsAA-LCR+10.7points ahead
Where Kimi K2 Instruct pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+17.3points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+13.7points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+6.4points ahead
Every test, side by side
All 34 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 Next 80B A3B
- LiveCodeBench v6Qwen3 Next 80B A3B by 1553.768.7+15
- Terminal-Bench HardKimi K2 Instruct by 13.723.59.8+13.7
- SciCodeQwen3 Next 80B A3B by 8.130.738.8+8.1
ReasoningQwen3 Next 80B A3B
- Humanity's Last ExamQwen3 Next 80B A3B by 6.26.412.6+6.2
- GPQA Diamondtie76.775.9tie
- CritPttie00tie
FactsKimi K2 Instruct
- AA-Omniscience · Non-hallucinationKimi K2 Instruct by 17.330.313+17.3
- Vectara HHEM hallucination ratelower is betterQwen3 Next 80B A3B by 8.617.99.3+8.6
- AA-Omniscience · AccuracyKimi K2 Instruct by 6.425.419+6.4
Long documentsQwen3 Next 80B A3B
- AA-LCRQwen3 Next 80B A3B by 10.75363.7+10.7
Following instructionsQwen3 Next 80B A3B
- IFBenchQwen3 Next 80B A3B by 1941.760.7+19
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIME 2025Qwen3 Next 80B A3B by 38.349.587.8+38.3
- AA Agentic IndexKimi K2 Instruct by 35.637.72.1+35.6
- HMMT 2025Qwen3 Next 80B A3B by 35.138.873.9+35.1
- τ²-Bench Telecom (AA run)Kimi K2 Instruct by 31.973.441.5+31.9
- AA-OmniscienceKimi K2 Instruct by 24.8-26.6-51.4+24.8
- HarmBenchKimi K2 Instruct by 12.897.484.6+12.8
- AIR-Bench 2024Qwen3 Next 80B A3B by 12.674.186.7+12.6
- vectara_avg_summary_lengthQwen3 Next 80B A3B by 11.759.270.9+11.7
- vectara_factual_consistencyQwen3 Next 80B A3B by 8.682.190.7+8.6
- Artificial Analysis Coding IndexKimi K2 Instruct by 8.525.917.4+8.5
- vectara_answer_rateKimi K2 Instruct by 4.298.694.4+4.2
- AA IntelligenceKimi K2 Instruct by 4.115.311.2+4.1
- Tau2 airlineQwen3 Next 80B A3B by 456.560.5+4
- SuperGPQAQwen3 Next 80B A3B by 3.657.260.8+3.6
- τ²-Bench (Retail)Kimi K2 Instruct by 2.870.667.8+2.8
- OJBenchQwen3 Next 80B A3B by 2.627.129.7+2.6
- MMLU-ProQwen3 Next 80B A3B by 1.681.182.7+1.6
- IFEvaltie89.888.9tie
- anthropic_red_teamtie99.399.9tie
- simple_safety_teststie10099.5tie
- MMLU-Reduxtie92.792.5tie
- XSTesttie98.298.3tie
Questions people ask
Which is better, Kimi K2 Instruct or Qwen3 Next 80B A3B?
Qwen3 Next 80B A3B wins four of the six areas where both have results: coding, reasoning, long documents and following instructions. Kimi K2 Instruct wins agents and facts.
Which is better for coding?
Qwen3 Next 80B A3B. It wins 2 of the 3 coding tests both models report; Kimi K2 Instruct wins 1.
Which is cheaper?
Kimi K2 Instruct costs $0.60 per million input tokens and $2.50 per million output tokens; Qwen3 Next 80B A3B costs $0.15 and $1.20. That makes Qwen3 Next 80B A3B about 56% cheaper for the same work.
How do you compare the two?
We use the 34 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.