Gemini 3 Pro vs Kimi K3
Wins 1 of 7 areas
Facts
Wins 6 of 7 areas
Coding · Agents · Reasoning · Images and charts · Math · Long documents
Kimi K3 is the stronger all-rounder.Gemini 3 Pro is cheaper and better at facts.
Scores updated · 31 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own04Kimi K34 of 4 tests
- ReasoningHard problems that need careful thinking14Kimi K34 of 5 tests
- CodingWriting and fixing software02Kimi K32 of 2 tests
- MathCompetition and research-level math02Kimi K32 of 2 tests
- Images and chartsUnderstanding pictures, charts and video01Kimi K31 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Kimi K31 of 1 test
- FactsGetting facts right instead of making them up21Gemini 3 Pro2 of 3 tests
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3 Pro costs 22% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Kimi K3 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+38.3points ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+32points ahead
- Multi-step tasks using many real software toolsMCP Atlas+30.1points ahead
Where Gemini 3 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+22.3points ahead
- Common-sense trick questionsSimpleBench+15.7points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+8.2points ahead
Every test, side by side
All 31 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingKimi K3
- LMArena · WebDevKimi K3 by 215 rating points14401655+215 rating
- SciCodeKimi K3 by 3.456.159.5+3.4
AgentsKimi K3
- BrowseCompKimi K3 by 3259.291.2+32
- MCP AtlasKimi K3 by 30.154.184.2+30.1
- AA ApexAgentsKimi K3 by 22.918.441.3+22.9
- GDPValKimi K3 by 17.534.251.7+17.5
ReasoningKimi K3
- ARC-AGI-2Kimi K3 by 29.331.160.4+29.3
- SimpleBenchGemini 3 Pro by 15.776.460.7+15.7
- CritPtKimi K3 by 14.39.123.4+14.3
- Humanity's Last ExamKimi K3 by 7.239.746.9+7.2
- GPQA DiamondKimi K3 by 2.790.893.5+2.7
FactsGemini 3 Pro
- AA-Omniscience · Non-hallucinationKimi K3 by 38.38.546.8+38.3
- SimpleQA VerifiedGemini 3 Pro by 22.372.950.6+22.3
- AA-Omniscience · AccuracyGemini 3 Pro by 8.255.847.6+8.2
Images and chartsKimi K3
- CharXiv (RQ)Kimi K3 by 9.981.491.3+9.9
- MMMU-Protie80.280.5tie
Other results12 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ToolathlonKimi K3 by 36.836.473.2+36.8
- ZeroBenchKimi K3 by 33841+33
- DeepSearchQA (F1)Kimi K3 by 31.863.295+31.8
- Artificial Analysis Coding IndexKimi K3 by 29.746.576.2+29.7
- ARC-AGI-1Kimi K3 by 19.57594.5+19.5
- AA IntelligenceKimi K3 by 15.62843.6+15.6
- HLE (with tools)Kimi K3 by 1445.859.8+14
- MATH-VisionKimi K3 by 11.786.197.8+11.7
- matharena_visual_math_overallKimi K3 by 7.184.291.3+7.1
- AA-OmniscienceKimi K3 by 4.415.319.7+4.4
- WorldVQAKimi K3 by 3.647.451+3.6
- AA Agentic IndexGemini 3 Pro by 1.45250.6+1.4
Questions people ask
Which is better, Gemini 3 Pro or Kimi K3?
Kimi K3 wins six of the seven areas where both have results: coding, agents, reasoning, images and charts, math and long documents. Gemini 3 Pro wins facts, and costs 22% less.
Which is better for coding?
Kimi K3. It wins 2 of the 2 coding tests both models report; Gemini 3 Pro wins none.
Which is cheaper?
Gemini 3 Pro costs $2.00 per million input tokens and $12.00 per million output tokens; Kimi K3 costs $3.00 and $15.00. That makes Gemini 3 Pro about 22% cheaper for the same work.
How do you compare the two?
We use the 31 benchmark tests both models have published scores on. The verdict counts the 19 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 12 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.