Gemini 3.1 Pro vs Kimi K3
Wins 2 of 8 areas
Facts · Following instructions
Wins 4 of 8 areas
Coding · Agents · Math · Long documents
Kimi K3 wins more areas, narrowly.Gemini 3.1 Pro is cheaper and better at facts.
Scores updated · 55 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own08Kimi K38 of 8 tests
- CodingWriting and fixing software04Kimi K34 of 5 tests · 1 tie
- MathCompetition and research-level math23Kimi K33 of 5 tests
- Long documentsFinding answers in very long texts01Kimi K31 of 1 test
- FactsGetting facts right instead of making them up30Gemini 3.1 Pro3 of 3 tests
- Following instructionsDoing exactly what it is asked10Gemini 3.1 Pro1 of 1 test
- ReasoningHard problems that need careful thinking22Even2 each · 2 ties
- Images and chartsUnderstanding pictures, charts and video11Even1 each
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3.1 Pro costs 22% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Kimi K3 pulls ahead
- Real work tasks from 44 professionsGDPVal+37.1points ahead
- Customer-service tasks in a simulated bankτ-Bench V3 · Banking+24.6points ahead
- Fixes real GitHub issues, working as an agentLiveBench · Agentic Coding+18.1points ahead
Where Gemini 3.1 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+22.9points ahead
- Common-sense trick questionsSimpleBench+18.9points ahead
- Abstract visual puzzles that people can solveARC-AGI-2+16.7points ahead
Every test, side by side
All 55 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingKimi K3
- LMArena · WebDevKimi K3 by 212 rating points14461658+212 rating
- LiveBench · Agentic CodingKimi K3 by 18.144.162.2+18.1
- Terminal-Bench 2.1Kimi K3 by 11.273.885+11.2
- LiveBench · CodingKimi K3 by 576.581.5+5
- SciCodetie58.759.5tie
AgentsKimi K3
- GDPValKimi K3 by 37.114.751.8+37.1
- τ-Bench V3 · BankingKimi K3 by 24.621.446+24.6
- AA IT-Bench SREKimi K3 by 17.430.347.7+17.4
- MCP AtlasKimi K3 by 1569.284.2+15
- AA ApexAgentsKimi K3 by 9.33241.3+9.3
- OSWorld-VerifiedKimi K3 by 8.676.284.8+8.6
- Terminal-Bench 4.0Kimi K3 by 8.6412.6+8.6
- BrowseCompKimi K3 by 5.385.991.2+5.3
ReasoningEven
- SimpleBenchGemini 3.1 Pro by 18.979.660.7+18.9
- ARC-AGI-2Gemini 3.1 Pro by 16.777.160.4+16.7
- LiveBench · ReasoningKimi K3 by 6.78490.7+6.7
- CritPtKimi K3 by 5.717.723.4+5.7
- GPQA Diamondtie94.193.5tie
- Humanity's Last Examtie4746.9tie
FactsGemini 3.1 Pro
- SimpleQA VerifiedGemini 3.1 Pro by 22.973.550.6+22.9
- AA-Omniscience · AccuracyGemini 3.1 Pro by 7.354.947.6+7.3
- AA-Omniscience · Non-hallucinationGemini 3.1 Pro by 2.349.146.8+2.3
Images and chartsEven
- CharXiv (RQ)Kimi K3 by 883.391.3+8
- MMMU-ProGemini 3.1 Pro by 1.982.480.5+1.9
MathKimi K3
- FrontierMath Tiers 1-3 (v2)Kimi K3 by 12.559.672.2+12.5
- FrontierMath Tier 4Kimi K3 by 12.226.839+12.2
- LiveBench · MathematicsGemini 3.1 Pro by 6.69184.4+6.6
- HMMT Feb. 2026Kimi K3 by 2.394.797+2.3
- AIME 2026Gemini 3.1 Pro by 1.698.396.7+1.6
Following instructionsGemini 3.1 Pro
- LiveBench · Instruction FollowingGemini 3.1 Pro by 7.779.171.4+7.7
Other results24 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDPval-AA v2 EloKimi K3 by 721 rating points9651686+721 rating
- GDPval-AA v2Kimi K3 by 720 rating points9621682+720 rating
- DeepSWEKimi K3 by 57.51067.5+57.5
- DeepSWE 1.1Kimi K3 by 571269+57
- CyberGymKimi K3 by 41.238.880+41.2
- AA Agentic IndexKimi K3 by 40.310.350.6+40.3
- Program BenchKimi K3 by 38.339.577.8+38.3
- SWE-MarathonKimi K3 by 38442+38
- BabyVisionKimi K3 by 31.354.485.7+31.3
- NL2RepoKimi K3 by 24.633.458+24.6
- ToolathlonKimi K3 by 24.448.873.2+24.4
- PostTrainBenchKimi K3 by 1521.636.6+15
- AA IntelligenceKimi K3 by 13.929.743.6+13.9
- DeepSearchQA (F1)Kimi K3 by 13.181.995+13.1
- AA-OmniscienceGemini 3.1 Pro by 12.231.919.7+12.2
- OfficeQA ProGemini 3.1 Pro by 9.272.563.3+9.2
- HLE (with tools)Kimi K3 by 8.451.459.8+8.4
- MATH-VisionKimi K3 by 889.897.8+8
- Artificial Analysis Coding IndexKimi K3 by 7.468.876.2+7.4
- WorldVQAKimi K3 by 6.744.351+6.7
- ARC-AGI-1Gemini 3.1 Pro by 3.59894.5+3.5
- matharena_visual_math_overallKimi K3 by 1.989.491.3+1.9
- livebench_data_analysistie78.578.7tie
- livebench_languagetie85.485.5tie
Questions people ask
Which is better, Gemini 3.1 Pro or Kimi K3?
Kimi K3 wins four of the eight areas we test: coding, agents, math and long documents. Gemini 3.1 Pro wins facts and following instructions, and costs 22% less. They are level on reasoning and images and charts.
Which is better for coding?
Kimi K3. It wins 4 of the 5 coding tests both models report; Gemini 3.1 Pro wins none, and 1 is a tie.
Which is cheaper?
Gemini 3.1 Pro costs $2.00 per million input tokens and $12.00 per million output tokens; Kimi K3 costs $3.00 and $15.00. That makes Gemini 3.1 Pro about 22% cheaper for the same work.
How do you compare the two?
We use the 55 benchmark tests both models have published scores on. The verdict counts the 31 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 24 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.