Gemini 3.1 Pro vs Kimi K2.5
Wins 8 of 8 areas
Coding · Agents · Reasoning · Facts · Images and charts · Math · Long documents · Following instructions
Wins 0 of 8 areas
—
Gemini 3.1 Pro is the stronger all-rounder.Kimi K2.5 is cheaper.
Both rank among the ten best models we track in images and charts.
Scores updated · 79 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software80Gemini 3.1 Pro8 of 8 tests
- AgentsCarrying out multi-step tasks on its own51Gemini 3.1 Pro5 of 6 tests
- ReasoningHard problems that need careful thinking40Gemini 3.1 Pro4 of 4 tests
- FactsGetting facts right instead of making them up40Gemini 3.1 Pro4 of 4 tests
- Images and chartsUnderstanding pictures, charts and video30Gemini 3.1 Pro3 of 5 tests · 2 ties
- MathCompetition and research-level math20Gemini 3.1 Pro2 of 3 tests · 1 tie
- Following instructionsDoing exactly what it is asked20Gemini 3.1 Pro2 of 2 tests
- Long documentsFinding answers in very long texts10Gemini 3.1 Pro1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Kimi K2.5 costs 74% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 3.1 Pro pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+65.3points ahead
- Short factual questions, answered correctlySimpleQA Verified+39.2points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+28.1points ahead
Where Kimi K2.5 pulls ahead
- Real work tasks from 44 professionsGDPVal+2.5points ahead
Every test, side by side
All 79 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 3.1 Pro
- Terminal-Bench 2.1Gemini 3.1 Pro by 28.173.845.7+28.1
- Terminal-Bench HardGemini 3.1 Pro by 1953.834.8+19
- SciCodeGemini 3.1 Pro by 9.758.749+9.7
- LiveCodeBench v6Gemini 3.1 Pro by 6.791.785+6.7
- SWE-bench MultilingualGemini 3.1 Pro by 3.976.973+3.9
- SWE-bench VerifiedGemini 3.1 Pro by 3.880.676.8+3.8
- SWE-bench ProGemini 3.1 Pro by 3.554.250.7+3.5
- LMArena · WebDevGemini 3.1 Pro by 11 rating points14471436+11 rating
AgentsGemini 3.1 Pro
- AA ApexAgentsGemini 3.1 Pro by 20.53211.5+20.5
- OSWorld-VerifiedGemini 3.1 Pro by 12.976.263.3+12.9
- BrowseCompGemini 3.1 Pro by 1185.974.9+11
- τ-Bench V3 · BankingGemini 3.1 Pro by 7.221.414.2+7.2
- MCP AtlasGemini 3.1 Pro by 5.469.263.8+5.4
- GDPValKimi K2.5 by 2.514.717.2+2.5
ReasoningGemini 3.1 Pro
- ARC-AGI-2Gemini 3.1 Pro by 65.377.111.8+65.3
- Humanity's Last ExamGemini 3.1 Pro by 16.34730.7+16.3
- CritPtGemini 3.1 Pro by 14.617.73.1+14.6
- GPQA DiamondGemini 3.1 Pro by 6.294.187.9+6.2
FactsGemini 3.1 Pro
- SimpleQA VerifiedGemini 3.1 Pro by 39.273.534.3+39.2
- AA-Omniscience · AccuracyGemini 3.1 Pro by 19.754.935.2+19.7
- AA-Omniscience · Non-hallucinationGemini 3.1 Pro by 14.849.134.3+14.8
- Vectara HHEM hallucination ratelower is betterGemini 3.1 Pro by 3.810.414.2+3.8
Images and chartsGemini 3.1 Pro
- MMMU-ProGemini 3.1 Pro by 782.475.4+7
- CharXiv (RQ)Gemini 3.1 Pro by 5.883.377.5+5.8
- LMArena · VisionGemini 3.1 Pro by 27 rating points12961269+27 rating
- Video-MMEtie86.787.4tie
- MathVistatie90.290.1tie
MathGemini 3.1 Pro
- HMMT Feb. 2026Gemini 3.1 Pro by 7.694.787.1+7.6
- AIME 2026Gemini 3.1 Pro by 2.598.395.8+2.5
- IMOAnswerBenchtie8181.8tie
Following instructionsGemini 3.1 Pro
- Multi-ChallengeGemini 3.1 Pro by 1071.461.4+10
- IFBenchGemini 3.1 Pro by 6.977.170.2+6.9
Other results46 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceGemini 3.1 Pro by 39.231.9-7.3+39.2
- ARC-AGI-1Gemini 3.1 Pro by 32.79865.3+32.7
- Vending-Bench 2Kimi K2.5 by 286.8 rating points911.21198+286.8 rating
- baby_vision_with_pythonGemini 3.1 Pro by 27.868.340.5+27.8
- MCPMarkGemini 3.1 Pro by 26.455.929.5+26.4
- Artificial Analysis Coding IndexGemini 3.1 Pro by 2268.846.8+22
- Tool-DecathlonGemini 3.1 Pro by 2148.827.8+21
- ToolathlonGemini 3.1 Pro by 2148.827.8+21
- BabyVisionGemini 3.1 Pro by 17.954.436.5+17.9
- Terminal-Bench 2.0Gemini 3.1 Pro by 17.768.550.8+17.7
- DeepSearchQA (accuracy)Kimi K2.5 by 16.960.277.1+16.9
- OJBench (python)Gemini 3.1 Pro by 1670.754.7+16
- frontiermath_tier_4_v1Gemini 3.1 Pro by 12.516.74.2+12.5
- AA Agentic IndexKimi K2.5 by 11.410.321.7+11.4
- charxiv_rq_with_pythonGemini 3.1 Pro by 11.289.978.7+11.2
- BrowseComp (w/ Ctx)Gemini 3.1 Pro by 1185.974.9+11
- browsecomp_with_context_managerGemini 3.1 Pro by 1185.974.9+11
- LiveBenchGemini 3.1 Pro by 10.879.969.1+10.8
- mathvision_with_pythonGemini 3.1 Pro by 10.795.785+10.7
- V* (w/ python)Gemini 3.1 Pro by 1096.986.9+10
- LVBenchKimi K2.5 by 9.766.275.9+9.7
- matharena_visual_math_overallGemini 3.1 Pro by 8.889.480.6+8.8
- Global-MMLU-LiteGemini 3.1 Pro by 8.792.784+8.7
- mmmu_pro_with_pythonGemini 3.1 Pro by 7.685.377.7+7.6
- BrowseComp (Agent Swarm)Gemini 3.1 Pro by 7.585.978.4+7.5
- vectara_answer_rateGemini 3.1 Pro by 7.299.492.2+7.2
- AA IntelligenceGemini 3.1 Pro by 6.229.723.5+6.2
- HMMT Nov. 2025Gemini 3.1 Pro by 5.694.889.2+5.6
- MATH-VisionGemini 3.1 Pro by 5.689.884.2+5.6
- DeepSearchQA (F1)Gemini 3.1 Pro by 4.881.977.1+4.8
- GDPval-AA v2Kimi K2.5 by 47 rating points9621009+47 rating
- MMLU-ProGemini 3.1 Pro by 3.99187.1+3.9
- vectara_factual_consistencyGemini 3.1 Pro by 3.889.685.8+3.8
- SWEBench Pro PublicGemini 3.1 Pro by 3.554.250.7+3.5
- LongVideoBenchKimi K2.5 by 3.376.579.8+3.3
- CyberGymKimi K2.5 by 2.538.841.3+2.5
- τ³-Bench BankingGemini 3.1 Pro by 2.316.514.2+2.3
- WorldVQAKimi K2.5 by 244.346.3+2
- StrongREJECTKimi K2.5 by 1.59899.5+1.5
- NL2RepoGemini 3.1 Pro by 1.433.432+1.4
- SimpleVQAKimi K2.5 by 1.369.971.2+1.3
- HLE (with tools)Gemini 3.1 Pro by 1.251.450.2+1.2
- τ³-BenchGemini 3.1 Pro by 1.167.166+1.1
- MotionBenchtie69.970.4tie
- vectara_avg_summary_lengthtie107.7112tie
- τ²-Bench Telecom (AA run)tie95.695.9tie
Questions people ask
Which is better, Gemini 3.1 Pro or Kimi K2.5?
Gemini 3.1 Pro wins all eight areas we test: coding, agents, reasoning, facts, images and charts, math, long documents and following instructions. Kimi K2.5 wins none, but costs 74% less.
Which is better for coding?
Gemini 3.1 Pro. It wins 8 of the 8 coding tests both models report; Kimi K2.5 wins none.
Which is cheaper?
Gemini 3.1 Pro costs $2.00 per million input tokens and $12.00 per million output tokens; Kimi K2.5 costs $0.60 and $3.00. That makes Kimi K2.5 about 74% cheaper for the same work.
How do you compare the two?
We use the 79 benchmark tests both models have published scores on. The verdict counts the 33 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 46 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.