GPT-5.5 vs Kimi K2.7 Code
Wins 7 of 7 areas
Coding · Agents · Reasoning · Facts · Math · Long documents · Following instructions
Wins 0 of 7 areas
—
GPT-5.5 is the stronger all-rounder.Kimi K2.7 Code is cheaper.
Scores updated · 36 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software60GPT-5.56 of 6 tests
- ReasoningHard problems that need careful thinking50GPT-5.55 of 5 tests
- AgentsCarrying out multi-step tasks on its own30GPT-5.53 of 4 tests · 1 tie
- MathCompetition and research-level math30GPT-5.53 of 3 tests
- Following instructionsDoing exactly what it is asked20GPT-5.52 of 2 tests
- FactsGetting facts right instead of making them up21GPT-5.52 of 3 tests
- Long documentsFinding answers in very long texts10GPT-5.51 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Kimi K2.7 Code costs 86% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-5.5 pulls ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+60.3points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+31.3points ahead
- Short factual questions, answered correctlySimpleQA Verified+26.5points ahead
Where Kimi K2.7 Code pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+6.6points ahead
Every test, side by side
All 36 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGPT-5.5
- Terminal-Bench 2.1GPT-5.5 by 16.984.367.4+16.9
- Terminal-Bench HardGPT-5.5 by 15.960.644.7+15.9
- LiveBench · Agentic CodingGPT-5.5 by 8.35445.7+8.3
- LiveBench · CodingGPT-5.5 by 8.282.274+8.2
- SciCodeGPT-5.5 by 855.847.8+8
- LMArena · WebDevGPT-5.5 by 39 rating points15121473+39 rating
AgentsGPT-5.5
- τ-Bench V3 · BankingGPT-5.5 by 18.83920.2+18.8
- GDPValGPT-5.5 by 15.742.727+15.7
- Terminal-Bench 4.0GPT-5.5 by 13.614.61+13.6
- MCP Atlastie75.376tie
ReasoningGPT-5.5
- CritPtGPT-5.5 by 17.127.110+17.1
- SimpleBenchGPT-5.5 by 11.16957.9+11.1
- Humanity's Last ExamGPT-5.5 by 10.845.835+10.8
- LiveBench · ReasoningGPT-5.5 by 6.989.782.8+6.9
- GPQA DiamondGPT-5.5 by 3.993.589.6+3.9
FactsGPT-5.5
- SimpleQA VerifiedGPT-5.5 by 26.56336.5+26.5
- AA-Omniscience · AccuracyGPT-5.5 by 18.45839.6+18.4
- AA-Omniscience · Non-hallucinationKimi K2.7 Code by 6.61117.6+6.6
MathGPT-5.5
- FrontierMath Tier 4GPT-5.5 by 60.372.512.2+60.3
- FrontierMath Tiers 1-3 (v2)GPT-5.5 by 31.385.354+31.3
- LiveBench · MathematicsGPT-5.5 by 16.395.979.6+16.3
Following instructionsGPT-5.5
- LiveBench · Instruction FollowingGPT-5.5 by 14.470.756.3+14.4
- IFBenchGPT-5.5 by 12.875.963.1+12.8
Other results12 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- DeepSWE 1.1GPT-5.5 by 366731+36
- AA-OmniscienceGPT-5.5 by 30.720.5-10.2+30.7
- livebench_data_analysisGPT-5.5 by 18.981.662.7+18.9
- Program BenchGPT-5.5 by 17.270.853.6+17.2
- AA Agentic IndexGPT-5.5 by 14.837.322.5+14.8
- Artificial Analysis Coding IndexGPT-5.5 by 14.174.960.8+14.1
- AA IntelligenceGPT-5.5 by 12.638.425.8+12.6
- MCP Mark VerifiedGPT-5.5 by 11.892.981.1+11.8
- livebench_languageGPT-5.5 by 9.587.477.9+9.5
- LiveBenchGPT-5.5 by 8.880.771.9+8.8
- τ²-Bench Telecom (AA run)GPT-5.5 by 3.893.990.1+3.8
- MLS Bench Litelower is bettertie35.535.1tie
Questions people ask
Which is better, GPT-5.5 or Kimi K2.7 Code?
GPT-5.5 wins all seven areas where both have results: coding, agents, reasoning, facts, math, long documents and following instructions. Kimi K2.7 Code wins none, but costs 86% less.
Which is better for coding?
GPT-5.5. It wins 6 of the 6 coding tests both models report; Kimi K2.7 Code wins none.
Which is cheaper?
GPT-5.5 costs $5.00 per million input tokens and $30.00 per million output tokens; Kimi K2.7 Code costs $0.95 and $4.00. That makes Kimi K2.7 Code about 86% cheaper for the same work.
How do you compare the two?
We use the 36 benchmark tests both models have published scores on. The verdict counts the 24 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 12 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.