GPT-4.1 vs GPT-4.5 Preview
Wins 2 of 4 areas
Coding · Images and charts
Wins 2 of 4 areas
Reasoning · Following instructions
The two are evenly matched.GPT-4.1 is better at images and charts and coding; GPT-4.5 Preview at reasoning and following instructions.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4.1 pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+16.6points ahead
- Reasoning about charts from research papersCharXiv (RQ)+1.3points ahead
Where GPT-4.5 Preview pulls ahead
- Common-sense trick questionsSimpleBench+7.5points ahead
- Keeps track of context across a multi-turn chatMulti-Challenge+5.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+2.9points ahead
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningGPT-4.5 Preview
- SimpleBenchGPT-4.5 Preview by 7.52734.5+7.5
- GPQA DiamondGPT-4.5 Preview by 2.966.669.5+2.9
Images and chartsGPT-4.1
- LMArena · VisionGPT-4.1 by 15 rating points12101195+15 rating
- CharXiv (RQ)GPT-4.1 by 1.356.755.4+1.3
- MMMUtie74.875.2tie
- MathVistatie72.272.3tie
Following instructionsGPT-4.5 Preview
- Multi-ChallengeGPT-4.5 Preview by 5.538.343.8+5.5
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIR-Bench 2024GPT-4.5 Preview by 9.464.874.2+9.4
- Aider-PolyglotGPT-4.1 by 6.751.644.9+6.7
- ARC-AGI-1GPT-4.5 Preview by 4.85.510.3+4.8
- HarmBenchGPT-4.5 Preview by 4.191.795.8+4.1
- AA IntelligenceGPT-4.1 by 3.112.79.6+3.1
- XSTestGPT-4.1 by 2.897.995.1+2.8
- MMMLUGPT-4.1 by 2.287.385.1+2.2
- IFEvaltie87.488.2tie
- bbqtie92.692tie
- TAU-bench (airline)tie49.450tie
- TAU-bench (retail)tie6868.4tie
- anthropic_red_teamtie99.399.4tie
- simple_safety_teststie100100tie
Questions people ask
Which is better, GPT-4.1 or GPT-4.5 Preview?
GPT-4.1 and GPT-4.5 Preview each win two of the four areas where both have results. GPT-4.1 wins coding and images and charts; GPT-4.5 Preview wins reasoning and following instructions.
Which is better for coding?
GPT-4.1. It wins the one coding test both models report.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.