GPT-4.1 mini vs GPT-4.5 Preview
Wins 0 of 4 areas
—
Wins 4 of 4 areas
Coding · Reasoning · Images and charts · Following instructions
GPT-4.5 Preview is the stronger all-rounder.
Scores updated · 20 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Images and chartsUnderstanding pictures, charts and video12GPT-4.5 Preview2 of 4 tests · 1 tie
- CodingWriting and fixing software01GPT-4.5 Preview1 of 1 test
- ReasoningHard problems that need careful thinking01GPT-4.5 Preview1 of 1 test
- Following instructionsDoing exactly what it is asked01GPT-4.5 Preview1 of 1 test
Coding, reasoning and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4.5 Preview pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+14.4points ahead
- Keeps track of context across a multi-turn chatMulti-Challenge+8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+3.1points ahead
Where GPT-4.1 mini pulls ahead
- Reasoning about charts from research papersCharXiv (RQ)+1.4points ahead
Every test, side by side
All 20 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsGPT-4.5 Preview
- MMMUGPT-4.5 Preview by 2.572.775.2+2.5
- CharXiv (RQ)GPT-4.1 mini by 1.456.855.4+1.4
- LMArena · VisionGPT-4.5 Preview by 14 rating points11811195+14 rating
- MathVistatie73.172.3tie
Following instructionsGPT-4.5 Preview
- Multi-ChallengeGPT-4.5 Preview by 835.843.8+8
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- TAU-bench (airline)GPT-4.5 Preview by 143650+14
- AIR-Bench 2024GPT-4.5 Preview by 13.860.474.2+13.8
- TAU-bench (retail)GPT-4.5 Preview by 12.655.868.4+12.6
- Aider-PolyglotGPT-4.5 Preview by 10.234.744.9+10.2
- HarmBenchGPT-4.5 Preview by 10.285.695.8+10.2
- ARC-AGI-1GPT-4.5 Preview by 6.83.510.3+6.8
- MMMLUGPT-4.5 Preview by 6.678.585.1+6.6
- IFEvalGPT-4.5 Preview by 4.184.188.2+4.1
- XSTestGPT-4.1 mini by 2.397.495.1+2.3
- AA Intelligencetie10.29.6tie
- anthropic_red_teamtie99.399.4tie
- bbqtie92.192tie
- simple_safety_teststie100100tie
Questions people ask
Which is better, GPT-4.1 mini or GPT-4.5 Preview?
GPT-4.5 Preview wins all four areas where both have results: coding, reasoning, images and charts and following instructions. GPT-4.1 mini wins none.
Which is better for coding?
GPT-4.5 Preview. It wins the one coding test both models report.
How do you compare the two?
We use the 20 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.