GPT-4.1 nano vs GPT-4.5 Preview
Wins 0 of 3 areas
—
Wins 3 of 3 areas
Reasoning · Images and charts · Following instructions
GPT-4.5 Preview is the stronger all-rounder.
Scores updated · 18 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4.5 Preview pulls ahead
- Keeps track of context across a multi-turn chatMulti-Challenge+28.8points ahead
- College exam questions with charts, maps and diagramsMMMU+19.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+18.3points ahead
Where GPT-4.1 nano pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 18 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsGPT-4.5 Preview
- MMMUGPT-4.5 Preview by 19.855.475.2+19.8
- MathVistaGPT-4.5 Preview by 16.156.272.3+16.1
- CharXiv (RQ)GPT-4.5 Preview by 14.940.555.4+14.9
Following instructionsGPT-4.5 Preview
- Multi-ChallengeGPT-4.5 Preview by 28.81543.8+28.8
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- TAU-bench (retail)GPT-4.5 Preview by 45.822.668.4+45.8
- TAU-bench (airline)GPT-4.5 Preview by 361450+36
- Aider-PolyglotGPT-4.5 Preview by 35.19.844.9+35.1
- MMMLUGPT-4.5 Preview by 18.266.985.1+18.2
- IFEvalGPT-4.5 Preview by 13.774.588.2+13.7
- AIR-Bench 2024GPT-4.5 Preview by 12.761.574.2+12.7
- ARC-AGI-1GPT-4.5 Preview by 10.3010.3+10.3
- HarmBenchGPT-4.5 Preview by 986.895.8+9
- bbqGPT-4.5 Preview by 4.587.592+4.5
- AA IntelligenceGPT-4.5 Preview by 1.87.89.6+1.8
- simple_safety_testsGPT-4.5 Preview by 199100+1
- XSTesttie9695.1tie
- anthropic_red_teamtie99.699.4tie
Questions people ask
Which is better, GPT-4.1 nano or GPT-4.5 Preview?
GPT-4.5 Preview wins all three areas where both have results: reasoning, images and charts and following instructions. GPT-4.1 nano wins none.
How do you compare the two?
We use the 18 benchmark tests both models have published scores on. The verdict counts the 5 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.