Claude Sonnet 4 vs Gemini 2.5 Flash (Sep) (Non-Reasoning)
Wins 2 of 7 areas
Coding · Following instructions
Wins 3 of 7 areas
Agents · Reasoning · Images and charts
Gemini 2.5 Flash (Sep) (Non-Reasoning) wins more areas, narrowly.Claude Sonnet 4 is better at coding.
Scores updated · 17 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking02Gemini 2.5 Flash (Sep) (Non-Reasoning)2 of 3 tests · 1 tie
- Images and chartsUnderstanding pictures, charts and video02Gemini 2.5 Flash (Sep) (Non-Reasoning)2 of 2 tests
- AgentsCarrying out multi-step tasks on its own01Gemini 2.5 Flash (Sep) (Non-Reasoning)1 of 1 test
- CodingWriting and fixing software10Claude Sonnet 41 of 2 tests · 1 tie
- Following instructionsDoing exactly what it is asked10Claude Sonnet 41 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
- Long documentsFinding answers in very long texts00Even0 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 2.5 Flash (Sep) (Non-Reasoning) pulls ahead
- Real work tasks from 44 professionsGDPVal+18.8points ahead
- Harder college exam questions with imagesMMMU-Pro+11.3points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+5.6points ahead
Where Claude Sonnet 4 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+60.8points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+14.4points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+2.4points ahead
Every test, side by side
All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Sonnet 4
- Terminal-Bench HardClaude Sonnet 4 by 14.431.116.7+14.4
- SciCodetie4040.5tie
AgentsGemini 2.5 Flash (Sep) (Non-Reasoning)
- GDPValGemini 2.5 Flash (Sep) (Non-Reasoning) by 18.89.728.5+18.8
ReasoningGemini 2.5 Flash (Sep) (Non-Reasoning)
- Humanity's Last ExamGemini 2.5 Flash (Sep) (Non-Reasoning) by 3.110.713.8+3.1
- GPQA DiamondGemini 2.5 Flash (Sep) (Non-Reasoning) by 1.677.779.3+1.6
- CritPttie0.30.3tie
FactsEven
- AA-Omniscience · Non-hallucinationClaude Sonnet 4 by 60.870.910.1+60.8
- AA-Omniscience · AccuracyGemini 2.5 Flash (Sep) (Non-Reasoning) by 5.622.728.3+5.6
Images and chartsGemini 2.5 Flash (Sep) (Non-Reasoning)
- MMMU-ProGemini 2.5 Flash (Sep) (Non-Reasoning) by 11.361.873.1+11.3
- LMArena · VisionGemini 2.5 Flash (Sep) (Non-Reasoning) by 63 rating points11911254+63 rating
Following instructionsClaude Sonnet 4
- IFBenchClaude Sonnet 4 by 2.454.752.3+2.4
Other results5 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceClaude Sonnet 4 by 36.30.2-36.1+36.3
- τ²-Bench Telecom (AA run)Claude Sonnet 4 by 1964.645.6+19
- AA Agentic IndexGemini 2.5 Flash (Sep) (Non-Reasoning) by 17.117.634.7+17.1
- Artificial Analysis Coding IndexClaude Sonnet 4 by 1337.624.6+13
- AA IntelligenceClaude Sonnet 4 by 3.418.915.5+3.4
Questions people ask
Which is better, Claude Sonnet 4 or Gemini 2.5 Flash (Sep) (Non-Reasoning)?
Gemini 2.5 Flash (Sep) (Non-Reasoning) wins three of the seven areas where both have results: agents, reasoning and images and charts. Claude Sonnet 4 wins coding and following instructions. They are level on facts and long documents.
Which is better for coding?
Claude Sonnet 4. It wins 1 of the 2 coding tests both models report; Gemini 2.5 Flash (Sep) (Non-Reasoning) wins none, and 1 is a tie.
How do you compare the two?
We use the 17 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 5 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.