Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Sonnet 4 vs Gemini 2.5 Flash (Sep) (Non-Reasoning)

Anthropic · released

Wins 2 of 7 areas

Coding · Following instructions

Google · released

Wins 3 of 7 areas

Agents · Reasoning · Images and charts

Gemini 2.5 Flash (Sep) (Non-Reasoning) wins more areas, narrowly.Claude Sonnet 4 is better at coding.

Scores updated · 17 tests both models report · How we compare

What it costs

Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.

Claude Sonnet 4$3.00 to read · $15.00 to write$18Price from anthropic
Gemini 2.5 Flash (Sep) (Non-Reasoning)No current price is tracked.

The biggest differences

The three tests each model wins by the widest margin. Scores are out of 100.

Where Gemini 2.5 Flash (Sep) (Non-Reasoning) pulls ahead

  • Real work tasks from 44 professionsGDPVal+18.8points aheadGemini 2.5 Flash (Sep) (Non-Reasoning)28.5Claude Sonnet 49.7
  • Harder college exam questions with imagesMMMU-Pro+11.3points aheadGemini 2.5 Flash (Sep) (Non-Reasoning)73.1Claude Sonnet 461.8
  • Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+5.6points aheadGemini 2.5 Flash (Sep) (Non-Reasoning)28.3Claude Sonnet 422.7

Where Claude Sonnet 4 pulls ahead

  • Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+60.8points aheadClaude Sonnet 470.9Gemini 2.5 Flash (Sep) (Non-Reasoning)10.1
  • Hard command-line tasks in a real terminalTerminal-Bench Hard+14.4points aheadClaude Sonnet 431.1Gemini 2.5 Flash (Sep) (Non-Reasoning)16.7
  • Follows unfamiliar, precisely checkable instructionsIFBench+2.4points aheadClaude Sonnet 454.7Gemini 2.5 Flash (Sep) (Non-Reasoning)52.3

Every test, side by side

All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.

CodingClaude Sonnet 4

Full coding ranking

AgentsGemini 2.5 Flash (Sep) (Non-Reasoning)
  • GDPValGemini 2.5 Flash (Sep) (Non-Reasoning) by 18.89.728.5+18.8

Full agents ranking

ReasoningGemini 2.5 Flash (Sep) (Non-Reasoning)

Full reasoning ranking

FactsEven

Full facts ranking

Images and chartsGemini 2.5 Flash (Sep) (Non-Reasoning)
  • MMMU-ProGemini 2.5 Flash (Sep) (Non-Reasoning) by 11.361.873.1+11.3
  • LMArena · VisionGemini 2.5 Flash (Sep) (Non-Reasoning) by 63 rating points11911254+63 rating

Full images and charts ranking

Long documentsEven

Full long documents ranking

Following instructionsClaude Sonnet 4
  • IFBenchClaude Sonnet 4 by 2.454.752.3+2.4

Full following instructions ranking

Other results5 tests, not counted

Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.

Questions people ask

Which is better, Claude Sonnet 4 or Gemini 2.5 Flash (Sep) (Non-Reasoning)?

Gemini 2.5 Flash (Sep) (Non-Reasoning) wins three of the seven areas where both have results: agents, reasoning and images and charts. Claude Sonnet 4 wins coding and following instructions. They are level on facts and long documents.

Which is better for coding?

Claude Sonnet 4. It wins 1 of the 2 coding tests both models report; Gemini 2.5 Flash (Sep) (Non-Reasoning) wins none, and 1 is a tie.

How do you compare the two?

We use the 17 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 5 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.