Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Gemini 2.5 Flash

Score basis

Gemini 2.5 Flash is behind the leaders in factuality, multimodal tasks, long context, instruction following, agentic tasks, reasoning, and coding. Too few results yet to rate safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited7 capabilities
  1. Factuality−27.7%3 of 4
  2. Multimodal−28.0%2 of 6
  3. Long Context−28.6%1 of 3
  4. Instruction Following−38.9%1 of 3
  5. Agentic−41.8%1 of 7
  6. Reasoning−42.1%5 of 6
  7. Coding−42.8%3 of 10
Not rated3 capabilities

Too few results yet to rate Safety, Math or Multilingual.

Price

$0.30input$2.50outputper million tokens

From Artificial Analysis · 4 providers tracked · All prices

Costs more than 55% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

102results on64benchmarks

  • 38 independently verified
  • 31 aggregator
  • 10 vendor-reported
  • 23 cross-referenced

From 23 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Gemini 2.5 Flash benchmark results

102 results on 64 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

LimitedFull factuality ranking

27.7% behind the leader3 of 4 ranked benchmarks measured

Show 2 more factuality resultsHide 2 factuality results

Multimodal

LimitedFull multimodal ranking

28.0% behind the leader2 of 6 ranked benchmarks measured

Show 3 more multimodal resultsHide 3 multimodal results

Long Context

LimitedFull long context ranking

28.6% behind the leader1 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Instruction Following

LimitedFull instruction following ranking

38.9% behind the leader1 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

Agentic

LimitedFull agentic ranking

41.8% behind the leader1 of 7 ranked benchmarks measured

Show 1 more agentic resultHide 1 agentic result

Reasoning

LimitedFull reasoning ranking

42.1% behind the leader5 of 6 ranked benchmarks measured

Show 10 more reasoning resultsHide 10 reasoning results

Coding

LimitedFull coding ranking

42.8% behind the leader3 of 10 ranked benchmarks measured

Show 2 more coding resultsHide 2 coding results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 60 more resultsHide 60 results

Gemini 2.5 Flash: common questions

Who makes Gemini 2.5 Flash?

Gemini 2.5 Flash is made by Google.

When was Gemini 2.5 Flash released?

Gemini 2.5 Flash was released on May 20, 2025, according to Artificial Analysis.

What is Gemini 2.5 Flash good at?

Gemini 2.5 Flash is behind the leaders in factuality, multimodal tasks, long context, instruction following, agentic tasks, reasoning, and coding. Too few results yet to rate safety, math, or multilingual tasks.

How much does Gemini 2.5 Flash cost?

Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 55% of the 331 priced models we track.

How many benchmarks has Gemini 2.5 Flash been tested on?

We track 102 results for Gemini 2.5 Flash on 64 benchmarks from 23 sources, 38 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Gemini 2.5 Flash support?

OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 2.5 Flash.

About this record

Where Gemini 2.5 Flash's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Sep 10, 2026

Where the results come from

Verification: 102 scores · 38 independently verified · 31 aggregator-attributed · 23 vendor cross-reference · 10 vendor-reported. How these tiers are assigned

From 23 sources on 15 sites. Artificial Analysis supplies 32 of them; the 38 independently verified results come from 12 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai32
  • arxiv.org18
  • api.llm-stats.com10
  • arcprize.org9
  • livecodebench.github.io8
  • storage.googleapis.com6
  • huggingface.co5
  • raw.githubusercontent.com4
  • aider.chat2
  • matharena.ai2
  • swebench.com2
  • datasets-server.huggingface.co1
  • epoch.ai1
  • lmarena.ai1
  • simple-bench.com1

Also known as

gemini 2.5 flash (04-17 preview)gemini 2.5 flash (2025-04-17)Gemini 2.5 Flash (Apr)gemini 2.5 flash (apr) (non-reasoning)gemini 2.5 flash (april 2025)gemini 2.5 flash (latest)gemini 2.5 flash (non-reasoning)gemini 2.5 flash (preview)gemini 2.5 flash (preview) (thinking 16k)gemini 2.5 flash (preview) (thinking 1k)gemini 2.5 flash (preview) (thinking 24k)gemini 2.5 flash (preview) (thinking 8k)gemini 2.5 flash (reasoning)gemini 2.5 flash (thinking)gemini 2.5 flash preview (may 2025)gemini 2.5 flash preview (non-reasoning)and 12 more