Gemini 2.5 Flash
Gemini 2.5 Flash is behind the leaders in factuality, multimodal tasks, long context, instruction following, agentic tasks, reasoning, and coding. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$0.30input$2.50outputper million tokens
From Artificial Analysis · 4 providers tracked · All prices
Evidence
102results on64benchmarks
- 38 independently verified
- 31 aggregator
- 10 vendor-reported
- 23 cross-referenced
From 23 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Gemini 2.5 Flash benchmark results
102 results on 64 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
27.7% behind the leader3 of 4 ranked benchmarks measured
- 7.80May 2, 2026
- AA-Omniscience · Accuracy25.95Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination24.71Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy26.13Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination6.99Oct 8, 2026omniscienceNonHallucination
28.0% behind the leader2 of 6 ranked benchmarks measured
- 79.70Oct 7, 2026
- MMMU-Pro69.08Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- 1235.48Aug 25, 2026
- 1214Jun 3, 2026
- MMMU-Pro65.49Oct 8, 2026aa_mmmu_pro
28.6% behind the leader1 of 3 ranked benchmarks measured
- 65.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 49.90Oct 8, 2026
38.9% behind the leader1 of 3 ranked benchmarks measured
- IFBench50.27Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench38.98Oct 8, 2026aa_ifbench
41.8% behind the leader1 of 7 ranked benchmarks measured
- 9.88Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 11.93Jun 15, 2026
42.1% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond78.99Oct 8, 2026gpqa
- 41.20May 10, 2026
- Humanity's Last Exam12.14Oct 8, 2026aa_hle
- 1.14Oct 8, 2026
- 2.54Sep 22, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 1.98Sep 22, 2026
- 2.12Sep 22, 2026
- 2.16Sep 22, 2026
- 1.69May 10, 2026
- 1.43Oct 8, 2026
- GPQA Diamond68.28Oct 8, 2026gpqa
- GPQA Diamond82.80Oct 7, 2026GPQA
- 82.80Jun 4, 2026
- Humanity's Last Exam4.73Oct 8, 2026aa_hle
- 11.00Oct 7, 2026
42.8% behind the leader3 of 10 ranked benchmarks measured
- 60.40Oct 7, 2026
- SciCode29.05Sep 4, 2026aa_scicode
- Terminal-Bench Hard13.64Oct 8, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- 28.73Sep 1, 2026
- Terminal-Bench Hard12.12Oct 8, 2026aa_terminalbench_hard
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 32.33Sep 22, 2026
- 33.33Sep 22, 2026
- 25.83Sep 22, 2026
- 16.00Sep 22, 2026
- 44.00Aug 29, 2026
- 70.83Jul 5, 2026
Show 60 more resultsHide 60 results
- 15.01Jun 18, 2026
- 18.76Jun 18, 2026
- AA Intelligence14.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence9.85Oct 8, 2026aa_intelligence_index
- AA Intelligence13.11Oct 8, 2026aa_intelligence_index
- -42.57Oct 8, 2026
- -29.80Oct 8, 2026
- 47.10May 1, 2026
- 61.90Oct 7, 2026
- AIME 202482.30Jun 4, 2026AIME 24
- 72.00Oct 7, 2026
- AIME 202572.00Jun 4, 2026AIME 25
- 68.67May 19, 2026
- 98.76May 10, 2026
- 17.76Jun 18, 2026
- 22.21Jun 18, 2026
- 97.70May 10, 2026
- CHiME-4 (C-4) WERlower is better14.79Jun 18, 2026
- Earnings-21 10m (E21 10m) WERlower is better8.09Jun 18, 2026
- Earnings-22 10m (E22 10m) WERlower is better10.80Jun 18, 2026
- FLEURS Arabic WERlower is better25.25Jun 18, 2026
- FLEURS Dutch WERlower is better6.20Jun 18, 2026
- FLEURS English WERlower is better4.64Jun 18, 2026
- FLEURS French WERlower is better6.17Jun 18, 2026
- FLEURS German WERlower is better4.74Jun 18, 2026
- FLEURS Hindi WERlower is better6.76Jun 18, 2026
- FLEURS Italian WERlower is better2.21Jun 18, 2026
- FLEURS Portuguese WERlower is better4.23Jun 18, 2026
- FLEURS Spanish WERlower is better3.17Jun 18, 2026
- frontiermath_tier_4_v14.17May 20, 2026frontiermath_tier_4
- GigaSpeech (GS) WERlower is better10.99Jun 18, 2026
- 88.40Oct 7, 2026
- 62.63May 10, 2026
- 64.17May 29, 2026
- HMMT Feb. 202564.20Jun 4, 2026HMMT Feb 25
- LibriSpeech Test Clean (LS-C) WERlower is better2.97Jun 18, 2026
- LibriSpeech Test Other (LS-O) WERlower is better6.15Jun 18, 2026
- 76.21Jun 20, 2026
- 75.07May 3, 2026
- LiveCodeBench62.30Jun 4, 2026LiveCodeBench (2408-2505)
- LiveCodeBench (v5)63.90Oct 7, 2026LiveCodeBench v5
- 99.07Jun 20, 2026
- 98.45May 3, 2026
- 48.29Jun 20, 2026
- 46.86May 3, 2026
- 82.51Jun 20, 2026
- 81.20May 3, 2026
- 32.00Oct 7, 2026
- 97.99May 10, 2026
- 26.90Oct 7, 2026
- SPGISpeech (SPGI) WERlower is better4.00Jun 18, 2026
- 28.73May 1, 2026
- SwitchBoard (SB) WERlower is better9.57Jun 18, 2026
- vectara_answer_rate99.00May 2, 2026Answer Rate
- vectara_avg_summary_length101.50May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency92.20May 2, 2026Factual Consistency Rate
- VoxPopuli (VP) WERlower is better7.84Jun 18, 2026
- 98.83May 10, 2026
- τ²-Bench Telecom (AA run)14.91Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)31.58Oct 8, 2026aa_tau2
Gemini 2.5 Flash: common questions
Who makes Gemini 2.5 Flash?
Gemini 2.5 Flash is made by Google.
When was Gemini 2.5 Flash released?
Gemini 2.5 Flash was released on May 20, 2025, according to Artificial Analysis.
What is Gemini 2.5 Flash good at?
Gemini 2.5 Flash is behind the leaders in factuality, multimodal tasks, long context, instruction following, agentic tasks, reasoning, and coding. Too few results yet to rate safety, math, or multilingual tasks.
How much does Gemini 2.5 Flash cost?
Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 55% of the 331 priced models we track.
How many benchmarks has Gemini 2.5 Flash been tested on?
We track 102 results for Gemini 2.5 Flash on 64 benchmarks from 23 sources, 38 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Gemini 2.5 Flash support?
OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 2.5 Flash.
About this record
Where Gemini 2.5 Flash's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 102 scores · 38 independently verified · 31 aggregator-attributed · 23 vendor cross-reference · 10 vendor-reported. How these tiers are assigned
From 23 sources on 15 sites. Artificial Analysis supplies 32 of them; the 38 independently verified results come from 12 sites. Bars are coloured by trust tier.
- artificialanalysis.ai32
- arxiv.org18
- api.llm-stats.com10
- arcprize.org9
- livecodebench.github.io8
- storage.googleapis.com6
- huggingface.co5
- raw.githubusercontent.com4
- aider.chat2
- matharena.ai2
- swebench.com2
- datasets-server.huggingface.co1
- epoch.ai1
- lmarena.ai1
- simple-bench.com1