Gemini 2.5 Flash (Sep) (Non-Reasoning)
Gemini 2.5 Flash (Sep) (Non-Reasoning) is capable in long context; and behind the leaders in multimodal tasks, factuality, agentic tasks, instruction following, and reasoning. Too few results yet to rate coding, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Coding, Safety, Math or Multilingual.
Price
No current price is tracked for this model. See the rate card
Evidence
34results on17benchmarks
- 2 independently verified
- 32 aggregator
From 3 sources · latest Oct 8, 2026 · How verification works
Gemini 2.5 Flash (Sep) (Non-Reasoning) benchmark results
34 results on 17 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
24.4% behind the leader1 of 3 ranked benchmarks measured
- 71.00Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 60.00Oct 8, 2026
31.1% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro73.12Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- 1254.02Sep 21, 2026
- 1226Jun 17, 2026
- MMMU-Pro70.23Oct 8, 2026aa_mmmu_pro
37.5% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy28.33Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination10.09Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy26.83Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.79Oct 8, 2026omniscienceNonHallucination
37.6% behind the leader1 of 7 ranked benchmarks measured
- 28.51Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 17.55Jun 15, 2026
38.0% behind the leader1 of 3 ranked benchmarks measured
- IFBench52.31Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench43.54Oct 8, 2026aa_ifbench
40.7% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond79.29Oct 8, 2026gpqa
- Humanity's Last Exam13.81Oct 8, 2026aa_hle
- 0.29Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond76.57Oct 8, 2026gpqa
- Humanity's Last Exam8.71Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard14.39Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard16.67Oct 8, 2026aa_terminalbench_hard
- SciCode40.51Sep 4, 2026aa_scicode
Show 1 more coding resultHide 1 coding result
- SciCode37.50Sep 4, 2026aa_scicode
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -39.90Oct 8, 2026
- AA Intelligence12.38Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)28.36Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)45.61Oct 8, 2026aa_tau2
- AA Intelligence15.47Oct 8, 2026aa_intelligence_index
- -36.10Oct 8, 2026
Show 4 more resultsHide 4 results
- 34.67Jun 18, 2026
- 22.99Jun 18, 2026
- Artificial Analysis Coding Index24.61Jun 18, 2026aa_coding_index
- Artificial Analysis Coding Index22.10Jun 18, 2026aa_coding_index
Gemini 2.5 Flash (Sep) (Non-Reasoning): common questions
Who makes Gemini 2.5 Flash (Sep) (Non-Reasoning)?
Gemini 2.5 Flash (Sep) (Non-Reasoning) is made by Google.
When was Gemini 2.5 Flash (Sep) (Non-Reasoning) released?
Gemini 2.5 Flash (Sep) (Non-Reasoning) was released on Sep 25, 2025, according to Artificial Analysis.
What is Gemini 2.5 Flash (Sep) (Non-Reasoning) good at?
Gemini 2.5 Flash (Sep) (Non-Reasoning) is capable in long context; and behind the leaders in multimodal tasks, factuality, agentic tasks, instruction following, and reasoning. Too few results yet to rate coding, safety, math, or multilingual tasks.
How many benchmarks has Gemini 2.5 Flash (Sep) (Non-Reasoning) been tested on?
We track 34 results for Gemini 2.5 Flash (Sep) (Non-Reasoning) on 17 benchmarks from 3 sources, 2 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Gemini 2.5 Flash (Sep) (Non-Reasoning)'s numbers come from, and every name it appears under.
- Tracked since
- Sep 11, 2026
- Newest source mention
- Sep 12, 2026
Where the results come from
Verification: 34 scores · 2 independently verified · 32 aggregator-attributed. How these tiers are assigned
From 3 sources on 3 sites. Artificial Analysis supplies 32 of them; the 2 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai32
- datasets-server.huggingface.co1
- lmarena.ai1