Gemini 3.5 Flash
Gemini 3.5 Flash is strong in factuality; capable in instruction following, reasoning, multimodal tasks, and agentic tasks; and behind the leaders in coding, long context, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$1.50input$9.00outputper million tokens
From Artificial Analysis · 4 providers tracked · All prices
Evidence
132results on75benchmarks
- 29 independently verified
- 53 aggregator
- 29 vendor-reported
- 21 cross-referenced
From 28 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Gemini 3.5 Flash benchmark results
132 results on 75 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
7.8% behind the leader3 of 4 ranked benchmarks measured
- 66.20May 29, 2026
- AA-Omniscience · Accuracy51.40Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination37.83Oct 8, 2026omniscienceNonHallucination
Show 4 more factuality resultsHide 4 factuality results
- AA-Omniscience · Accuracy51.05Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy43.02Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination38.24Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination25.68Oct 8, 2026omniscienceNonHallucination
11.1% behind the leader2 of 3 ranked benchmarks measured
- LiveBench · Instruction Following75.60Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench76.33Oct 8, 2026aa_ifbench
13.6% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond92.22Oct 8, 2026gpqa
- LiveBench · Reasoning82.00Oct 8, 2026livebench_reasoning@2026-06-25
- 76.70May 21, 2026
- ARC-AGI-272.10Oct 7, 2026ARC-AGI v2
- Humanity's Last Exam42.68Oct 8, 2026aa_hle
- 13.14Oct 8, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 8.89Sep 22, 2026
- 72.08May 19, 2026
- 10.86Oct 8, 2026
- 1.43Oct 8, 2026
- GPQA Diamond92.12Oct 8, 2026gpqa
- GPQA Diamond82.83Oct 8, 2026gpqa
- Humanity's Last Exam41.33Oct 8, 2026aa_hle
- Humanity's Last Exam24.10Oct 8, 2026aa_hle
- 40.20Oct 7, 2026
- 40.20Jul 7, 2026
15.3% behind the leader3 of 6 ranked benchmarks measured
- 87.20Sep 28, 2026
- MMMU-Pro84.28Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)84.20Oct 7, 2026CharXiv-R
Show 5 more multimodal resultsHide 5 multimodal results
18.5% behind the leader6 of 7 ranked benchmarks measured
- 83.60Oct 7, 2026
- 78.40Oct 7, 2026
- τ-Bench V3 · Banking32.16Oct 8, 2026tauBanking
- 35.24Oct 8, 2026
- Terminal-Bench 4.06.57Oct 8, 2026
Show 7 more agentic resultsHide 7 agentic results
- 47.05Oct 8, 2026
- 40.35Oct 8, 2026
- 57.79Jun 15, 2026
- 47.00Jun 15, 2026
- 83.60Oct 8, 2026
- 83.60Jun 1, 2026
- 78.40Aug 12, 2026
25.1% behind the leader7 of 10 ranked benchmarks measured
- Terminal-Bench 2.178.65Oct 8, 2026terminalbenchV21
- LiveBench · Coding78.18Oct 8, 2026livebench_coding@2026-06-25
- SciCode53.94Oct 8, 2026aa_scicode
- 1498.97Sep 21, 2026
- LiveBench · Agentic Coding48.99Oct 8, 2026livebench_agentic_coding@2026-06-25
- Terminal-Bench Hard40.91Oct 8, 2026aa_terminalbench_hard
- 55.10Oct 7, 2026
Show 10 more coding resultsHide 10 coding results
- 1490.91Sep 21, 2026
- 1491.55May 22, 2026
- SciCode53.01Sep 4, 2026aa_scicode
- SciCode48.84Sep 4, 2026aa_scicode
- SWE-bench Pro53.90Aug 24, 2026SWE-Bench Pro (Public)
- 55.10Aug 12, 2026
- 76.20Aug 24, 2026
- 76.20Aug 12, 2026
- Terminal-Bench Hard39.39Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard46.21Oct 8, 2026aa_terminalbench_hard
25.7% behind the leader2 of 3 ranked benchmarks measured
- 73.33Oct 8, 2026
- MRCR v2 (8-needle, 128K)26.60Aug 24, 2026MRCR v2 (8-needle)
Show 3 more long context resultsHide 3 long context results
- 74.33Oct 8, 2026
- 61.33Oct 8, 2026
- MRCR v2 (8-needle, 128K)77.30Jul 6, 2026MRCR v2 (8-needle) Long context performance
46.0% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics88.24Oct 8, 2026livebench_math@2026-06-25
- 62.81Sep 9, 2026
- 26.83Sep 8, 2026
Show 3 more math resultsHide 3 math results
- 95.00Sep 2, 2026
- 95.45Sep 2, 2026
- 95.45Jun 24, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language84.58Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis64.86Oct 8, 2026livebench_data_analysis@2026-06-25
- 27.67Oct 8, 2026
- 1477Oct 6, 2026
- 48.83Sep 22, 2026
- 89.86Sep 2, 2026
Show 51 more resultsHide 51 results
- 27.25Sep 9, 2026
- 70.43Jun 18, 2026
- 50.96Jun 18, 2026
- AA Intelligence50.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence33.63Oct 8, 2026aa_intelligence_index
- AA Intelligence23.85Oct 8, 2026aa_intelligence_index
- AA Intelligence32.60Oct 8, 2026aa_intelligence_index
- 20.82Oct 8, 2026
- 21.18Oct 8, 2026
- 0.67Oct 8, 2026
- 92.50May 19, 2026
- Artificial Analysis Coding Index70.14Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index43.93Jun 18, 2026aa_coding_index
- Artificial Analysis Coding Index47.09Jun 18, 2026aa_coding_index
- 33.60Oct 7, 2026
- 33.60Jun 12, 2026
- 84.90Jul 22, 2026
- 37.00Aug 4, 2026
- 37.00Oct 7, 2026
- 57.90Oct 7, 2026
- Finance Agent57.90Jun 10, 2026Finance Agent v2
- 57.86Oct 7, 2026
- 57.90Aug 24, 2026
- 57.90Sep 18, 2026
- frontiermath_tier_4_v114.58May 29, 2026frontiermath_tier_4
- 26.60Jul 22, 2026
- GDPval-AA (Elo)1656.00Aug 24, 2026GDpval-AA
- 1349.00Aug 4, 2026
- 1357.00Aug 12, 2026
- 1348.00Jun 30, 2026
- GDPval-AA v2 Elo1349.00Jul 22, 2026GDPVal-AA v2 (Elo)
- 40.20Jul 7, 2026
- 0.80Oct 7, 2026
- 0.80Jul 7, 2026
- 75.02Oct 7, 2026
- 77.60Sep 28, 2026
- 76.30Sep 28, 2026
- 68.60Sep 28, 2026
- 49.70Jul 22, 2026
- 70.60Sep 28, 2026
- mrcr_v2_8needle_1m_pointwise22.10Jul 6, 2026MRCR v2 (8-needle) Long context performance 1M (pointwise)
- 56.50Sep 28, 2026
- 76.20Oct 7, 2026
- 71.90Sep 28, 2026
- 56.50Oct 7, 2026
- 76.40Sep 28, 2026
- 67.10Sep 28, 2026
- 76.00Sep 28, 2026
- τ²-Bench Telecom (AA run)95.61Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)95.32Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)58.77Oct 8, 2026aa_tau2
Gemini 3.5 Flash: common questions
Who makes Gemini 3.5 Flash?
Gemini 3.5 Flash is made by Google.
When was Gemini 3.5 Flash released?
Gemini 3.5 Flash was released on May 19, 2026, according to Artificial Analysis.
What is Gemini 3.5 Flash good at?
Gemini 3.5 Flash is strong in factuality; capable in instruction following, reasoning, multimodal tasks, and agentic tasks; and behind the leaders in coding, long context, and math. Too few results yet to rate safety or multilingual tasks.
How much does Gemini 3.5 Flash cost?
Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 78% of the 330 priced models we track.
How many benchmarks has Gemini 3.5 Flash been tested on?
We track 132 results for Gemini 3.5 Flash on 75 benchmarks from 28 sources, 29 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Gemini 3.5 Flash support?
OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 3.5 Flash.
About this record
Where Gemini 3.5 Flash's numbers come from, and every name it appears under.
- Tracked since
- May 19, 2026
- Newest source mention
- Aug 24, 2026
Where the results come from
Verification: 132 scores · 29 independently verified · 53 aggregator-attributed · 21 vendor cross-reference · 29 vendor-reported. How these tiers are assigned
From 28 sources on 17 sites. Artificial Analysis supplies 54 of them; the 29 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai54
- api.llm-stats.com15
- lf3-static.bytednsdoc.com10
- www-cdn.anthropic.com8
- deepmind.google7
- livebench.ai7
- datasets-server.huggingface.co5
- storage.googleapis.com5
- arcprize.org4
- epoch.ai4
- matharena.ai4
- anthropic.com2
- blog.google2
- labs.scale.com2
- cdn.sanity.io1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name Gemini 3.5 Flash at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| minimal | gemini 3.5 flash (minimal) | gemini-3-5-flash-minimal |
| medium | gemini 3.5 flash (medium) | gemini-3-5-flash-medium gemini-3.5-flash-medium |
| high | gemini 3.5 flash (high) | gemini-3.5-flash-high |