Gemini 3 Flash Preview
Gemini 3 Flash Preview is capable in long context, reasoning, instruction following, and multimodal tasks; and behind the leaders in coding, factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.50input$3.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
105results on66benchmarks
- 41 independently verified
- 34 aggregator
- 30 vendor-reported
From 24 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Gemini 3 Flash Preview benchmark results
105 results on 66 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
15.1% behind the leader2 of 3 ranked benchmarks measured
- 78.00Oct 7, 2026
- MRCR v2 (8-needle, 128K)67.20Jul 6, 2026MRCR v2 (8-needle) Long context performance
Show 1 more long context resultHide 1 long context result
- 55.33Oct 7, 2026
21.1% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond89.80Oct 7, 2026gpqa
- 61.10May 15, 2026
- Humanity's Last Exam36.56Oct 7, 2026aa_hle
- ARC-AGI-233.60Oct 6, 2026ARC-AGI v2
- 8.57Oct 7, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 3.33Sep 22, 2026
- 33.61May 10, 2026
- 12.78May 10, 2026
- 1.25May 10, 2026
- 1.43Oct 7, 2026
- GPQA Diamond81.21Oct 7, 2026gpqa
- GPQA Diamond90.40Oct 6, 2026GPQA
- Humanity's Last Exam14.97Oct 7, 2026aa_hle
- 43.50Oct 6, 2026
- Humanity's Last Exam33.70Jul 6, 2026Humanity’s Last Exam Academic reasoning (full set, text + MM)
21.7% behind the leader1 of 3 ranked benchmarks measured
- IFBench77.96Oct 7, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench55.10Oct 7, 2026aa_ifbench
23.9% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro79.94Oct 7, 2026aa_mmmu_pro
- CharXiv (reasoning)80.30Oct 6, 2026CharXiv-R
Show 6 more multimodal resultsHide 6 multimodal results
- 1265.52Sep 22, 2026
- 1285.36Aug 25, 2026
- 1256Jun 17, 2026
- 1260May 25, 2026
- MMMU-Pro78.55Oct 7, 2026aa_mmmu_pro
- 81.20Oct 6, 2026
26.6% behind the leader6 of 10 ranked benchmarks measured
- 78.00Oct 6, 2026
- SciCode50.58Sep 4, 2026aa_scicode
- Terminal-Bench 2.158.00Jul 6, 2026Terminal-bench 2.1 Agentic terminal coding
- 1438.86May 22, 2026
- Terminal-Bench Hard38.64Oct 7, 2026aa_terminalbench_hard
- 49.60Aug 4, 2026
Show 6 more coding resultsHide 6 coding results
- 1383.99Sep 22, 2026
- SciCode49.88Sep 4, 2026aa_scicode
- 72.70May 1, 2026
- 34.63Oct 7, 2026
- 75.80Sep 25, 2026
- Terminal-Bench Hard31.82Oct 7, 2026aa_terminalbench_hard
27.6% behind the leader4 of 4 ranked benchmarks measured
- 66.80May 20, 2026
- 13.50May 2, 2026
- AA-Omniscience · Accuracy53.43Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.02Oct 7, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy45.78Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.59Oct 7, 2026omniscienceNonHallucination
36.1% behind the leader4 of 7 ranked benchmarks measured
- 65.10Aug 4, 2026
- 57.40Oct 6, 2026
- τ-Bench V3 · Banking20.82Oct 7, 2026tauBanking
- 35.19Jun 15, 2026
Show 4 more agentic resultsHide 4 agentic results
- 27.73Oct 7, 2026
- 30.72Jun 15, 2026
- 62.00Oct 7, 2026
- MCP Atlas62.00Jul 6, 2026MCP Atlas Multi-step workflows using MCP
44.3% behind the leader2 of 5 ranked benchmarks measured
- 51.23Sep 9, 2026
- 17.07Sep 8, 2026
Show 4 more math resultsHide 4 math results
- 96.67Sep 2, 2026
- 96.67May 2, 2026
- 89.39Sep 2, 2026
- 89.39May 10, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1473Oct 5, 2026
- 21.50Sep 22, 2026
- 85.83Sep 2, 2026
- 97.50Jul 5, 2026
- 75.80May 30, 2026
- 82.80May 22, 2026
Show 39 more resultsHide 39 results
- 49.66Jun 18, 2026
- 35.01Jun 18, 2026
- AA Intelligence17.93Oct 7, 2026aa_intelligence_index
- AA Intelligence26.33Oct 7, 2026aa_intelligence_index
- 10.13Oct 7, 2026
- -4.32Oct 7, 2026
- 99.70Oct 6, 2026
- 84.67May 10, 2026
- 57.67May 10, 2026
- 29.00May 10, 2026
- 42.62Jun 18, 2026
- 37.84Jun 18, 2026
- Blueprint-Bench 20.00Jul 6, 2026Blueprint-Bench 2 Agentic spatial reasoning
- 42.55Oct 6, 2026
- Finance Agent v242.60Jul 6, 2026Finance Agent v2 Financial analysis and decision-making
- frontiermath_tier_4_v14.17May 20, 2026frontiermath_tier_4
- 66.04May 22, 2026
- 48.98May 22, 2026
- 68.44May 22, 2026
- GDPval-AA (Elo)1204.00Jul 6, 2026GDPval-AA Economically valuable knowledge work
- 97.50May 11, 2026
- 93.33May 10, 2026
- 33.70Jul 7, 2026
- 0.00Oct 6, 2026
- 72.40Oct 6, 2026
- 91.80Oct 6, 2026
- mrcr_v2_8needle_1m_pointwise26.30Jul 6, 2026MRCR v2 (8-needle) Long context performance 1M (pointwise)
- ScreenSpot-Pro (No tools)69.10Oct 6, 2026ScreenSpot Pro
- 68.70Oct 6, 2026
- 90.20Oct 6, 2026
- 47.60Oct 6, 2026
- 49.40Oct 6, 2026
- vectara_answer_rate99.80May 2, 2026Answer Rate
- vectara_avg_summary_length90.20May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency86.50May 2, 2026Factual Consistency Rate
- 363500.00Oct 6, 2026
- 86.90Oct 6, 2026
- τ²-Bench Telecom (AA run)43.27Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)80.41Oct 7, 2026aa_tau2
Gemini 3 Flash Preview: common questions
Who makes Gemini 3 Flash Preview?
Gemini 3 Flash Preview is made by Google.
When was Gemini 3 Flash Preview released?
Gemini 3 Flash Preview was released on Dec 17, 2025, according to Artificial Analysis.
What is Gemini 3 Flash Preview good at?
Gemini 3 Flash Preview is capable in long context, reasoning, instruction following, and multimodal tasks; and behind the leaders in coding, factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
How much does Gemini 3 Flash Preview cost?
Gemini 3 Flash Preview costs $0.50 per million input tokens and $3.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 61% of the 329 priced models we track.
How many benchmarks has Gemini 3 Flash Preview been tested on?
We track 105 results for Gemini 3 Flash Preview on 66 benchmarks from 24 sources, 41 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does Gemini 3 Flash Preview support?
OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 3 Flash Preview.
About this record
Where Gemini 3 Flash Preview's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 105 scores · 41 independently verified · 34 aggregator-attributed · 30 vendor-reported. How these tiers are assigned
From 24 sources on 14 sites. Artificial Analysis supplies 34 of them; the 41 independently verified results come from 10 sites. Bars are coloured by trust tier.
- artificialanalysis.ai34
- api.llm-stats.com19
- deepmind.google9
- arcprize.org8
- matharena.ai8
- datasets-server.huggingface.co4
- epoch.ai4
- huggingface.co4
- raw.githubusercontent.com4
- lmarena.ai3
- swebench.com3
- blog.google2
- labs.scale.com2
- simple-bench.com1
Also known as
How our sources name Gemini 3 Flash Preview at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| minimal | gemini 3 flash preview (minimal) | gemini-3-flash (thinking-minimal) |
| low | gemini 3 flash preview (low) | — |
| medium | gemini 3 flash preview (medium) | — |
| high | gemini 3 flash (high reasoning) gemini 3 flash preview (high) gemini 3 flash (high) | — |