Gemini 3.7 Flash
Gemini 3.7 Flash is at the frontier in long context; strong in factuality; capable in reasoning, multimodal tasks, agentic tasks, and coding; and behind the leaders in instruction following and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.75input$3.75outputper million tokens
From Artificial Analysis · 4 providers tracked · All prices
Evidence
89results on51benchmarks
- 20 independently verified
- 46 aggregator
- 23 vendor-reported
From 14 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Gemini 3.7 Flash benchmark results
89 results on 51 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
1.3% behind the leader2 of 3 ranked benchmarks measured
- MRCR v2 (8-needle, 128K)97.00Aug 14, 2026GDM-MRCR v2 (8-needle) (128k (average))
- 81.67Oct 8, 2026
6.4% behind the leader3 of 4 ranked benchmarks measured
- 69.20Sep 21, 2026
- AA-Omniscience · Accuracy55.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination35.47Oct 8, 2026omniscienceNonHallucination
Show 4 more factuality resultsHide 4 factuality results
- AA-Omniscience · Accuracy54.00Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy53.57Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination32.30Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination34.13Oct 8, 2026omniscienceNonHallucination
11.1% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond94.55Oct 8, 2026gpqa
- LiveBench · Reasoning87.80Oct 8, 2026livebench_reasoning@2026-06-25
- 84.58Sep 21, 2026
- Humanity's Last Exam47.87Oct 8, 2026aa_hle
- 14.29Oct 8, 2026
Show 8 more reasoning resultsHide 8 reasoning results
- 52.92Sep 21, 2026
- 63.75Sep 21, 2026
- 9.43Oct 8, 2026
- 5.71Oct 8, 2026
- GPQA Diamond92.12Oct 8, 2026gpqa
- GPQA Diamond90.10Oct 8, 2026gpqa
- Humanity's Last Exam35.13Oct 8, 2026aa_hle
- Humanity's Last Exam38.97Oct 8, 2026aa_hle
15.6% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro85.49Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)88.70Oct 7, 2026CharXiv-R
Show 4 more multimodal resultsHide 4 multimodal results
- CharXiv (reasoning)84.50Aug 14, 2026CharXiv Reasoning (No tools)
- 1314.98Sep 29, 2026
- MMMU-Pro84.86Oct 8, 2026aa_mmmu_pro
- MMMU-Pro84.74Oct 8, 2026aa_mmmu_pro
21.3% behind the leader4 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking32.78Oct 8, 2026tauBanking
- 44.62Oct 8, 2026
- Terminal-Bench 4.013.64Oct 8, 2026
Show 4 more agentic resultsHide 4 agentic results
- 41.36Oct 8, 2026
- 43.05Oct 8, 2026
- τ-Bench V3 · Banking29.48Oct 8, 2026tauBanking
- τ-Bench V3 · Banking35.46Oct 8, 2026tauBanking
23.2% behind the leader5 of 10 ranked benchmarks measured
- Terminal-Bench 2.185.77Oct 8, 2026terminalbenchV21
- LiveBench · Coding78.89Oct 8, 2026livebench_coding@2026-06-25
- SciCode57.18Oct 8, 2026aa_scicode
- 1591.95Sep 21, 2026
- LiveBench · Agentic Coding58.28Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 5 more coding resultsHide 5 coding results
- SciCode55.67Oct 8, 2026aa_scicode
- SciCode59.84Oct 8, 2026aa_scicode
- Terminal-Bench 2.178.28Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.179.78Oct 8, 2026terminalbenchV21
- 85.80Oct 7, 2026
27.4% behind the leader1 of 3 ranked benchmarks measured
- LiveBench · Instruction Following79.93Oct 8, 2026livebench_instruction_following@2026-06-25
36.3% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics93.47Oct 8, 2026livebench_math@2026-06-25
- 71.58Sep 21, 2026
- 36.59Sep 21, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1488Oct 8, 2026
- livebench_language85.46Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis67.96Oct 8, 2026livebench_data_analysis@2026-06-25
- AA Intelligence39.00Sep 21, 2026Artificial Analysis Intelligence Index
- 85.17Sep 21, 2026
- 91.17Sep 21, 2026
Show 32 more resultsHide 32 results
- 36.41Sep 9, 2026
- 41.57Sep 4, 2026
- 45.10Sep 4, 2026
- AA Intelligence39.62Oct 8, 2026aa_intelligence_index
- AA Intelligence36.95Oct 8, 2026aa_intelligence_index
- AA Intelligence39.06Oct 8, 2026aa_intelligence_index
- 22.13Oct 8, 2026
- 23.70Oct 8, 2026
- 26.48Oct 8, 2026
- 26.30Aug 14, 2026
- 26.30Oct 7, 2026
- 95.50Sep 21, 2026
- 56.00Oct 7, 2026
- Artificial Analysis Coding Index71.05Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index71.47Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index76.12Sep 9, 2026aa_coding_index
- 88.70Aug 14, 2026
- 1588.00Aug 24, 2026
- 1588.00Aug 14, 2026
- 65.30Oct 7, 2026
- 97.00Aug 24, 2026
- GDP (Surge AI)34.00Oct 7, 2026GDP.pdf
- GDP.PDF (All pass rate)34.00Sep 9, 2026
- 1525.00Aug 24, 2026
- GDPval-AA v2 Elo1482.00Aug 14, 2026GDPVal-AA v2 (Elo)
- 53.60Oct 7, 2026
- 85.40Oct 7, 2026
- 47.90Oct 7, 2026
- OSWorld 2.0 (partial)50.60Sep 9, 2026OSWorld-2.0 (Partial score)
- OSWorld 2.0 (partial)47.90Aug 24, 2026OSWorld-2.0
- 14.90Oct 7, 2026
- Terminal-Bench 4.011.20Sep 9, 2026
Gemini 3.7 Flash: common questions
Who makes Gemini 3.7 Flash?
Gemini 3.7 Flash is made by Google.
When was Gemini 3.7 Flash released?
Gemini 3.7 Flash was released on Aug 13, 2026, according to Artificial Analysis.
What is Gemini 3.7 Flash good at?
Gemini 3.7 Flash is at the frontier in long context; strong in factuality; capable in reasoning, multimodal tasks, agentic tasks, and coding; and behind the leaders in instruction following and math. Too few results yet to rate safety or multilingual tasks.
How much does Gemini 3.7 Flash cost?
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 65% of the 330 priced models we track.
How many benchmarks has Gemini 3.7 Flash been tested on?
We track 89 results for Gemini 3.7 Flash on 51 benchmarks from 14 sources, 20 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Gemini 3.7 Flash support?
OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 3.7 Flash.
About this record
Where Gemini 3.7 Flash's numbers come from, and every name it appears under.
- Tracked since
- Aug 13, 2026
- Newest source mention
- Aug 15, 2026
Where the results come from
Verification: 89 scores · 20 independently verified · 46 aggregator-attributed · 23 vendor-reported. How these tiers are assigned
From 14 sources on 9 sites. Artificial Analysis supplies 47 of them; the 20 independently verified results come from 6 sites. Bars are coloured by trust tier.
- artificialanalysis.ai47
- api.llm-stats.com10
- deepmind.google9
- livebench.ai7
- arcprize.org6
- storage.googleapis.com4
- epoch.ai3
- datasets-server.huggingface.co2
- lmarena.ai1
Also known as
How our sources name Gemini 3.7 Flash at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | gemini 3.7 flash (low) | gemini-3-7-flash-low |
| medium | gemini 3.7 flash (medium) | gemini-3-7-flash-medium |
| high | gemini 3.7 flash (high) | gemini-3.7-flash-high |