Step 3.7 Flash
Step 3.7 Flash is capable in long context; and behind the leaders in multimodal tasks, coding, instruction following, reasoning, factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.20input$1.15outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
26results on23benchmarks
- 3 independently verified
- 20 aggregator
- 3 vendor-reported
From 4 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Step 3.7 Flash benchmark results
26 results on 23 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
20.8% behind the leader1 of 3 ranked benchmarks measured
- 73.67Oct 8, 2026
28.7% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro75.32Oct 8, 2026aa_mmmu_pro
30.3% behind the leader4 of 10 ranked benchmarks measured
- SciCode43.87Oct 8, 2026aa_scicode
- 56.30Oct 7, 2026
- Terminal-Bench Hard35.61Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.139.33Oct 8, 2026terminalbenchV21
Show 1 more coding resultHide 1 coding result
- 59.50Oct 7, 2026
30.7% behind the leader1 of 3 ranked benchmarks measured
- IFBench67.28Oct 8, 2026aa_ifbench
32.6% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond80.91Oct 8, 2026gpqa
- Humanity's Last Exam21.41Oct 8, 2026aa_hle
- 2.29Oct 8, 2026
34.9% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy25.78Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination15.02Oct 8, 2026omniscienceNonHallucination
40.5% behind the leader2 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking11.96Oct 8, 2026tauBanking
- 18.00Oct 8, 2026
Show 3 more agentic resultsHide 3 agentic results
- 14.82Oct 8, 2026
- 30.27Oct 8, 2026
- GDPval (win rate)45.80Oct 7, 2026GDPval
46.1% behind the leader2 of 5 ranked benchmarks measured
- 87.88Sep 2, 2026
- 95.00Sep 2, 2026
- 94.23Aug 17, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- τ²-Bench Telecom (AA run)98.54Oct 8, 2026aa_tau2
- AA Intelligence19.48Oct 8, 2026aa_intelligence_index
- -37.28Oct 8, 2026
- Artificial Analysis Coding Index39.57Sep 9, 2026aa_coding_index
- 21.74Sep 4, 2026
Step 3.7 Flash: common questions
Who makes Step 3.7 Flash?
Step 3.7 Flash is made by StepFun.
When was Step 3.7 Flash released?
Step 3.7 Flash was released on May 29, 2026, according to Artificial Analysis.
What is Step 3.7 Flash good at?
Step 3.7 Flash is capable in long context; and behind the leaders in multimodal tasks, coding, instruction following, reasoning, factuality, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
How much does Step 3.7 Flash cost?
Step 3.7 Flash costs $0.20 per million input tokens and $1.15 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 63% of the 331 priced models we track.
How many benchmarks has Step 3.7 Flash been tested on?
We track 26 results for Step 3.7 Flash on 23 benchmarks from 4 sources, 3 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Step 3.7 Flash support?
OpenRouter lists tool calling, structured outputs, and reasoning for Step 3.7 Flash.
About this record
Where Step 3.7 Flash's numbers come from, and every name it appears under.
- Tracked since
- May 29, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 26 scores · 3 independently verified · 20 aggregator-attributed · 3 vendor-reported. How these tiers are assigned
From 4 sources on 3 sites. Artificial Analysis supplies 20 of them; the 3 independently verified results come from 1 site. Bars are coloured by trust tier.
- artificialanalysis.ai20
- api.llm-stats.com3
- matharena.ai3