Step3 VL 10B
Step3 VL 10B is capable in multimodal tasks and behind the leaders in instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
40results on36benchmarks
- 16 aggregator
- 24 vendor-reported
From 3 sources · latest Oct 8, 2026 · How verification works
Step3 VL 10B benchmark results
40 results on 36 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
21.9% behind the leader4 of 6 ranked benchmarks measured
- 86.75Jun 5, 2026
- 84.00Oct 7, 2026
- 78.10Oct 7, 2026
- MMMU-Pro63.99Oct 8, 2026aa_mmmu_pro
34.0% behind the leader2 of 3 ranked benchmarks measured
- 62.60Oct 7, 2026
- IFBench50.20Oct 8, 2026aa_ifbench
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond68.99Oct 8, 2026gpqa
- Humanity's Last Exam10.84Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard5.30Oct 8, 2026aa_terminalbench_hard
- SciCode31.13Sep 4, 2026aa_scicode
0 of 7 ranked benchmarks measured
- 0.00Jun 15, 2026
0 of 3 ranked benchmarks measured
- 0.00Oct 8, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy13.27Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination16.68Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -59.00Oct 8, 2026
- τ²-Bench Telecom (AA run)16.08Oct 8, 2026aa_tau2
- AA Intelligence7.69Oct 8, 2026aa_intelligence_index
- 5.36Jun 18, 2026
- 13.91Jun 18, 2026
- 87.70Oct 7, 2026
Show 15 more resultsHide 15 results
- 89.35Jun 5, 2026
- 87.66Jun 5, 2026
- 57.21Jun 5, 2026
- 78.18Jun 5, 2026
- 66.05Jun 5, 2026
- 75.77Jun 5, 2026
- 70.80Oct 7, 2026
- 70.81Jun 5, 2026
- 91.80Oct 7, 2026
- 92.05Jun 5, 2026
- 59.02Jun 5, 2026
- 59.45Jun 5, 2026
- 67.29Jun 5, 2026
- ScreenSpot-Pro (No tools)51.55Jun 5, 2026ScreenSpot-Pro
- 92.61Jun 5, 2026
Step3 VL 10B: common questions
Who makes Step3 VL 10B?
Step3 VL 10B is made by StepFun.
When was Step3 VL 10B released?
Step3 VL 10B was released on Jan 20, 2026, according to Artificial Analysis.
What is Step3 VL 10B good at?
Step3 VL 10B is capable in multimodal tasks and behind the leaders in instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or factuality.
How many benchmarks has Step3 VL 10B been tested on?
We track 40 results for Step3 VL 10B on 36 benchmarks from 3 sources. The latest was recorded on Oct 8, 2026.
About this record
Where Step3 VL 10B's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- May 16, 2026
Where the results come from
Verification: 40 scores · 0 independently verified · 16 aggregator-attributed · 24 vendor-reported. How these tiers are assigned
From 3 sources on 3 sites. Hugging Face supplies 18 of them. Bars are coloured by trust tier.
- huggingface.co18
- artificialanalysis.ai16
- api.llm-stats.com6