Devstral Small 2
Devstral Small 2 is behind the leaders in multimodal tasks and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
24results on23benchmarks
- 2 independently verified
- 19 aggregator
- 3 vendor-reported
From 3 sources · latest Oct 7, 2026 · How verification works
Devstral Small 2 benchmark results
24 results on 23 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
41.2% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro44.62Oct 7, 2026aa_mmmu_pro
42.8% behind the leader5 of 10 ranked benchmarks measured
- 68.00Jun 15, 2026
- 55.70Jun 15, 2026
- SciCode32.41Oct 7, 2026aa_scicode
- Terminal-Bench 2.129.59Oct 7, 2026terminalbenchV21
- Terminal-Bench Hard16.67Oct 7, 2026aa_terminalbench_hard
Show 1 more coding resultHide 1 coding result
- 56.40May 1, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 7, 2026
- GPQA Diamond53.23Oct 7, 2026gpqa
- Humanity's Last Exam3.48Oct 7, 2026aa_hle
0 of 7 ranked benchmarks measured
- 1.58Oct 7, 2026
- Terminal-Bench 4.00.00Oct 7, 2026
- τ-Bench V3 · Banking10.72Oct 7, 2026tauBanking
0 of 3 ranked benchmarks measured
- 28.00Oct 7, 2026
0 of 3 ranked benchmarks measured
- IFBench31.16Oct 7, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy15.80Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination12.71Oct 7, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 53.80May 1, 2026
- τ²-Bench Telecom (AA run)23.39Oct 7, 2026aa_tau2
- AA Intelligence7.54Oct 7, 2026aa_intelligence_index
- -57.70Oct 7, 2026
- 4.78Sep 9, 2026
- Artificial Analysis Coding Index29.33Sep 9, 2026aa_coding_index
Show 1 more resultHide 1 result
- Terminal-Bench 2.022.50Jun 15, 2026Terminal Bench 2
Devstral Small 2: common questions
Who makes Devstral Small 2?
Devstral Small 2 is made by Mistral.
When was Devstral Small 2 released?
Devstral Small 2 was released on Dec 9, 2025, according to Artificial Analysis.
What is Devstral Small 2 good at?
Devstral Small 2 is behind the leaders in multimodal tasks and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How many benchmarks has Devstral Small 2 been tested on?
We track 24 results for Devstral Small 2 on 23 benchmarks from 3 sources, 2 of them independently verified. The latest was recorded on Oct 7, 2026.
About this record
Where Devstral Small 2's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- May 2, 2026
Where the results come from
Verification: 24 scores · 2 independently verified · 19 aggregator-attributed · 3 vendor-reported. How these tiers are assigned
From 3 sources on 3 sites. Artificial Analysis supplies 19 of them; the 2 independently verified results come from 1 site. Bars are coloured by trust tier.
- artificialanalysis.ai19
- huggingface.co3
- swebench.com2