Devstral Small
Devstral Small is behind the leaders in agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multimodal, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
29results on15benchmarks
- 29 aggregator
From 1 source · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
1 paper reference Devstral SmallDevstral Small benchmark results
29 results on 15 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
40.1% behind the leader1 of 7 ranked benchmarks measured
- 16.43Jun 15, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 7, 2026
- GPQA Diamond43.43Oct 7, 2026gpqa
- Humanity's Last Exam4.03Oct 7, 2026aa_hle
Show 3 more reasoning resultsHide 3 reasoning results
- 0.00Oct 7, 2026
- GPQA Diamond41.41Oct 7, 2026gpqa
- Humanity's Last Exam3.75Oct 7, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard6.06Oct 7, 2026aa_terminalbench_hard
- Terminal-Bench Hard6.06Oct 7, 2026aa_terminalbench_hard
- SciCode24.54Sep 4, 2026aa_scicode
Show 1 more coding resultHide 1 coding result
- SciCode24.31Sep 4, 2026aa_scicode
0 of 3 ranked benchmarks measured
- 32.00Oct 7, 2026
- 18.67Oct 7, 2026
0 of 3 ranked benchmarks measured
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy15.87Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination13.65Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy14.73Oct 7, 2026omniscienceAccuracy
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination23.10Oct 7, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence8.74Oct 7, 2026aa_intelligence_index
- -56.78Oct 7, 2026
- τ²-Bench Telecom (AA run)38.01Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)28.36Oct 7, 2026aa_tau2
- AA Intelligence7.60Oct 7, 2026aa_intelligence_index
- -50.83Oct 7, 2026
Show 4 more resultsHide 4 results
- 14.26Jun 18, 2026
- 25.13Jun 18, 2026
- Artificial Analysis Coding Index12.14Jun 18, 2026aa_coding_index
- 12.22Jun 18, 2026
Devstral Small: common questions
Who makes Devstral Small?
Devstral Small is made by Mistral.
When was Devstral Small released?
Devstral Small was released on May 21, 2025, according to Artificial Analysis.
What is Devstral Small good at?
Devstral Small is behind the leaders in agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
How many benchmarks has Devstral Small been tested on?
We track 29 results for Devstral Small on 15 benchmarks from 1 source. The latest was recorded on Oct 7, 2026.
Which API features does Devstral Small support?
OpenRouter lists tool calling and structured outputs for Devstral Small.
About this record
Where Devstral Small's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 29 scores · 0 independently verified · 29 aggregator-attributed. How these tiers are assigned
From 1 source on 1 site. Artificial Analysis supplies all of them. Bars are coloured by trust tier.
- artificialanalysis.ai29