Mistral Small 4
Mistral Small 4 is behind the leaders in factuality, multimodal tasks, coding, instruction following, and reasoning. Too few results yet to rate agentic tasks, safety, long context, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Long Context, Math or Multilingual.
Price
$0.15input$0.60outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
43results on23benchmarks
- 1 independently verified
- 34 aggregator
- 8 vendor-reported
From 3 sources · latest Oct 8, 2026 · How verification works
Mistral Small 4 benchmark results
43 results on 23 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
31.3% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination33.50Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy21.68Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy16.57Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination21.99Oct 8, 2026omniscienceNonHallucination
37.6% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro56.82Oct 8, 2026aa_mmmu_pro
39.0% behind the leader3 of 10 ranked benchmarks measured
- SciCode38.77Oct 8, 2026aa_scicode
- Terminal-Bench Hard17.42Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.120.97Oct 8, 2026terminalbenchV21
Show 2 more coding resultsHide 2 coding results
- SciCode28.13Sep 4, 2026aa_scicode
- Terminal-Bench Hard10.61Oct 8, 2026aa_terminalbench_hard
40.0% behind the leader1 of 3 ranked benchmarks measured
- IFBench48.16Oct 8, 2026aa_ifbench
44.4% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond76.87Oct 8, 2026gpqa
- Humanity's Last Exam9.87Oct 8, 2026aa_hle
- 0.29Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond57.07Oct 8, 2026gpqa
- GPQA Diamond71.20Oct 7, 2026GPQA
- Humanity's Last Exam3.80Oct 8, 2026aa_hle
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking4.95Oct 8, 2026tauBanking
Show 1 more agentic resultHide 1 agentic result
- 17.22Jun 15, 2026
0 of 3 ranked benchmarks measured
- 49.67Oct 8, 2026
- 28.33Oct 8, 2026
- 71.20Oct 7, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence9.00Jun 21, 2026Artificial Analysis Intelligence Index
- AA Intelligence8.99Oct 8, 2026aa_intelligence_index
- -48.52Oct 8, 2026
- τ²-Bench Telecom (AA run)18.42Oct 8, 2026aa_tau2
- AA Intelligence11.27Oct 8, 2026aa_intelligence_index
- -30.40Oct 8, 2026
Show 9 more resultsHide 9 results
- 1.44Sep 9, 2026
- 18.55Jun 18, 2026
- 83.80Oct 7, 2026
- 58.30Sep 8, 2026
- Artificial Analysis Coding Index26.64Sep 9, 2026aa_coding_index
- 16.45Jun 18, 2026
- 63.60Aug 23, 2026
- 78.00Oct 7, 2026
- τ²-Bench Telecom (AA run)41.23Oct 8, 2026aa_tau2
Mistral Small 4: common questions
Who makes Mistral Small 4?
Mistral Small 4 is made by Mistral.
When was Mistral Small 4 released?
Mistral Small 4 was released on Mar 16, 2026, according to Mistral's own announcement.
What is Mistral Small 4 good at?
Mistral Small 4 is behind the leaders in factuality, multimodal tasks, coding, instruction following, and reasoning. Too few results yet to rate agentic tasks, safety, long context, math, or multilingual tasks.
How much does Mistral Small 4 cost?
Mistral Small 4 costs $0.15 per million input tokens and $0.60 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 77% of the 331 priced models we track.
How many benchmarks has Mistral Small 4 been tested on?
We track 43 results for Mistral Small 4 on 23 benchmarks from 3 sources, 1 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Mistral Small 4's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Jul 20, 2026
Where the results come from
Verification: 43 scores · 1 independently verified · 34 aggregator-attributed · 8 vendor-reported. How these tiers are assigned
From 3 sources on 2 sites. Artificial Analysis supplies 35 of them; the 1 independently verified result comes from 1 site. Bars are coloured by trust tier.
- artificialanalysis.ai35
- api.llm-stats.com8