Mistral Large 3
Mistral Large 3 is behind the leaders in multimodal tasks and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
$0.50input$1.50outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
37results on34benchmarks
- 8 independently verified
- 19 aggregator
- 10 vendor-reported
From 6 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Mistral Large 3 benchmark results
37 results on 34 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
38.0% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro55.66Oct 8, 2026aa_mmmu_pro
Show 1 more multimodal resultHide 1 multimodal result
- 1221.82Aug 25, 2026
43.9% behind the leader4 of 10 ranked benchmarks measured
- SciCode36.57Oct 8, 2026aa_scicode
- 1230.02May 22, 2026
- Terminal-Bench Hard15.91Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.111.99Oct 8, 2026terminalbenchV21
Show 1 more coding resultHide 1 coding result
- 1223.06Jun 19, 2026
0 of 6 ranked benchmarks measured
- 20.40May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond67.98Oct 8, 2026gpqa
Show 2 more reasoning resultsHide 2 reasoning results
- GPQA Diamond43.90Oct 7, 2026GPQA
- Humanity's Last Exam4.17Oct 8, 2026aa_hle
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking5.77Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 36.00Oct 8, 2026
0 of 3 ranked benchmarks measured
- IFBench36.19Oct 8, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- 14.50May 2, 2026
- AA-Omniscience · Accuracy24.95Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination14.04Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- vectara_avg_summary_length112.70May 2, 2026Average Summary Length (Words)
- vectara_answer_rate98.80May 2, 2026Answer Rate
- vectara_factual_consistency85.50May 2, 2026Factual Consistency Rate
- AA Intelligence9.27Oct 8, 2026aa_intelligence_index
- -39.57Oct 8, 2026
- τ²-Bench Telecom (AA run)24.56Oct 8, 2026aa_tau2
Show 11 more resultsHide 11 results
- 2.44Sep 9, 2026
- 55.10Sep 8, 2026
- Artificial Analysis Coding Index20.07Sep 9, 2026aa_coding_index
- 34.40Aug 23, 2026
- 84.90Oct 7, 2026
- 82.00Oct 7, 2026
- 85.50Oct 7, 2026
- 74.20Oct 7, 2026
- 23.80Oct 7, 2026
- 74.90Oct 7, 2026
- 68.50Oct 7, 2026
Mistral Large 3: common questions
Who makes Mistral Large 3?
Mistral Large 3 is made by Mistral.
When was Mistral Large 3 released?
Mistral Large 3 was released on Dec 2, 2025, according to Artificial Analysis.
What is Mistral Large 3 good at?
Mistral Large 3 is behind the leaders in multimodal tasks and coding. Too few results yet to rate reasoning, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How much does Mistral Large 3 cost?
Mistral Large 3 costs $0.50 per million input tokens and $1.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 51% of the 331 priced models we track.
How many benchmarks has Mistral Large 3 been tested on?
We track 37 results for Mistral Large 3 on 34 benchmarks from 6 sources, 8 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Mistral Large 3 support?
OpenRouter lists tool calling and structured outputs for Mistral Large 3.
About this record
Where Mistral Large 3's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Oct 6, 2026
Where the results come from
Verification: 37 scores · 8 independently verified · 19 aggregator-attributed · 10 vendor-reported. How these tiers are assigned
From 6 sources on 5 sites. Artificial Analysis supplies 19 of them; the 8 independently verified results come from 3 sites. Bars are coloured by trust tier.
- artificialanalysis.ai19
- api.llm-stats.com10
- raw.githubusercontent.com4
- datasets-server.huggingface.co3
- simple-bench.com1