o4-mini
o4-mini is behind the leaders in multimodal tasks, long context, coding, instruction following, agentic tasks, reasoning, and math. Too few results yet to rate safety, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Factuality.
Price
$1.10input$4.40outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
97results on71benchmarks
- 50 independently verified
- 16 aggregator
- 12 vendor-reported
- 19 cross-referenced
From 30 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
12 papers reference o4-minio4-mini benchmark results
97 results on 71 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
25.0% behind the leader4 of 6 ranked benchmarks measured
- 81.60Oct 7, 2026
- 84.30Oct 7, 2026
- MMMU-Pro69.25Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)72.00Oct 7, 2026CharXiv-R
Show 2 more multimodal resultsHide 2 multimodal results
- 1193.97Aug 25, 2026
- 1201May 6, 2026
31.3% behind the leader1 of 3 ranked benchmarks measured
- 61.00Oct 8, 2026
34.7% behind the leader3 of 10 ranked benchmarks measured
- 68.10Oct 7, 2026
- SciCode46.53Sep 4, 2026aa_scicode
- Terminal-Bench Hard15.15Oct 8, 2026aa_terminalbench_hard
Show 1 more coding resultHide 1 coding result
- 45.00Sep 1, 2026
34.9% behind the leader2 of 3 ranked benchmarks measured
- IFBench68.71Oct 8, 2026aa_ifbench
- 43.00Oct 7, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 44.90Oct 8, 2026
39.4% behind the leader2 of 7 ranked benchmarks measured
- 51.50Oct 7, 2026
- 25.41Jun 15, 2026
39.4% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond78.38Oct 8, 2026gpqa
- Humanity's Last Exam16.54Oct 8, 2026aa_hle
- 6.11May 10, 2026
- 0.57Oct 8, 2026
Show 5 more reasoning resultsHide 5 reasoning results
- 2.36May 10, 2026
- 1.67May 10, 2026
- GPQA Diamond81.40Oct 7, 2026GPQA
- 14.70Oct 7, 2026
- 38.70May 10, 2026
50.5% behind the leader2 of 5 ranked benchmarks measured
- 36.14Sep 21, 2026
- 4.88Sep 21, 2026
Show 2 more math resultsHide 2 math results
- 28.77Sep 21, 2026
- 16.14Sep 21, 2026
0 of 4 ranked benchmarks measured
- 19.60Sep 21, 2026
- 18.60Aug 29, 2026
- 18.80May 20, 2026
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy24.80Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination19.55Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 72.00Sep 11, 2026
- 14.29Sep 2, 2026
- 19.05Sep 2, 2026
- vectara_avg_summary_length127.70Aug 29, 2026Average Summary Length (Words)
- vectara_answer_rate99.20Aug 29, 2026Answer Rate
- vectara_factual_consistency81.40Aug 29, 2026Factual Consistency Rate
Show 57 more resultsHide 57 results
- 36.09Jun 18, 2026
- AA Intelligence17.00Aug 29, 2026Artificial Analysis Intelligence Index
- AA Intelligence26.00Aug 9, 2026Artificial Analysis Intelligence Index
- AA Intelligence16.65Oct 8, 2026aa_intelligence_index
- -35.70Oct 8, 2026
- 68.90Oct 7, 2026
- 61.67May 20, 2026
- 91.67May 20, 2026
- 84.17May 20, 2026
- 92.70Oct 7, 2026
- 78.49May 19, 2026
- 98.24May 10, 2026
- 58.67May 10, 2026
- 41.83May 10, 2026
- 21.33May 10, 2026
- 25.61Jun 18, 2026
- 94.00May 10, 2026
- 0.82Jun 13, 2026
- 0.95Jun 13, 2026
- 0.26Jun 13, 2026
- frontiermath_tier_4_v16.25Aug 29, 2026frontiermath_tier_4
- frontiermath_tier_4_v12.08May 20, 2026frontiermath_tier_4
- 97.00May 10, 2026
- 0.00May 10, 2026
- 77.44May 3, 2026
- 84.45May 3, 2026
- 98.14May 3, 2026
- 98.76May 3, 2026
- 48.86May 3, 2026
- 62.86May 3, 2026
- 86.16May 3, 2026
- 92.17May 3, 2026
- 0.86Jun 13, 2026
- 0.85Jun 13, 2026
- 0.87Jun 13, 2026
- 0.87Jun 13, 2026
- 0.87Jun 13, 2026
- 0.86Jun 13, 2026
- 0.87Jun 13, 2026
- 0.88Jun 13, 2026
- 0.87Jun 13, 2026
- 0.87Jun 13, 2026
- 0.88Jun 13, 2026
- 0.81Jun 13, 2026
- 0.71Jun 13, 2026
- 69.30May 26, 2025
- PersonQA hallucination ratelower is better0.48Jun 13, 2026
- 100.00May 10, 2026
- SimpleQA hallucination ratelower is better0.79Jun 13, 2026
- 96.0Jun 13, 2026
- 45.00May 1, 2026
- TAU-bench (airline)49.20Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)71.80Oct 7, 2026TAU-bench Retail
- 16.67May 10, 2026
- 130.90May 2, 2026
- 97.39May 10, 2026
- τ²-Bench Telecom (AA run)55.56Oct 8, 2026aa_tau2
o4-mini: common questions
Who makes o4-mini?
o4-mini is made by OpenAI.
When was o4-mini released?
o4-mini was released on Apr 16, 2025, according to Artificial Analysis.
What is o4-mini good at?
o4-mini is behind the leaders in multimodal tasks, long context, coding, instruction following, agentic tasks, reasoning, and math. Too few results yet to rate safety, multilingual tasks, or factuality.
How much does o4-mini cost?
o4-mini costs $1.10 per million input tokens and $4.40 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 71% of the 331 priced models we track.
How many benchmarks has o4-mini been tested on?
We track 97 results for o4-mini on 71 benchmarks from 30 sources, 50 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does o4-mini support?
OpenRouter lists tool calling, structured outputs, and reasoning for o4-mini.
About this record
Where o4-mini's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 2, 2026
Where the results come from
Verification: 97 scores · 50 independently verified · 16 aggregator-attributed · 19 vendor cross-reference · 12 vendor-reported. How these tiers are assigned
From 30 sources on 16 sites. cdn.openai.com supplies 19 of them; the 50 independently verified results come from 14 sites. Bars are coloured by trust tier.
- cdn.openai.com19
- artificialanalysis.ai18
- api.llm-stats.com12
- epoch.ai8
- livecodebench.github.io8
- matharena.ai7
- arcprize.org6
- storage.googleapis.com6
- raw.githubusercontent.com5
- swebench.com2
- aider.chat1
- arxiv.org1
- datasets-server.huggingface.co1
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name o4-mini at each reasoning setting.
| Setting | API id |
|---|---|
| low | o4-mini (low) o4-mini-low-2025-04-16 |
| medium | o4-mini (medium) o4-mini (medium) (april 2025) |
| high | o4-mini (high) o4-mini-2025-04-16-reasoning-high o4-mini-high o4-mini-high-2025-04-16 o4-mini (high) (april 2025) |