o3-mini
o3-mini is behind the leaders in instruction following and math. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Multimodal, Multilingual or Factuality.
Price
$1.10input$4.40outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
104results on67benchmarks
- 48 independently verified
- 25 aggregator
- 12 vendor-reported
- 19 cross-referenced
From 25 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
6 papers reference o3-minio3-mini benchmark results
104 results on 67 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
37.4% behind the leader2 of 3 ranked benchmarks measured
- IFBench67.14Oct 8, 2026aa_ifbench
- 39.90Oct 7, 2026
Show 3 more instruction following resultsHide 3 instruction following results
- LiveBench · Instruction Following82.50Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following78.58Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following79.85Aug 23, 2026livebench_instruction_following@2025-04-07
57.4% behind the leader2 of 5 ranked benchmarks measured
- 18.60Sep 21, 2026
- 0.00Sep 21, 2026
Show 2 more math resultsHide 2 math results
- 10.53Sep 21, 2026
- 3.86Sep 9, 2026
0 of 6 ranked benchmarks measured
- 2.08Sep 21, 2026
- 22.80May 10, 2026
- 2.99May 10, 2026
Show 8 more reasoning resultsHide 8 reasoning results
- 0.00May 10, 2026
- 0.29Oct 8, 2026
- GPQA Diamond77.27Oct 8, 2026gpqa
- GPQA Diamond74.85Oct 8, 2026gpqa
- GPQA Diamond77.20Oct 7, 2026GPQA
- 76.80Jun 4, 2026
- Humanity's Last Exam7.91Oct 8, 2026aa_hle
- Humanity's Last Exam12.04Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- LiveBench · Coding82.81Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding64.84Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding69.53Aug 23, 2026livebench_coding@2025-04-07
Show 6 more coding resultsHide 6 coding results
- SciCode42.82Oct 8, 2026aa_scicode
- SciCode39.93Sep 4, 2026aa_scicode
- 49.30Oct 7, 2026
- Terminal-Bench 2.14.49Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard6.06Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard6.82Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking5.15Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 43.00Oct 8, 2026
0 of 4 ranked benchmarks measured
- 15.30Sep 21, 2026
- AA-Omniscience · Accuracy21.27Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination18.92Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 22.33Sep 21, 2026
- 53.80Sep 11, 2026
- 2.08Sep 2, 2026
- AA Intelligence11.00Aug 29, 2026Artificial Analysis Intelligence Index
- frontiermath_tier_4_v14.17Aug 29, 2026frontiermath_tier_4
- livebench_language49.46Aug 23, 2026livebench_language@2025-04-07
Show 62 more resultsHide 62 results
- 0.86Sep 9, 2026
- AA Intelligence12.00Aug 7, 2026Artificial Analysis Intelligence Index
- AA Intelligence10.96Oct 8, 2026aa_intelligence_index
- AA Intelligence12.47Oct 8, 2026aa_intelligence_index
- -42.57Oct 8, 2026
- 66.70Oct 7, 2026
- AIME 202479.60Jun 4, 2026AIME 24
- 48.33May 20, 2026
- 76.67May 19, 2026
- AIME 202576.70Jun 4, 2026AIME 25
- 74.89May 19, 2026
- 98.64May 10, 2026
- 34.50May 10, 2026
- 14.50May 10, 2026
- Artificial Analysis Coding Index16.26Sep 9, 2026aa_coding_index
- 17.86Jun 18, 2026
- 94.10May 10, 2026
- FrontierMath (overall)9.20Oct 7, 2026FrontierMath
- 95.19May 10, 2026
- 28.33May 13, 2026
- HMMT Feb. 202553.30Jun 4, 2026HMMT Feb 25
- 93.90Aug 31, 2026
- 84.60Oct 7, 2026
- livebench_language39.45Aug 23, 2026livebench_language@2025-04-07
- livebench_language44.98Aug 23, 2026livebench_language@2025-04-07
- 75.36Jun 18, 2026
- 77.73May 3, 2026
- 75.36May 3, 2026
- LiveCodeBench65.90Jun 4, 2026LiveCodeBench (2408-2505)
- 98.76Jun 18, 2026
- 99.07May 3, 2026
- 98.76May 3, 2026
- 43.14Jun 18, 2026
- 47.43May 3, 2026
- 43.14May 3, 2026
- 85.12Jun 18, 2026
- 87.47May 3, 2026
- 85.12May 3, 2026
- 92.00Oct 7, 2026
- 0.82Jun 13, 2026
- 0.81Jun 13, 2026
- 0.80Jun 13, 2026
- 0.84Jun 13, 2026
- 0.84Jun 13, 2026
- 0.81Jun 13, 2026
- 0.81Jun 13, 2026
- 0.83Jun 13, 2026
- 0.84Jun 13, 2026
- 0.83Jun 13, 2026
- 0.83Jun 13, 2026
- 0.84Jun 13, 2026
- 0.74Jun 13, 2026
- 0.64Jun 13, 2026
- 80.70Oct 7, 2026
- 99.00May 10, 2026
- 15.00Oct 7, 2026
- TAU-bench (airline)32.40Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)57.60Oct 7, 2026TAU-bench Retail
- 0.00May 10, 2026
- 94.06May 10, 2026
- τ²-Bench Telecom (AA run)31.29Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)28.65Oct 8, 2026aa_tau2
o3-mini: common questions
Who makes o3-mini?
o3-mini is made by OpenAI.
When was o3-mini released?
o3-mini was released on Jan 31, 2025, according to Artificial Analysis.
What is o3-mini good at?
o3-mini is behind the leaders in instruction following and math. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, multimodal tasks, multilingual tasks, or factuality.
How much does o3-mini cost?
o3-mini costs $1.10 per million input tokens and $4.40 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 71% of the 330 priced models we track.
How many benchmarks has o3-mini been tested on?
We track 104 results for o3-mini on 67 benchmarks from 25 sources, 48 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does o3-mini support?
OpenRouter lists tool calling, structured outputs, and reasoning for o3-mini.
About this record
Where o3-mini's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 28, 2026
Where the results come from
Verification: 104 scores · 48 independently verified · 25 aggregator-attributed · 19 vendor cross-reference · 12 vendor-reported. How these tiers are assigned
From 25 sources on 11 sites. Artificial Analysis supplies 27 of them; the 48 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai27
- cdn.openai.com14
- huggingface.co14
- api.llm-stats.com12
- livecodebench.github.io12
- arcprize.org6
- epoch.ai6
- storage.googleapis.com6
- matharena.ai5
- aider.chat1
- simple-bench.com1
Also known as
How our sources name o3-mini at each reasoning setting.
| Setting | API id |
|---|---|
| low | o3-mini (low) o3-mini-2025-01-31 (low) o3-mini-2025-01-31-low |
| medium | o3-mini (medium) o3-mini-2025-01-31-medium |
| high | o3-mini (high) o3-mini-2025-01-31 (high) o3-mini-2025-01-31-high o3-mini-high |