Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

o4-mini

Score basis

o4-mini is behind the leaders in multimodal tasks, long context, coding, instruction following, agentic tasks, reasoning, and math. Too few results yet to rate safety, multilingual tasks, or factuality.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited7 capabilities
  1. Multimodal−25.0%4 of 6
  2. Long Context−31.3%1 of 3
  3. Coding−34.7%3 of 10
  4. Instruction Following−34.9%2 of 3
  5. Agentic−39.4%2 of 7
  6. Reasoning−39.4%4 of 6
  7. Math−50.5%2 of 5
Not rated3 capabilities

Too few results yet to rate Safety, Multilingual or Factuality.

Price

$1.10input$4.40outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 71% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

97results on71benchmarks

  • 50 independently verified
  • 16 aggregator
  • 12 vendor-reported
  • 19 cross-referenced

From 30 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

o4-mini benchmark results

97 results on 71 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Multimodal

LimitedFull multimodal ranking

25.0% behind the leader4 of 6 ranked benchmarks measured

Show 2 more multimodal resultsHide 2 multimodal results

Long Context

LimitedFull long context ranking

31.3% behind the leader1 of 3 ranked benchmarks measured

Coding

LimitedFull coding ranking

34.7% behind the leader3 of 10 ranked benchmarks measured

Show 1 more coding resultHide 1 coding result

Instruction Following

LimitedFull instruction following ranking

34.9% behind the leader2 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

Agentic

LimitedFull agentic ranking

39.4% behind the leader2 of 7 ranked benchmarks measured

Reasoning

LimitedFull reasoning ranking

39.4% behind the leader4 of 6 ranked benchmarks measured

Show 5 more reasoning resultsHide 5 reasoning results

50.5% behind the leader2 of 5 ranked benchmarks measured

Show 2 more math resultsHide 2 math results

Factuality

Not enough dataFull factuality ranking

0 of 4 ranked benchmarks measured

Show 2 more factuality resultsHide 2 factuality results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 57 more resultsHide 57 results

o4-mini: common questions

Who makes o4-mini?

o4-mini is made by OpenAI.

When was o4-mini released?

o4-mini was released on Apr 16, 2025, according to Artificial Analysis.

What is o4-mini good at?

o4-mini is behind the leaders in multimodal tasks, long context, coding, instruction following, agentic tasks, reasoning, and math. Too few results yet to rate safety, multilingual tasks, or factuality.

How much does o4-mini cost?

o4-mini costs $1.10 per million input tokens and $4.40 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 71% of the 331 priced models we track.

How many benchmarks has o4-mini been tested on?

We track 97 results for o4-mini on 71 benchmarks from 30 sources, 50 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does o4-mini support?

OpenRouter lists tool calling, structured outputs, and reasoning for o4-mini.

About this record

Where o4-mini's numbers come from, and every name it appears under.

Tracked since
May 1, 2026
Newest source mention
Sep 2, 2026

Where the results come from

Verification: 97 scores · 50 independently verified · 16 aggregator-attributed · 19 vendor cross-reference · 12 vendor-reported. How these tiers are assigned

From 30 sources on 16 sites. cdn.openai.com supplies 19 of them; the 50 independently verified results come from 14 sites. Bars are coloured by trust tier.

  • cdn.openai.com19
  • artificialanalysis.ai18
  • api.llm-stats.com12
  • epoch.ai8
  • livecodebench.github.io8
  • matharena.ai7
  • arcprize.org6
  • storage.googleapis.com6
  • raw.githubusercontent.com5
  • swebench.com2
  • aider.chat1
  • arxiv.org1
  • datasets-server.huggingface.co1
  • labs.scale.com1
  • lmarena.ai1
  • simple-bench.com1

Also known as

How our sources name o4-mini at each reasoning setting.

SettingAPI id
lowo4-mini (low) o4-mini-low-2025-04-16
mediumo4-mini (medium) o4-mini (medium) (april 2025)
higho4-mini (high) o4-mini-2025-04-16-reasoning-high o4-mini-high o4-mini-high-2025-04-16 o4-mini (high) (april 2025)
Also listed aso4-mini (2025-04-16)o4-mini-2025-04-16o4-mini:batchopenai o4-mini