GPT-4o mini
GPT-4o mini is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
$0.15input$0.60outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
93results on74benchmarks
- 32 independently verified
- 11 aggregator
- 7 vendor-reported
- 43 cross-referenced
From 28 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsWeb search
As listed by OpenRouter
Research
104 papers reference GPT-4o miniGPT-4o mini benchmark results
93 results on 74 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
52.3% behind the leader3 of 6 ranked benchmarks measured
- 59.40Oct 7, 2026
- 56.70Oct 7, 2026
- MMMU-Pro41.50Oct 8, 2026aa_mmmu_pro
0 of 6 ranked benchmarks measured
- 10.70May 10, 2026
- 0.00May 10, 2026
- GPQA Diamond42.63Oct 8, 2026gpqa
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond40.20Oct 7, 2026GPQA
- 39.40Aug 25, 2026
- GPQA Diamond40.90May 30, 2026GPQA
- Humanity's Last Exam4.20Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- LiveBench · Coding42.97Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding42.97Jun 17, 2026livebench_coding@2025-04-07
- Terminal-Bench 2.15.62Oct 8, 2026terminalbenchV21
Show 2 more coding resultsHide 2 coding results
- SciCode22.92Sep 4, 2026aa_scicode
- 8.70Oct 7, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- τ-Bench V3 · Banking2.89Oct 8, 2026tauBanking
0 of 5 ranked benchmarks measured
- 0.70Sep 9, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following65.34Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following65.80Jun 17, 2026livebench_instruction_following@2025-04-07
- IFBench30.95Oct 8, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- 8.30Sep 1, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language32.90Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence7.00Aug 7, 2026Artificial Analysis Intelligence Index
- 92.00Jul 2, 2026
- 84.30Jul 2, 2026
- livebench_language32.90Jun 17, 2026livebench_language@2025-04-07
- 4.65Jun 9, 2026
Show 61 more resultsHide 61 results
- 0.96Aug 18, 2026
- AA Intelligence6.66Oct 8, 2026aa_intelligence_index
- AI2D75.20Aug 24, 2026AI2D (test)
- 77.80Jun 5, 2026
- 3.60May 1, 2026
- 56.26May 19, 2026
- 98.32May 10, 2026
- Artificial Analysis Coding Index11.38Sep 9, 2026aa_coding_index
- 88.20May 10, 2026
- 53.60Jun 5, 2026
- 79.70Oct 7, 2026
- 79.30May 30, 2026
- 7.53May 22, 2026
- 4.40May 22, 2026
- 0.00May 22, 2026
- 84.94May 10, 2026
- 87.20Oct 7, 2026
- 87.20Aug 25, 2026
- 86.20May 30, 2026
- 57.90Jun 5, 2026
- 39.90Aug 24, 2026
- 35.52May 3, 2026
- 81.71May 3, 2026
- 4.34May 3, 2026
- 25.17May 3, 2026
- 58.20Jun 5, 2026
- 80.15May 10, 2026
- 70.20Aug 25, 2026
- 73.00May 30, 2026
- 84.80Aug 25, 2026
- 87.00Oct 7, 2026
- 86.50May 30, 2026
- 48.10Jun 5, 2026
- 83.80Aug 24, 2026
- 77.10Jun 5, 2026
- 29.00Jun 5, 2026
- 66.78May 10, 2026
- 82.00Aug 25, 2026
- 81.80May 30, 2026
- 61.70Aug 25, 2026
- MMMU (val) (Pass@1)52.10Aug 24, 2026MMMU (val)
- 60.00Jun 5, 2026
- 54.80Jun 5, 2026
- 66.90Jun 5, 2026
- 61.60Jun 5, 2026
- 76.79May 10, 2026
- 38.55May 10, 2026
- 785.00Jun 5, 2026
- 83.60Aug 24, 2026
- 67.10Jun 5, 2026
- 84.00Aug 24, 2026
- ScreenSpot-V26.90Jun 5, 2026ScreenSpot-V2 (Acc)
- 97.75May 10, 2026
- 9.90May 30, 2026
- TextVQA-val70.90Aug 24, 2026TextVQA (val)
- 28.80Jun 5, 2026
- 64.80Jun 5, 2026
- 61.20Aug 24, 2026
- 68.90Jun 5, 2026
- 2.70Jun 5, 2026
- 96.00May 10, 2026
GPT-4o mini: common questions
Who makes GPT-4o mini?
GPT-4o mini is made by OpenAI.
When was GPT-4o mini released?
GPT-4o mini was released on Jul 18, 2024, according to Artificial Analysis.
What is GPT-4o mini good at?
GPT-4o mini is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How much does GPT-4o mini cost?
GPT-4o mini costs $0.15 per million input tokens and $0.60 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 77% of the 330 priced models we track.
How many benchmarks has GPT-4o mini been tested on?
We track 93 results for GPT-4o mini on 74 benchmarks from 28 sources, 32 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-4o mini support?
OpenRouter lists tool calling, structured outputs, and web search for GPT-4o mini.
About this record
Where GPT-4o mini's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Aug 28, 2026
Where the results come from
Verification: 93 scores · 32 independently verified · 11 aggregator-attributed · 43 vendor cross-reference · 7 vendor-reported. How these tiers are assigned
From 28 sources on 9 sites. Hugging Face supplies 53 of them; the 32 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co53
- artificialanalysis.ai12
- storage.googleapis.com12
- api.llm-stats.com7
- livecodebench.github.io4
- epoch.ai2
- aider.chat1
- arcprize.org1
- simple-bench.com1