o3
o3 is capable in long context, multimodal tasks, and instruction following; and behind the leaders in factuality, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$2.00input$8.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
103results on81benchmarks
- 36 independently verified
- 16 aggregator
- 17 vendor-reported
- 34 cross-referenced
From 29 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
7 papers reference o3o3 benchmark results
103 results on 81 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
18.7% behind the leader2 of 3 ranked benchmarks measured
- 58.80Jun 5, 2026
- 74.67Oct 8, 2026
19.8% behind the leader4 of 6 ranked benchmarks measured
- 82.90Oct 7, 2026
- 86.80Oct 7, 2026
- CharXiv (reasoning)78.60Oct 7, 2026CharXiv-R
- MMMU-Pro70.06Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- 1214.19Aug 25, 2026
- 1217Jun 17, 2026
- 76.40Oct 7, 2026
24.3% behind the leader2 of 3 ranked benchmarks measured
- IFBench71.43Oct 8, 2026aa_ifbench
- 60.40Oct 7, 2026
Show 2 more instruction following resultsHide 2 instruction following results
- 56.62Oct 8, 2026
- 56.50Jun 5, 2026
26.6% behind the leader3 of 4 ranked benchmarks measured
- 49.40May 20, 2026
- AA-Omniscience · Accuracy38.57Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination11.91Oct 8, 2026omniscienceNonHallucination
31.5% behind the leader3 of 10 ranked benchmarks measured
- 69.10Oct 7, 2026
- SciCode40.97Sep 4, 2026aa_scicode
- Terminal-Bench Hard37.12Oct 8, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- 58.40Sep 1, 2026
- 69.10Jun 5, 2026
35.6% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond82.73Oct 8, 2026gpqa
- 53.10May 10, 2026
- Humanity's Last Exam20.05Oct 8, 2026aa_hle
- ARC-AGI-26.50Oct 7, 2026ARC-AGI v2
- 1.14Oct 8, 2026
Show 7 more reasoning resultsHide 7 reasoning results
- 6.53Sep 21, 2026
- 2.98Sep 21, 2026
- 1.99Sep 21, 2026
- GPQA Diamond83.30Oct 7, 2026GPQA
- 83.30Jun 5, 2026
- 14.70Oct 7, 2026
- Humanity's Last Exam20.30Jun 5, 2026HLE (no tools)
42.3% behind the leader2 of 7 ranked benchmarks measured
- 49.70Oct 7, 2026
- 12.84Jun 15, 2026
0 of 5 ranked benchmarks measured
- 33.33Sep 21, 2026
- 29.82Sep 21, 2026
- 19.30Sep 21, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 60.83Sep 21, 2026
- 53.83Sep 21, 2026
- 41.50Sep 21, 2026
- 76.90Sep 11, 2026
- 16.67Sep 2, 2026
- frontiermath_tier_4_v12.08Aug 29, 2026frontiermath_tier_4
Show 59 more resultsHide 59 results
- 36.09Jun 18, 2026
- AA Intelligence20.00Aug 7, 2026Artificial Analysis Intelligence Index
- AA Intelligence20.20Oct 8, 2026aa_intelligence_index
- -15.55Oct 8, 2026
- 81.30Oct 7, 2026
- 91.60Jun 5, 2026
- 89.17May 2, 2026
- 86.40Oct 7, 2026
- 88.90Jun 5, 2026
- 84.47May 19, 2026
- 98.25May 10, 2026
- 38.40Jun 18, 2026
- 97.90May 10, 2026
- 0.94Jun 13, 2026
- 0.93Jun 13, 2026
- 0.25Jun 13, 2026
- 64.00Oct 7, 2026
- FrontierMath (overall)15.80Oct 7, 2026FrontierMath
- 69.30Jun 5, 2026
- 98.38May 10, 2026
- 77.50May 13, 2026
- 0.00May 10, 2026
- 84.74May 3, 2026
- 75.80Jun 5, 2026
- 99.07May 3, 2026
- 66.00May 3, 2026
- 89.82May 3, 2026
- MATH-500 (EM)98.10Jun 5, 2026MATH-500
- 64.65May 10, 2026
- 15.00May 10, 2026
- 99.75May 10, 2026
- 0.90Jun 13, 2026
- 0.89Jun 13, 2026
- 0.88Jun 13, 2026
- 0.89Jun 13, 2026
- 0.91Jun 13, 2026
- 0.91Jun 13, 2026
- 0.90Jun 13, 2026
- 0.90Jun 13, 2026
- 0.91Jun 13, 2026
- 0.89Jun 13, 2026
- 0.91Jun 13, 2026
- 85.00Jun 5, 2026
- 56.50Jun 5, 2026
- PersonQA hallucination ratelower is better0.33Jun 13, 2026
- 99.00May 10, 2026
- 49.40Jun 5, 2026
- 0.49Jun 13, 2026
- SimpleQA hallucination ratelower is better0.51Jun 13, 2026
- 97.0Jun 13, 2026
- 58.40May 1, 2026
- 52.00Jun 5, 2026
- 73.90Jun 5, 2026
- 64.80Oct 7, 2026
- 83.30Oct 7, 2026
- 97.28May 10, 2026
- 95.80Jun 5, 2026
- τ²-Bench (Retail)80.20Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)80.70Oct 8, 2026aa_tau2
o3: common questions
Who makes o3?
o3 is made by OpenAI.
When was o3 released?
o3 was released on Apr 16, 2025, according to Artificial Analysis.
What is o3 good at?
o3 is capable in long context, multimodal tasks, and instruction following; and behind the leaders in factuality, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does o3 cost?
o3 costs $2.00 per million input tokens and $8.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 81% of the 330 priced models we track.
How many benchmarks has o3 been tested on?
We track 103 results for o3 on 81 benchmarks from 29 sources, 36 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does o3 support?
OpenRouter lists tool calling, structured outputs, and reasoning for o3.
About this record
Where o3's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 103 scores · 36 independently verified · 16 aggregator-attributed · 34 vendor cross-reference · 17 vendor-reported. How these tiers are assigned
From 29 sources on 16 sites. cdn.openai.com supplies 18 of them; the 36 independently verified results come from 13 sites. Bars are coloured by trust tier.
- cdn.openai.com18
- api.llm-stats.com17
- artificialanalysis.ai17
- huggingface.co16
- arcprize.org6
- storage.googleapis.com6
- epoch.ai5
- livecodebench.github.io4
- matharena.ai4
- raw.githubusercontent.com3
- swebench.com2
- aider.chat1
- datasets-server.huggingface.co1
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name o3 at each reasoning setting.
| Setting | API id |
|---|---|
| low | o3 (low) o3 (preview, low) ¹ |
| medium | o3 (medium) o3 (medium) (april 2025) |
| high | o3 (high) o3 (high) (April 2025) o3-2025-04-16-reasoning-high o3-high |