Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

o3

Score basis

o3 is capable in long context, multimodal tasks, and instruction following; and behind the leaders in factuality, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable3 capabilities
  1. Long Context−18.7%2 of 3
  2. Multimodal−19.8%4 of 6
  3. Instruction Following−24.3%2 of 3
Limited4 capabilities
  1. Factuality−26.6%3 of 4
  2. Coding−31.5%3 of 10
  3. Reasoning−35.6%5 of 6
  4. Agentic−42.3%2 of 7
Not rated3 capabilities

Too few results yet to rate Safety, Math or Multilingual.

Price

$2.00input$8.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 81% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

103results on81benchmarks

  • 36 independently verified
  • 16 aggregator
  • 17 vendor-reported
  • 34 cross-referenced

From 29 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

o3 benchmark results

103 results on 81 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

CapableFull long context ranking

18.7% behind the leader2 of 3 ranked benchmarks measured

Multimodal

CapableFull multimodal ranking

19.8% behind the leader4 of 6 ranked benchmarks measured

Show 3 more multimodal resultsHide 3 multimodal results

Instruction Following

CapableFull instruction following ranking

24.3% behind the leader2 of 3 ranked benchmarks measured

Show 2 more instruction following resultsHide 2 instruction following results

Factuality

LimitedFull factuality ranking

26.6% behind the leader3 of 4 ranked benchmarks measured

Coding

LimitedFull coding ranking

31.5% behind the leader3 of 10 ranked benchmarks measured

Show 2 more coding resultsHide 2 coding results

Reasoning

LimitedFull reasoning ranking

35.6% behind the leader5 of 6 ranked benchmarks measured

Show 7 more reasoning resultsHide 7 reasoning results

Agentic

LimitedFull agentic ranking

42.3% behind the leader2 of 7 ranked benchmarks measured

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 59 more resultsHide 59 results

o3: common questions

Who makes o3?

o3 is made by OpenAI.

When was o3 released?

o3 was released on Apr 16, 2025, according to Artificial Analysis.

What is o3 good at?

o3 is capable in long context, multimodal tasks, and instruction following; and behind the leaders in factuality, coding, reasoning, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

How much does o3 cost?

o3 costs $2.00 per million input tokens and $8.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 81% of the 330 priced models we track.

How many benchmarks has o3 been tested on?

We track 103 results for o3 on 81 benchmarks from 29 sources, 36 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does o3 support?

OpenRouter lists tool calling, structured outputs, and reasoning for o3.

About this record

Where o3's numbers come from, and every name it appears under.

Tracked since
Apr 25, 2026
Newest source mention
Sep 1, 2026

Where the results come from

Verification: 103 scores · 36 independently verified · 16 aggregator-attributed · 34 vendor cross-reference · 17 vendor-reported. How these tiers are assigned

From 29 sources on 16 sites. cdn.openai.com supplies 18 of them; the 36 independently verified results come from 13 sites. Bars are coloured by trust tier.

  • cdn.openai.com18
  • api.llm-stats.com17
  • artificialanalysis.ai17
  • huggingface.co16
  • arcprize.org6
  • storage.googleapis.com6
  • epoch.ai5
  • livecodebench.github.io4
  • matharena.ai4
  • raw.githubusercontent.com3
  • swebench.com2
  • aider.chat1
  • datasets-server.huggingface.co1
  • labs.scale.com1
  • lmarena.ai1
  • simple-bench.com1

Also known as

How our sources name o3 at each reasoning setting.

SettingAPI id
lowo3 (low) o3 (preview, low) ¹
mediumo3 (medium) o3 (medium) (april 2025)
higho3 (high) o3 (high) (April 2025) o3-2025-04-16-reasoning-high o3-high
Also listed aso3 (2025-04-16)o3 (inspect)o3-2025-04-16o3:batchopenai o3