Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

o1

Score basis

o1 is behind the leaders in factuality, instruction following, long context, agentic tasks, and coding. Too few results yet to rate reasoning, safety, math, multimodal tasks, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited5 capabilities
  1. Factuality−25.6%3 of 4
  2. Instruction Following−28.5%1 of 3
  3. Long Context−29.0%1 of 3
  4. Agentic−41.6%1 of 7
  5. Coding−42.6%3 of 10
Not rated5 capabilities

Too few results yet to rate Reasoning, Safety, Math, Multimodal or Multilingual.

Price

$15.00input$60.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 97% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

96results on74benchmarks

  • 35 independently verified
  • 15 aggregator
  • 13 vendor-reported
  • 33 cross-referenced

From 22 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

o1 benchmark results

96 results on 74 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

LimitedFull factuality ranking

25.6% behind the leader3 of 4 ranked benchmarks measured

Instruction Following

LimitedFull instruction following ranking

28.5% behind the leader1 of 3 ranked benchmarks measured

Show 5 more instruction following resultsHide 5 instruction following results

Long Context

LimitedFull long context ranking

29.0% behind the leader1 of 3 ranked benchmarks measured

Agentic

LimitedFull agentic ranking

41.6% behind the leader1 of 7 ranked benchmarks measured

Coding

LimitedFull coding ranking

42.6% behind the leader3 of 10 ranked benchmarks measured

Show 6 more coding resultsHide 6 coding results

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 4 more reasoning resultsHide 4 reasoning results

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

Multimodal

Not enough dataFull multimodal ranking

0 of 6 ranked benchmarks measured

Show 1 more multimodal resultHide 1 multimodal result

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 56 more resultsHide 56 results

o1: common questions

Who makes o1?

o1 is made by OpenAI.

When was o1 released?

o1 was released on Dec 5, 2024, according to Artificial Analysis.

What is o1 good at?

o1 is behind the leaders in factuality, instruction following, long context, agentic tasks, and coding. Too few results yet to rate reasoning, safety, math, multimodal tasks, or multilingual tasks.

How much does o1 cost?

o1 costs $15.00 per million input tokens and $60.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 97% of the 331 priced models we track.

How many benchmarks has o1 been tested on?

We track 96 results for o1 on 74 benchmarks from 22 sources, 35 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does o1 support?

OpenRouter lists tool calling, structured outputs, and reasoning for o1.

About this record

Where o1's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Sep 9, 2026

Where the results come from

Verification: 96 scores · 35 independently verified · 15 aggregator-attributed · 33 vendor cross-reference · 13 vendor-reported. How these tiers are assigned

From 22 sources on 12 sites. Hugging Face supplies 22 of them; the 35 independently verified results come from 10 sites. Bars are coloured by trust tier.

  • huggingface.co22
  • cdn.openai.com20
  • artificialanalysis.ai16
  • api.llm-stats.com13
  • raw.githubusercontent.com9
  • storage.googleapis.com6
  • epoch.ai4
  • simple-bench.com2
  • aider.chat1
  • datasets-server.huggingface.co1
  • lmarena.ai1
  • matharena.ai1

Also known as

How our sources name o1 at each reasoning setting.

SettingAPI id
lowo1 (low) o1-2024-12-17-low
mediumo1 (medium) o1-2024-12-17-medium
higho1 (high) o1-2024-12-17 (high) o1-2024-12-17-high
Also listed aso1 (2024-12-17)o1 (december 2024)o1 (inspect)o1 previewo1-2024-12-17o1-preview-2024-09-12o1:batchopenai o1-1217