o1
o1 is behind the leaders in factuality, instruction following, long context, agentic tasks, and coding. Too few results yet to rate reasoning, safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Safety, Math, Multimodal or Multilingual.
Price
$15.00input$60.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
96results on74benchmarks
- 35 independently verified
- 15 aggregator
- 13 vendor-reported
- 33 cross-referenced
From 22 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
o1 benchmark results
96 results on 74 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
25.6% behind the leader3 of 4 ranked benchmarks measured
- 41.10Sep 21, 2026
- AA-Omniscience · Accuracy34.53Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination30.37Oct 8, 2026omniscienceNonHallucination
28.5% behind the leader1 of 3 ranked benchmarks measured
- IFBench70.34Oct 8, 2026aa_ifbench
Show 5 more instruction following resultsHide 5 instruction following results
- LiveBench · Instruction Following83.15Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following82.90Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following82.94Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following80.87Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following83.17Jun 17, 2026livebench_instruction_following@2025-04-07
29.0% behind the leader1 of 3 ranked benchmarks measured
- 65.00Oct 8, 2026
41.6% behind the leader1 of 7 ranked benchmarks measured
- 11.48Jun 15, 2026
42.6% behind the leader3 of 10 ranked benchmarks measured
- SciCode35.76Sep 4, 2026aa_scicode
- 41.30Oct 7, 2026
- Terminal-Bench Hard12.88Oct 8, 2026aa_terminalbench_hard
Show 6 more coding resultsHide 6 coding results
- LiveBench · Coding68.75Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding55.47Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding52.34Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding53.13Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding52.34Jun 17, 2026livebench_coding@2025-04-07
- SWE-bench Verified48.90Jun 4, 2026SWE Verified (Resolved)
0 of 6 ranked benchmarks measured
- 40.10May 10, 2026
- 41.70May 10, 2026
- 0.29Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond74.75Oct 8, 2026gpqa
- GPQA Diamond73.30Oct 7, 2026GPQA
- GPQA Diamond75.70Jun 4, 2026GPQA-Diamond (Pass@1)
- Humanity's Last Exam7.01Oct 8, 2026aa_hle
0 of 5 ranked benchmarks measured
- 14.74Sep 21, 2026
- 10.18Sep 21, 2026
- 8.42Sep 21, 2026
0 of 6 ranked benchmarks measured
- 1168.23Aug 25, 2026
- 1193May 6, 2026
- 77.60Oct 7, 2026
Show 1 more multimodal resultHide 1 multimodal result
- 71.80Oct 7, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 61.70Aug 29, 2026
- livebench_language64.75Aug 23, 2026livebench_language@2025-04-07
- livebench_language63.54Aug 23, 2026livebench_language@2025-04-07
- livebench_language77.36Aug 23, 2026livebench_language@2025-04-07
- livebench_language61.13Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence15.00Aug 7, 2026Artificial Analysis Intelligence Index
Show 56 more resultsHide 56 results
- 31.08Jun 18, 2026
- AA Intelligence15.23Oct 8, 2026aa_intelligence_index
- -11.05Oct 8, 2026
- Aider-Polyglot61.70Aug 31, 2026Aider-Polyglot (Acc.)
- AIME 202479.20Jun 4, 2026AIME 2024 (Pass@1)
- 81.67May 20, 2026
- 79.96May 19, 2026
- 98.29May 10, 2026
- Artificial Analysis Coding Index39.72Sep 9, 2026aa_coding_index
- 97.30May 10, 2026
- 0.96Jun 13, 2026
- 0.93Jun 13, 2026
- 0.05Jun 13, 2026
- 2061.00May 1, 2026
- Codeforces (Percentile)96.60Jul 3, 2026code_codeforces_percentile
- Codeforces (Rating)2061.00Jul 3, 2026code_codeforces_rating
- 90.20Jun 4, 2026
- FrontierMath (overall)5.50Oct 7, 2026FrontierMath
- 97.10Oct 7, 2026
- 96.31May 10, 2026
- 88.10Oct 7, 2026
- 52.30Oct 7, 2026
- livebench_language77.36Jun 17, 2026livebench_language@2025-04-07
- 63.40Jun 4, 2026
- 96.40May 1, 2026
- MATH-500 (EM)96.40Sep 11, 2026MATH-500 (Pass@1)
- 42.57May 10, 2026
- 0.00May 10, 2026
- 100.00May 10, 2026
- 90.80Oct 7, 2026
- MMLU91.80Jun 4, 2026MMLU (Pass@1)
- 0.88Jun 13, 2026
- 0.87Jun 13, 2026
- 0.89Jun 13, 2026
- 0.89Jun 13, 2026
- 0.88Jun 13, 2026
- 0.89Jun 13, 2026
- 0.90Jun 13, 2026
- 0.89Jun 13, 2026
- 0.88Jun 13, 2026
- 0.90Jun 13, 2026
- 0.90Jun 13, 2026
- 0.85Jun 13, 2026
- 0.75Jun 13, 2026
- 87.70Oct 7, 2026
- PersonQA hallucination ratelower is better0.16Jun 13, 2026
- 99.00May 10, 2026
- 42.40Oct 7, 2026
- SimpleQA47.00Jun 4, 2026SimpleQA (Correct)
- 0.47Jun 13, 2026
- SimpleQA hallucination ratelower is better0.44Jun 13, 2026
- 97.0Jun 13, 2026
- TAU-bench (airline)50.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)70.80Oct 7, 2026TAU-bench Retail
- 97.00May 10, 2026
- τ²-Bench Telecom (AA run)62.57Oct 8, 2026aa_tau2
o1: common questions
Who makes o1?
o1 is made by OpenAI.
When was o1 released?
o1 was released on Dec 5, 2024, according to Artificial Analysis.
What is o1 good at?
o1 is behind the leaders in factuality, instruction following, long context, agentic tasks, and coding. Too few results yet to rate reasoning, safety, math, multimodal tasks, or multilingual tasks.
How much does o1 cost?
o1 costs $15.00 per million input tokens and $60.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 97% of the 331 priced models we track.
How many benchmarks has o1 been tested on?
We track 96 results for o1 on 74 benchmarks from 22 sources, 35 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does o1 support?
OpenRouter lists tool calling, structured outputs, and reasoning for o1.
About this record
Where o1's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 9, 2026
Where the results come from
Verification: 96 scores · 35 independently verified · 15 aggregator-attributed · 33 vendor cross-reference · 13 vendor-reported. How these tiers are assigned
From 22 sources on 12 sites. Hugging Face supplies 22 of them; the 35 independently verified results come from 10 sites. Bars are coloured by trust tier.
- huggingface.co22
- cdn.openai.com20
- artificialanalysis.ai16
- api.llm-stats.com13
- raw.githubusercontent.com9
- storage.googleapis.com6
- epoch.ai4
- simple-bench.com2
- aider.chat1
- datasets-server.huggingface.co1
- lmarena.ai1
- matharena.ai1
Also known as
How our sources name o1 at each reasoning setting.
| Setting | API id |
|---|---|
| low | o1 (low) o1-2024-12-17-low |
| medium | o1 (medium) o1-2024-12-17-medium |
| high | o1 (high) o1-2024-12-17 (high) o1-2024-12-17-high |