Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-4o

Score basis

GPT-4o is behind the leaders in factuality and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or instruction following.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited2 capabilities
  1. Factuality−35.2%4 of 4
  2. Multimodal−51.8%5 of 6
Not rated8 capabilities

Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual or Instruction Following.

Price

$5.00input$15.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 91% of 328 priced models · 3:1 input-to-output blend, log scale

Evidence

261results on179benchmarks

  • 52 independently verified
  • 41 aggregator
  • 24 vendor-reported
  • 144 cross-referenced

From 39 sources · latest Oct 7, 2026 · How verification works

API features

Tool callingStructured outputsWeb search

As listed by OpenRouter

GPT-4o benchmark results

261 results on 179 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

LimitedFull factuality ranking

35.2% behind the leader4 of 4 ranked benchmarks measured

Show 3 more factuality resultsHide 3 factuality results

Multimodal

LimitedFull multimodal ranking

51.8% behind the leader5 of 6 ranked benchmarks measured

Show 11 more multimodal resultsHide 11 multimodal results

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 9 more reasoning resultsHide 9 reasoning results

Coding

Not enough dataFull coding ranking

0 of 10 ranked benchmarks measured

Show 11 more coding resultsHide 11 coding results

Agentic

Not enough dataFull agentic ranking

0 of 7 ranked benchmarks measured

Long Context

Not enough dataFull long context ranking

0 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

Instruction Following

Not enough dataFull instruction following ranking

0 of 3 ranked benchmarks measured

Show 4 more instruction following resultsHide 4 instruction following results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 191 more resultsHide 191 results

GPT-4o: common questions

Who makes GPT-4o?

GPT-4o is made by OpenAI.

When was GPT-4o released?

GPT-4o was released on Aug 6, 2024, according to Artificial Analysis.

What is GPT-4o good at?

GPT-4o is behind the leaders in factuality and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or instruction following.

How much does GPT-4o cost?

GPT-4o costs $5.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 91% of the 328 priced models we track.

How many benchmarks has GPT-4o been tested on?

We track 261 results for GPT-4o on 179 benchmarks from 39 sources, 52 of them independently verified. The latest was recorded on Oct 7, 2026.

Which API features does GPT-4o support?

OpenRouter lists tool calling, structured outputs, and web search for GPT-4o.

About this record

Where GPT-4o's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Sep 10, 2026

Where the results come from

Verification: 261 scores · 52 independently verified · 41 aggregator-attributed · 144 vendor cross-reference · 24 vendor-reported. How these tiers are assigned

From 39 sources on 13 sites. Hugging Face supplies 135 of them; the 52 independently verified results come from 10 sites. Bars are coloured by trust tier.

  • huggingface.co135
  • artificialanalysis.ai41
  • api.llm-stats.com24
  • raw.githubusercontent.com18
  • storage.googleapis.com15
  • www-cdn.anthropic.com9
  • arxiv.org6
  • livecodebench.github.io4
  • epoch.ai3
  • arcprize.org2
  • swebench.com2
  • lmarena.ai1
  • simple-bench.com1

Also known as

GPT-4o (11-20)gpt-4o (2024-05-13)gpt-4o (2024-08-06)gpt-4o (2024-11-20)gpt-4o (aug '24)gpt-4o (chatgpt)gpt-4o (inspect)gpt-4o (may '24)gpt-4o (Non-Reasoning)gpt-4o (nov '24)gpt-4o (Reasoning)gpt-4o 0513gpt-4o 08-06gpt-4o-2024-05-13gpt-4o-2024-08-06gpt-4o-2024-11-20and 11 more