GPT-4o
GPT-4o is behind the leaders in factuality and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual or Instruction Following.
Price
$5.00input$15.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
261results on179benchmarks
- 52 independently verified
- 41 aggregator
- 24 vendor-reported
- 144 cross-referenced
From 39 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsWeb search
As listed by OpenRouter
Research
318 papers reference GPT-4oGPT-4o benchmark results
261 results on 179 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
35.2% behind the leader4 of 4 ranked benchmarks measured
- 9.60May 2, 2026
- AA-Omniscience · Non-hallucination62.08Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy23.70Oct 7, 2026omniscienceAccuracy
- 26.00Sep 9, 2026
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy19.87Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination41.59Oct 7, 2026omniscienceNonHallucination
- 26.00Sep 1, 2026
51.8% behind the leader5 of 6 ranked benchmarks measured
- 806.00Jun 12, 2026
- 72.20Oct 6, 2026
- 61.40Oct 6, 2026
- MMMU-Pro56.30Oct 7, 2026aa_mmmu_pro
- CharXiv (reasoning)58.80Oct 6, 2026CharXiv-R
Show 11 more multimodal resultsHide 11 multimodal results
0 of 6 ranked benchmarks measured
- 17.80May 10, 2026
- 0.00May 10, 2026
- GPQA Diamond52.63Oct 7, 2026gpqa
Show 9 more reasoning resultsHide 9 reasoning results
- 0.00Oct 7, 2026
- 0.00Oct 7, 2026
- GPQA Diamond52.12Oct 7, 2026gpqa
- GPQA Diamond51.11Oct 7, 2026gpqa
- GPQA Diamond54.34Oct 7, 2026gpqa
- Humanity's Last Exam1.76Oct 7, 2026aa_hle
- Humanity's Last Exam2.42Oct 7, 2026aa_hle
- Humanity's Last Exam2.33Oct 7, 2026aa_hle
- Humanity's Last Exam2.73Oct 7, 2026aa_hle
0 of 10 ranked benchmarks measured
- 21.62Sep 1, 2026
- LiveBench · Coding50.00Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding50.00Jun 17, 2026livebench_coding@2025-04-07
Show 11 more coding resultsHide 11 coding results
- LiveBench · Coding46.09Jun 17, 2026livebench_coding@2025-04-07
- LiveBench · Coding50.00Jun 17, 2026livebench_coding@2025-04-07
- LiveCodeBench v630.90Jun 5, 2026LiveCodeBench v6 (Pass@1)
- SciCode33.33Sep 4, 2026aa_scicode
- SciCode30.90Sep 4, 2026aa_scicode
- SciCode33.10Sep 4, 2026aa_scicode
- SciCode33.45Sep 4, 2026aa_scicode
- 33.20Oct 6, 2026
- SWE-bench Verified38.80Jun 4, 2026SWE Verified (Resolved)
- Terminal-Bench Hard8.33Oct 7, 2026aa_terminalbench_hard
- Terminal-Bench Hard8.33Oct 7, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 7, 2026
- 0.00Jun 15, 2026
- OSWorld-Verified5.03Jun 5, 2026OSWorld (Pass@1)
0 of 3 ranked benchmarks measured
- 56.00Oct 7, 2026
- 41.00Oct 7, 2026
- 49.33Oct 7, 2026
Show 1 more long context resultHide 1 long context result
- 48.10May 3, 2026
0 of 5 ranked benchmarks measured
- 0.35Sep 9, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following69.72Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following69.58Jun 17, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following72.89Jun 17, 2026livebench_instruction_following@2025-04-07
Show 4 more instruction following resultsHide 4 instruction following results
- IFBench34.29Oct 7, 2026aa_ifbench
- IFBench35.99Oct 7, 2026aa_ifbench
- LiveBench · Instruction Following70.12Jun 17, 2026livebench_instruction_following@2025-04-07
- 40.30Oct 6, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language51.20Aug 23, 2026livebench_language@2025-04-07
- 96.80Jul 2, 2026
- 96.60Jul 2, 2026
- 90.90Jul 2, 2026
- 90.50Jul 2, 2026
- livebench_language50.17Jun 17, 2026livebench_language@2025-04-07
Show 191 more resultsHide 191 results
- 8.38Jun 18, 2026
- 9.69Jun 18, 2026
- AA Intelligence7.33Oct 7, 2026aa_intelligence_index
- AA Intelligence7.74Oct 7, 2026aa_intelligence_index
- AA Intelligence8.44Oct 7, 2026aa_intelligence_index
- AA Intelligence7.19Oct 7, 2026aa_intelligence_index
- -20.87Oct 7, 2026
- -10.52Oct 7, 2026
- 94.20Oct 6, 2026
- AI2D94.20Sep 7, 2026AI2D (test)
- 83.10Jun 12, 2026
- 83.10Jun 5, 2026
- 84.60Jun 5, 2026
- 72.90May 3, 2026
- 30.70Oct 6, 2026
- Aider-Polyglot16.00Aug 31, 2026Aider-Polyglot (Acc.)
- AIME 20249.30Jun 5, 2026AIME 2024 (Pass@1)
- AIME 202413.40Jun 4, 2026AIME 2024 cons@64
- AIME 202511.60Jun 5, 2026AIME 2025 (Pass@1)
- 62.35May 19, 2026
- 52.79May 19, 2026
- AlpacaEval 2 LC57.50Sep 10, 2026
- 51.10Jun 4, 2026
- 99.14May 10, 2026
- 4.50May 10, 2026
- 79.30Sep 10, 2026
- 92.40Jun 6, 2026
- Arena-Hard (GPT-4-1106 judge)80.40Jun 4, 2026ArenaHard (GPT-4-1106)
- Artificial Analysis Coding Index24.24Sep 9, 2026aa_coding_index
- 16.67Jun 18, 2026
- Artificial Analysis Coding Index16.59Jun 18, 2026aa_coding_index
- Avg.27.80Aug 24, 2026AVG
- 95.10May 10, 2026
- 68.00Jun 5, 2026
- C-Eval76.00Aug 31, 2026C-Eval (EM)
- 85.70Oct 6, 2026
- 88.10Jun 5, 2026
- 88.10Jun 12, 2026
- Chinese SimpleQA (C-SimpleQA)58.70Sep 11, 2026C-SimpleQA (Correct)
- Chinese SimpleQA (C-SimpleQA)64.60Jun 6, 2026C-SimpleQA
- CLUEWSC87.90Aug 31, 2026CLUEWSC (EM)
- CNMO 202410.80Jul 3, 2026math_cnmo_2024_pass1
- 23.60May 3, 2026
- 759.00May 1, 2026
- Codeforces (Percentile)23.60Jul 3, 2026code_codeforces_percentile
- Codeforces (Rating)759.00Jul 3, 2026code_codeforces_rating
- 92.80Oct 6, 2026
- 91.10Jun 12, 2026
- 91.10Jun 5, 2026
- DocVQA (test, ANLS score)92.80Sep 7, 2026
- 92.80Aug 24, 2026
- 83.40Oct 6, 2026
- 80.90May 30, 2026
- 83.70Jun 5, 2026
- 89.20Jun 6, 2026
- DROP F1 Score83.40Sep 7, 2026
- 72.20Oct 6, 2026
- 72.20Jun 5, 2026
- 72.20Aug 24, 2026
- 35.20Oct 6, 2026
- 80.50Jun 4, 2026
- 13.98May 22, 2026
- 8.81May 22, 2026
- 2.04May 22, 2026
- 9.30Jun 9, 2026
- GPQA (Diamond) 0-shot CoT53.60Sep 7, 2026
- 95.60Jun 6, 2026
- 55.00Aug 24, 2026
- 82.94May 10, 2026
- 90.20Oct 6, 2026
- 90.20Jun 6, 2026
- 90.60May 30, 2026
- HumanEval 0-shot90.20Sep 7, 2026
- 80.50May 3, 2026
- 81.00Aug 31, 2026
- 84.10Jun 12, 2026
- IFEval84.30Jun 5, 2026IF-Eval (Prompt Strict)
- 84.10Jun 6, 2026
- 80.70Jun 5, 2026
- livebench_language49.02Jun 17, 2026livebench_language@2025-04-07
- livebench_language51.20Jun 17, 2026livebench_language@2025-04-07
- 38.30May 3, 2026
- LiveCodeBench32.90Jun 4, 2026LiveCodeBench pass@1
- LiveCodeBench34.20May 3, 2026LiveCodeBench (Pass@1)
- 34.20Jun 4, 2026
- LiveCodeBench (v5)32.90Jun 5, 2026LiveCodeBench v5 (Pass@1)
- 82.52May 3, 2026
- 4.46May 3, 2026
- 32.06May 3, 2026
- 54.20Jun 6, 2026
- 57.40Jun 6, 2026
- 49.70Jun 6, 2026
- 45.60Jun 6, 2026
- 43.50Jun 6, 2026
- 40.20Jun 12, 2026
- 48.60Jun 6, 2026
- 52.40Jun 6, 2026
- 51.40Jun 6, 2026
- 50.10Jun 6, 2026
- 59.60Jun 6, 2026
- 53.30Jun 6, 2026
- 66.70Jun 5, 2026
- M-LongDoc (multimodal long-document benchmark)41.40Jun 12, 2026M-LongDoc_acc
- 41.40Jun 5, 2026
- 76.60Sep 7, 2026
- 76.60Jun 6, 2026
- 74.60May 30, 2026
- MATH-500 (EM)74.60Sep 11, 2026MATH-500 (Pass@1)
- 30.40Aug 24, 2026
- 74.60May 16, 2026
- 30.40Jun 5, 2026
- 76.20Jun 6, 2026
- 49.40Jun 12, 2026
- MEGA-Bench_macro49.40Jun 5, 2026MEGA-Bench
- 25.00May 10, 2026
- 0.00May 10, 2026
- 99.24May 10, 2026
- 90.50Oct 6, 2026
- 90.40May 30, 2026
- MGSM 0-shot CoT90.50Sep 7, 2026
- 64.60Jun 5, 2026
- 82.10Aug 24, 2026
- 83.40Aug 24, 2026
- 83.10Jun 5, 2026
- 82.20Aug 24, 2026
- 2328.70Aug 24, 2026
- 42.80Jun 5, 2026
- 85.70Jun 6, 2026
- MMLU87.20Jun 4, 2026MMLU (Pass@1)
- 88.70Sep 7, 2026
- 74.70Oct 6, 2026
- MMLU-Pro72.60Aug 31, 2026MMLU-Pro (EM)
- 74.40Jun 6, 2026
- 61.10May 26, 2025
- MMLU-Redux88.00Aug 31, 2026MMLU-Redux (EM)
- 81.40Oct 6, 2026
- MMMU (val) (Pass@1)69.10Aug 24, 2026MMMU_val
- 69.10Sep 7, 2026
- 64.70Jun 5, 2026
- 65.50Aug 24, 2026
- 69.10Jun 5, 2026
- 69.10Jun 5, 2026
- MMVU67.40Jun 5, 2026MMVU (Pass@1)
- 67.40Jun 5, 2026
- MT-Bench8.74Sep 10, 2026MT-Bench (GPT-4-Turbo)
- 54.30Jun 6, 2026
- 9.90Jun 6, 2026
- 58.30Jun 6, 2026
- 33.20Jun 6, 2026
- 79.54May 10, 2026
- 80.39May 10, 2026
- 49.58May 10, 2026
- 50.13May 10, 2026
- 815.00Jun 5, 2026
- OlympiadBench25.20Jun 12, 2026OlympiadBench_full
- 25.20Jun 5, 2026
- 75.40Aug 24, 2026
- 75.40Jun 5, 2026
- 31.60Aug 24, 2026
- 97.00Jun 6, 2026
- 92.10Jun 6, 2026
- 89.00Jun 6, 2026
- 88.80Jun 6, 2026
- 88.40Jun 6, 2026
- 0.80Jun 5, 2026
- ScreenSpot-V218.10Jun 5, 2026ScreenSpot-V2 (Acc)
- 98.50May 10, 2026
- 38.20Oct 6, 2026
- 39.00Jun 6, 2026
- SimpleQA38.20Jun 4, 2026SimpleQA (Correct)
- SuperGPQA42.40Jun 5, 2026SuperGPQA (Pass@1)
- 21.62May 1, 2026
- TAU-bench (airline)42.80Oct 6, 2026TAU-bench Airline
- TAU-bench (retail)60.30Oct 6, 2026TAU-bench Retail
- 45.50Oct 6, 2026
- 37.70Jun 5, 2026
- 91.55Aug 24, 2026
- vectara_answer_rate93.80May 2, 2026Answer Rate
- vectara_avg_summary_length86.60May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency90.40May 2, 2026Factual Consistency Rate
- 71.90Jun 5, 2026
- 71.90Aug 24, 2026
- VideoMME (w sub.)77.20Jun 5, 2026Video-MME (w/ sub.)
- 61.20Oct 6, 2026
- VideoMMMU61.20Jun 5, 2026VideoMMMU (Pass@1)
- 34.00Jun 5, 2026
- 9.40Jun 5, 2026
- 97.28May 10, 2026
- τ²-Bench (Retail)63.40Oct 6, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)25.15Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)28.95Oct 7, 2026aa_tau2
GPT-4o: common questions
Who makes GPT-4o?
GPT-4o is made by OpenAI.
When was GPT-4o released?
GPT-4o was released on Aug 6, 2024, according to Artificial Analysis.
What is GPT-4o good at?
GPT-4o is behind the leaders in factuality and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, or instruction following.
How much does GPT-4o cost?
GPT-4o costs $5.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 91% of the 328 priced models we track.
How many benchmarks has GPT-4o been tested on?
We track 261 results for GPT-4o on 179 benchmarks from 39 sources, 52 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does GPT-4o support?
OpenRouter lists tool calling, structured outputs, and web search for GPT-4o.
About this record
Where GPT-4o's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 261 scores · 52 independently verified · 41 aggregator-attributed · 144 vendor cross-reference · 24 vendor-reported. How these tiers are assigned
From 39 sources on 13 sites. Hugging Face supplies 135 of them; the 52 independently verified results come from 10 sites. Bars are coloured by trust tier.
- huggingface.co135
- artificialanalysis.ai41
- api.llm-stats.com24
- raw.githubusercontent.com18
- storage.googleapis.com15
- www-cdn.anthropic.com9
- arxiv.org6
- livecodebench.github.io4
- epoch.ai3
- arcprize.org2
- swebench.com2
- lmarena.ai1
- simple-bench.com1