GPT-4
GPT-4 is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$30.00input$60.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
48results on30benchmarks
- 30 independently verified
- 8 aggregator
- 6 vendor-reported
- 4 cross-referenced
From 14 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
164 papers reference GPT-4GPT-4 benchmark results
48 results on 30 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
32.5% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination30.59Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy21.00Oct 8, 2026omniscienceAccuracy
0 of 6 ranked benchmarks measured
- GPQA Diamond34.95Oct 8, 2026gpqa
- GPQA Diamond35.70Oct 7, 2026GPQA
0 of 10 ranked benchmarks measured
- LiveBench · Coding46.09Aug 23, 2026livebench_coding@2025-04-07
- LiveBench · Coding42.19Jun 17, 2026livebench_coding@2025-04-07
- LiveBench · Coding36.72Jun 17, 2026livebench_coding@2025-04-07
Show 1 more coding resultHide 1 coding result
- LiveBench · Coding46.09Jun 17, 2026livebench_coding@2025-04-07
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following76.17Aug 23, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following64.70Jun 17, 2026livebench_instruction_following@2025-04-07
- LiveBench · Instruction Following77.08Jun 17, 2026livebench_instruction_following@2025-04-07
Show 2 more instruction following resultsHide 2 instruction following results
- IFBench33.20Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following71.79Jun 17, 2026livebench_instruction_following@2025-04-07
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language48.58Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence7.00Aug 5, 2026Artificial Analysis Intelligence Index
- 96.00Jul 2, 2026
- 93.20Jul 2, 2026
- livebench_language40.03Jun 17, 2026livebench_language@2025-04-07
- livebench_language48.58Jun 17, 2026livebench_language@2025-04-07
Show 28 more resultsHide 28 results
- AA Intelligence6.70Oct 8, 2026aa_intelligence_index
- -33.83Oct 8, 2026
- 64.17May 19, 2026
- Artificial Analysis Coding Index13.14Sep 9, 2026aa_coding_index
- 80.90Oct 7, 2026
- 3.23May 22, 2026
- 0.00May 22, 2026
- 0.00May 22, 2026
- 1.00Jun 9, 2026
- 92.00Aug 12, 2026
- 95.30Oct 7, 2026
- 67.00Oct 7, 2026
- 67.00Aug 12, 2026
- livebench_language49.86Jun 17, 2026livebench_language@2025-04-07
- 80.24May 10, 2026
- 42.50Aug 12, 2026
- 19.71May 10, 2026
- 18.86May 10, 2026
- 0.00May 10, 2026
- 0.00May 10, 2026
- 98.48May 10, 2026
- 98.48May 10, 2026
- 74.50Oct 7, 2026
- 73.48May 10, 2026
- 86.40Aug 12, 2026
- 76.78May 10, 2026
- 45.69May 10, 2026
- 87.50Oct 7, 2026
GPT-4: common questions
Who makes GPT-4?
GPT-4 is made by OpenAI.
When was GPT-4 released?
GPT-4 was released on Mar 14, 2023, according to Artificial Analysis.
What is GPT-4 good at?
GPT-4 is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does GPT-4 cost?
GPT-4 costs $30.00 per million input tokens and $60.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 98% of the 330 priced models we track.
How many benchmarks has GPT-4 been tested on?
We track 48 results for GPT-4 on 30 benchmarks from 14 sources, 30 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-4 support?
OpenRouter lists tool calling and structured outputs for GPT-4.
About this record
Where GPT-4's numbers come from, and every name it appears under.
- Tracked since
- May 10, 2026
- Newest source mention
- Aug 18, 2026
Where the results come from
Verification: 48 scores · 30 independently verified · 8 aggregator-attributed · 4 vendor cross-reference · 6 vendor-reported. How these tiers are assigned
From 14 sources on 6 sites. Hugging Face supplies 16 of them; the 30 independently verified results come from 4 sites. Bars are coloured by trust tier.
- huggingface.co16
- artificialanalysis.ai9
- storage.googleapis.com7
- api.llm-stats.com6
- raw.githubusercontent.com6
- x.ai4