Grok4.5
Grok4.5 is capable in long context, factuality, reasoning, multimodal tasks, agentic tasks, and coding; and behind the leaders in instruction following and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$2.00input$6.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
56results on47benchmarks
- 22 independently verified
- 16 aggregator
- 12 vendor-reported
- 6 cross-referenced
From 15 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
4 papers reference Grok4.5Grok4.5 benchmark results
56 results on 47 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
10.5% behind the leader2 of 3 ranked benchmarks measured
- 79.33Oct 8, 2026
- MRCR v2 (8-needle, 128K)81.40Jul 22, 2026GDM-MRCR v2 (8-needle) (128k (average))
13.5% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy51.55Oct 8, 2026omniscienceAccuracy
- 48.30Sep 21, 2026
- AA-Omniscience · Non-hallucination45.85Oct 8, 2026omniscienceNonHallucination
13.7% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond93.13Oct 8, 2026gpqa
- LiveBench · Reasoning87.17Oct 8, 2026livebench_reasoning@2026-06-25
- 70.00Jul 9, 2026
- Humanity's Last Exam42.68Oct 8, 2026aa_hle
- 52.64Sep 21, 2026
- 15.43Oct 8, 2026
Show 5 more reasoning resultsHide 5 reasoning results
- 33.06Sep 21, 2026
- 0.26Sep 21, 2026
- 0.32Sep 21, 2026
- 0.30Sep 21, 2026
- GPQA Diamond93.00Jul 17, 2026GPQA
21.8% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro80.40Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)81.60Jul 22, 2026CharXiv Reasoning (No tools)
Show 1 more multimodal resultHide 1 multimodal result
- 1287.56Aug 25, 2026
22.5% behind the leader4 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking42.06Oct 8, 2026tauBanking
- 44.46Oct 8, 2026
- Terminal-Bench 4.010.61Oct 8, 2026
Show 1 more agentic resultHide 1 agentic result
- AA ApexAgents47.10Aug 24, 2026APEX-Agents
24.2% behind the leader6 of 10 ranked benchmarks measured
- Terminal-Bench 2.181.65Oct 8, 2026terminalbenchV21
- SciCode54.98Oct 8, 2026aa_scicode
- 1552.85Jul 10, 2026
- LiveBench · Coding68.59Oct 8, 2026livebench_coding@2026-06-25
- LiveBench · Agentic Coding56.46Oct 8, 2026livebench_agentic_coding@2026-06-25
- 64.70Oct 7, 2026
Show 2 more coding resultsHide 2 coding results
- SWE-bench Pro64.70Jul 22, 2026SWE-Bench Pro (Public)
- 83.30Oct 7, 2026
36.7% behind the leader1 of 3 ranked benchmarks measured
- LiveBench · Instruction Following71.53Oct 8, 2026livebench_instruction_following@2026-06-25
41.7% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics90.82Oct 8, 2026livebench_math@2026-06-25
- 57.19Sep 21, 2026
- 24.39Sep 21, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language82.80Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis73.04Oct 8, 2026livebench_data_analysis@2026-06-25
- 1466Oct 5, 2026
- 79.17Sep 21, 2026
- 87.17Sep 21, 2026
- 85.67Sep 21, 2026
Show 15 more resultsHide 15 results
- 42.09Sep 9, 2026
- AA Intelligence38.81Oct 8, 2026aa_intelligence_index
- 25.32Oct 8, 2026
- 53.60Aug 24, 2026
- Artificial Analysis Coding Index72.45Sep 9, 2026aa_coding_index
- 53.00Jul 17, 2026
- 54.00Oct 7, 2026
- DeepSWE 1.154.00Jul 22, 2026DeepSWE v1.1
- 1526.00Aug 24, 2026
- GDPval-AA v2 Elo1535.00Jul 22, 2026GDPVal-AA v2 (Elo)
- 43.20Jul 22, 2026
- 29.00Oct 7, 2026
- Terminal-bench 3.015.70Aug 24, 2026Terminal-Bench v3.0
- Terminal-Bench 4.012.40Oct 7, 2026
- τ³-Bench Banking33.00Jul 17, 2026Tau3 Banking
Grok4.5: common questions
Who makes Grok4.5?
Grok4.5 is made by SpaceXAI.
When was Grok4.5 released?
Grok4.5 was released on Jul 8, 2026, according to Artificial Analysis.
What is Grok4.5 good at?
Grok4.5 is capable in long context, factuality, reasoning, multimodal tasks, agentic tasks, and coding; and behind the leaders in instruction following and math. Too few results yet to rate safety or multilingual tasks.
How much does Grok4.5 cost?
Grok4.5 costs $2.00 per million input tokens and $6.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 75% of the 330 priced models we track.
How many benchmarks has Grok4.5 been tested on?
We track 56 results for Grok4.5 on 47 benchmarks from 15 sources, 22 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Grok4.5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Grok4.5.
About this record
Where Grok4.5's numbers come from, and every name it appears under.
- Tracked since
- Jun 29, 2026
- Newest source mention
- Sep 26, 2026
Where the results come from
Verification: 56 scores · 22 independently verified · 16 aggregator-attributed · 6 vendor cross-reference · 12 vendor-reported. How these tiers are assigned
From 15 sources on 10 sites. Artificial Analysis supplies 16 of them; the 22 independently verified results come from 6 sites. Bars are coloured by trust tier.
- artificialanalysis.ai16
- api.llm-stats.com8
- arcprize.org8
- livebench.ai7
- deepmind.google6
- x.ai4
- epoch.ai3
- datasets-server.huggingface.co2
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name Grok4.5 at each reasoning setting.
| Setting | Short form |
|---|---|
| low | grok 4.5 (low) |
| medium | grok 4.5 (medium) |
| high | grok 4.5 (high) grok 4.5 high |