Grok 4
Grok 4 is capable in factuality; and behind the leaders in long context, reasoning, multimodal tasks, instruction following, and agentic tasks. Too few results yet to rate coding, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Coding, Safety, Math or Multilingual.
Price
$3.00input$15.00outputper million tokens
From Artificial Analysis · All prices
Evidence
54results on43benchmarks
- 19 independently verified
- 16 aggregator
- 6 vendor-reported
- 13 cross-referenced
From 21 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
26 papers reference Grok 4Grok 4 benchmark results
54 results on 43 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
18.3% behind the leader3 of 4 ranked benchmarks measured
- 47.90May 20, 2026
- AA-Omniscience · Accuracy40.48Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination35.51Oct 8, 2026omniscienceNonHallucination
27.2% behind the leader1 of 3 ranked benchmarks measured
- 68.00Oct 8, 2026
30.4% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond87.68Oct 8, 2026gpqa
- Humanity's Last Exam26.69Oct 8, 2026aa_hle
- ARC-AGI-215.90Oct 7, 2026ARC-AGI v2
- 2.00Oct 8, 2026
Show 6 more reasoning resultsHide 6 reasoning results
- 15.97May 10, 2026
- GPQA Diamond87.50Oct 7, 2026GPQA
- GPQA Diamond87.50May 15, 2026GPQA no tools
- 40.00Oct 7, 2026
- Humanity's Last Exam25.40Jun 12, 2026HLE (Text-only) no tools
- 60.50May 10, 2026
32.9% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro68.84Oct 8, 2026aa_mmmu_pro
Show 2 more multimodal resultsHide 2 multimodal results
- 1210.48Aug 25, 2026
- 1182May 1, 2026
37.2% behind the leader1 of 3 ranked benchmarks measured
- IFBench53.67Oct 8, 2026aa_ifbench
38.8% behind the leader1 of 7 ranked benchmarks measured
- 24.59Jun 15, 2026
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard37.88Oct 8, 2026aa_terminalbench_hard
- SciCode45.72Sep 4, 2026aa_scicode
0 of 5 ranked benchmarks measured
- IMO-AnswerBench73.10May 15, 2026IMO-AnswerBench no tools
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 11.90Sep 2, 2026
- 70.03Sep 2, 2026
- 79.60Aug 29, 2026
- AA Intelligence33.00Jul 3, 2026Artificial Analysis Intelligence Index
- frontiermath_tier_4_v12.08May 20, 2026frontiermath_tier_4
- 44.38May 19, 2026
Show 26 more resultsHide 26 results
- 41.50Jun 18, 2026
- AA Intelligence22.49Oct 8, 2026aa_intelligence_index
- 2.10Oct 8, 2026
- 91.70Oct 7, 2026
- AIME 202591.70May 15, 2026AIME25
- 100.00May 15, 2026
- 91.70Jun 12, 2026
- 98.80May 15, 2026
- 95.65May 10, 2026
- 66.67May 10, 2026
- 40.49Jun 18, 2026
- 93.69May 10, 2026
- 39.69May 10, 2026
- HLE (with tools)50.70May 15, 2026HLE (Text-only) heavy
- HLE (with tools)41.00May 15, 2026HLE (Text-only) w/ tools
- HMMT 202590.00Oct 7, 2026HMMT25
- HMMT 202590.00May 15, 2026HMMT25
- 88.33May 10, 2026
- 96.70May 15, 2026
- 90.00Jun 12, 2026
- 93.90May 15, 2026
- 0.00May 10, 2026
- 79.00Aug 23, 2026
- 92.21May 10, 2026
- 96.61May 10, 2026
- τ²-Bench Telecom (AA run)74.85Oct 8, 2026aa_tau2
Grok 4: common questions
Who makes Grok 4?
Grok 4 is made by SpaceXAI.
When was Grok 4 released?
Grok 4 was released on Jul 10, 2025, according to Artificial Analysis.
What is Grok 4 good at?
Grok 4 is capable in factuality; and behind the leaders in long context, reasoning, multimodal tasks, instruction following, and agentic tasks. Too few results yet to rate coding, safety, math, or multilingual tasks.
How much does Grok 4 cost?
Grok 4 costs $3.00 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 88% of the 331 priced models we track.
How many benchmarks has Grok 4 been tested on?
We track 54 results for Grok 4 on 43 benchmarks from 21 sources, 19 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Grok 4 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.
About this record
Where Grok 4's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Jun 18, 2026
Where the results come from
Verification: 54 scores · 19 independently verified · 16 aggregator-attributed · 13 vendor cross-reference · 6 vendor-reported. How these tiers are assigned
From 21 sources on 11 sites. Artificial Analysis supplies 17 of them; the 19 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai17
- huggingface.co13
- api.llm-stats.com6
- storage.googleapis.com6
- matharena.ai4
- arcprize.org2
- epoch.ai2
- aider.chat1
- datasets-server.huggingface.co1
- lmarena.ai1
- simple-bench.com1