Grok 4.20
Grok 4.20 is capable in instruction following and reasoning; and behind the leaders in factuality, long context, multimodal tasks, agentic tasks, and math. Too few results yet to rate coding, safety, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Coding, Safety or Multilingual.
Price
$1.25input$2.50outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
23results on23benchmarks
- 6 independently verified
- 17 aggregator
From 7 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
1 paper reference Grok 4.20Grok 4.20 benchmark results
23 results on 23 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
19.3% behind the leader1 of 3 ranked benchmarks measured
- IFBench82.93Oct 8, 2026aa_ifbench
22.4% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond88.48Oct 8, 2026gpqa
- 65.14May 10, 2026
- Humanity's Last Exam32.39Oct 8, 2026aa_hle
- 6.00Oct 8, 2026
26.7% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination77.60Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy28.88Oct 8, 2026omniscienceAccuracy
- 30.20Jul 14, 2026
27.5% behind the leader1 of 3 ranked benchmarks measured
- 67.67Oct 8, 2026
30.9% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro73.18Oct 8, 2026aa_mmmu_pro
38.0% behind the leader1 of 7 ranked benchmarks measured
- 27.03Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 14.23Oct 8, 2026
45.2% behind the leader2 of 5 ranked benchmarks measured
- 44.91Sep 9, 2026
- 17.07Sep 8, 2026
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard40.91Oct 8, 2026aa_terminalbench_hard
- SciCode44.68Sep 4, 2026aa_scicode
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 20.00Oct 8, 2026
- 89.50May 10, 2026
- 12.95Oct 8, 2026
- τ²-Bench Telecom (AA run)96.49Oct 8, 2026aa_tau2
- AA Intelligence25.22Oct 8, 2026aa_intelligence_index
- 50.88Jun 18, 2026
Show 1 more resultHide 1 result
- 42.16Jun 18, 2026
Grok 4.20: common questions
Who makes Grok 4.20?
Grok 4.20 is made by SpaceXAI.
When was Grok 4.20 released?
Grok 4.20 was released on Mar 10, 2026, according to Artificial Analysis.
What is Grok 4.20 good at?
Grok 4.20 is capable in instruction following and reasoning; and behind the leaders in factuality, long context, multimodal tasks, agentic tasks, and math. Too few results yet to rate coding, safety, or multilingual tasks.
How much does Grok 4.20 cost?
Grok 4.20 costs $1.25 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 67% of the 330 priced models we track.
How many benchmarks has Grok 4.20 been tested on?
We track 23 results for Grok 4.20 on 23 benchmarks from 7 sources, 6 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Grok 4.20 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.20.
About this record
Where Grok 4.20's numbers come from, and every name it appears under.
- Tracked since
- May 10, 2026
- Newest source mention
- Jun 23, 2026
Where the results come from
Verification: 23 scores · 6 independently verified · 17 aggregator-attributed. How these tiers are assigned
From 7 sources on 4 sites. Artificial Analysis supplies 17 of them; the 6 independently verified results come from 3 sites. Bars are coloured by trust tier.
- artificialanalysis.ai17
- epoch.ai3
- arcprize.org2
- labs.scale.com1