Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Grok 4.20

Score basis

Grok 4.20 is capable in instruction following and reasoning; and behind the leaders in factuality, long context, multimodal tasks, agentic tasks, and math. Too few results yet to rate coding, safety, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable2 capabilities
  1. Instruction Following−19.3%1 of 3
  2. Reasoning−22.4%4 of 6
Limited5 capabilities
  1. Factuality−26.7%3 of 4
  2. Long Context−27.5%1 of 3
  3. Multimodal−30.9%1 of 6
  4. Agentic−38.0%1 of 7
  5. Math−45.2%2 of 5
Not rated3 capabilities

Too few results yet to rate Coding, Safety or Multilingual.

Price

$1.25input$2.50outputper million tokens

From Artificial Analysis · 3 providers tracked · All prices

Costs more than 67% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

23results on23benchmarks

  • 6 independently verified
  • 17 aggregator

From 7 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Grok 4.20 benchmark results

23 results on 23 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Instruction Following

CapableFull instruction following ranking

19.3% behind the leader1 of 3 ranked benchmarks measured

Reasoning

CapableFull reasoning ranking

22.4% behind the leader4 of 6 ranked benchmarks measured

Factuality

LimitedFull factuality ranking

26.7% behind the leader3 of 4 ranked benchmarks measured

Long Context

LimitedFull long context ranking

27.5% behind the leader1 of 3 ranked benchmarks measured

Multimodal

LimitedFull multimodal ranking

30.9% behind the leader1 of 6 ranked benchmarks measured

Agentic

LimitedFull agentic ranking

38.0% behind the leader1 of 7 ranked benchmarks measured

Show 1 more agentic resultHide 1 agentic result

45.2% behind the leader2 of 5 ranked benchmarks measured

Coding

Not enough dataFull coding ranking

0 of 10 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 1 more resultHide 1 result

Grok 4.20: common questions

Who makes Grok 4.20?

Grok 4.20 is made by SpaceXAI.

When was Grok 4.20 released?

Grok 4.20 was released on Mar 10, 2026, according to Artificial Analysis.

What is Grok 4.20 good at?

Grok 4.20 is capable in instruction following and reasoning; and behind the leaders in factuality, long context, multimodal tasks, agentic tasks, and math. Too few results yet to rate coding, safety, or multilingual tasks.

How much does Grok 4.20 cost?

Grok 4.20 costs $1.25 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 67% of the 330 priced models we track.

How many benchmarks has Grok 4.20 been tested on?

We track 23 results for Grok 4.20 on 23 benchmarks from 7 sources, 6 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Grok 4.20 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.20.

About this record

Where Grok 4.20's numbers come from, and every name it appears under.

Tracked since
May 10, 2026
Newest source mention
Jun 23, 2026

Where the results come from

Verification: 23 scores · 6 independently verified · 17 aggregator-attributed. How these tiers are assigned

From 7 sources on 4 sites. Artificial Analysis supplies 17 of them; the 6 independently verified results come from 3 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai17
  • epoch.ai3
  • arcprize.org2
  • labs.scale.com1

Also known as

grok-4.20-0309-reasoninggrok-4-20grok 4.20 0309 (reasoning)grok 4.20 (reasoning)