Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Grok 4.6

Score basis

Grok 4.6 is strong in reasoning, factuality, and long context; capable in agentic tasks; and behind the leaders in coding, math, and instruction following. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong3 capabilities
  1. Reasoning−8.1%6 of 6
  2. Factuality−8.5%3 of 4
  3. Long Context−9.6%1 of 3
Capable1 capability
  1. Agentic−18.6%4 of 7
Limited3 capabilities
  1. Coding−26.0%5 of 10
  2. Math−31.5%3 of 5
  3. Instruction Following−35.7%1 of 3
Not rated3 capabilities

Too few results yet to rate Safety, Multimodal or Multilingual.

Price

$2.00input$6.00outputper million tokens

From Artificial Analysis · 3 providers tracked · All prices

Costs more than 75% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

87results on38benchmarks

  • 23 independently verified
  • 58 aggregator
  • 6 vendor-reported

From 14 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Grok 4.6 benchmark results

87 results on 38 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Reasoning

StrongFull reasoning ranking

8.1% behind the leader6 of 6 ranked benchmarks measured

Show 12 more reasoning resultsHide 12 reasoning results

Factuality

StrongFull factuality ranking

8.5% behind the leader3 of 4 ranked benchmarks measured

Show 7 more factuality resultsHide 7 factuality results

Long Context

StrongFull long context ranking

9.6% behind the leader1 of 3 ranked benchmarks measured

Show 2 more long context resultsHide 2 long context results

Agentic

CapableFull agentic ranking

18.6% behind the leader4 of 7 ranked benchmarks measured

Show 10 more agentic resultsHide 10 agentic results

Coding

LimitedFull coding ranking

26.0% behind the leader5 of 10 ranked benchmarks measured

Show 6 more coding resultsHide 6 coding results

31.5% behind the leader3 of 5 ranked benchmarks measured

Instruction Following

LimitedFull instruction following ranking

35.7% behind the leader1 of 3 ranked benchmarks measured

Multimodal

Not enough dataFull multimodal ranking

0 of 6 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 21 more resultsHide 21 results

Grok 4.6: common questions

Who makes Grok 4.6?

Grok 4.6 is made by SpaceXAI.

When was Grok 4.6 released?

Grok 4.6 was released on Sep 2, 2026, according to SpaceXAI's own announcement.

What is Grok 4.6 good at?

Grok 4.6 is strong in reasoning, factuality, and long context; capable in agentic tasks; and behind the leaders in coding, math, and instruction following. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.

How much does Grok 4.6 cost?

Grok 4.6 costs $2.00 per million input tokens and $6.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 75% of the 331 priced models we track.

How many benchmarks has Grok 4.6 been tested on?

We track 87 results for Grok 4.6 on 38 benchmarks from 14 sources, 23 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Grok 4.6 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.6.

About this record

Where Grok 4.6's numbers come from, and every name it appears under.

Tracked since
Jul 29, 2026
Newest source mention
Sep 22, 2026

Where the results come from

Verification: 87 scores · 23 independently verified · 58 aggregator-attributed · 6 vendor-reported. How these tiers are assigned

From 14 sources on 9 sites. Artificial Analysis supplies 58 of them; the 23 independently verified results come from 6 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai58
  • arcprize.org8
  • livebench.ai7
  • api.llm-stats.com5
  • epoch.ai4
  • datasets-server.huggingface.co2
  • lmarena.ai1
  • simple-bench.com1
  • x.ai1

Also known as

How our sources name Grok 4.6 at each reasoning setting.

SettingShort formAPI id
lowgrok 4.6 (low)grok-4-6-low
mediumgrok 4.6 (medium)grok-4-6-medium
highgrok 4.6 (high)grok-4.6-high
xhighgrok 4.6 (xhigh)grok-4-6-xhigh
Also listed asgrok-4-6introducing grok 4.6