Grok 4.6
Grok 4.6 is strong in reasoning, factuality, and long context; capable in agentic tasks; and behind the leaders in coding, math, and instruction following. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multimodal or Multilingual.
Price
$2.00input$6.00outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
87results on38benchmarks
- 23 independently verified
- 58 aggregator
- 6 vendor-reported
From 14 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
5 papers reference Grok 4.6Grok 4.6 benchmark results
87 results on 38 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.1% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning90.51Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond93.54Oct 8, 2026gpqa
- 75.90Aug 14, 2026
- 67.08Sep 21, 2026
- Humanity's Last Exam44.07Oct 8, 2026aa_hle
- 19.71Oct 8, 2026
Show 12 more reasoning resultsHide 12 reasoning results
- 27.64Sep 21, 2026
- 61.25Sep 21, 2026
- 65.14Sep 21, 2026
- 2.11Sep 21, 2026
- 17.14Oct 8, 2026
- 5.71Oct 8, 2026
- 17.71Oct 8, 2026
- GPQA Diamond94.95Oct 8, 2026gpqa
- GPQA Diamond87.88Oct 8, 2026gpqa
- Humanity's Last Exam42.91Oct 8, 2026aa_hle
- Humanity's Last Exam27.62Oct 8, 2026aa_hle
- Humanity's Last Exam42.12Oct 8, 2026aa_hle
8.5% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination75.99Oct 8, 2026omniscienceNonHallucination
- 48.90Sep 21, 2026
- AA-Omniscience · Accuracy43.00Oct 8, 2026omniscienceAccuracy
Show 7 more factuality resultsHide 7 factuality results
- AA-Omniscience · Accuracy48.23Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy43.28Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy41.93Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination65.71Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination69.35Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination76.00Oct 8, 2026omniscienceNonHallucination
- 49.30Sep 21, 2026
9.6% behind the leader1 of 3 ranked benchmarks measured
- 81.00Oct 8, 2026
18.6% behind the leader4 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking43.30Oct 8, 2026tauBanking
- 57.33Oct 8, 2026
- Terminal-Bench 4.017.17Oct 8, 2026
Show 10 more agentic resultsHide 10 agentic results
- AA ApexAgents57.50Oct 7, 2026APEX-Agents
- 56.08Oct 8, 2026
- 45.46Oct 8, 2026
- 55.86Oct 8, 2026
- Terminal-Bench 4.021.21Oct 8, 2026
- Terminal-Bench 4.03.03Oct 8, 2026
- Terminal-Bench 4.013.13Oct 8, 2026
- τ-Bench V3 · Banking50.72Oct 8, 2026tauBanking
- τ-Bench V3 · Banking38.14Oct 8, 2026tauBanking
- τ-Bench V3 · Banking44.33Oct 8, 2026tauBanking
26.0% behind the leader5 of 10 ranked benchmarks measured
- Terminal-Bench 2.188.01Oct 8, 2026terminalbenchV21
- 1616.96Sep 21, 2026
- LiveBench · Coding76.78Oct 8, 2026livebench_coding@2026-06-25
- SciCode53.01Oct 8, 2026aa_scicode
- LiveBench · Agentic Coding57.02Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 6 more coding resultsHide 6 coding results
- SciCode56.48Oct 8, 2026aa_scicode
- SciCode49.42Oct 8, 2026aa_scicode
- SciCode55.90Oct 8, 2026aa_scicode
- Terminal-Bench 2.188.39Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.175.28Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.184.27Oct 8, 2026terminalbenchV21
31.5% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics92.57Oct 8, 2026livebench_math@2026-06-25
- 65.96Sep 21, 2026
- 31.71Sep 21, 2026
35.7% behind the leader1 of 3 ranked benchmarks measured
- LiveBench · Instruction Following71.87Oct 8, 2026livebench_instruction_following@2026-06-25
0 of 6 ranked benchmarks measured
- 1264.80Sep 21, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language83.70Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis73.86Oct 8, 2026livebench_data_analysis@2026-06-25
- 74.83Sep 21, 2026
- 87.50Sep 21, 2026
- 87.00Sep 21, 2026
- 1461Sep 1, 2026
Show 21 more resultsHide 21 results
- 53.36Sep 9, 2026
- 52.73Sep 9, 2026
- 42.21Sep 9, 2026
- 51.20Sep 9, 2026
- AA Intelligence44.20Oct 8, 2026aa_intelligence_index
- AA Intelligence44.31Oct 8, 2026aa_intelligence_index
- AA Intelligence35.12Oct 8, 2026aa_intelligence_index
- AA Intelligence42.84Oct 8, 2026aa_intelligence_index
- 29.32Oct 8, 2026
- 30.48Oct 8, 2026
- 25.90Oct 8, 2026
- 28.00Oct 8, 2026
- 56.40Oct 7, 2026
- Artificial Analysis Coding Index76.79Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index75.88Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index66.31Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index74.39Sep 9, 2026aa_coding_index
- 65.90Oct 7, 2026
- 1753.00Aug 24, 2026
- 26.00Oct 7, 2026
- Terminal-Bench 4.020.30Oct 7, 2026
Grok 4.6: common questions
Who makes Grok 4.6?
Grok 4.6 is made by SpaceXAI.
When was Grok 4.6 released?
Grok 4.6 was released on Sep 2, 2026, according to SpaceXAI's own announcement.
What is Grok 4.6 good at?
Grok 4.6 is strong in reasoning, factuality, and long context; capable in agentic tasks; and behind the leaders in coding, math, and instruction following. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
How much does Grok 4.6 cost?
Grok 4.6 costs $2.00 per million input tokens and $6.00 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 75% of the 331 priced models we track.
How many benchmarks has Grok 4.6 been tested on?
We track 87 results for Grok 4.6 on 38 benchmarks from 14 sources, 23 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Grok 4.6 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.6.
About this record
Where Grok 4.6's numbers come from, and every name it appears under.
- Tracked since
- Jul 29, 2026
- Newest source mention
- Sep 22, 2026
Where the results come from
Verification: 87 scores · 23 independently verified · 58 aggregator-attributed · 6 vendor-reported. How these tiers are assigned
From 14 sources on 9 sites. Artificial Analysis supplies 58 of them; the 23 independently verified results come from 6 sites. Bars are coloured by trust tier.
- artificialanalysis.ai58
- arcprize.org8
- livebench.ai7
- api.llm-stats.com5
- epoch.ai4
- datasets-server.huggingface.co2
- lmarena.ai1
- simple-bench.com1
- x.ai1
Also known as
How our sources name Grok 4.6 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | grok 4.6 (low) | grok-4-6-low |
| medium | grok 4.6 (medium) | grok-4-6-medium |
| high | grok 4.6 (high) | grok-4.6-high |
| xhigh | grok 4.6 (xhigh) | grok-4-6-xhigh |