Grok 3
Grok 3 is behind the leaders in factuality, long context, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Math, Multimodal or Multilingual.
Price
$4.00input$20.00outputper million tokens
From Artificial Analysis · All prices
Evidence
36results on32benchmarks
- 14 independently verified
- 15 aggregator
- 7 vendor-reported
From 13 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
4 papers reference Grok 3Grok 3 benchmark results
36 results on 32 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
30.2% behind the leader3 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better5.80Jun 19, 2026
- AA-Omniscience · Accuracy28.90Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination12.68Oct 8, 2026omniscienceNonHallucination
32.5% behind the leader1 of 3 ranked benchmarks measured
- 58.00Oct 8, 2026
40.4% behind the leader1 of 3 ranked benchmarks measured
- IFBench46.94Oct 8, 2026aa_ifbench
42.2% behind the leader1 of 7 ranked benchmarks measured
- 8.33Jun 15, 2026
0 of 6 ranked benchmarks measured
- 36.10May 10, 2026
- 0.00May 10, 2026
- 0.00Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond69.29Oct 8, 2026gpqa
- GPQA Diamond84.60Oct 7, 2026GPQA
- Humanity's Last Exam4.12Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- LiveBench · Coding53.91Aug 23, 2026livebench_coding@2025-04-07
- Terminal-Bench Hard11.36Oct 8, 2026aa_terminalbench_hard
- SciCode36.81Sep 4, 2026aa_scicode
0 of 6 ranked benchmarks measured
- 78.00Oct 7, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence12.00Sep 4, 2026Artificial Analysis Intelligence Index
- 4.76Aug 24, 2026
- livebench_language51.44Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence19.00Jun 20, 2026Artificial Analysis Intelligence Index
- 95.90Jun 19, 2026
- 93.00Jun 19, 2026
Show 14 more resultsHide 14 results
- 24.64Jun 18, 2026
- AA Intelligence12.11Oct 8, 2026aa_intelligence_index
- -33.18Oct 8, 2026
- 93.30Oct 7, 2026
- 5.50May 10, 2026
- Arena ELO1402.00Jul 13, 2026Chatbot Arena
- 19.84Jun 18, 2026
- BrowseComp-ZH10.80Jul 13, 2026BrowseComp (zh)
- frontiermath_tier_4_v10.00Jun 19, 2026frontiermath_tier_4
- 79.40Aug 23, 2026
- 82.00Jul 13, 2026
- 4.17May 10, 2026
- 94.20Jun 19, 2026
- τ²-Bench Telecom (AA run)48.83Oct 8, 2026aa_tau2
Grok 3: common questions
Who makes Grok 3?
Grok 3 is made by SpaceXAI.
When was Grok 3 released?
Grok 3 was released on Feb 19, 2025, according to Artificial Analysis.
What is Grok 3 good at?
Grok 3 is behind the leaders in factuality, long context, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, or multilingual tasks.
How much does Grok 3 cost?
Grok 3 costs $4.00 per million input tokens and $20.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 92% of the 331 priced models we track.
How many benchmarks has Grok 3 been tested on?
We track 36 results for Grok 3 on 32 benchmarks from 13 sources, 14 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Grok 3 support?
OpenRouter lists tool calling and structured outputs for Grok 3.
About this record
Where Grok 3's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Jul 13, 2026
Where the results come from
Verification: 36 scores · 14 independently verified · 15 aggregator-attributed · 7 vendor-reported. How these tiers are assigned
From 13 sources on 9 sites. Artificial Analysis supplies 17 of them; the 14 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai17
- api.llm-stats.com4
- raw.githubusercontent.com4
- x.ai3
- arcprize.org2
- huggingface.co2
- matharena.ai2
- epoch.ai1
- simple-bench.com1