Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Grok 3

Score basis

Grok 3 is behind the leaders in factuality, long context, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited4 capabilities
  1. Factuality−30.2%3 of 4
  2. Long Context−32.5%1 of 3
  3. Instruction Following−40.4%1 of 3
  4. Agentic−42.2%1 of 7
Not rated6 capabilities

Too few results yet to rate Reasoning, Coding, Safety, Math, Multimodal or Multilingual.

Price

$4.00input$20.00outputper million tokens

From Artificial Analysis · All prices

Costs more than 92% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

36results on32benchmarks

  • 14 independently verified
  • 15 aggregator
  • 7 vendor-reported

From 13 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputs

As listed by OpenRouter

Grok 3 benchmark results

36 results on 32 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

LimitedFull factuality ranking

30.2% behind the leader3 of 4 ranked benchmarks measured

Long Context

LimitedFull long context ranking

32.5% behind the leader1 of 3 ranked benchmarks measured

Instruction Following

LimitedFull instruction following ranking

40.4% behind the leader1 of 3 ranked benchmarks measured

Agentic

LimitedFull agentic ranking

42.2% behind the leader1 of 7 ranked benchmarks measured

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 3 more reasoning resultsHide 3 reasoning results

Coding

Not enough dataFull coding ranking

0 of 10 ranked benchmarks measured

Multimodal

Not enough dataFull multimodal ranking

0 of 6 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 14 more resultsHide 14 results

Grok 3: common questions

Who makes Grok 3?

Grok 3 is made by SpaceXAI.

When was Grok 3 released?

Grok 3 was released on Feb 19, 2025, according to Artificial Analysis.

What is Grok 3 good at?

Grok 3 is behind the leaders in factuality, long context, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, math, multimodal tasks, or multilingual tasks.

How much does Grok 3 cost?

Grok 3 costs $4.00 per million input tokens and $20.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 92% of the 331 priced models we track.

How many benchmarks has Grok 3 been tested on?

We track 36 results for Grok 3 on 32 benchmarks from 13 sources, 14 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Grok 3 support?

OpenRouter lists tool calling and structured outputs for Grok 3.

About this record

Where Grok 3's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Jul 13, 2026

Where the results come from

Verification: 36 scores · 14 independently verified · 15 aggregator-attributed · 7 vendor-reported. How these tiers are assigned

From 13 sources on 9 sites. Artificial Analysis supplies 17 of them; the 14 independently verified results come from 7 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai17
  • api.llm-stats.com4
  • raw.githubusercontent.com4
  • x.ai3
  • arcprize.org2
  • huggingface.co2
  • matharena.ai2
  • epoch.ai1
  • simple-bench.com1

Also known as

Grok 3 (Non-Reasoning)Grok 3 (Reasoning)grok 3 (think)grok 3 (no reasoning)