Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Grok 4.3

Score basis

Grok 4.3 is capable in factuality and long context; and behind the leaders in reasoning, multimodal tasks, instruction following, math, coding, and agentic tasks. Too few results yet to rate safety or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable2 capabilities
  1. Factuality−17.3%2 of 4
  2. Long Context−21.8%1 of 3
Limited6 capabilities
  1. Reasoning−25.4%4 of 6
  2. Multimodal−26.5%1 of 6
  3. Instruction Following−27.3%2 of 3
  4. Math−41.1%1 of 5
  5. Coding−42.4%6 of 10
  6. Agentic−46.2%4 of 7
Not rated2 capabilities

Too few results yet to rate Safety or Multilingual.

Price

$1.25input$2.50outputper million tokens

From Artificial Analysis · 3 providers tracked · All prices

Costs more than 67% of 329 priced models · 3:1 input-to-output blend, log scale

Evidence

90results on34benchmarks

  • 11 independently verified
  • 76 aggregator
  • 1 vendor-reported
  • 2 cross-referenced

From 7 sources · latest Oct 7, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Grok 4.3 benchmark results

90 results on 34 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

CapableFull factuality ranking

17.3% behind the leader2 of 4 ranked benchmarks measured

Show 6 more factuality resultsHide 6 factuality results

Long Context

CapableFull long context ranking

21.8% behind the leader1 of 3 ranked benchmarks measured

Show 4 more long context resultsHide 4 long context results

Reasoning

LimitedFull reasoning ranking

25.4% behind the leader4 of 6 ranked benchmarks measured

Show 10 more reasoning resultsHide 10 reasoning results

Multimodal

LimitedFull multimodal ranking

26.5% behind the leader1 of 6 ranked benchmarks measured

Show 5 more multimodal resultsHide 5 multimodal results

Instruction Following

LimitedFull instruction following ranking

27.3% behind the leader2 of 3 ranked benchmarks measured

Show 3 more instruction following resultsHide 3 instruction following results

41.1% behind the leader1 of 5 ranked benchmarks measured

Coding

LimitedFull coding ranking

42.4% behind the leader6 of 10 ranked benchmarks measured

Show 7 more coding resultsHide 7 coding results

Agentic

LimitedFull agentic ranking

46.2% behind the leader4 of 7 ranked benchmarks measured

Show 7 more agentic resultsHide 7 agentic results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 22 more resultsHide 22 results

Grok 4.3: common questions

Who makes Grok 4.3?

Grok 4.3 is made by SpaceXAI.

When was Grok 4.3 released?

Grok 4.3 was released on Apr 30, 2026, according to Artificial Analysis.

What is Grok 4.3 good at?

Grok 4.3 is capable in factuality and long context; and behind the leaders in reasoning, multimodal tasks, instruction following, math, coding, and agentic tasks. Too few results yet to rate safety or multilingual tasks.

How much does Grok 4.3 cost?

Grok 4.3 costs $1.25 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 67% of the 329 priced models we track.

How many benchmarks has Grok 4.3 been tested on?

We track 90 results for Grok 4.3 on 34 benchmarks from 7 sources, 11 of them independently verified. The latest was recorded on Oct 7, 2026.

Which API features does Grok 4.3 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.3.

About this record

Where Grok 4.3's numbers come from, and every name it appears under.

Tracked since
May 1, 2026
Newest source mention
Sep 28, 2026

Where the results come from

Verification: 90 scores · 11 independently verified · 76 aggregator-attributed · 2 vendor cross-reference · 1 vendor-reported. How these tiers are assigned

From 7 sources on 6 sites. Artificial Analysis supplies 76 of them; the 11 independently verified results come from 3 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai76
  • livebench.ai7
  • datasets-server.huggingface.co2
  • lmarena.ai2
  • thinkingmachines.ai2
  • api.llm-stats.com1

Also known as

How our sources name Grok 4.3 at each reasoning setting.

SettingShort formAPI id
lowgrok 4.3 (low)grok-4-3-low
mediumgrok 4.3 (medium)grok-4-3-medium
highGrok 4.3 (high)—
Also listed asgrok-4-3 (reasoning)grok-4-3-non-reasoninggrok 4.3 (non-reasoning)grok-4-3grok-4-3 (Non-Reasoning)grok-4.3:batch