Grok 4.3
Grok 4.3 is capable in factuality and long context; and behind the leaders in reasoning, multimodal tasks, instruction following, math, coding, and agentic tasks. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$1.25input$2.50outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
90results on34benchmarks
- 11 independently verified
- 76 aggregator
- 1 vendor-reported
- 2 cross-referenced
From 7 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
7 papers reference Grok 4.3Grok 4.3 benchmark results
90 results on 34 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
17.3% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination74.24Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy34.78Oct 7, 2026omniscienceAccuracy
Show 6 more factuality resultsHide 6 factuality results
- AA-Omniscience · Accuracy26.68Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy23.72Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy28.77Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination82.61Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination25.54Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination83.06Oct 7, 2026omniscienceNonHallucination
21.8% behind the leader1 of 3 ranked benchmarks measured
- 73.00Oct 7, 2026
25.4% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond90.10Oct 7, 2026gpqa
- LiveBench · Reasoning70.82Oct 7, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam37.21Oct 7, 2026aa_hle
- 8.00Oct 7, 2026
Show 10 more reasoning resultsHide 10 reasoning results
- 0.57Oct 7, 2026
- 0.00Oct 7, 2026
- 4.86Oct 7, 2026
- GPQA Diamond84.34Oct 7, 2026gpqa
- GPQA Diamond65.76Oct 7, 2026gpqa
- GPQA Diamond88.99Oct 7, 2026gpqa
- Humanity's Last Exam18.35Oct 7, 2026aa_hle
- Humanity's Last Exam6.77Oct 7, 2026aa_hle
- Humanity's Last Exam29.98Oct 7, 2026aa_hle
- 17.33May 2, 2026
26.5% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro78.09Oct 7, 2026aa_mmmu_pro
Show 5 more multimodal resultsHide 5 multimodal results
27.3% behind the leader2 of 3 ranked benchmarks measured
- IFBench81.29Oct 7, 2026aa_ifbench
- LiveBench · Instruction Following62.75Oct 7, 2026livebench_instruction_following@2026-06-25
41.1% behind the leader1 of 5 ranked benchmarks measured
- LiveBench · Mathematics84.34Oct 7, 2026livebench_math@2026-06-25
42.4% behind the leader6 of 10 ranked benchmarks measured
- LiveBench · Coding69.93Oct 7, 2026livebench_coding@2026-06-25
- SciCode48.26Oct 7, 2026aa_scicode
- Terminal-Bench Hard37.88Oct 7, 2026aa_terminalbench_hard
- 1357.02May 22, 2026
- Terminal-Bench 2.139.70Oct 7, 2026terminalbenchV21
- LiveBench · Agentic Coding18.54Oct 7, 2026livebench_agentic_coding@2026-06-25
Show 7 more coding resultsHide 7 coding results
- SciCode39.35Oct 7, 2026aa_scicode
- SciCode41.90Sep 4, 2026aa_scicode
- SciCode44.56Sep 4, 2026aa_scicode
- Terminal-Bench 2.134.08Oct 7, 2026terminalbenchV21
- Terminal-Bench Hard26.52Oct 7, 2026aa_terminalbench_hard
- Terminal-Bench Hard18.94Oct 7, 2026aa_terminalbench_hard
- Terminal-Bench Hard30.30Oct 7, 2026aa_terminalbench_hard
46.2% behind the leader4 of 7 ranked benchmarks measured
- 22.04Oct 7, 2026
- τ-Bench V3 · Banking12.37Oct 7, 2026tauBanking
- Terminal-Bench 4.00.00Oct 7, 2026
Show 7 more agentic resultsHide 7 agentic results
- 17.04Oct 7, 2026
- 32.72Oct 7, 2026
- 22.26Oct 7, 2026
- 31.33Jun 15, 2026
- 40.59Jun 15, 2026
- 31.24May 2, 2026
- τ-Bench V3 · Banking8.04Oct 7, 2026tauBanking
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis55.77Oct 7, 2026livebench_data_analysis@2026-06-25
- livebench_language73.58Oct 7, 2026livebench_language@2026-06-25
- 1442Jul 23, 2026
- τ²-Bench Telecom (AA run)88.89Oct 7, 2026aa_tau2
- 13.93Oct 7, 2026
- AA Intelligence24.30Oct 7, 2026aa_intelligence_index
Show 22 more resultsHide 22 results
- 15.80Sep 9, 2026
- 17.24Sep 9, 2026
- 50.42Jun 18, 2026
- 57.47Jun 18, 2026
- AA Intelligence13.99Oct 7, 2026aa_intelligence_index
- AA Intelligence24.88Oct 7, 2026aa_intelligence_index
- AA Intelligence24.78Oct 7, 2026aa_intelligence_index
- 43.89May 2, 2026
- -33.08Oct 7, 2026
- 17.98Oct 7, 2026
- 16.70Oct 7, 2026
- 13.67May 2, 2026
- Artificial Analysis Coding Index42.25Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index35.18Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index31.64Jun 18, 2026aa_coding_index
- Artificial Analysis Coding Index35.06Jun 18, 2026aa_coding_index
- 37.71Oct 6, 2026
- 61.30Jul 31, 2026
- 0.13Jul 31, 2026
- τ²-Bench Telecom (AA run)97.66Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)65.79Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)91.23Oct 7, 2026aa_tau2
Grok 4.3: common questions
Who makes Grok 4.3?
Grok 4.3 is made by SpaceXAI.
When was Grok 4.3 released?
Grok 4.3 was released on Apr 30, 2026, according to Artificial Analysis.
What is Grok 4.3 good at?
Grok 4.3 is capable in factuality and long context; and behind the leaders in reasoning, multimodal tasks, instruction following, math, coding, and agentic tasks. Too few results yet to rate safety or multilingual tasks.
How much does Grok 4.3 cost?
Grok 4.3 costs $1.25 per million input tokens and $2.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 67% of the 329 priced models we track.
How many benchmarks has Grok 4.3 been tested on?
We track 90 results for Grok 4.3 on 34 benchmarks from 7 sources, 11 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does Grok 4.3 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Grok 4.3.
About this record
Where Grok 4.3's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Sep 28, 2026
Where the results come from
Verification: 90 scores · 11 independently verified · 76 aggregator-attributed · 2 vendor cross-reference · 1 vendor-reported. How these tiers are assigned
From 7 sources on 6 sites. Artificial Analysis supplies 76 of them; the 11 independently verified results come from 3 sites. Bars are coloured by trust tier.
- artificialanalysis.ai76
- livebench.ai7
- datasets-server.huggingface.co2
- lmarena.ai2
- thinkingmachines.ai2
- api.llm-stats.com1
Also known as
How our sources name Grok 4.3 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | grok 4.3 (low) | grok-4-3-low |
| medium | grok 4.3 (medium) | grok-4-3-medium |
| high | Grok 4.3 (high) | — |