Llama 4 Maverick
Llama 4 Maverick is behind the leaders in multimodal tasks, long context, factuality, and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Math or Multilingual.
Price
$0.25input$0.87outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
41results on39benchmarks
- 10 independently verified
- 19 aggregator
- 12 vendor-reported
From 12 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Llama 4 Maverick benchmark results
41 results on 39 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
35.4% behind the leader3 of 6 ranked benchmarks measured
- 73.40Oct 7, 2026
- 73.70Oct 7, 2026
- MMMU-Pro62.14Oct 8, 2026aa_mmmu_pro
Show 2 more multimodal resultsHide 2 multimodal results
- 1147May 1, 2026
- 59.60Oct 7, 2026
36.7% behind the leader1 of 3 ranked benchmarks measured
- 50.00Oct 8, 2026
38.7% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy24.90Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination11.12Oct 8, 2026omniscienceNonHallucination
42.8% behind the leader1 of 3 ranked benchmarks measured
- IFBench42.99Oct 8, 2026aa_ifbench
0 of 6 ranked benchmarks measured
- 27.70Aug 29, 2026
- 0.00Aug 29, 2026
- 0.00Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond67.07Oct 8, 2026gpqa
- GPQA Diamond69.80Oct 7, 2026GPQA
- Humanity's Last Exam4.91Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- 5.24Oct 8, 2026
- 21.04Sep 1, 2026
- SciCode31.71Oct 8, 2026aa_scicode
Show 2 more coding resultsHide 2 coding results
- Terminal-Bench 2.17.87Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard6.82Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking3.71Oct 8, 2026tauBanking
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 4.38Aug 29, 2026
- 15.60Aug 29, 2026
- 22.50Jul 5, 2026
- 8.33May 11, 2026
- 21.04May 1, 2026
- AA Intelligence9.99Oct 8, 2026aa_intelligence_index
Show 12 more resultsHide 12 results
- 0.62Sep 9, 2026
- -41.85Oct 8, 2026
- Artificial Analysis Coding Index16.28Sep 9, 2026aa_coding_index
- 90.00Oct 7, 2026
- 94.40Oct 7, 2026
- 73.70May 1, 2026
- 73.40May 1, 2026
- LiveCodeBench43.40Sep 11, 2026code_livecodebench_1001202402012025
- 77.60Oct 7, 2026
- 92.30Oct 7, 2026
- 80.50Oct 7, 2026
- τ²-Bench Telecom (AA run)17.84Oct 8, 2026aa_tau2
Llama 4 Maverick: common questions
Who makes Llama 4 Maverick?
Llama 4 Maverick is made by Meta.
When was Llama 4 Maverick released?
Llama 4 Maverick was released on Apr 5, 2025, according to Artificial Analysis.
What is Llama 4 Maverick good at?
Llama 4 Maverick is behind the leaders in multimodal tasks, long context, factuality, and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, or multilingual tasks.
How much does Llama 4 Maverick cost?
Llama 4 Maverick costs $0.25 per million input tokens and $0.87 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 65% of the 331 priced models we track.
How many benchmarks has Llama 4 Maverick been tested on?
We track 41 results for Llama 4 Maverick on 39 benchmarks from 12 sources, 10 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Llama 4 Maverick support?
OpenRouter lists tool calling and structured outputs for Llama 4 Maverick.
About this record
Where Llama 4 Maverick's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Aug 24, 2026
Where the results come from
Verification: 41 scores · 10 independently verified · 19 aggregator-attributed · 12 vendor-reported. How these tiers are assigned
From 12 sources on 10 sites. Artificial Analysis supplies 19 of them; the 10 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai19
- api.llm-stats.com9
- raw.githubusercontent.com3
- arcprize.org2
- matharena.ai2
- swebench.com2
- aider.chat1
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1