Qwen3 Max (Reasoning)
Qwen3 Max (Reasoning) is capable in long context and instruction following and behind the leaders in reasoning and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Math, Multimodal, Multilingual or Factuality.
Price
$1.20input$6.00outputper million tokens
From Alibaba's own price page · 4 providers tracked · All prices
Evidence
39results on36benchmarks
- 11 aggregator
- 28 vendor-reported
From 2 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Qwen3 Max (Reasoning) benchmark results
39 results on 36 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
17.1% behind the leader2 of 3 ranked benchmarks measured
- 60.60Oct 7, 2026
- 74.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 68.70Oct 7, 2026
22.3% behind the leader2 of 3 ranked benchmarks measured
- IFBench70.75Oct 8, 2026aa_ifbench
- 63.30Oct 7, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 70.90Oct 7, 2026
30.2% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond86.06Oct 8, 2026gpqa
- Humanity's Last Exam27.99Oct 8, 2026aa_hle
- 1.71Oct 8, 2026
Show 1 more reasoning resultHide 1 reasoning result
- GPQA Diamond87.40Oct 7, 2026GPQA
32.9% behind the leader4 of 10 ranked benchmarks measured
- 85.90Oct 7, 2026
- 75.30Oct 7, 2026
- 66.70Oct 7, 2026
- Terminal-Bench Hard24.24Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 53.90Oct 7, 2026
0 of 5 ranked benchmarks measured
- 83.90Oct 7, 2026
- 93.30Oct 7, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy29.82Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination5.34Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- τ²-Bench Telecom (AA run)83.63Oct 8, 2026aa_tau2
- -36.62Oct 8, 2026
- AA Intelligence21.26Oct 8, 2026aa_intelligence_index
- 57.90Oct 7, 2026
- 40.90Oct 7, 2026
- 18.80Oct 7, 2026
Show 14 more resultsHide 14 results
- 81.60Oct 7, 2026
- 67.70Oct 7, 2026
- 60.90Oct 7, 2026
- 93.70Oct 7, 2026
- 28.70Oct 7, 2026
- 37.60Oct 7, 2026
- 33.50Oct 7, 2026
- 85.70Oct 7, 2026
- 92.80Oct 7, 2026
- 84.40Oct 7, 2026
- 46.90Oct 7, 2026
- 67.30Oct 7, 2026
- 74.80Oct 7, 2026
- 22.50Oct 7, 2026
Qwen3 Max (Reasoning): common questions
Who makes Qwen3 Max (Reasoning)?
Qwen3 Max (Reasoning) is made by Alibaba.
When was Qwen3 Max (Reasoning) released?
Qwen3 Max (Reasoning) was released on Nov 3, 2025, according to Artificial Analysis.
What is Qwen3 Max (Reasoning) good at?
Qwen3 Max (Reasoning) is capable in long context and instruction following and behind the leaders in reasoning and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
How much does Qwen3 Max (Reasoning) cost?
Qwen3 Max (Reasoning) costs $1.20 per million input tokens and $6.00 per million output tokens, according to Alibaba's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 74% of the 331 priced models we track.
How many benchmarks has Qwen3 Max (Reasoning) been tested on?
We track 39 results for Qwen3 Max (Reasoning) on 36 benchmarks from 2 sources. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 Max (Reasoning) support?
OpenRouter lists tool calling and structured outputs for Qwen3 Max (Reasoning).
About this record
Where Qwen3 Max (Reasoning)'s numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Jul 20, 2026
Where the results come from
Verification: 39 scores · 0 independently verified · 11 aggregator-attributed · 28 vendor-reported. How these tiers are assigned
From 2 sources on 2 sites. api.llm-stats.com supplies 28 of them. Bars are coloured by trust tier.
- api.llm-stats.com28
- artificialanalysis.ai11