Qwen3.7 Max
Qwen3.7 Max is strong in factuality; capable in instruction following, long context, reasoning, and coding; and behind the leaders in agentic tasks and math. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multimodal or Multilingual.
Price
$2.50input$7.50outputper million tokens
From Alibaba's own price page · 4 providers tracked · All prices
Evidence
86results on70benchmarks
- 13 independently verified
- 19 aggregator
- 41 vendor-reported
- 13 cross-referenced
From 11 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
6 papers reference Qwen3.7 MaxQwen3.7 Max benchmark results
86 results on 70 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.8% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination74.37Oct 8, 2026omniscienceNonHallucination
- 55.77Sep 20, 2026
- AA-Omniscience · Accuracy31.13Oct 8, 2026omniscienceAccuracy
12.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench80.54Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following74.04Oct 8, 2026livebench_instruction_following@2026-06-25
Show 1 more instruction following resultHide 1 instruction following result
- 79.10Oct 7, 2026
13.2% behind the leader1 of 3 ranked benchmarks measured
- 79.00Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 65.30Aug 24, 2026
17.2% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond92.32Oct 8, 2026gpqa
- LiveBench · Reasoning83.34Oct 8, 2026livebench_reasoning@2026-06-25
- 70.40Sep 20, 2026
- Humanity's Last Exam40.50Oct 8, 2026aa_hle
- 13.43Oct 8, 2026
Show 6 more reasoning resultsHide 6 reasoning results
- 11.40Oct 7, 2026
- 13.40Sep 20, 2026
- GPQA Diamond92.40Oct 7, 2026GPQA
- 90.00Sep 20, 2026
- 41.40Oct 7, 2026
- 41.40Sep 20, 2026
19.5% behind the leader10 of 10 ranked benchmarks measured
- 91.60Oct 7, 2026
- 80.40Oct 7, 2026
- 78.30Oct 7, 2026
- Terminal-Bench 2.174.53Oct 8, 2026terminalbenchV21
- LiveBench · Coding74.22Oct 8, 2026livebench_coding@2026-06-25
- Terminal-Bench Hard50.76Oct 8, 2026aa_terminalbench_hard
- SciCode49.54Oct 8, 2026aa_scicode
- 1515.08Sep 20, 2026
- 60.60Oct 7, 2026
- LiveBench · Agentic Coding43.59Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 4 more coding resultsHide 4 coding results
- 53.50Oct 7, 2026
- 60.60Sep 20, 2026
- 74.50Aug 12, 2026
- Terminal-Bench 2.175.00Sep 20, 2026Terminal Bench 2.1 (Terminus-2)
34.9% behind the leader5 of 7 ranked benchmarks measured
- 76.40Oct 7, 2026
- 31.56Oct 8, 2026
- τ-Bench V3 · Banking11.75Oct 8, 2026tauBanking
- Terminal-Bench 4.01.52Oct 8, 2026
Show 2 more agentic resultsHide 2 agentic results
- 42.47Oct 8, 2026
- MCP Atlas76.40Sep 20, 2026MCP-Atlas (Public Set)
35.5% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics85.25Oct 8, 2026livebench_math@2026-06-25
- 64.56Sep 20, 2026
- 34.15Sep 20, 2026
Show 5 more math resultsHide 5 math results
- 97.00Sep 20, 2026
- HMMT Feb 202697.10Oct 7, 2026HMMT Feb 26
- HMMT Feb 202697.10Sep 20, 2026HMMT Feb. 2026
- 90.00Oct 7, 2026
- 90.00Sep 20, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language79.74Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis71.79Oct 8, 2026livebench_data_analysis@2026-06-25
- 1475Oct 5, 2026
- AA Intelligence29.46Oct 8, 2026aa_intelligence_index
- 13.48Oct 8, 2026
- τ²-Bench Telecom (AA run)94.74Oct 8, 2026aa_tau2
Show 33 more resultsHide 33 results
- 23.90Sep 9, 2026
- 31.10Aug 12, 2026
- 56.50Aug 12, 2026
- Artificial Analysis Coding Index65.97Sep 9, 2026aa_coding_index
- 75.00Oct 7, 2026
- Claw Eval (pass@3)65.20Oct 7, 2026Claw-Eval
- 18.00Sep 20, 2026
- 21.60Aug 12, 2026
- 48.35Oct 7, 2026
- 54.50Aug 12, 2026
- HLE (with tools)53.50Aug 12, 2026HLE w/ tools
- HLE (with tools)53.50Sep 20, 2026HLE (w/ Tools)
- 95.00Sep 20, 2026
- 94.30Aug 31, 2026
- 86.20Oct 7, 2026
- 31.30Aug 12, 2026
- 74.29Oct 7, 2026
- MathArena Apex (Pass@1)44.50Oct 7, 2026MathArena Apex
- 60.80Oct 7, 2026
- MLS Bench Litelower is better31.70Aug 12, 2026
- 89.60Oct 7, 2026
- 87.00Oct 7, 2026
- 95.00Oct 7, 2026
- 90.30Oct 7, 2026
- 86.70Aug 24, 2026
- 47.20Oct 7, 2026
- 47.20Sep 20, 2026
- 64.80Aug 12, 2026
- 73.60Oct 7, 2026
- 69.70Oct 7, 2026
- Toolathlon Verified49.70Aug 12, 2026Toolathlon Verified (Pass@1)
- 47.90Oct 7, 2026
- 75.20Aug 12, 2026
Qwen3.7 Max: common questions
Who makes Qwen3.7 Max?
Qwen3.7 Max is made by Alibaba.
When was Qwen3.7 Max released?
Qwen3.7 Max was released on May 19, 2026, according to Artificial Analysis.
What is Qwen3.7 Max good at?
Qwen3.7 Max is strong in factuality; capable in instruction following, long context, reasoning, and coding; and behind the leaders in agentic tasks and math. Too few results yet to rate safety, multimodal tasks, or multilingual tasks.
How much does Qwen3.7 Max cost?
Qwen3.7 Max costs $2.50 per million input tokens and $7.50 per million output tokens, according to Alibaba's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 82% of the 331 priced models we track.
How many benchmarks has Qwen3.7 Max been tested on?
We track 86 results for Qwen3.7 Max on 70 benchmarks from 11 sources, 13 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3.7 Max support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3.7 Max.
About this record
Where Qwen3.7 Max's numbers come from, and every name it appears under.
- Tracked since
- May 20, 2026
- Newest source mention
- Aug 31, 2026
Where the results come from
Verification: 86 scores · 13 independently verified · 19 aggregator-attributed · 13 vendor cross-reference · 41 vendor-reported. How these tiers are assigned
From 11 sources on 8 sites. api.llm-stats.com supplies 28 of them; the 13 independently verified results come from 5 sites. Bars are coloured by trust tier.
- api.llm-stats.com28
- huggingface.co26
- artificialanalysis.ai19
- livebench.ai7
- epoch.ai3
- datasets-server.huggingface.co1
- lmarena.ai1
- simple-bench.com1