Qwen3.6 Plus
Qwen3.6 Plus is strong in long context; capable in multimodal tasks and factuality; and behind the leaders in coding, reasoning, agentic tasks, instruction following, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.50input$3.00outputper million tokens
From Alibaba's own price page · 3 providers tracked · All prices
Evidence
102results on87benchmarks
- 12 independently verified
- 18 aggregator
- 59 vendor-reported
- 13 cross-referenced
From 10 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
3 papers reference Qwen3.6 PlusQwen3.6 Plus benchmark results
102 results on 87 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
7.5% behind the leader2 of 3 ranked benchmarks measured
- 62.00Oct 7, 2026
- 78.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 68.30Oct 7, 2026
17.0% behind the leader4 of 6 ranked benchmarks measured
- 86.00Oct 7, 2026
- 84.20Oct 7, 2026
- MMMU-Pro77.98Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)81.50Oct 7, 2026CharXiv-R
20.5% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination65.36Oct 8, 2026omniscienceNonHallucination
- 44.14May 20, 2026
- AA-Omniscience · Accuracy26.38Oct 8, 2026omniscienceAccuracy
27.4% behind the leader10 of 10 ranked benchmarks measured
- 87.10Oct 7, 2026
- LiveBench · Coding78.18Oct 8, 2026livebench_coding@2026-06-25
- 78.80Oct 7, 2026
- 73.80Oct 7, 2026
- Terminal-Bench 2.161.42Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard43.94Oct 8, 2026aa_terminalbench_hard
- 1461.20May 22, 2026
- 56.60Oct 7, 2026
- SciCode40.74Sep 4, 2026aa_scicode
- LiveBench · Agentic Coding41.36Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 1 more coding resultHide 1 coding result
- 56.60Oct 8, 2026
30.9% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond88.18Oct 8, 2026gpqa
- LiveBench · Reasoning75.83Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam27.85Oct 8, 2026aa_hle
- 2.86Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond90.40Oct 7, 2026GPQA
- 90.40Oct 8, 2026
- 28.80Oct 7, 2026
- 28.80Oct 8, 2026
35.4% behind the leader4 of 7 ranked benchmarks measured
- 74.10Oct 7, 2026
- 62.50Oct 7, 2026
- τ-Bench V3 · Banking20.82Oct 8, 2026tauBanking
- 24.65Oct 8, 2026
Show 1 more agentic resultHide 1 agentic result
- MCP Atlas74.10Oct 8, 2026MCP-Atlas (Public Set)
35.8% behind the leader2 of 3 ranked benchmarks measured
- IFBench75.17Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following58.34Oct 8, 2026livebench_instruction_following@2026-06-25
Show 1 more instruction following resultHide 1 instruction following result
- 74.20Oct 7, 2026
46.4% behind the leader2 of 5 ranked benchmarks measured
- LiveBench · Mathematics83.72Oct 8, 2026livebench_math@2026-06-25
- 38.25Sep 9, 2026
Show 6 more math resultsHide 6 math results
- 95.30Oct 7, 2026
- 95.10Oct 8, 2026
- HMMT Feb 202687.80Oct 7, 2026HMMT Feb 26
- HMMT Feb 202687.80Oct 8, 2026HMMT Feb. 2026
- 83.80Oct 7, 2026
- 83.80Oct 8, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language74.99Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis69.91Oct 8, 2026livebench_data_analysis@2026-06-25
- 1444Jul 23, 2026
- frontiermath_tier_4_v18.33May 20, 2026frontiermath_tier_4
- 0.88Oct 8, 2026
- AA Intelligence27.01Oct 8, 2026aa_intelligence_index
Show 49 more resultsHide 49 results
- 29.00Sep 4, 2026
- 57.65Jun 24, 2026
- 55.28Jun 24, 2026
- 60.33Jun 24, 2026
- 50.81Jun 24, 2026
- 21.94Jun 24, 2026
- 59.08Jun 24, 2026
- 50.58Jun 24, 2026
- 50.78Jun 24, 2026
- 94.40Oct 7, 2026
- Artificial Analysis Coding Index54.53Sep 9, 2026aa_coding_index
- 93.30Oct 7, 2026
- 83.40Oct 7, 2026
- Claw Eval (pass@3)58.70Oct 7, 2026Claw-Eval
- 41.50Oct 7, 2026
- 88.00Oct 7, 2026
- 65.70Oct 7, 2026
- 40.85Oct 7, 2026
- HLE (with tools)50.60Oct 8, 2026HLE (w/ Tools)
- 94.60Oct 7, 2026
- 94.60Oct 8, 2026
- 94.30Aug 31, 2026
- 85.10Oct 7, 2026
- 70.85Oct 7, 2026
- 88.00Oct 7, 2026
- 48.20Oct 7, 2026
- 86.70Oct 7, 2026
- 62.00Oct 7, 2026
- 88.50Oct 7, 2026
- 84.70Oct 7, 2026
- 94.50Oct 7, 2026
- 89.50Oct 7, 2026
- 37.90Oct 7, 2026
- 37.90Oct 8, 2026
- 91.20Oct 7, 2026
- 85.40Oct 7, 2026
- ScreenSpot-Pro (No tools)68.20Oct 7, 2026ScreenSpot Pro
- 67.30Oct 7, 2026
- 71.60Oct 7, 2026
- 61.60Oct 7, 2026
- Terminal-Bench 2.061.60Oct 8, 2026Terminal-Bench 2.0 (Terminus-2)
- 39.80Oct 8, 2026
- 39.80Oct 7, 2026
- 84.00Oct 7, 2026
- 44.30Oct 7, 2026
- 74.30Oct 7, 2026
- τ²-Bench Telecom (AA run)97.66Oct 8, 2026aa_tau2
- τ³-Bench70.70Oct 7, 2026TAU3-Bench
- 70.70Oct 8, 2026
Qwen3.6 Plus: common questions
Who makes Qwen3.6 Plus?
Qwen3.6 Plus is made by Alibaba.
When was Qwen3.6 Plus released?
Qwen3.6 Plus was released on Apr 2, 2026, according to Artificial Analysis.
What is Qwen3.6 Plus good at?
Qwen3.6 Plus is strong in long context; capable in multimodal tasks and factuality; and behind the leaders in coding, reasoning, agentic tasks, instruction following, and math. Too few results yet to rate safety or multilingual tasks.
How much does Qwen3.6 Plus cost?
Qwen3.6 Plus costs $0.50 per million input tokens and $3.00 per million output tokens, according to Alibaba's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 61% of the 330 priced models we track.
How many benchmarks has Qwen3.6 Plus been tested on?
We track 102 results for Qwen3.6 Plus on 87 benchmarks from 10 sources, 12 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3.6 Plus support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3.6 Plus.
About this record
Where Qwen3.6 Plus's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- May 30, 2026
Where the results come from
Verification: 102 scores · 12 independently verified · 18 aggregator-attributed · 13 vendor cross-reference · 59 vendor-reported. How these tiers are assigned
From 10 sources on 7 sites. api.llm-stats.com supplies 51 of them; the 12 independently verified results come from 4 sites. Bars are coloured by trust tier.
- api.llm-stats.com51
- huggingface.co21
- artificialanalysis.ai18
- livebench.ai7
- epoch.ai3
- datasets-server.huggingface.co1
- lmarena.ai1