Qwen3.7 Plus Preview
Qwen3.7 Plus Preview is capable in multimodal tasks, coding, instruction following, long context, reasoning, and factuality; and behind the leaders in agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$0.40input$1.60outputper million tokens
From Alibaba's own price page · 3 providers tracked · All prices
Evidence
85results on73benchmarks
- 4 independently verified
- 20 aggregator
- 61 vendor-reported
From 7 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Qwen3.7 Plus Preview benchmark results
85 results on 73 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
14.7% behind the leader3 of 6 ranked benchmarks measured
- 88.00Oct 6, 2026
- CharXiv (reasoning)85.90Oct 6, 2026CharXiv-R
- MMMU-Pro80.46Oct 7, 2026aa_mmmu_pro
Show 4 more multimodal resultsHide 4 multimodal results
- 1280.38Aug 25, 2026
- 1263Jun 17, 2026
- 79.00Oct 6, 2026
- 67.10Oct 6, 2026
19.6% behind the leader7 of 10 ranked benchmarks measured
- 89.60Oct 6, 2026
- 77.70Oct 6, 2026
- 75.80Oct 6, 2026
- Terminal-Bench Hard46.97Oct 7, 2026aa_terminalbench_hard
- SciCode46.06Oct 7, 2026aa_scicode
- Terminal-Bench 2.161.05Oct 7, 2026terminalbenchV21
- 57.60Oct 6, 2026
Show 3 more coding resultsHide 3 coding results
- 51.30Oct 6, 2026
- 55.80Aug 26, 2026
- Terminal-Bench 2.164.00Aug 14, 2026Terminal Bench 2.1 (Terminus)
21.7% behind the leader1 of 3 ranked benchmarks measured
- IFBench77.96Oct 7, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- 79.10Oct 6, 2026
21.8% behind the leader1 of 3 ranked benchmarks measured
- 73.00Oct 7, 2026
22.5% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond90.00Oct 7, 2026gpqa
- Humanity's Last Exam35.63Oct 7, 2026aa_hle
- 9.14Oct 7, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- 6.00Oct 6, 2026
- GPQA Diamond90.30Oct 6, 2026GPQA
- 34.70Oct 6, 2026
23.5% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination72.34Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy22.52Oct 7, 2026omniscienceAccuracy
35.0% behind the leader5 of 7 ranked benchmarks measured
- 73.30Oct 6, 2026
- 73.20Oct 6, 2026
- τ-Bench V3 · Banking17.53Oct 7, 2026tauBanking
- 13.52Oct 7, 2026
- Terminal-Bench 4.01.01Oct 7, 2026
Show 1 more agentic resultHide 1 agentic result
- 22.42Oct 7, 2026
0 of 5 ranked benchmarks measured
- 34.39Sep 9, 2026
- 86.00Oct 6, 2026
- HMMT Feb 202692.90Oct 6, 2026HMMT Feb 26
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1458Aug 11, 2026
- AA Intelligence25.16Oct 7, 2026aa_intelligence_index
- 1.08Oct 7, 2026
- τ²-Bench Telecom (AA run)92.98Oct 7, 2026aa_tau2
- 19.67Sep 9, 2026
- Artificial Analysis Coding Index55.86Sep 9, 2026aa_coding_index
Show 42 more resultsHide 42 results
- 13.20Aug 26, 2026
- 33.60Aug 24, 2026
- 81.00Oct 6, 2026
- Apex (Pass@1)22.70Oct 6, 2026Apex
- 70.40Oct 6, 2026
- 72.90Oct 6, 2026
- Claw Eval (pass@3)62.70Oct 6, 2026Claw-Eval
- 77.00Oct 6, 2026
- 62.30Oct 6, 2026
- 16.50Aug 26, 2026
- 14.20Aug 14, 2026
- 69.80Oct 6, 2026
- 38.22Oct 6, 2026
- 94.60Aug 31, 2026
- 83.00Oct 6, 2026
- 27.60Aug 26, 2026
- 76.20Oct 6, 2026
- 90.30Oct 6, 2026
- 88.70Aug 26, 2026
- 58.70Oct 6, 2026
- 71.00Oct 6, 2026
- 87.40Oct 6, 2026
- 88.50Oct 6, 2026
- 85.40Oct 6, 2026
- 94.50Oct 6, 2026
- 89.00Oct 6, 2026
- 41.10Oct 6, 2026
- 91.40Oct 6, 2026
- OSWorld 2.0 (partial)21.50Aug 26, 2026OSWorld 2.0
- 61.80Oct 6, 2026
- 86.90Oct 6, 2026
- ScreenSpot-Pro (No tools)79.00Oct 6, 2026ScreenSpot Pro
- 81.70Oct 6, 2026
- 71.40Oct 6, 2026
- 30.00Aug 24, 2026
- 70.30Oct 6, 2026
- Toolathlon Verified50.60Aug 26, 2026Toolathlon Verified (Pass@1)
- 78.20Oct 6, 2026
- 85.40Oct 6, 2026
- 45.60Oct 6, 2026
- 55.30Aug 24, 2026
- 61.10Oct 6, 2026
Qwen3.7 Plus Preview: common questions
Who makes Qwen3.7 Plus Preview?
Qwen3.7 Plus Preview is made by Alibaba.
When was Qwen3.7 Plus Preview released?
Qwen3.7 Plus Preview was released on Jun 1, 2026, according to Artificial Analysis.
What is Qwen3.7 Plus Preview good at?
Qwen3.7 Plus Preview is capable in multimodal tasks, coding, instruction following, long context, reasoning, and factuality; and behind the leaders in agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does Qwen3.7 Plus Preview cost?
Qwen3.7 Plus Preview costs $0.40 per million input tokens and $1.60 per million output tokens, according to Alibaba's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 50% of the 328 priced models we track.
How many benchmarks has Qwen3.7 Plus Preview been tested on?
We track 85 results for Qwen3.7 Plus Preview on 73 benchmarks from 7 sources, 4 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does Qwen3.7 Plus Preview support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3.7 Plus Preview.
About this record
Where Qwen3.7 Plus Preview's numbers come from, and every name it appears under.
- Tracked since
- May 19, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 85 scores · 4 independently verified · 20 aggregator-attributed · 61 vendor-reported. How these tiers are assigned
From 7 sources on 6 sites. api.llm-stats.com supplies 49 of them; the 4 independently verified results come from 3 sites. Bars are coloured by trust tier.
- api.llm-stats.com49
- artificialanalysis.ai20
- huggingface.co12
- lmarena.ai2
- datasets-server.huggingface.co1
- epoch.ai1