Qwen3.5 27B
Qwen3.5 27B is capable in multimodal tasks, long context, and instruction following; and behind the leaders in reasoning, coding, and agentic tasks. Too few results yet to rate safety, math, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multilingual or Factuality.
Price
$0.30input$2.40outputper million tokens
From Alibaba's own price page · 5 providers tracked · All prices
Evidence
121results on91benchmarks
- 11 independently verified
- 33 aggregator
- 77 vendor-reported
From 11 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
24 papers reference Qwen3.5 27BQwen3.5 27B benchmark results
121 results on 91 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
10.1% behind the leader5 of 6 ranked benchmarks measured
- 89.40Oct 7, 2026
- MathVista87.80Oct 7, 2026MathVista-Mini
- 82.30Oct 7, 2026
- MMMU-Pro75.03Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)79.50Oct 7, 2026CharXiv-R
Show 5 more multimodal resultsHide 5 multimodal results
- 1240.11Aug 25, 2026
- 1221Jun 17, 2026
- MMMU-Pro70.00Oct 8, 2026aa_mmmu_pro
- 75.00Oct 7, 2026
- 81.00Oct 7, 2026
12.8% behind the leader2 of 3 ranked benchmarks measured
- 60.60Oct 7, 2026
- 77.67Oct 8, 2026
20.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench75.58Oct 8, 2026aa_ifbench
- 60.80Oct 7, 2026
33.3% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond85.76Oct 8, 2026gpqa
- Humanity's Last Exam23.91Oct 8, 2026aa_hle
- 0.85Oct 8, 2026
Show 6 more reasoning resultsHide 6 reasoning results
- 0.29Oct 8, 2026
- GPQA Diamond84.24Oct 8, 2026gpqa
- GPQA Diamond85.50Oct 7, 2026GPQA
- Humanity's Last Exam13.95Oct 8, 2026aa_hle
- 48.50Oct 7, 2026
- 24.30Jun 15, 2026
35.4% behind the leader7 of 10 ranked benchmarks measured
- 80.70Oct 7, 2026
- 72.40Oct 7, 2026
- 69.30Jun 15, 2026
- SciCode39.47Sep 4, 2026aa_scicode
- 51.20Jun 15, 2026
- 1357.48May 22, 2026
- Terminal-Bench Hard32.58Oct 8, 2026aa_terminalbench_hard
Show 3 more coding resultsHide 3 coding results
- SciCode36.69Sep 4, 2026aa_scicode
- 75.00Jun 15, 2026
- Terminal-Bench Hard31.82Oct 8, 2026aa_terminalbench_hard
39.1% behind the leader4 of 7 ranked benchmarks measured
- 68.40Jun 6, 2026
- 61.00Oct 7, 2026
- 56.20Oct 7, 2026
- 33.01Jun 15, 2026
Show 2 more agentic resultsHide 2 agentic results
- 35.50Oct 8, 2026
- 32.96Jun 15, 2026
0 of 5 ranked benchmarks measured
- 81.06Sep 2, 2026
- 91.67Sep 2, 2026
- 78.72May 10, 2026
Show 4 more math resultsHide 4 math results
- 91.67May 2, 2026
- AIME 202692.60Jun 15, 2026AIME26
- HMMT Feb 202684.30Jun 15, 2026HMMT Feb 26
- 79.90Jun 15, 2026
0 of 4 ranked benchmarks measured
- 12.10May 2, 2026
- AA-Omniscience · Accuracy15.68Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination24.87Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy20.67Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination18.51Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- vectara_avg_summary_length94.40May 2, 2026Average Summary Length (Words)
- vectara_answer_rate99.80May 2, 2026Answer Rate
- vectara_factual_consistency87.90May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)87.13Oct 8, 2026aa_tau2
- AA Intelligence19.37Oct 8, 2026aa_intelligence_index
- -47.67Oct 8, 2026
Show 60 more resultsHide 60 results
- 51.49Jun 18, 2026
- 54.61Jun 18, 2026
- AA Intelligence22.90Oct 8, 2026aa_intelligence_index
- -43.98Oct 8, 2026
- 92.90Oct 7, 2026
- 33.44Jun 18, 2026
- 34.87Jun 18, 2026
- 44.60Oct 7, 2026
- 68.50Oct 7, 2026
- 62.10Oct 7, 2026
- 90.50Oct 7, 2026
- 81.00Oct 7, 2026
- Claw Eval (pass@3)46.20Aug 24, 2026Claw-Eval Pass^3
- 64.30Jun 15, 2026
- 22.60Oct 7, 2026
- 87.70Oct 7, 2026
- 84.50Oct 7, 2026
- 60.50Oct 7, 2026
- 70.00Oct 7, 2026
- 89.80Oct 7, 2026
- HMMT Feb. 202592.00Jun 15, 2026HMMT Feb 25
- HMMT Nov. 202589.80Jun 15, 2026HMMT Nov 25
- 95.00Aug 31, 2026
- 81.60Oct 7, 2026
- 73.60Oct 7, 2026
- 86.00Oct 7, 2026
- 36.30Jun 6, 2026
- 85.90Oct 7, 2026
- 92.60Aug 24, 2026
- 60.20Oct 7, 2026
- 86.10Oct 7, 2026
- 82.20Oct 7, 2026
- 93.20Oct 7, 2026
- 85.90Oct 7, 2026
- 73.30Oct 7, 2026
- 74.60Oct 7, 2026
- 27.30Jun 15, 2026
- 40.10Oct 7, 2026
- 88.90Oct 7, 2026
- 52.20Jun 15, 2026
- 1068.00Jun 15, 2026
- 83.70Oct 7, 2026
- 67.70Oct 7, 2026
- ScreenSpot-Pro (No tools)70.30Oct 7, 2026ScreenSpot Pro
- 47.20Oct 7, 2026
- 56.00Oct 7, 2026
- 27.20Jun 15, 2026
- 65.60Oct 7, 2026
- 79.00Oct 7, 2026
- 41.60Oct 7, 2026
- 31.50Jun 6, 2026
- 87.00Aug 24, 2026
- 82.30Oct 7, 2026
- 41.90Oct 7, 2026
- 41.80Jun 6, 2026
- 61.10Oct 7, 2026
- 66.40Jun 6, 2026
- 10.00Oct 7, 2026
- τ²-Bench Telecom (AA run)93.86Oct 8, 2026aa_tau2
- τ³-Bench68.40Jun 6, 2026TAU3-Bench
Qwen3.5 27B: common questions
Who makes Qwen3.5 27B?
Qwen3.5 27B is made by Alibaba.
When was Qwen3.5 27B released?
Qwen3.5 27B was released on Feb 24, 2026, according to Artificial Analysis.
What is Qwen3.5 27B good at?
Qwen3.5 27B is capable in multimodal tasks, long context, and instruction following; and behind the leaders in reasoning, coding, and agentic tasks. Too few results yet to rate safety, math, multilingual tasks, or factuality.
How much does Qwen3.5 27B cost?
Qwen3.5 27B costs $0.30 per million input tokens and $2.40 per million output tokens, according to Alibaba's own price page. We track its price at 5 providers. At a mix of three input tokens to one output token, it costs more than 54% of the 331 priced models we track.
How many benchmarks has Qwen3.5 27B been tested on?
We track 121 results for Qwen3.5 27B on 91 benchmarks from 11 sources, 11 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3.5 27B support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3.5 27B.
About this record
Where Qwen3.5 27B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 25, 2026
Where the results come from
Verification: 121 scores · 11 independently verified · 33 aggregator-attributed · 77 vendor-reported. How these tiers are assigned
From 11 sources on 7 sites. api.llm-stats.com supplies 54 of them; the 11 independently verified results come from 4 sites. Bars are coloured by trust tier.
- api.llm-stats.com54
- artificialanalysis.ai33
- huggingface.co23
- matharena.ai4
- raw.githubusercontent.com4
- datasets-server.huggingface.co2
- lmarena.ai1