Qwen3.8 27B
Qwen3.8 27B is strong in long context; capable in instruction following, agentic tasks, coding, multimodal tasks, and reasoning; and behind the leaders in factuality and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$0.50input$3.00outputper million tokens
From Alibaba's own price page · 7 providers tracked · All prices
Evidence
108results on52benchmarks
- 20 independently verified
- 62 aggregator
- 26 vendor-reported
From 12 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
23 papers reference Qwen3.8 27BQwen3.8 27B benchmark results
108 results on 52 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.4% behind the leader1 of 3 ranked benchmarks measured
- 82.00Oct 8, 2026
14.1% behind the leader2 of 3 ranked benchmarks measured
- 79.50Oct 7, 2026
- LiveBench · Instruction Following72.66Oct 8, 2026livebench_instruction_following@2026-06-25
16.4% behind the leader4 of 7 ranked benchmarks measured
- 84.30Oct 7, 2026
- τ-Bench V3 · Banking48.04Oct 8, 2026tauBanking
- 46.16Oct 8, 2026
- Terminal-Bench 4.05.56Oct 8, 2026
Show 9 more agentic resultsHide 9 agentic results
- 43.17Oct 8, 2026
- 44.17Oct 8, 2026
- 28.71Oct 8, 2026
- Terminal-Bench 4.02.53Oct 8, 2026
- Terminal-Bench 4.05.05Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking47.42Oct 8, 2026tauBanking
- τ-Bench V3 · Banking32.16Oct 8, 2026tauBanking
- τ-Bench V3 · Banking20.00Oct 8, 2026tauBanking
18.5% behind the leader8 of 10 ranked benchmarks measured
- 90.30Oct 7, 2026
- Terminal-Bench 2.179.78Oct 8, 2026terminalbenchV21
- 1592.03Aug 25, 2026
- LiveBench · Coding75.69Oct 8, 2026livebench_coding@2026-06-25
- LiveBench · Agentic Coding61.36Oct 8, 2026livebench_agentic_coding@2026-06-25
- 73.80Aug 26, 2026
- SciCode46.64Oct 8, 2026aa_scicode
- 61.70Oct 7, 2026
Show 7 more coding resultsHide 7 coding results
- SciCode40.05Oct 8, 2026aa_scicode
- SciCode39.00Oct 8, 2026aa_scicode
- SciCode36.23Oct 8, 2026aa_scicode
- Terminal-Bench 2.167.42Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.165.17Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.149.06Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.173.00Aug 14, 2026Terminal Bench 2.1 (Terminus)
19.2% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)90.20Oct 7, 2026CharXiv-R
- MMMU-Pro76.30Oct 8, 2026aa_mmmu_pro
Show 5 more multimodal resultsHide 5 multimodal results
24.2% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond90.51Oct 8, 2026gpqa
- LiveBench · Reasoning80.03Oct 8, 2026livebench_reasoning@2026-06-25
- 60.20Aug 21, 2026
- Humanity's Last Exam33.92Oct 8, 2026aa_hle
- 42.36Oct 2, 2026
- 5.43Oct 8, 2026
Show 12 more reasoning resultsHide 12 reasoning results
- 1.53Oct 2, 2026
- 22.78Oct 2, 2026
- 13.19Oct 2, 2026
- 0.00Oct 8, 2026
- 0.29Oct 8, 2026
- GPQA Diamond84.55Oct 8, 2026gpqa
- GPQA Diamond81.82Oct 8, 2026gpqa
- GPQA Diamond89.20Oct 7, 2026GPQA
- Humanity's Last Exam14.04Oct 8, 2026aa_hle
- Humanity's Last Exam14.09Oct 8, 2026aa_hle
- Humanity's Last Exam12.14Oct 8, 2026aa_hle
- 30.80Oct 7, 2026
31.8% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination69.71Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy15.58Oct 8, 2026omniscienceAccuracy
Show 6 more factuality resultsHide 6 factuality results
- AA-Omniscience · Accuracy18.35Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy17.13Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy8.58Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination47.14Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination33.25Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination81.91Oct 8, 2026omniscienceNonHallucination
40.5% behind the leader1 of 5 ranked benchmarks measured
- LiveBench · Mathematics86.21Oct 8, 2026livebench_math@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis76.59Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language74.35Oct 8, 2026livebench_language@2026-06-25
- 56.90Oct 4, 2026
- 34.00Oct 2, 2026
- 69.17Oct 2, 2026
- 68.67Oct 2, 2026
Show 34 more resultsHide 34 results
- 46.51Sep 9, 2026
- 39.86Sep 9, 2026
- 49.81Sep 4, 2026
- 30.36Sep 4, 2026
- AA Intelligence33.70Oct 8, 2026aa_intelligence_index
- AA Intelligence26.20Oct 8, 2026aa_intelligence_index
- AA Intelligence27.55Oct 8, 2026aa_intelligence_index
- AA Intelligence20.15Oct 8, 2026aa_intelligence_index
- -9.98Oct 8, 2026
- -26.67Oct 8, 2026
- -36.15Oct 8, 2026
- -7.95Oct 8, 2026
- 42.90Oct 7, 2026
- 20.40Aug 26, 2026
- 81.90Oct 7, 2026
- 87.50Oct 2, 2026
- Artificial Analysis Coding Index68.08Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index58.22Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index56.14Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index44.59Sep 9, 2026aa_coding_index
- 85.60Oct 7, 2026
- 42.20Oct 7, 2026
- 65.50Oct 7, 2026
- 33.40Oct 7, 2026
- 72.40Aug 26, 2026
- 94.60Oct 7, 2026
- 42.30Oct 7, 2026
- 91.10Oct 7, 2026
- OSWorld 2.0 (partial)48.00Aug 26, 2026OSWorld 2.0
- 85.90Oct 7, 2026
- 38.60Aug 14, 2026
- 38.60Oct 7, 2026
- Toolathlon Verified67.10Aug 26, 2026Toolathlon Verified (Pass@1)
- 64.80Oct 7, 2026
Qwen3.8 27B: common questions
Who makes Qwen3.8 27B?
Qwen3.8 27B is made by Alibaba.
When was Qwen3.8 27B released?
Qwen3.8 27B was released on Aug 14, 2026, according to Artificial Analysis.
What is Qwen3.8 27B good at?
Qwen3.8 27B is strong in long context; capable in instruction following, agentic tasks, coding, multimodal tasks, and reasoning; and behind the leaders in factuality and math. Too few results yet to rate safety or multilingual tasks.
How much does Qwen3.8 27B cost?
Qwen3.8 27B costs $0.50 per million input tokens and $3.00 per million output tokens, according to Alibaba's own price page. We track its price at 7 providers. At a mix of three input tokens to one output token, it costs more than 61% of the 331 priced models we track.
How many benchmarks has Qwen3.8 27B been tested on?
We track 108 results for Qwen3.8 27B on 52 benchmarks from 12 sources, 20 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3.8 27B support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3.8 27B.
About this record
Where Qwen3.8 27B's numbers come from, and every name it appears under.
- Tracked since
- Aug 3, 2026
- Newest source mention
- Oct 2, 2026
Where the results come from
Verification: 108 scores · 20 independently verified · 62 aggregator-attributed · 26 vendor-reported. How these tiers are assigned
From 12 sources on 8 sites. Artificial Analysis supplies 62 of them; the 20 independently verified results come from 5 sites. Bars are coloured by trust tier.
- artificialanalysis.ai62
- api.llm-stats.com20
- arcprize.org8
- livebench.ai7
- huggingface.co6
- 99franklin.github.io2
- datasets-server.huggingface.co2
- simple-bench.com1
Also known as
How our sources name Qwen3.8 27B at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | qwen3.8 27b (low) | qwen3-8-27b-low |
| medium | qwen3.8 27b (medium) | qwen3-8-27b-medium |
| xhigh | qwen3.8 27b (xhigh) | — |