Qwen3 VL Thinking (8B)
Qwen3 VL Thinking (8B) is behind the leaders in multimodal tasks and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
$0.18input$2.10outputper million tokens
From Alibaba's own price page · 3 providers tracked · All prices
Evidence
74results on63benchmarks
- 16 aggregator
- 39 vendor-reported
- 19 cross-referenced
From 4 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Qwen3 VL Thinking (8B) benchmark results
74 results on 63 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
45.8% behind the leader6 of 6 ranked benchmarks measured
- MathVista81.40Oct 7, 2026MathVista-Mini
- 81.90Oct 7, 2026
- 73.53Jun 5, 2026
- 71.80Oct 7, 2026
- MMMU-Pro56.65Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)53.00Oct 7, 2026CharXiv-R
46.0% behind the leader2 of 7 ranked benchmarks measured
- OSWorld-Verified33.90Oct 7, 2026OSWorld
- 8.41Jun 15, 2026
0 of 6 ranked benchmarks measured
- 0.29Oct 8, 2026
- GPQA Diamond57.88Oct 8, 2026gpqa
- Humanity's Last Exam3.85Oct 8, 2026aa_hle
Show 2 more reasoning resultsHide 2 reasoning results
- GPQA Diamond69.90Oct 7, 2026GPQA
- 67.10Jun 5, 2026
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard3.79Oct 8, 2026aa_terminalbench_hard
- SciCode21.88Sep 4, 2026aa_scicode
- 58.60Oct 7, 2026
0 of 3 ranked benchmarks measured
- 33.33Oct 8, 2026
0 of 3 ranked benchmarks measured
- IFBench39.86Oct 8, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy20.43Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.59Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -52.30Oct 8, 2026
- τ²-Bench Telecom (AA run)22.51Oct 8, 2026aa_tau2
- AA Intelligence8.17Oct 8, 2026aa_intelligence_index
- 15.57Jun 18, 2026
- 9.82Jun 18, 2026
- 69.50Oct 7, 2026
Show 41 more resultsHide 41 results
- 84.90Oct 7, 2026
- 83.32Jun 5, 2026
- AIME 202486.00Jun 5, 2026AIME24
- 80.30Oct 7, 2026
- AIME 202579.80Jun 5, 2026AIME25
- 45.88Jun 5, 2026
- 51.10Oct 7, 2026
- 63.00Oct 7, 2026
- 76.30Oct 7, 2026
- 59.90Oct 7, 2026
- 95.30Oct 7, 2026
- 46.80Oct 7, 2026
- 65.40Oct 7, 2026
- HMMT 202560.60Oct 7, 2026HMMT25
- 26.94Jun 5, 2026
- 83.20Aug 31, 2026
- 86.00Oct 7, 2026
- 58.00Jun 5, 2026
- 55.80Oct 7, 2026
- 62.70Oct 7, 2026
- 59.60Jun 5, 2026
- 8.00Oct 7, 2026
- 90.55Jun 5, 2026
- 77.30Oct 7, 2026
- 70.70Oct 7, 2026
- 88.80Oct 7, 2026
- MMMU (val) (Pass@1)74.10Oct 7, 2026MMMU (val)
- 69.00Oct 7, 2026
- OCRBench v2 (Chinese)59.20Oct 7, 2026OCRBench-V2 (zh)
- OCRBench v2_en63.90Oct 7, 2026OCRBench-V2 (en)
- 56.70Jun 5, 2026
- 57.67Jun 5, 2026
- 73.50Oct 7, 2026
- 57.17Jun 5, 2026
- 93.60Oct 7, 2026
- ScreenSpot-Pro (No tools)46.60Oct 7, 2026ScreenSpot Pro
- ScreenSpot-Pro (No tools)46.60Jun 5, 2026ScreenSpot-Pro
- 93.60Jun 5, 2026
- 49.60Oct 7, 2026
- 51.20Oct 7, 2026
- 72.80Oct 7, 2026
Qwen3 VL Thinking (8B): common questions
Who makes Qwen3 VL Thinking (8B)?
Qwen3 VL Thinking (8B) is made by Alibaba.
When was Qwen3 VL Thinking (8B) released?
Qwen3 VL Thinking (8B) was released on Oct 14, 2025, according to Artificial Analysis.
What is Qwen3 VL Thinking (8B) good at?
Qwen3 VL Thinking (8B) is behind the leaders in multimodal tasks and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, instruction following, or factuality.
How much does Qwen3 VL Thinking (8B) cost?
Qwen3 VL Thinking (8B) costs $0.18 per million input tokens and $2.10 per million output tokens, according to Alibaba's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 52% of the 331 priced models we track.
How many benchmarks has Qwen3 VL Thinking (8B) been tested on?
We track 74 results for Qwen3 VL Thinking (8B) on 63 benchmarks from 4 sources. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 VL Thinking (8B) support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3 VL Thinking (8B).
About this record
Where Qwen3 VL Thinking (8B)'s numbers come from, and every name it appears under.
- Tracked since
- May 16, 2026
- Newest source mention
- Sep 29, 2026
Where the results come from
Verification: 74 scores · 0 independently verified · 16 aggregator-attributed · 19 vendor cross-reference · 39 vendor-reported. How these tiers are assigned
From 4 sources on 3 sites. api.llm-stats.com supplies 39 of them. Bars are coloured by trust tier.
- api.llm-stats.com39
- huggingface.co19
- artificialanalysis.ai16