Qwen3 VL 4B (Reasoning)
Qwen3 VL 4B (Reasoning) is behind the leaders in agentic tasks and multimodal tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
118results on102benchmarks
- 16 aggregator
- 37 vendor-reported
- 65 cross-referenced
From 6 sources · latest Oct 8, 2026 · How verification works
Qwen3 VL 4B (Reasoning) benchmark results
118 results on 102 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
45.1% behind the leader2 of 7 ranked benchmarks measured
- OSWorld-Verified31.40Oct 7, 2026OSWorld
- 13.78Jun 15, 2026
54.8% behind the leader5 of 6 ranked benchmarks measured
- MathVista79.50Oct 7, 2026MathVista-Mini
- 80.80Oct 7, 2026
- 59.70Jul 26, 2026
- MMMU-Pro52.02Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)50.30Oct 7, 2026CharXiv-R
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond49.39Oct 8, 2026gpqa
- Humanity's Last Exam4.59Oct 8, 2026aa_hle
Show 2 more reasoning resultsHide 2 reasoning results
- GPQA Diamond64.10Oct 7, 2026GPQA
- 60.10Jun 5, 2026
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard1.52Oct 8, 2026aa_terminalbench_hard
- SciCode17.13Sep 4, 2026aa_scicode
- 51.30Oct 7, 2026
0 of 3 ranked benchmarks measured
- 21.33Oct 8, 2026
0 of 3 ranked benchmarks measured
- IFBench36.60Oct 8, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy12.22Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.18Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence7.05Oct 8, 2026aa_intelligence_index
- -68.38Oct 8, 2026
- τ²-Bench Telecom (AA run)15.50Oct 8, 2026aa_tau2
- 6.72Jun 18, 2026
- 14.35Jun 18, 2026
- 64.60Oct 7, 2026
Show 86 more resultsHide 86 results
- 43.90Jun 15, 2026
- 84.90Oct 7, 2026
- 81.54Jul 26, 2026
- AIME 202472.90Jun 5, 2026AIME24
- 74.50Oct 7, 2026
- AIME 202569.70Jun 5, 2026AIME25
- 46.70Jun 15, 2026
- 36.80Oct 7, 2026
- 67.30Oct 7, 2026
- 53.10Jul 26, 2026
- 73.80Oct 7, 2026
- CCOCR-multilan74.20Sep 2, 2026
- CCOCR-overall76.50Sep 2, 2026
- 59.00Oct 7, 2026
- 74.90Sep 2, 2026
- 83.30Sep 2, 2026
- 83.96Jul 26, 2026
- 36.20Sep 2, 2026
- 26.79Jul 26, 2026
- 85.37Jul 26, 2026
- 85.70Jun 15, 2026
- 76.50Jun 15, 2026
- 94.90Sep 2, 2026
- 94.20Oct 7, 2026
- 94.69Jul 26, 2026
- 38.80Jun 15, 2026
- 80.70Jun 15, 2026
- 47.30Oct 7, 2026
- 47.30Jun 15, 2026
- 64.10Oct 7, 2026
- HMMT 202553.10Oct 7, 2026HMMT25
- 82.60Aug 31, 2026
- InfoVQA (val)79.50Jul 26, 2026InfoVQA-val
- 83.00Oct 7, 2026
- 51.30Jun 5, 2026
- 57.70Jul 26, 2026
- 53.50Oct 7, 2026
- 39.20Jul 26, 2026
- 60.00Oct 7, 2026
- 31.00Jun 15, 2026
- 61.50Jul 26, 2026
- 7.70Oct 7, 2026
- 80.58Jul 26, 2026
- 83.25Jul 26, 2026
- 1703.50Jul 26, 2026
- 63.20Jul 26, 2026
- 73.60Oct 7, 2026
- 65.00Oct 7, 2026
- 86.00Oct 7, 2026
- MMMU (val) (Pass@1)70.80Oct 7, 2026MMMU (val)
- 25.10Jun 15, 2026
- 87.21Jul 26, 2026
- 69.30Oct 7, 2026
- 66.70Jul 26, 2026
- 79.80Jul 26, 2026
- OCRBench Score873.00Sep 2, 2026
- OCRBench v2 (Chinese)55.80Oct 7, 2026OCRBench-V2 (zh)
- OCRBench v2_en61.80Oct 7, 2026OCRBench-V2 (en)
- OCRBenchv2 (en/zh) Chinese59.13Sep 2, 2026
- OCRBenchv2 (en/zh) English60.68Sep 2, 2026
- 64.70Sep 2, 2026
- 73.20Oct 7, 2026
- 7.48Jul 26, 2026
- 45.30Jun 15, 2026
- 45.80Jun 15, 2026
- 36.40Jun 15, 2026
- 63.20Jun 15, 2026
- 72.80Jul 26, 2026
- 69.30Jul 26, 2026
- 56.70Jun 15, 2026
- 92.90Oct 7, 2026
- ScreenSpot-Pro (No tools)49.20Oct 7, 2026ScreenSpot Pro
- 25.50Jun 15, 2026
- 62.20Jun 15, 2026
- 50.90Jun 15, 2026
- 61.00Jun 15, 2026
- 58.00Jun 15, 2026
- 46.80Oct 7, 2026
- 81.80Sep 2, 2026
- 80.55Jul 26, 2026
- 28.40Jul 26, 2026
- 69.40Oct 7, 2026
- 41.60Jun 15, 2026
- 53.30Jul 26, 2026
- 55.20Jun 15, 2026
- 59.00Jun 15, 2026
Qwen3 VL 4B (Reasoning): common questions
Who makes Qwen3 VL 4B (Reasoning)?
Qwen3 VL 4B (Reasoning) is made by Alibaba.
When was Qwen3 VL 4B (Reasoning) released?
Qwen3 VL 4B (Reasoning) was released on Oct 14, 2025, according to Artificial Analysis.
What is Qwen3 VL 4B (Reasoning) good at?
Qwen3 VL 4B (Reasoning) is behind the leaders in agentic tasks and multimodal tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, instruction following, or factuality.
How many benchmarks has Qwen3 VL 4B (Reasoning) been tested on?
We track 118 results for Qwen3 VL 4B (Reasoning) on 102 benchmarks from 6 sources. The latest was recorded on Oct 8, 2026.
About this record
Where Qwen3 VL 4B (Reasoning)'s numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 118 scores · 0 independently verified · 16 aggregator-attributed · 65 vendor cross-reference · 37 vendor-reported. How these tiers are assigned
From 6 sources on 3 sites. Hugging Face supplies 65 of them. Bars are coloured by trust tier.
- huggingface.co65
- api.llm-stats.com37
- artificialanalysis.ai16