Qwen3 VL 4B Instruct
Qwen3 VL 4B Instruct is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
57results on55benchmarks
- 16 aggregator
- 36 vendor-reported
- 5 cross-referenced
From 4 sources · latest Oct 8, 2026 · How verification works
Qwen3 VL 4B Instruct benchmark results
57 results on 55 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
44.8% behind the leader4 of 6 ranked benchmarks measured
- 88.10Oct 7, 2026
- MathVista73.70Oct 7, 2026MathVista-Mini
- MMMU-Pro43.87Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)39.70Oct 7, 2026CharXiv-R
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond37.07Oct 8, 2026gpqa
- Humanity's Last Exam3.61Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard0.00Oct 8, 2026aa_terminalbench_hard
- SciCode13.66Sep 4, 2026aa_scicode
- 37.90Oct 7, 2026
0 of 7 ranked benchmarks measured
- 0.00Jun 15, 2026
- OSWorld-Verified26.20Oct 7, 2026OSWorld
0 of 3 ranked benchmarks measured
- 14.00Oct 8, 2026
0 of 3 ranked benchmarks measured
- IFBench31.84Oct 8, 2026aa_ifbench
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy10.87Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination2.58Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence5.66Oct 8, 2026aa_intelligence_index
- -75.97Oct 8, 2026
- τ²-Bench Telecom (AA run)23.39Oct 8, 2026aa_tau2
- 4.55Jun 18, 2026
- 7.80Jun 18, 2026
- 61.40Oct 7, 2026
Show 32 more resultsHide 32 results
- 84.10Oct 7, 2026
- 46.60Oct 7, 2026
- 43.80Jun 5, 2026
- 63.30Oct 7, 2026
- 76.20Oct 7, 2026
- 55.50Oct 7, 2026
- 95.30Oct 7, 2026
- 41.30Oct 7, 2026
- 57.60Oct 7, 2026
- HMMT 202530.70Oct 7, 2026HMMT25
- 82.30Aug 31, 2026
- 80.30Oct 7, 2026
- 56.20Oct 7, 2026
- 90.00Jun 5, 2026
- 51.60Oct 7, 2026
- 7.50Oct 7, 2026
- 8.01Jun 5, 2026
- 67.10Oct 7, 2026
- 59.40Oct 7, 2026
- 81.50Oct 7, 2026
- MMMU (val) (Pass@1)67.40Oct 7, 2026MMMU (val)
- 68.90Oct 7, 2026
- OCRBench v2 (Chinese)57.60Oct 7, 2026OCRBench-V2 (zh)
- OCRBench v2_en63.70Oct 7, 2026OCRBench-V2 (en)
- 70.90Oct 7, 2026
- 94.00Oct 7, 2026
- ScreenSpot-Pro (No tools)59.50Oct 7, 2026ScreenSpot Pro
- 48.00Oct 7, 2026
- 40.30Oct 7, 2026
- 56.20Oct 7, 2026
- 92.00Aug 9, 2026
- 56.80Jun 5, 2026
Qwen3 VL 4B Instruct: common questions
Who makes Qwen3 VL 4B Instruct?
Qwen3 VL 4B Instruct is made by Alibaba.
When was Qwen3 VL 4B Instruct released?
Qwen3 VL 4B Instruct was released on Oct 14, 2025, according to Artificial Analysis.
What is Qwen3 VL 4B Instruct good at?
Qwen3 VL 4B Instruct is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How many benchmarks has Qwen3 VL 4B Instruct been tested on?
We track 57 results for Qwen3 VL 4B Instruct on 55 benchmarks from 4 sources. The latest was recorded on Oct 8, 2026.
About this record
Where Qwen3 VL 4B Instruct's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Jun 5, 2026
Where the results come from
Verification: 57 scores · 0 independently verified · 16 aggregator-attributed · 5 vendor cross-reference · 36 vendor-reported. How these tiers are assigned
From 4 sources on 3 sites. api.llm-stats.com supplies 36 of them. Bars are coloured by trust tier.
- api.llm-stats.com36
- artificialanalysis.ai16
- huggingface.co5