Qwen3.5 4B
Qwen3.5 4B is behind the leaders in multimodal tasks, long context, instruction following, and coding. Too few results yet to rate reasoning, agentic tasks, safety, math, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Math, Multilingual or Factuality.
Price
$0.03input$0.15outputper million tokens
From Artificial Analysis · All prices
Evidence
141results on85benchmarks
- 5 independently verified
- 35 aggregator
- 70 vendor-reported
- 31 cross-referenced
From 12 sources · latest Oct 8, 2026 · How verification works
Research
62 papers reference Qwen3.5 4BQwen3.5 4B benchmark results
141 results on 85 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
33.3% behind the leader4 of 6 ranked benchmarks measured
- 86.90Sep 3, 2026
- 73.40Sep 3, 2026
- MMMU-Pro65.38Oct 8, 2026aa_mmmu_pro
- MathVista63.60Sep 10, 2026MathVista (mini)
Show 10 more multimodal resultsHide 10 multimodal results
- 58.70Sep 17, 2026
- 65.00Sep 25, 2026
- 69.70Sep 25, 2026
- MMMU-Pro62.08Oct 8, 2026aa_mmmu_pro
- 36.00Sep 10, 2026
- 60.90Sep 25, 2026
- 59.30Sep 10, 2026
- 75.30Sep 3, 2026
- 73.30Sep 25, 2026
- 58.80Sep 25, 2026
34.1% behind the leader2 of 3 ranked benchmarks measured
- 50.00Oct 7, 2026
- 63.00Oct 8, 2026
40.9% behind the leader2 of 3 ranked benchmarks measured
- IFBench51.97Oct 8, 2026aa_ifbench
- 49.00Oct 7, 2026
45.0% behind the leader4 of 10 ranked benchmarks measured
- 55.80Oct 7, 2026
- Terminal-Bench 2.125.84Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard18.18Oct 8, 2026aa_terminalbench_hard
- SciCode16.10Sep 4, 2026aa_scicode
Show 3 more coding resultsHide 3 coding results
- SciCode18.29Sep 4, 2026aa_scicode
- Terminal-Bench 2.121.35Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard11.36Oct 8, 2026aa_terminalbench_hard
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond77.07Oct 8, 2026gpqa
- Humanity's Last Exam9.92Oct 8, 2026aa_hle
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond71.21Oct 8, 2026gpqa
- GPQA Diamond76.20Oct 7, 2026GPQA
- Humanity's Last Exam7.97Oct 8, 2026aa_hle
0 of 7 ranked benchmarks measured
- τ-Bench V3 · Banking6.80Oct 8, 2026tauBanking
- τ-Bench V3 · Banking4.33Oct 8, 2026tauBanking
- 0.49Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 8.32Jun 15, 2026
0 of 5 ranked benchmarks measured
- 72.87Sep 2, 2026
- 89.67Sep 2, 2026
- 84.09May 10, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy15.12Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination13.45Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy11.22Oct 8, 2026omniscienceAccuracy
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination3.19Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence20.00Jun 17, 2026Artificial Analysis Intelligence Index
- AA Intelligence13.12Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)92.11Oct 8, 2026aa_tau2
- -58.35Oct 8, 2026
- τ²-Bench Telecom (AA run)87.72Oct 8, 2026aa_tau2
- AA Intelligence10.81Oct 8, 2026aa_intelligence_index
Show 83 more resultsHide 83 results
- 32.46Jun 18, 2026
- 36.26Jun 18, 2026
- -74.73Oct 8, 2026
- AIME25 no tools49.33Aug 10, 2026AIME25
- AIME25 no tools54.28Aug 9, 2026AIME25
- AIME25 no tools49.33Aug 9, 2026AIME25
- Artificial Analysis Coding Index22.60Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index20.33Sep 9, 2026aa_coding_index
- 71.06Aug 9, 2026
- 50.30Oct 7, 2026
- 53.60Sep 17, 2026
- 50.56Aug 10, 2026
- 53.60Sep 25, 2026
- 54.01Aug 9, 2026
- 50.56Aug 9, 2026
- 24.46Aug 10, 2026
- 24.46Aug 9, 2026
- 85.10Oct 7, 2026
- 84.20Sep 25, 2026
- ChartQA Test84.20Sep 17, 2026ChartQA (test)
- 65.10Sep 3, 2026
- Claw-Eval Avg62.28Aug 10, 2026Claw-Eval average (EN)
- Claw-Eval Avg62.28Aug 9, 2026Claw-Eval average (EN)
- 35.90Sep 3, 2026
- 17.60Oct 7, 2026
- DocVQA-val94.80Sep 17, 2026DocVQA (val)
- Ego3D RMSE ↓13.17Sep 3, 2026
- 46.30Sep 3, 2026
- 51.70Sep 17, 2026
- HMMT 202576.80Oct 7, 2026HMMT25
- 86.20Sep 17, 2026
- 89.80Aug 31, 2026
- 87.80Aug 9, 2026
- 71.00Oct 7, 2026
- 80.30Sep 17, 2026
- LiveCodeBenchV6 no tools60.85Aug 10, 2026LiveCodeBenchv6
- LiveCodeBenchV6 no tools60.85Aug 9, 2026LiveCodeBenchv6
- MATH500 (Pass@1)80.76Aug 9, 2026MATH500
- 63.10Sep 10, 2026
- 63.80Sep 25, 2026
- MMBench78.40Sep 10, 2026MMBench (dev EN v1.1)
- 87.10Sep 3, 2026
- 79.50Sep 10, 2026
- 79.50Sep 25, 2026
- 79.10Oct 7, 2026
- 71.50Oct 7, 2026
- 88.80Oct 7, 2026
- 82.00Sep 19, 2026
- 83.40Sep 25, 2026
- 76.10Oct 7, 2026
- MMMU (val) (Pass@1)50.30Sep 10, 2026MMMU (val)
- MMMU-Pro std64.90Sep 3, 2026
- MMMU-Pro vis61.30Sep 3, 2026
- 77.00Sep 10, 2026
- 82.00Sep 10, 2026
- 85.60Sep 17, 2026
- OCRBench v2_en58.70Sep 17, 2026OCRBench v2 (En)
- 40.80Sep 3, 2026
- 47.40Sep 3, 2026
- 71.26Aug 10, 2026
- 71.26Aug 9, 2026
- 86.00Sep 17, 2026
- 86.00Sep 25, 2026
- 67.10Sep 10, 2026
- 76.30Sep 3, 2026
- 76.20Sep 25, 2026
- 86.60Sep 25, 2026
- 78.50Sep 25, 2026
- 76.30Sep 17, 2026
- 81.40Sep 17, 2026
- 77.80Sep 17, 2026
- 76.10Sep 10, 2026
- 40.70Sep 10, 2026
- 47.80Sep 3, 2026
- 52.90Oct 7, 2026
- 79.90Oct 7, 2026
- 87.72Aug 9, 2026
- TextVQA-val81.20Sep 17, 2026TextVQA (val)
- 22.00Oct 7, 2026
- WaymoQA all67.10Sep 3, 2026
- τ²-Bench (Retail)71.93Aug 9, 2026Tau² Retail
- 5.45Aug 10, 2026
- 5.45Aug 9, 2026
Qwen3.5 4B: common questions
Who makes Qwen3.5 4B?
Qwen3.5 4B is made by Alibaba.
When was Qwen3.5 4B released?
Qwen3.5 4B was released on Mar 2, 2026, according to Artificial Analysis.
What is Qwen3.5 4B good at?
Qwen3.5 4B is behind the leaders in multimodal tasks, long context, instruction following, and coding. Too few results yet to rate reasoning, agentic tasks, safety, math, multilingual tasks, or factuality.
How much does Qwen3.5 4B cost?
Qwen3.5 4B costs $0.03 per million input tokens and $0.15 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it is cheaper than 96% of the 331 priced models we track.
How many benchmarks has Qwen3.5 4B been tested on?
We track 141 results for Qwen3.5 4B on 85 benchmarks from 12 sources, 5 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Qwen3.5 4B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Oct 3, 2026
Where the results come from
Verification: 141 scores · 5 independently verified · 35 aggregator-attributed · 31 vendor cross-reference · 70 vendor-reported. How these tiers are assigned
From 12 sources on 4 sites. Hugging Face supplies 82 of them; the 5 independently verified results come from 2 sites. Bars are coloured by trust tier.
- huggingface.co82
- artificialanalysis.ai36
- api.llm-stats.com19
- matharena.ai4