Qwen3.5 2B
Qwen3.5 2B is behind the leaders in multimodal tasks and math. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Multilingual, Instruction Following or Factuality.
Price
No current price is tracked for this model. See the rate card
Evidence
126results on76benchmarks
- 33 aggregator
- 66 vendor-reported
- 27 cross-referenced
From 8 sources · latest Oct 8, 2026 · How verification works
Research
12 papers reference Qwen3.5 2BQwen3.5 2B benchmark results
126 results on 76 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
46.3% behind the leader3 of 6 ranked benchmarks measured
Show 10 more multimodal resultsHide 10 multimodal results
- 48.60Sep 10, 2026
- 59.30Sep 25, 2026
- 68.80Sep 25, 2026
- MMMU-Pro42.66Oct 8, 2026aa_mmmu_pro
- 24.90Sep 10, 2026
- 43.50Sep 25, 2026
- 55.10Sep 10, 2026
- 67.90Sep 25, 2026
- 0.86Aug 12, 2026
- 48.00Sep 25, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond45.56Oct 8, 2026gpqa
- Humanity's Last Exam2.59Oct 8, 2026aa_hle
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond43.84Oct 8, 2026gpqa
- GPQA Diamond51.60Oct 7, 2026GPQA
- 41.97Aug 15, 2026
- Humanity's Last Exam4.96Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench 2.13.00Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard3.79Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.10.00Oct 8, 2026terminalbenchV21
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- τ-Bench V3 · Banking2.47Aug 10, 2026tauBanking
- τ-Bench V3 · Banking2.27Aug 10, 2026tauBanking
0 of 3 ranked benchmarks measured
- 21.00Oct 8, 2026
- 14.00Oct 8, 2026
- 38.70Oct 7, 2026
Show 1 more long context resultHide 1 long context result
- 25.60Oct 7, 2026
0 of 3 ranked benchmarks measured
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy8.07Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination25.76Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy7.38Oct 8, 2026omniscienceAccuracy
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination2.27Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -60.18Oct 8, 2026
- τ²-Bench Telecom (AA run)69.01Oct 8, 2026aa_tau2
- AA Intelligence6.94Oct 8, 2026aa_intelligence_index
- AA Intelligence6.19Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)81.58Oct 8, 2026aa_tau2
- -83.13Oct 8, 2026
Show 79 more resultsHide 79 results
- 0.82Aug 10, 2026
- 0.76Aug 10, 2026
- AIME 202417.00Aug 15, 2026AIME 24
- AIME 202519.00Aug 15, 2026AIME 25
- Artificial Analysis Coding Index2.92Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index2.39Sep 9, 2026aa_coding_index
- 68.91Aug 15, 2026
- 43.60Oct 7, 2026
- 33.90Sep 10, 2026
- 33.90Sep 25, 2026
- 73.20Oct 7, 2026
- 78.30Sep 25, 2026
- ChartQA Test78.40Sep 10, 2026ChartQA (test)
- 0.78Aug 12, 2026
- 0.78Aug 12, 2026
- DocVQA-val92.60Sep 10, 2026DocVQA (val)
- 0.93Aug 12, 2026
- 64.45Aug 15, 2026
- 0.54Aug 12, 2026
- 0.54Aug 12, 2026
- 49.30Sep 10, 2026
- 65.50Aug 12, 2026
- 0.66Aug 12, 2026
- 66.00Aug 15, 2026
- 73.60Sep 10, 2026
- 78.60Aug 31, 2026
- 75.06Aug 15, 2026
- 55.40Oct 7, 2026
- 73.50Sep 10, 2026
- MATH-500 (EM)81.52Aug 15, 2026Math-500
- 57.62Aug 15, 2026
- 86.52Aug 15, 2026
- 55.40Sep 10, 2026
- 52.10Sep 25, 2026
- MMBench73.10Sep 10, 2026MMBench (dev EN v1.1)
- 0.76Aug 12, 2026
- 0.76Aug 12, 2026
- MMBench_DEV0.67Sep 17, 2026
- 76.20Sep 10, 2026
- 76.40Sep 25, 2026
- 0.54Aug 12, 2026
- 0.54Aug 12, 2026
- 66.50Oct 7, 2026
- 0.30Aug 12, 2026
- 0.30Aug 12, 2026
- 52.30Oct 7, 2026
- 55.24Aug 15, 2026
- 79.60Oct 7, 2026
- 75.90Sep 19, 2026
- 74.50Sep 3, 2026
- 73.60Sep 25, 2026
- 0.74Aug 12, 2026
- 63.10Oct 7, 2026
- MMMU (val) (Pass@1)44.10Sep 10, 2026MMMU (val)
- 0.47Aug 12, 2026
- 0.47Aug 12, 2026
- 0.67Aug 12, 2026
- 0.67Aug 12, 2026
- 69.90Sep 10, 2026
- 75.90Sep 10, 2026
- 84.40Sep 10, 2026
- OCRBench v2_en47.70Sep 10, 2026OCRBench v2 (En)
- 48.10Aug 12, 2026
- 0.48Aug 12, 2026
- 88.60Sep 10, 2026
- 88.70Sep 25, 2026
- 65.10Sep 10, 2026
- 71.40Sep 25, 2026
- 78.50Sep 25, 2026
- 66.50Sep 25, 2026
- 63.80Sep 10, 2026
- 69.70Sep 10, 2026
- 65.90Sep 10, 2026
- 75.80Sep 10, 2026
- 35.20Sep 10, 2026
- 37.50Oct 7, 2026
- 48.80Oct 7, 2026
- TextVQA-val77.30Sep 10, 2026TextVQA (val)
- 89.70Aug 9, 2026
Qwen3.5 2B: common questions
Who makes Qwen3.5 2B?
Qwen3.5 2B is made by Alibaba.
When was Qwen3.5 2B released?
Qwen3.5 2B was released on Mar 2, 2026, according to Artificial Analysis.
What is Qwen3.5 2B good at?
Qwen3.5 2B is behind the leaders in multimodal tasks and math. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, multilingual tasks, instruction following, or factuality.
How many benchmarks has Qwen3.5 2B been tested on?
We track 126 results for Qwen3.5 2B on 76 benchmarks from 8 sources. The latest was recorded on Oct 8, 2026.
About this record
Where Qwen3.5 2B's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 126 scores · 0 independently verified · 33 aggregator-attributed · 27 vendor cross-reference · 66 vendor-reported. How these tiers are assigned
From 8 sources on 3 sites. Hugging Face supplies 78 of them. Bars are coloured by trust tier.
- huggingface.co78
- artificialanalysis.ai33
- api.llm-stats.com15