Qwen3 VL 4B (Reasoning) is the stronger all-rounder.
Scores updated · 29 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Images and charts rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3 VL 4B (Reasoning) pulls ahead
- Reasoning about charts from research papersCharXiv (RQ)+9points ahead
Where Qwen3 VL 2B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 29 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsQwen3 VL 4B (Reasoning)
- CharXiv (RQ)Qwen3 VL 4B (Reasoning) by 941.350.3+9
Other results28 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ShareRobot-Traj.Qwen3 VL 4B (Reasoning) by 20.641.662.2+20.6
- RoboSpatial-HomeQwen3 VL 4B (Reasoning) by 17.945.363.2+17.9
- RefSpatial-BenchQwen3 VL 4B (Reasoning) by 16.428.945.3+16.4
- Where2PlaceQwen3 VL 4B (Reasoning) by 144559+14
- ChartQAProQwen3 VL 4B (Reasoning) by 11.724.536.2+11.7
- SATQwen3 VL 4B (Reasoning) by 11.445.356.7+11.4
- RoboBench-MCQQwen3 VL 4B (Reasoning) by 8.936.945.8+8.9
- SIBench-miniQwen3 VL 4B (Reasoning) by 8.94250.9+8.9
- SITE-Bench-ImageQwen3 VL 4B (Reasoning) by 8.752.361+8.7
- VSI-BenchQwen3 VL 4B (Reasoning) by 7.24855.2+7.2
- DA-2KQwen3 VL 4B (Reasoning) by 769.576.5+7
- SITE-Bench-VideoQwen3 VL 4B (Reasoning) by 5.852.258+5.8
- CV-BenchQwen3 VL 4B (Reasoning) by 5.78085.7+5.7
- ShareRobot-Aff.Qwen3 VL 4B (Reasoning) by 5.719.825.5+5.7
- ERQAQwen3 VL 4B (Reasoning) by 5.541.847.3+5.5
- OCRVQAQwen3 VL 4B (Reasoning) by 5.459.364.7+5.4
- ChartQAQwen3 VL 4B (Reasoning) by 578.383.3+5
- EmbSpatial-BenchQwen3 VL 4B (Reasoning) by 4.875.980.7+4.8
- All-Angles-BenchQwen3 VL 4B (Reasoning) by 4.442.346.7+4.4
- ViewSpatialQwen3 VL 4B (Reasoning) by 4.437.241.6+4.4
- 3DSRBenchQwen3 VL 4B (Reasoning) by 439.943.9+4
- Ego-Plan2Qwen3 VL 4B (Reasoning) by 3.335.538.8+3.3
- MindCubeQwen3 VL 4B (Reasoning) by 2.628.431+2.6
- DocVQAQwen3 VL 4B (Reasoning) by 2.292.794.9+2.2
- TextVQAQwen3 VL 4B (Reasoning) by 1.979.981.8+1.9
- ChartBenchQwen3 VL 4B (Reasoning) by 1.773.274.9+1.7
- MMSI-BenchQwen3 VL 4B (Reasoning) by 1.523.625.1+1.5
- RoboBench-Planningtie36.236.4tie
Questions people ask
Which is better, Qwen3 VL 2B or Qwen3 VL 4B (Reasoning)?
Qwen3 VL 4B (Reasoning) wins the one area where both have results: images and charts. Qwen3 VL 2B wins none.
How do you compare the two?
We use the 29 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 28 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.