InternVL3_5-2B vs Qwen3.5 4B
Wins 0 of 2 areas
—
Wins 2 of 2 areas
Images and charts · Following instructions
Qwen3.5 4B is the stronger all-rounder.
Scores updated · 28 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 4B pulls ahead
Where InternVL3_5-2B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 28 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsQwen3.5 4B
Following instructionsQwen3.5 4B
- IFBenchQwen3.5 4B by 27.624.452+27.6
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- IFEvalQwen3.5 4B by 53.832.486.2+53.8
- MM-IFEvalQwen3.5 4B by 1647.163.1+16
- OCRBench v2_enQwen3.5 4B by 13.245.558.7+13.2
- InfographicVQA (val)Qwen3.5 4B by 1169.380.3+11
- SimpleVQAQwen3.5 4B by 10.230.540.7+10.2
- DocVQA-valQwen3.5 4B by 6.488.494.8+6.4
- Multilingual MMBenchQwen3.5 4B by 6.170.977+6.1
- MMEQwen3.5 4B by 5.973.679.5+5.9
- MMMBQwen3.5 4B by 5.776.382+5.7
- Multilingual MMMBQwen3.5 4B by 5.776.382+5.7
- RealWorldQAQwen3.5 4B by 5.561.667.1+5.5
- ScreenSpot-v2 MobileInternVL3_5-2B by 4.886.281.4+4.8
- TextVQA-valQwen3.5 4B by 4.676.681.2+4.6
- HallusionBenchQwen3.5 4B by 4.147.651.7+4.1
- ScreenSpot-v2 DesktopInternVL3_5-2B by 3.679.976.3+3.6
- ChartQA TestQwen3.5 4B by 2.581.784.2+2.5
- MMBenchQwen3.5 4B by 2.276.278.4+2.2
- ScreenSpot-v2 WebInternVL3_5-2B by 2.179.977.8+2.1
- POPEInternVL3_5-2B by 28886+2
- MMMU (val) (Pass@1)InternVL3_5-2B by 1.75250.3+1.7
- OCRBench v1Qwen3.5 4B by 1.783.985.6+1.7
- SEED-Bench (image)tie75.476.1tie
Questions people ask
Which is better, InternVL3_5-2B or Qwen3.5 4B?
Qwen3.5 4B wins all two areas where both have results: images and charts and following instructions. InternVL3_5-2B wins none.
How do you compare the two?
We use the 28 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.