Qwen3 VL 4B (Reasoning) vs Qwen3 VL 8B Instruct
Wins 5 of 7 areas
Coding · Reasoning · Images and charts · Long documents · Following instructions
Wins 1 of 7 areas
Facts
Qwen3 VL 4B (Reasoning) is the stronger all-rounder.Qwen3 VL 8B Instruct is better at facts.
Scores updated · 51 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking20Qwen3 VL 4B (Reasoning)2 of 3 tests · 1 tie
- Images and chartsUnderstanding pictures, charts and video43Qwen3 VL 4B (Reasoning)4 of 7 tests
- CodingWriting and fixing software10Qwen3 VL 4B (Reasoning)1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts10Qwen3 VL 4B (Reasoning)1 of 1 test
- Following instructionsDoing exactly what it is asked10Qwen3 VL 4B (Reasoning)1 of 1 test
- FactsGetting facts right instead of making them up01Qwen3 VL 8B Instruct1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own11Even1 each
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Qwen3 VL 4B (Reasoning) pulls ahead
- Recent programming contest problemsLiveCodeBench v6+12points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+6.7points ahead
- Real work tasks from 44 professionsGDPVal+4.7points ahead
Where Qwen3 VL 8B Instruct pulls ahead
- Questions about short and long videosVideo-MME+11.7points ahead
- Reads text in imagesOCRBench+8.8points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+8.3points ahead
Every test, side by side
All 51 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 VL 4B (Reasoning)
- LiveCodeBench v6Qwen3 VL 4B (Reasoning) by 1251.339.3+12
- Terminal-Bench Hardtie1.52.3tie
- SciCodetie17.117.4tie
AgentsEven
- GDPValQwen3 VL 4B (Reasoning) by 4.713.89.1+4.7
- OSWorld-VerifiedQwen3 VL 8B Instruct by 2.531.433.9+2.5
ReasoningQwen3 VL 4B (Reasoning)
- GPQA DiamondQwen3 VL 4B (Reasoning) by 6.749.442.7+6.7
- Humanity's Last ExamQwen3 VL 4B (Reasoning) by 1.94.62.7+1.9
- CritPttie00tie
FactsQwen3 VL 8B Instruct
- AA-Omniscience · AccuracyQwen3 VL 8B Instruct by 8.312.220.5+8.3
- AA-Omniscience · Non-hallucinationtie8.29tie
Images and chartsQwen3 VL 4B (Reasoning)
- Video-MMEQwen3 VL 8B Instruct by 11.759.771.4+11.7
- OCRBenchQwen3 VL 8B Instruct by 8.880.889.6+8.8
- BLINKQwen3 VL 8B Instruct by 5.763.469.1+5.7
- MMMU-ProQwen3 VL 4B (Reasoning) by 4.75247.3+4.7
- CharXiv (RQ)Qwen3 VL 4B (Reasoning) by 3.950.346.4+3.9
- MathVistaQwen3 VL 4B (Reasoning) by 2.379.577.2+2.3
- MMStarQwen3 VL 4B (Reasoning) by 2.373.270.9+2.3
Long documentsQwen3 VL 4B (Reasoning)
- AA-LCRQwen3 VL 4B (Reasoning) by 4.621.316.7+4.6
Following instructionsQwen3 VL 4B (Reasoning)
- IFBenchQwen3 VL 4B (Reasoning) by 4.336.632.3+4.3
Other results32 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIME 2025Qwen3 VL 4B (Reasoning) by 28.674.545.9+28.6
- HMMT 2025Qwen3 VL 4B (Reasoning) by 20.653.132.5+20.6
- AA-OmniscienceQwen3 VL 8B Instruct by 16.5-68.4-51.9+16.5
- τ²-Bench Telecom (AA run)Qwen3 VL 8B Instruct by 13.715.529.2+13.7
- CC-OCRQwen3 VL 8B Instruct by 6.173.879.9+6.1
- MATH-VisionQwen3 VL 4B (Reasoning) by 6.16053.9+6.1
- OCRBench v2 (Chinese)Qwen3 VL 8B Instruct by 5.455.861.2+5.4
- ScreenSpot-Pro (No tools)Qwen3 VL 8B Instruct by 5.449.254.6+5.4
- LVBenchQwen3 VL 8B Instruct by 4.553.558+4.5
- VideoMMMUQwen3 VL 4B (Reasoning) by 4.169.465.3+4.1
- AA Agentic IndexQwen3 VL 8B Instruct by 414.418.4+4
- OCRBench v2_enQwen3 VL 8B Instruct by 3.661.865.4+3.6
- CharadesSTAQwen3 VL 4B (Reasoning) by 35956+3
- HallusionBenchQwen3 VL 4B (Reasoning) by 364.161.1+3
- IncludeQwen3 VL 8B Instruct by 2.464.667+2.4
- SuperGPQAQwen3 VL 4B (Reasoning) by 2.346.844.5+2.3
- MMLU-ProQwen3 VL 4B (Reasoning) by 273.671.6+2
- DocVQA_testQwen3 VL 8B Instruct by 1.994.296.1+1.9
- RealWorldQAQwen3 VL 4B (Reasoning) by 1.773.271.5+1.7
- ERQAQwen3 VL 4B (Reasoning) by 1.547.345.8+1.5
- ScreenSpotQwen3 VL 8B Instruct by 1.592.994.4+1.5
- MMMU (val) (Pass@1)Qwen3 VL 4B (Reasoning) by 1.270.869.6+1.2
- IFEvalQwen3 VL 8B Instruct by 1.182.683.7+1.1
- MMLU-ReduxQwen3 VL 4B (Reasoning) by 1.18684.9+1.1
- BFCL v3Qwen3 VL 4B (Reasoning) by 167.366.3+1
- AI2Dtie84.985.7tie
- Artificial Analysis Coding Indextie6.77.3tie
- MV-Benchtie69.368.7tie
- MMLU-ProXtie6565.4tie
- AA Intelligencetie77.3tie
- InfoVQA_testtie8383.1tie
- MM MTBenchtie7.77.7tie
Questions people ask
Which is better, Qwen3 VL 4B (Reasoning) or Qwen3 VL 8B Instruct?
Qwen3 VL 4B (Reasoning) wins five of the seven areas where both have results: coding, reasoning, images and charts, long documents and following instructions. Qwen3 VL 8B Instruct wins facts. They are level on agents.
Which is better for coding?
Qwen3 VL 4B (Reasoning). It wins 1 of the 3 coding tests both models report; Qwen3 VL 8B Instruct wins none, and 2 are ties.
How do you compare the two?
We use the 51 benchmark tests both models have published scores on. The verdict counts the 19 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 32 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.