Ministral 3 3B vs Qwen3 VL 4B Instruct
Wins 2 of 7 areas
Coding · Long documents
Wins 2 of 7 areas
Images and charts · Following instructions
The two are evenly matched.Ministral 3 3B is better at coding and long documents; Qwen3 VL 4B Instruct at images and charts and following instructions.
Scores updated · 22 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Images and chartsUnderstanding pictures, charts and video02Qwen3 VL 4B Instruct2 of 2 tests
- Following instructionsDoing exactly what it is asked01Qwen3 VL 4B Instruct1 of 1 test
- CodingWriting and fixing software10Ministral 3 3B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts10Ministral 3 3B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- ReasoningHard problems that need careful thinking11Even1 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Qwen3 VL 4B Instruct pulls ahead
Where Ministral 3 3B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+17.2points ahead
- Reasons across sets of long documentsAA-LCR+3points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.8points ahead
Every test, side by side
All 22 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingMinistral 3 3B
- SciCodeMinistral 3 3B by 1.615.313.7+1.6
- Terminal-Bench Hardtie00tie
ReasoningEven
- Humanity's Last ExamMinistral 3 3B by 1.85.43.6+1.8
- GPQA DiamondQwen3 VL 4B Instruct by 1.335.837.1+1.3
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationMinistral 3 3B by 17.219.82.6+17.2
- AA-Omniscience · AccuracyQwen3 VL 4B Instruct by 1.9910.9+1.9
Images and chartsQwen3 VL 4B Instruct
Following instructionsQwen3 VL 4B Instruct
- IFBenchQwen3 VL 4B Instruct by 526.831.8+5
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIME 2025Ministral 3 3B by 25.572.146.6+25.5
- OCRBench v2_enQwen3 VL 4B Instruct by 22.341.463.7+22.3
- Arena HardQwen3 VL 4B Instruct by 13.330.543.8+13.3
- AA-OmniscienceMinistral 3 3B by 12-64-76+12
- HallusionBenchMinistral 3 3B by 7.665.257.6+7.6
- τ²-Bench Telecom (AA run)Ministral 3 3B by 1.524.923.4+1.5
- AA Intelligencetie4.85.7tie
- MM MTBenchtie7.87.5tie
- Artificial Analysis Coding Indextie4.84.6tie
- WildBenchtie56.856.8tie
Questions people ask
Which is better, Ministral 3 3B or Qwen3 VL 4B Instruct?
Ministral 3 3B and Qwen3 VL 4B Instruct each win two of the seven areas where both have results. Ministral 3 3B wins coding and long documents; Qwen3 VL 4B Instruct wins images and charts and following instructions. They are level on agents, reasoning and facts.
Which is better for coding?
Ministral 3 3B. It wins 1 of the 2 coding tests both models report; Qwen3 VL 4B Instruct wins none, and 1 is a tie.
How do you compare the two?
We use the 22 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.