Qwen3.5 122B A10B vs Qwen3.5 4B
Wins 7 of 7 areas
Coding · Agents · Reasoning · Facts · Images and charts · Long documents · Following instructions
Wins 0 of 7 areas
—
Qwen3.5 122B A10B is the stronger all-rounder.Qwen3.5 4B is cheaper.
Scores updated · 41 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Images and chartsUnderstanding pictures, charts and video50Qwen3.5 122B A10B5 of 5 tests
- CodingWriting and fixing software40Qwen3.5 122B A10B4 of 4 tests
- ReasoningHard problems that need careful thinking20Qwen3.5 122B A10B2 of 3 tests · 1 tie
- Long documentsFinding answers in very long texts20Qwen3.5 122B A10B2 of 2 tests
- Following instructionsDoing exactly what it is asked20Qwen3.5 122B A10B2 of 2 tests
- FactsGetting facts right instead of making them up10Qwen3.5 122B A10B1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own10Qwen3.5 122B A10B1 of 1 test
Agents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Qwen3.5 4B costs 95% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 122B A10B pulls ahead
Where Qwen3.5 4B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 41 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3.5 122B A10B
- SciCodeQwen3.5 122B A10B by 23.639.716.1+23.6
- LiveCodeBench v6Qwen3.5 122B A10B by 23.178.955.8+23.1
- Terminal-Bench 2.1Qwen3.5 122B A10B by 21.847.625.8+21.8
- Terminal-Bench HardQwen3.5 122B A10B by 12.931.118.2+12.9
ReasoningQwen3.5 122B A10B
- Humanity's Last ExamQwen3.5 122B A10B by 15.325.29.9+15.3
- GPQA DiamondQwen3.5 122B A10B by 8.685.777.1+8.6
- CritPttie0.60tie
FactsQwen3.5 122B A10B
- AA-Omniscience · AccuracyQwen3.5 122B A10B by 9.324.415.1+9.3
- AA-Omniscience · Non-hallucinationtie12.913.4tie
Images and chartsQwen3.5 122B A10B
Long documentsQwen3.5 122B A10B
- AA-LCRQwen3.5 122B A10B by 13.376.363+13.3
- LongBench v2Qwen3.5 122B A10B by 10.260.250+10.2
Following instructionsQwen3.5 122B A10B
- IFBenchQwen3.5 122B A10B by 23.775.752+23.7
- Multi-ChallengeQwen3.5 122B A10B by 12.561.549+12.5
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Artificial Analysis Coding IndexQwen3.5 122B A10B by 23.145.722.6+23.1
- AA Agentic IndexQwen3.5 4B by 22.99.632.5+22.9
- BFCLv4Qwen3.5 122B A10B by 21.672.250.6+21.6
- SimpleVQAQwen3.5 122B A10B by 2161.740.7+21
- RealWorldQAQwen3.5 122B A10B by 1885.167.1+18
- AA-OmniscienceQwen3.5 122B A10B by 16.8-41.5-58.4+16.8
- HallusionBenchQwen3.5 122B A10B by 15.967.651.7+15.9
- ERQAQwen3.5 122B A10B by 15.76246.3+15.7
- SuperGPQAQwen3.5 122B A10B by 14.267.152.9+14.2
- HMMT 2025Qwen3.5 122B A10B by 13.590.376.8+13.5
- IncludeQwen3.5 122B A10B by 11.882.871+11.8
- VITA-BenchQwen3.5 122B A10B by 11.633.622+11.6
- MMLU-ProXQwen3.5 122B A10B by 10.782.271.5+10.7
- MMMLUQwen3.5 122B A10B by 10.686.776.1+10.6
- MMLU-ProQwen3.5 122B A10B by 7.686.779.1+7.6
- IFEvalQwen3.5 122B A10B by 7.293.486.2+7.2
- C-EvalQwen3.5 122B A10B by 6.891.985.1+6.8
- DeepPlanningQwen3.5 122B A10B by 6.524.117.6+6.5
- MMLU-ReduxQwen3.5 122B A10B by 5.29488.8+5.2
- AA IntelligenceQwen3.5 122B A10B by 2.515.613.1+2.5
- τ²-Bench Telecom (AA run)Qwen3.5 122B A10B by 1.593.692.1+1.5
- t2-benchtie79.579.9tie
Questions people ask
Which is better, Qwen3.5 122B A10B or Qwen3.5 4B?
Qwen3.5 122B A10B wins all seven areas where both have results: coding, agents, reasoning, facts, images and charts, long documents and following instructions. Qwen3.5 4B wins none, but costs 95% less.
Which is better for coding?
Qwen3.5 122B A10B. It wins 4 of the 4 coding tests both models report; Qwen3.5 4B wins none.
Which is cheaper?
Qwen3.5 122B A10B costs $0.40 per million input tokens and $3.20 per million output tokens; Qwen3.5 4B costs $0.03 and $0.15. That makes Qwen3.5 4B about 95% cheaper for the same work.
How do you compare the two?
We use the 41 benchmark tests both models have published scores on. The verdict counts the 19 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.