Gemma 4 26B A4B vs GLM 5V Turbo
Wins 1 of 7 areas
Following instructions
Wins 5 of 7 areas
Coding · Agents · Facts · Images and charts · Long documents
GLM 5V Turbo is the stronger all-rounder.Gemma 4 26B A4B is cheaper and better at following instructions.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software03GLM 5V Turbo3 of 3 tests
- FactsGetting facts right instead of making them up02GLM 5V Turbo2 of 2 tests
- Images and chartsUnderstanding pictures, charts and video01GLM 5V Turbo1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own01GLM 5V Turbo1 of 1 test
- Long documentsFinding answers in very long texts01GLM 5V Turbo1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemma 4 26B A4B1 of 1 test
- ReasoningHard problems that need careful thinking11Even1 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemma 4 26B A4B costs 90% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GLM 5V Turbo pulls ahead
- Real work tasks from 44 professionsGDPVal+38points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+19points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+17.6points ahead
Where Gemma 4 26B A4B pulls ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+11.3points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+2.2points ahead
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGLM 5V Turbo
- Terminal-Bench HardGLM 5V Turbo by 1913.632.6+19
- LMArena · WebDevGLM 5V Turbo by 42 rating points13591401+42 rating
- SciCodeGLM 5V Turbo by 3.54043.5+3.5
ReasoningEven
- Humanity's Last ExamGemma 4 26B A4B by 2.219.317.1+2.2
- GPQA DiamondGLM 5V Turbo by 1.779.280.9+1.7
- CritPttie00.6tie
FactsGLM 5V Turbo
- AA-Omniscience · Non-hallucinationGLM 5V Turbo by 17.613.631.2+17.6
- AA-Omniscience · AccuracyGLM 5V Turbo by 10.219.129.3+10.2
Images and chartsGLM 5V Turbo
- MMMU-ProGLM 5V Turbo by 3.669.272.8+3.6
- LMArena · Visiontie12601264tie
Following instructionsGemma 4 26B A4B
- IFBenchGemma 4 26B A4B by 11.372.461.1+11.3
Other results8 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- τ²-Bench Telecom (AA run)GLM 5V Turbo by 54.943.698.5+54.9
- AA Agentic IndexGLM 5V Turbo by 50.11161.1+50.1
- Claw Eval (pass@3)GLM 5V Turbo by 472875+47
- AA-OmniscienceGLM 5V Turbo by 31.5-50.8-19.3+31.5
- SimpleVQAGLM 5V Turbo by 2652.278.2+26
- AA IntelligenceGLM 5V Turbo by 6.816.723.5+6.8
- PinchBenchGLM 5V Turbo by 674.780.7+6
- Artificial Analysis Coding IndexGemma 4 26B A4B by 3.139.336.2+3.1
Questions people ask
Which is better, Gemma 4 26B A4B or GLM 5V Turbo?
GLM 5V Turbo wins five of the seven areas where both have results: coding, agents, facts, images and charts and long documents. Gemma 4 26B A4B wins following instructions, and costs 90% less. They are level on reasoning.
Which is better for coding?
GLM 5V Turbo. It wins 3 of the 3 coding tests both models report; Gemma 4 26B A4B wins none.
Which is cheaper?
Gemma 4 26B A4B costs $0.13 per million input tokens and $0.40 per million output tokens; GLM 5V Turbo costs $1.20 and $4.00. That makes Gemma 4 26B A4B about 90% cheaper for the same work.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 13 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 8 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.