GPT-4o vs Qwen3 235B A22B
Wins 1 of 6 areas
Facts
Wins 4 of 6 areas
Coding · Agents · Reasoning · Following instructions
Qwen3 235B A22B is the stronger all-rounder.GPT-4o is better at facts.
Scores updated · 47 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software13Qwen3 235B A22B3 of 4 tests
- ReasoningHard problems that need careful thinking02Qwen3 235B A22B2 of 3 tests · 1 tie
- Following instructionsDoing exactly what it is asked01Qwen3 235B A22B1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own01Qwen3 235B A22B1 of 1 test
- FactsGetting facts right instead of making them up21GPT-4o2 of 4 tests · 1 tie
- Long documentsFinding answers in very long texts11Even1 each
Agents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Qwen3 235B A22B costs 83% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Qwen3 235B A22B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+17.4points ahead
- Short factual questions, answered correctlySimpleQA Verified+14.4points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+9.2points ahead
Where GPT-4o pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+39.7points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+5.2points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+2.2points ahead
Every test, side by side
All 47 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 235B A22B
- SciCodeQwen3 235B A22B by 6.633.339.9+6.6
- LiveCodeBench v6Qwen3 235B A22B by 6.130.937+6.1
- Terminal-Bench HardGPT-4o by 2.28.36.1+2.2
- SWE-bench VerifiedQwen3 235B A22B by 1.233.234.4+1.2
ReasoningQwen3 235B A22B
- GPQA DiamondQwen3 235B A22B by 17.452.670+17.4
- Humanity's Last ExamQwen3 235B A22B by 9.21.811+9.2
- CritPttie00tie
FactsGPT-4o
- AA-Omniscience · Non-hallucinationGPT-4o by 39.762.122.4+39.7
- SimpleQA VerifiedQwen3 235B A22B by 14.42640.4+14.4
- AA-Omniscience · AccuracyGPT-4o by 5.223.718.5+5.2
- Vectara HHEM hallucination ratelower is bettertie9.69.3tie
Long documentsEven
- AA-LCRGPT-4o by 56560+56
- LongBench v2Qwen3 235B A22B by 248.150.1+2
Following instructionsQwen3 235B A22B
- IFBenchQwen3 235B A22B by 2.73638.7+2.7
- Multi-Challengetie40.341.2tie
Other results31 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIME 2024Qwen3 235B A22B by 76.49.385.7+76.4
- AIME 2025Qwen3 235B A22B by 69.811.781.5+69.8
- livecodebench_mediumQwen3 235B A22B by 56.732.188.8+56.7
- livecodebench_hardQwen3 235B A22B by 49.54.554+49.5
- CNMO 2024Qwen3 235B A22B by 37.810.848.6+37.8
- AA-OmniscienceGPT-4o by 34.2-10.5-44.7+34.2
- LiveCodeBenchQwen3 235B A22B by 32.438.370.7+32.4
- SimpleQAGPT-4o by 2538.213.2+25
- Tau2 airlineGPT-4o by 1945.526.5+19
- livecodebench_easyQwen3 235B A22B by 16.682.599.1+16.6
- MATH-500 (EM)Qwen3 235B A22B by 16.674.691.2+16.6
- Arena HardQwen3 235B A22B by 16.379.395.6+16.3
- MMLUQwen3 235B A22B by 13.273.887+13.2
- AA Agentic IndexQwen3 235B A22B by 108.418.4+10
- TAU-bench (airline)GPT-4o by 8.142.834.7+8.1
- MGSMGPT-4o by 790.583.5+7
- Artificial Analysis Coding IndexGPT-4o by 6.824.217.4+6.8
- MMLU-ProGPT-4o by 6.574.768.2+6.5
- τ²-Bench (Retail)GPT-4o by 6.463.457+6.4
- MMLU-ProXQwen3 235B A22B by 5.661.166.7+5.6
- MMMLUQwen3 235B A22B by 5.381.486.7+5.3
- τ²-Bench Telecom (AA run)GPT-4o by 4.928.924+4.9
- GSM8KQwen3 235B A22B by 3.590.994.4+3.5
- AA IntelligenceQwen3 235B A22B by 2.27.39.5+2.2
- IFEvalQwen3 235B A22B by 2.28183.2+2.2
- vectara_avg_summary_lengthQwen3 235B A22B by 19 rating points86.6105.6+19 rating
- SuperGPQAQwen3 235B A22B by 1.742.444.1+1.7
- TAU-bench (retail)GPT-4o by 1.760.358.6+1.7
- vectara_answer_rateQwen3 235B A22B by 1.193.894.9+1.1
- MMLU-Reduxtie8887.4tie
- vectara_factual_consistencytie90.490.7tie
Questions people ask
Which is better, GPT-4o or Qwen3 235B A22B?
Qwen3 235B A22B wins four of the six areas where both have results: coding, agents, reasoning and following instructions. GPT-4o wins facts. They are level on long documents.
Which is better for coding?
Qwen3 235B A22B. It wins 3 of the 4 coding tests both models report; GPT-4o wins 1.
Which is cheaper?
GPT-4o costs $5.00 per million input tokens and $15.00 per million output tokens; Qwen3 235B A22B costs $0.70 and $2.80. That makes Qwen3 235B A22B about 83% cheaper for the same work.
How do you compare the two?
We use the 47 benchmark tests both models have published scores on. The verdict counts the 16 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 31 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.