Claude 3 Haiku vs EXAONE 4.0 32B
Wins 2 of 5 areas
Facts · Long documents
Wins 2 of 5 areas
Coding · Reasoning
The two are evenly matched.Claude 3 Haiku is better at facts and long documents; EXAONE 4.0 32B at reasoning and coding.
Scores updated · 18 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking02EXAONE 4.0 32B2 of 3 tests · 1 tie
- CodingWriting and fixing software02EXAONE 4.0 32B2 of 2 tests
- FactsGetting facts right instead of making them up20Claude 3 Haiku2 of 2 tests
- Long documentsFinding answers in very long texts10Claude 3 Haiku1 of 1 test
- Following instructionsDoing exactly what it is asked00Even0 each · 1 tie
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where EXAONE 4.0 32B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+36.5points ahead
- Code for real scientific research problemsSciCode+15.8points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+7.3points ahead
Where Claude 3 Haiku pulls ahead
- Reasons across sets of long documentsAA-LCR+13points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+5.3points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+3.7points ahead
Every test, side by side
All 18 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEXAONE 4.0 32B
- SciCodeEXAONE 4.0 32B by 15.818.634.4+15.8
- Terminal-Bench HardEXAONE 4.0 32B by 30.83.8+3
ReasoningEXAONE 4.0 32B
- GPQA DiamondEXAONE 4.0 32B by 36.537.473.9+36.5
- Humanity's Last ExamEXAONE 4.0 32B by 7.34.111.4+7.3
- CritPttie00tie
FactsClaude 3 Haiku
- AA-Omniscience · Non-hallucinationClaude 3 Haiku by 5.319.514.2+5.3
- AA-Omniscience · AccuracyClaude 3 Haiku by 3.717.613.9+3.7
Other results9 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- livecodebench_mediumEXAONE 4.0 32B by 79.3988.3+79.3
- LiveCodeBenchEXAONE 4.0 32B by 58.422.580.9+58.4
- livecodebench_hardEXAONE 4.0 32B by 54.41.956.3+54.4
- livecodebench_easyEXAONE 4.0 32B by 37.761.198.8+37.7
- AA-OmniscienceClaude 3 Haiku by 11.5-48.6-60.1+11.5
- Artificial Analysis Coding IndexEXAONE 4.0 32B by 7.36.714+7.3
- τ²-Bench Telecom (AA run)Claude 3 Haiku by 3.821.117.3+3.8
- AA IntelligenceEXAONE 4.0 32B by 2.65.68.2+2.6
- AA Agentic IndexEXAONE 4.0 32B by 2.579.5+2.5
Questions people ask
Which is better, Claude 3 Haiku or EXAONE 4.0 32B?
Claude 3 Haiku and EXAONE 4.0 32B each win two of the five areas where both have results. Claude 3 Haiku wins facts and long documents; EXAONE 4.0 32B wins coding and reasoning. They are level on following instructions.
Which is better for coding?
EXAONE 4.0 32B. It wins 2 of the 2 coding tests both models report; Claude 3 Haiku wins none.
How do you compare the two?
We use the 18 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 9 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.