Agnes 2.5 Pro Alpha vs Claude Haiku 4.5
Wins 4 of 5 areas
Coding · Agents · Reasoning · Long documents
Wins 0 of 5 areas
—
Agnes 2.5 Pro Alpha is the stronger all-rounder.
Scores updated · 15 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own30Agnes 2.5 Pro Alpha3 of 3 tests
- ReasoningHard problems that need careful thinking30Agnes 2.5 Pro Alpha3 of 3 tests
- CodingWriting and fixing software10Agnes 2.5 Pro Alpha1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts10Agnes 2.5 Pro Alpha1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Agnes 2.5 Pro Alpha pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+23.2points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+22.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+20.4points ahead
Where Claude Haiku 4.5 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+60.8points ahead
Every test, side by side
All 15 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingAgnes 2.5 Pro Alpha
- Terminal-Bench 2.1Agnes 2.5 Pro Alpha by 22.86744.2+22.8
- SciCodetie42.942.2tie
AgentsAgnes 2.5 Pro Alpha
- GDPValAgnes 2.5 Pro Alpha by 17.629.411.8+17.6
- τ-Bench V3 · BankingAgnes 2.5 Pro Alpha by 3.112.49.3+3.1
- Terminal-Bench 4.0Agnes 2.5 Pro Alpha by 330+3
ReasoningAgnes 2.5 Pro Alpha
- Humanity's Last ExamAgnes 2.5 Pro Alpha by 23.233.610.4+23.2
- GPQA DiamondAgnes 2.5 Pro Alpha by 20.487.667.2+20.4
- CritPtAgnes 2.5 Pro Alpha by 10.910.90+10.9
FactsEven
- AA-Omniscience · Non-hallucinationClaude Haiku 4.5 by 60.811.972.7+60.8
- AA-Omniscience · AccuracyAgnes 2.5 Pro Alpha by 15.533.518+15.5
Long documentsAgnes 2.5 Pro Alpha
- AA-LCRAgnes 2.5 Pro Alpha by 1.475.774.3+1.4
Other results4 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceClaude Haiku 4.5 by 20.7-25.1-4.4+20.7
- AA Agentic IndexAgnes 2.5 Pro Alpha by 19.129.410.3+19.1
- Artificial Analysis Coding IndexAgnes 2.5 Pro Alpha by 14.958.843.9+14.9
- AA IntelligenceAgnes 2.5 Pro Alpha by 9.926.816.9+9.9
Questions people ask
Which is better, Agnes 2.5 Pro Alpha or Claude Haiku 4.5?
Agnes 2.5 Pro Alpha wins four of the five areas where both have results: coding, agents, reasoning and long documents. Claude Haiku 4.5 wins none. They are level on facts.
Which is better for coding?
Agnes 2.5 Pro Alpha. It wins 1 of the 2 coding tests both models report; Claude Haiku 4.5 wins none, and 1 is a tie.
How do you compare the two?
We use the 15 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 4 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.