Claude Opus 4.8
Claude Opus 4.8 is strong in factuality, reasoning, and coding; capable in agentic tasks, long context, multimodal tasks, and math; and behind the leaders in instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$5.00input$25.00outputper million tokens
From Anthropic's own price page · 4 providers tracked · All prices
Evidence
199results on139benchmarks
- 32 independently verified
- 18 aggregator
- 67 vendor-reported
- 82 cross-referenced
From 36 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 4.8 benchmark results
199 results on 139 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.4% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy48.83Oct 8, 2026omniscienceAccuracy
- 53.00Jun 13, 2026
- AA-Omniscience · Non-hallucination60.75Oct 8, 2026omniscienceNonHallucination
8.6% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning89.19Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond92.02Oct 8, 2026gpqa
- 72.10Jul 24, 2026
- Humanity's Last Exam48.66Oct 8, 2026aa_hle
- 64.80May 29, 2026
- 20.86Oct 8, 2026
Show 14 more reasoning resultsHide 14 reasoning results
- 72.08Sep 21, 2026
- 71.67Sep 21, 2026
- 62.22Sep 21, 2026
- 1.52Sep 21, 2026
- 1.52Jul 24, 2026
- 1.50Jul 24, 2026
- 20.90Jul 27, 2026
- GPQA Diamond93.60Oct 7, 2026GPQA
- 92.00Aug 12, 2026
- 91.00Jul 27, 2026
- 57.90Oct 7, 2026
- Humanity's Last Exam49.80Aug 5, 2026Multidisciplinary reasoning Humanity's Last Exam no tools
- 45.70Aug 12, 2026
- Humanity's Last Exam49.80Jul 27, 2026HLE-Full
9.7% behind the leader9 of 10 ranked benchmarks measured
- Terminal-Bench 2.184.64Oct 8, 2026terminalbenchV21
- 88.60Oct 7, 2026
- 84.40Oct 7, 2026
- LiveBench · Coding81.83Oct 8, 2026livebench_coding@2026-06-25
- Terminal-Bench Hard58.33Oct 8, 2026aa_terminalbench_hard
- SciCode54.40Oct 8, 2026aa_scicode
- 1556.06Sep 21, 2026
- 69.20Oct 7, 2026
- LiveBench · Agentic Coding50.51Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 5 more coding resultsHide 5 coding results
- 1536.01Jun 9, 2026
- 53.50Jul 27, 2026
- 69.20Aug 12, 2026
- Terminal-Bench 2.178.90Jul 13, 2026Terminal Bench 2.1 (Best Reported Harness)
- Terminal-Bench 2.185.00Jul 13, 2026Terminal Bench 2.1 (Terminus-2)
10.3% behind the leader7 of 7 ranked benchmarks measured
- 83.40Oct 7, 2026
- 82.20Oct 7, 2026
- 84.30Oct 7, 2026
- τ-Bench V3 · Banking34.23Oct 8, 2026tauBanking
- 47.77Oct 8, 2026
- Terminal-Bench 4.021.72Oct 8, 2026
Show 5 more agentic resultsHide 5 agentic results
- AA ApexAgents39.40Jul 27, 2026APEX-Agents
- 84.30Jul 27, 2026
- 82.20Oct 8, 2026
- 83.60Jul 27, 2026
- 83.40Aug 24, 2026
15.5% behind the leader1 of 3 ranked benchmarks measured
- 77.67Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 69.10Aug 24, 2026
17.5% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)89.90Oct 7, 2026CharXiv-R
- 78.90Aug 24, 2026
Show 5 more multimodal resultsHide 5 multimodal results
- CharXiv (reasoning)80.50Jul 7, 2026CharXiv Reasoning (without tools)
- CharXiv (reasoning)80.50Aug 24, 2026CharXiv (RQ)
- 1288.99Sep 22, 2026
- 1294.97Sep 21, 2026
- 1280Jun 17, 2026
24.1% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics94.32Oct 8, 2026livebench_math@2026-06-25
- 80.00Sep 9, 2026
- 56.10Sep 8, 2026
Show 6 more math resultsHide 6 math results
- 100.00Sep 2, 2026
- 95.70Jul 13, 2026
- 95.45Sep 2, 2026
- HMMT Feb 202696.70Jul 13, 2026HMMT Feb. 2026
- 83.50Jul 13, 2026
- 96.70Jul 7, 2026
26.7% behind the leader2 of 3 ranked benchmarks measured
- LiveBench · Instruction Following72.03Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench62.24Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- 62.20Aug 12, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language79.66Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis66.03Oct 8, 2026livebench_data_analysis@2026-06-25
- 35.33Oct 8, 2026
- 1482Oct 6, 2026
- AA Intelligence42.00Oct 3, 2026Artificial Analysis Intelligence Index
- 92.50Sep 21, 2026
Show 124 more resultsHide 124 results
- 42.59Sep 9, 2026
- AA Intelligence41.79Oct 8, 2026aa_intelligence_index
- 28.75Oct 8, 2026
- 25.70Aug 31, 2026
- 27.00Jul 27, 2026
- 25.70Aug 28, 2026
- 45.10Aug 12, 2026
- 61.50Jun 24, 2026
- 54.93Jun 24, 2026
- 66.62Jun 24, 2026
- 56.59Jun 24, 2026
- 35.14Jun 24, 2026
- 64.10Jun 24, 2026
- 59.18Jun 24, 2026
- 54.66Jun 24, 2026
- 84.00Jul 13, 2026
- 69.80Aug 12, 2026
- 92.00Sep 21, 2026
- 91.50Sep 21, 2026
- 88.00Sep 21, 2026
- 92.50Jul 24, 2026
- Artificial Analysis Coding Index74.25Sep 9, 2026aa_coding_index
- AutomationBench Public27.20Aug 31, 2026AutomationBench (Public)
- baby_vision_with_python81.20Aug 24, 2026BabyVision w/ python
- 14.50Jun 12, 2026
- 88.50Jun 12, 2026
- 84.30Jun 12, 2026
- 89.70Jul 7, 2026
- 75.80Jul 7, 2026
- 69.40Jun 1, 2026
- 72.30Jun 1, 2026
- 89.90Jul 7, 2026
- 78.80Oct 7, 2026
- 78.30Aug 31, 2026
- 78.10Aug 28, 2026
- DeepSearchQA (F1)93.10Oct 7, 2026DeepSearchQA
- 93.10Jul 27, 2026
- 58.00Aug 31, 2026
- 59.00Oct 7, 2026
- 59.00Aug 12, 2026
- 71.60Aug 13, 2026
- 71.70Aug 31, 2026
- 53.90Oct 7, 2026
- 53.92Oct 7, 2026
- 53.90Jul 8, 2026
- 53.90Aug 24, 2026
- frontiermath_tier_4_v131.25Jun 13, 2026frontiermath_tier_4
- GDP (Surge AI)84.80Jul 24, 2026GDP.pdf (with tools)
- GDP (Surge AI)22.50Jun 12, 2026GDP.pdf
- GDPval-AA (Elo)1890.00Jul 8, 2026GDPval-AA
- 1593.00Jul 24, 2026
- 1588.00Aug 28, 2026
- GDPval-AA v2 Elo1593.00Jul 27, 2026GDPval-AA v2 (Elo)
- GraphWalks — Parents task, 256K-token context subset99.30Jun 12, 2026GraphWalks Parents 256K subset
- 68.10Jun 12, 2026
- GraphWalks BFS 256K85.90Jun 12, 2026GraphWalks BFS 256K subset
- 83.30Jun 12, 2026
- 59.30Jun 12, 2026
- 52.40Aug 12, 2026
- 58.80Jul 24, 2026
- 55.80Oct 7, 2026
- 57.40Jul 24, 2026
- 56.90Jun 12, 2026
- 60.30Jul 24, 2026
- HLE (with tools)57.90Aug 5, 2026Multidisciplinary reasoning Humanity's Last Exam with tools
- HLE (with tools)57.90Aug 28, 2026HLE w/ Tools
- 57.90Aug 13, 2026
- 26.50Jul 13, 2026
- 96.50Jul 13, 2026
- 87.60Oct 7, 2026
- 48.40Aug 12, 2026
- 10.00Jun 18, 2026
- Legal Agent Benchmark, Full Public Set9.60Jun 12, 2026Legal Agent Full Public Benchmark Set
- 77.22Oct 7, 2026
- 86.70Aug 24, 2026
- 81.60Sep 2, 2026
- 76.40Jul 27, 2026
- MLS Bench Litelower is better42.80Aug 12, 2026
- 79.20Aug 24, 2026
- 83.20Aug 24, 2026
- 69.70Aug 31, 2026
- 77.60Jul 24, 2026
- 66.20Oct 7, 2026
- 48.10Jun 12, 2026
- 63.90Jul 27, 2026
- 87.90Aug 24, 2026
- 84.00Jun 12, 2026
- 3.30Jul 13, 2026
- 55.70Jul 24, 2026
- 55.70Aug 24, 2026
- 80.30Aug 12, 2026
- 32.90Aug 28, 2026
- 34.10Jul 27, 2026
- 37.20Jul 13, 2026
- 71.90Jul 27, 2026
- 63.80Jun 14, 2026
- 15.50Aug 28, 2026
- 73.50Jul 27, 2026
- 34.00Jun 12, 2026
- ScreenSpot-Pro (No tools)87.90Oct 7, 2026ScreenSpot Pro
- 82.30Jun 1, 2026
- 87.90Jun 1, 2026
- 31.60Jul 27, 2026
- 38.40Oct 7, 2026
- SWE-Marathon48.80Aug 28, 2026SWE-Marathon (v1.1)
- 40.00Jul 27, 2026
- 26.00Jul 13, 2026
- 74.60Oct 7, 2026
- 21.10Aug 28, 2026
- Terminal-Bench 4.023.60Oct 7, 2026
- 59.90Jul 13, 2026
- 59.90Oct 7, 2026
- Toolathlon67.60Jun 30, 2026Toolathlon Pass@3
- 24.50Jun 30, 2026
- 48.10Jul 7, 2026
- Toolathlon Verified88.00Sep 22, 2026Toolathlon Verified Pass@3
- Toolathlon Verified79.90Sep 22, 2026Toolathlon Verified Pass@1
- 76.20Aug 31, 2026
- VideoMME (w sub.)86.00Aug 24, 2026Video-MME (w. sub)
- 72.90Aug 12, 2026
- ZeroBench34.00Aug 31, 2026ZeroBench (Pass@5)
- ZeroBench17.00Aug 24, 2026ZeroBench (pass@5)
- τ²-Bench Telecom (AA run)94.44Oct 8, 2026aa_tau2
- τ³-Bench Banking27.60Aug 24, 2026τ³-Banking
Claude Opus 4.8: common questions
Who makes Claude Opus 4.8?
Claude Opus 4.8 is made by Anthropic.
When was Claude Opus 4.8 released?
Claude Opus 4.8 was released on May 28, 2026, according to Artificial Analysis.
What is Claude Opus 4.8 good at?
Claude Opus 4.8 is strong in factuality, reasoning, and coding; capable in agentic tasks, long context, multimodal tasks, and math; and behind the leaders in instruction following. Too few results yet to rate safety or multilingual tasks.
How much does Claude Opus 4.8 cost?
Claude Opus 4.8 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 330 priced models we track.
How many benchmarks has Claude Opus 4.8 been tested on?
We track 199 results for Claude Opus 4.8 on 139 benchmarks from 36 sources, 32 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4.8 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.8.
About this record
Where Claude Opus 4.8's numbers come from, and every name it appears under.
- Tracked since
- May 29, 2026
- Newest source mention
- Oct 5, 2026
Where the results come from
Verification: 199 scores · 32 independently verified · 18 aggregator-attributed · 82 vendor cross-reference · 67 vendor-reported. How these tiers are assigned
From 36 sources on 15 sites. Hugging Face supplies 79 of them; the 32 independently verified results come from 9 sites. Bars are coloured by trust tier.
- huggingface.co79
- www-cdn.anthropic.com34
- api.llm-stats.com23
- artificialanalysis.ai19
- arcprize.org8
- livebench.ai7
- anthropic.com6
- cdn.sanity.io4
- datasets-server.huggingface.co4
- epoch.ai4
- ai.meta.com3
- matharena.ai3
- labs.scale.com2
- lmarena.ai2
- simple-bench.com1
Also known as
How our sources name Claude Opus 4.8 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude opus 4.8 (low) | — | — |
| medium | claude opus 4.8 (medium) | — | — |
| high | claude opus 4.8 (high) | — | claude-opus-4-8-high |
| max | claude opus 4.8 (max) claude opus 4.8 max | claude opus 4.8 (adaptive reasoning, max effort) | claude-opus-4-8 (max) |