Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.8

Score basis

Claude Opus 4.8 is strong in factuality, reasoning, and coding; capable in agentic tasks, long context, multimodal tasks, and math; and behind the leaders in instruction following. Too few results yet to rate safety or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong3 capabilities
  1. Factuality−8.4%3 of 4
  2. Reasoning−8.6%6 of 6
  3. Coding−9.7%9 of 10
Capable4 capabilities
  1. Agentic−10.3%7 of 7
  2. Long Context−15.5%1 of 3
  3. Multimodal−17.5%2 of 6
  4. Math−24.1%5 of 5
Limited1 capability
  1. Instruction Following−26.7%2 of 3
Not rated2 capabilities

Too few results yet to rate Safety or Multilingual.

Price

$5.00input$25.00outputper million tokens

From Anthropic's own price page · 4 providers tracked · All prices

Costs more than 93% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

199results on139benchmarks

  • 32 independently verified
  • 18 aggregator
  • 67 vendor-reported
  • 82 cross-referenced

From 36 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Claude Opus 4.8 benchmark results

199 results on 139 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Factuality

StrongFull factuality ranking

8.4% behind the leader3 of 4 ranked benchmarks measured

Reasoning

StrongFull reasoning ranking

8.6% behind the leader6 of 6 ranked benchmarks measured

Show 14 more reasoning resultsHide 14 reasoning results

9.7% behind the leader9 of 10 ranked benchmarks measured

Show 5 more coding resultsHide 5 coding results

Agentic

CapableFull agentic ranking

10.3% behind the leader7 of 7 ranked benchmarks measured

Show 5 more agentic resultsHide 5 agentic results

Long Context

CapableFull long context ranking

15.5% behind the leader1 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Multimodal

CapableFull multimodal ranking

17.5% behind the leader2 of 6 ranked benchmarks measured

Show 5 more multimodal resultsHide 5 multimodal results

24.1% behind the leader5 of 5 ranked benchmarks measured

Show 6 more math resultsHide 6 math results

Instruction Following

LimitedFull instruction following ranking

26.7% behind the leader2 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 124 more resultsHide 124 results

Claude Opus 4.8: common questions

Who makes Claude Opus 4.8?

Claude Opus 4.8 is made by Anthropic.

When was Claude Opus 4.8 released?

Claude Opus 4.8 was released on May 28, 2026, according to Artificial Analysis.

What is Claude Opus 4.8 good at?

Claude Opus 4.8 is strong in factuality, reasoning, and coding; capable in agentic tasks, long context, multimodal tasks, and math; and behind the leaders in instruction following. Too few results yet to rate safety or multilingual tasks.

How much does Claude Opus 4.8 cost?

Claude Opus 4.8 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 330 priced models we track.

How many benchmarks has Claude Opus 4.8 been tested on?

We track 199 results for Claude Opus 4.8 on 139 benchmarks from 36 sources, 32 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Claude Opus 4.8 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.8.

About this record

Where Claude Opus 4.8's numbers come from, and every name it appears under.

Tracked since
May 29, 2026
Newest source mention
Oct 5, 2026

Where the results come from

Verification: 199 scores · 32 independently verified · 18 aggregator-attributed · 82 vendor cross-reference · 67 vendor-reported. How these tiers are assigned

From 36 sources on 15 sites. Hugging Face supplies 79 of them; the 32 independently verified results come from 9 sites. Bars are coloured by trust tier.

  • huggingface.co79
  • www-cdn.anthropic.com34
  • api.llm-stats.com23
  • artificialanalysis.ai19
  • arcprize.org8
  • livebench.ai7
  • anthropic.com6
  • cdn.sanity.io4
  • datasets-server.huggingface.co4
  • epoch.ai4
  • ai.meta.com3
  • matharena.ai3
  • labs.scale.com2
  • lmarena.ai2
  • simple-bench.com1

Also known as

How our sources name Claude Opus 4.8 at each reasoning setting.

SettingShort formLong formAPI id
lowclaude opus 4.8 (low)——
mediumclaude opus 4.8 (medium)——
highclaude opus 4.8 (high)—claude-opus-4-8-high
maxclaude opus 4.8 (max) claude opus 4.8 maxclaude opus 4.8 (adaptive reasoning, max effort)claude-opus-4-8 (max)
Also listed asclaude 4.8claude-opus-4-8claude-opus-4-8-thinkingclaude-opus-4.8:batchclaude's opus 4.8opus 4.8anthropics opus 4.8