Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.6

Score basis

Claude Opus 4.6 is capable in long context, reasoning, coding, and factuality; and behind the leaders in multimodal tasks, agentic tasks, and math. Too few results yet to rate safety, multilingual tasks, or instruction following.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable4 capabilities
  1. Long Context−10.5%2 of 3
  2. Reasoning−13.7%6 of 6
  3. Coding−21.1%9 of 10
  4. Factuality−23.2%4 of 4
Limited3 capabilities
  1. Multimodal−30.0%2 of 6
  2. Agentic−31.5%4 of 7
  3. Math−36.4%3 of 5
Not rated3 capabilities

Too few results yet to rate Safety, Multilingual or Instruction Following.

Price

$5.00input$25.00outputper million tokens

From Anthropic's own price page · 3 providers tracked · All prices

Costs more than 93% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

231results on134benchmarks

  • 55 independently verified
  • 33 aggregator
  • 50 vendor-reported
  • 93 cross-referenced

From 42 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

Claude Opus 4.6 benchmark results

231 results on 134 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

CapableFull long context ranking

10.5% behind the leader2 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Reasoning

CapableFull reasoning ranking

13.7% behind the leader6 of 6 ranked benchmarks measured

Show 16 more reasoning resultsHide 16 reasoning results

Coding

CapableFull coding ranking

21.1% behind the leader9 of 10 ranked benchmarks measured

Show 20 more coding resultsHide 20 coding results

Factuality

CapableFull factuality ranking

23.2% behind the leader4 of 4 ranked benchmarks measured

Show 3 more factuality resultsHide 3 factuality results

Multimodal

LimitedFull multimodal ranking

30.0% behind the leader2 of 6 ranked benchmarks measured

Show 12 more multimodal resultsHide 12 multimodal results

Agentic

LimitedFull agentic ranking

31.5% behind the leader4 of 7 ranked benchmarks measured

Show 13 more agentic resultsHide 13 agentic results

36.4% behind the leader3 of 5 ranked benchmarks measured

Show 11 more math resultsHide 11 math results

Instruction Following

Not enough dataFull instruction following ranking

0 of 3 ranked benchmarks measured

Show 4 more instruction following resultsHide 4 instruction following results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 112 more resultsHide 112 results

Claude Opus 4.6: common questions

Who makes Claude Opus 4.6?

Claude Opus 4.6 is made by Anthropic.

When was Claude Opus 4.6 released?

Claude Opus 4.6 was released on Feb 5, 2026, according to Artificial Analysis.

What is Claude Opus 4.6 good at?

Claude Opus 4.6 is capable in long context, reasoning, coding, and factuality; and behind the leaders in multimodal tasks, agentic tasks, and math. Too few results yet to rate safety, multilingual tasks, or instruction following.

How much does Claude Opus 4.6 cost?

Claude Opus 4.6 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 331 priced models we track.

How many benchmarks has Claude Opus 4.6 been tested on?

We track 231 results for Claude Opus 4.6 on 134 benchmarks from 42 sources, 55 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does Claude Opus 4.6 support?

OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.6.

About this record

Where Claude Opus 4.6's numbers come from, and every name it appears under.

Tracked since
Apr 25, 2026
Newest source mention
Sep 26, 2026

Where the results come from

Verification: 231 scores · 55 independently verified · 33 aggregator-attributed · 93 vendor cross-reference · 50 vendor-reported. How these tiers are assigned

From 42 sources on 19 sites. Hugging Face supplies 82 of them; the 55 independently verified results come from 12 sites. Bars are coloured by trust tier.

  • huggingface.co82
  • artificialanalysis.ai35
  • www-cdn.anthropic.com22
  • api.llm-stats.com20
  • deepmind.google10
  • arcprize.org9
  • anthropic.com7
  • livebench.ai7
  • raw.githubusercontent.com7
  • matharena.ai6
  • datasets-server.huggingface.co5
  • epoch.ai5
  • labs.scale.com5
  • lmarena.ai3
  • swebench.com3
  • 99franklin.github.io2
  • cdn.sanity.io1
  • mistral.ai1
  • simple-bench.com1

Also known as

How our sources name Claude Opus 4.6 at each reasoning setting.

SettingShort formLong formAPI id
lowclaude opus 4.6 (120k, low)——
mediumclaude opus 4.6 (120k, medium) opus4.6 medium——
highclaude opus 4.6 (120k, high) claude opus 4.6 (non-reasoning, high)Claude Opus 4.6 (Non-reasoning, High Effort)claude-opus-4.6 (high) claude-opus-4-6-high
maxanthropic opus 4.6 (max) claude opus 4.6 (120k, max) Opus-4.6 Max Opus 4.6 Thinking (Max) claude opus 4.6 (max) opus4.6 max claude opus 4.6 maxclaude opus 4.6 (adaptive reasoning, max effort) claude opus 4.6 (max effort)claude-opus-4-6-thinking-max claude-opus-4-6 (max)
Also listed asclaude-opus-4-6-thinkingclaude opus~4.6claude opus 4.6 (inspect)claude-opus-4-6 (non-thinking)claude 4.6 opusclaude-opus-4-6-adaptiveclaude-opus-4-6 (Non-Reasoning)claude-opus-4.6-fastmini-swe-agent + claude opus 4.6claude opus 4.6 (no thinking)claude 4.6claude-opus-4-6claude opus 4.6 (64k thinking)opus 4.6claude-opus-4-6 (thinking)*claude-opus-4-6 (reasoning)and 5 more