Claude Opus 4.7
Claude Opus 4.7 is capable in coding, multimodal tasks, reasoning, factuality, long context, and agentic tasks; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$5.00input$25.00outputper million tokens
From Anthropic's own price page · 4 providers tracked · All prices
Evidence
168results on106benchmarks
- 39 independently verified
- 35 aggregator
- 51 vendor-reported
- 43 cross-referenced
From 33 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Claude Opus 4.7 benchmark results
168 results on 106 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
11.0% behind the leader9 of 10 ranked benchmarks measured
- 87.60Oct 7, 2026
- Terminal-Bench 2.183.15Oct 8, 2026terminalbenchV21
- LiveBench · Coding82.09Oct 8, 2026livebench_coding@2026-06-25
- 80.50Jun 21, 2026
- SciCode54.51Sep 4, 2026aa_scicode
- 1557.52Sep 21, 2026
- Terminal-Bench Hard51.52Oct 8, 2026aa_terminalbench_hard
- 64.30Oct 7, 2026
- LiveBench · Agentic Coding50.66Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 8 more coding resultsHide 8 coding results
- 1555.18Aug 29, 2026
- 1556.70May 22, 2026
- SciCode50.12Sep 4, 2026aa_scicode
- SWE-bench Pro64.30Jul 6, 2026SWE-Bench Pro (Public) Diverse agentic coding tasks
- 71.70Sep 28, 2026
- Terminal-Bench 2.188.00Sep 10, 2026Terminal-Bench 2.1 (Pass@1)
- Terminal-Bench 2.166.10Jul 6, 2026Terminal-bench 2.1 Agentic terminal coding
- Terminal-Bench Hard54.55Oct 8, 2026aa_terminalbench_hard
11.6% behind the leader3 of 6 ranked benchmarks measured
- CharXiv (reasoning)91.00Oct 7, 2026CharXiv-R
- 84.40Sep 28, 2026
- MMMU-Pro78.84Oct 8, 2026aa_mmmu_pro
Show 10 more multimodal resultsHide 10 multimodal results
- 70.40Sep 28, 2026
- CharXiv (reasoning)82.10Jun 21, 2026CharXiv Reasoning (No tools)
- CharXiv (reasoning)82.10Jul 6, 2026CharXiv Reasoning Information synthesis from complex charts
- 1312.93Sep 21, 2026
- 1316.43Aug 25, 2026
- 1304May 25, 2026
- MMMU-Pro76.36Oct 8, 2026aa_mmmu_pro
- 74.00Sep 28, 2026
- MMMU-Pro75.20Jul 6, 2026MMMU-Pro Multimodal understanding and reasoning
- 56.90Sep 28, 2026
13.4% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond91.41Oct 8, 2026gpqa
- LiveBench · Reasoning87.19Oct 8, 2026livebench_reasoning@2026-06-25
- 75.83Jul 24, 2026
- 61.70Aug 29, 2026
- Humanity's Last Exam42.31Oct 8, 2026aa_hle
- 12.00Oct 8, 2026
Show 13 more reasoning resultsHide 13 reasoning results
- 67.50Sep 21, 2026
- 75.83May 10, 2026
- 68.33May 10, 2026
- 62.08May 10, 2026
- ARC-AGI-275.80Jul 6, 2026ARC-AGI-2 Abstract reasoning puzzles
- 0.18May 10, 2026
- 5.14Oct 8, 2026
- GPQA Diamond88.48Oct 8, 2026gpqa
- GPQA Diamond94.20Oct 7, 2026GPQA
- Humanity's Last Exam33.32Oct 8, 2026aa_hle
- 54.70Oct 7, 2026
- Humanity's Last Exam46.90Jun 1, 2026Humanity's Last Exam (No tools)
- Humanity's Last Exam46.90Jul 6, 2026Humanity’s Last Exam Academic reasoning (full set, text + MM)
14.6% behind the leader4 of 4 ranked benchmarks measured
- 12.00Aug 29, 2026
- AA-Omniscience · Accuracy48.88Oct 8, 2026omniscienceAccuracy
- 51.70Sep 21, 2026
- AA-Omniscience · Non-hallucination57.71Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy44.73Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination45.90Oct 8, 2026omniscienceNonHallucination
15.5% behind the leader2 of 3 ranked benchmarks measured
- 78.67Oct 8, 2026
- MRCR v2 (8-needle, 128K)59.30Jul 6, 2026MRCR v2 (8-needle) Long context performance
Show 1 more long context resultHide 1 long context result
- 75.67Oct 8, 2026
20.0% behind the leader6 of 7 ranked benchmarks measured
- 78.00Oct 7, 2026
- 77.30Oct 7, 2026
- 79.30Oct 7, 2026
- τ-Bench V3 · Banking34.64Oct 8, 2026tauBanking
- 42.81Oct 8, 2026
Show 11 more agentic resultsHide 11 agentic results
- AA ApexAgents33.90Sep 28, 2026APEX Agents
- 46.66Oct 8, 2026
- 79.80Jun 1, 2026
- 58.57Jun 15, 2026
- GDPval (win rate)82.70Sep 28, 2026GDPval
- 79.10Oct 8, 2026
- 79.10Jun 1, 2026
- 79.10Sep 28, 2026
- 82.80Jun 1, 2026
- 82.30May 31, 2026
- OSWorld-Verified78.00Jul 6, 2026OSWorld-Verified Agentic computer use
30.5% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics92.85Oct 8, 2026livebench_math@2026-06-25
- 70.18Sep 21, 2026
- 31.71Sep 21, 2026
Show 5 more math resultsHide 5 math results
- 95.83Sep 2, 2026
- 95.83May 2, 2026
- 93.94Sep 2, 2026
- 93.94May 10, 2026
- 69.30May 3, 2026
35.2% behind the leader2 of 3 ranked benchmarks measured
- LiveBench · Instruction Following66.74Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench58.64Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench43.61Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1494Oct 8, 2026
- livebench_language77.91Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis78.26Oct 8, 2026livebench_data_analysis@2026-06-25
- 41.67Oct 8, 2026
- AA Intelligence31.00Sep 21, 2026Artificial Analysis Intelligence Index
- 91.00Sep 21, 2026
Show 77 more resultsHide 77 results
- 39.49Sep 9, 2026
- 64.64Jun 18, 2026
- AA Intelligence44.00Jun 21, 2026Artificial Analysis Intelligence Index
- AA Intelligence40.69Oct 8, 2026aa_intelligence_index
- AA Intelligence30.93Oct 8, 2026aa_intelligence_index
- 27.27Oct 8, 2026
- 14.83Oct 8, 2026
- 92.00May 10, 2026
- 93.50May 10, 2026
- 92.00Jun 21, 2026
- Artificial Analysis Coding Index73.60Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index53.07Jun 18, 2026aa_coding_index
- 22.20Sep 28, 2026
- 90.90Sep 9, 2026
- BigLaw Bench (substantive accuracy, high effort)90.90May 3, 2026BigLaw Bench (high effort)
- Blueprint-Bench 224.50Jul 6, 2026Blueprint-Bench 2 Agentic spatial reasoning
- 65.50Sep 28, 2026
- 67.60Jun 1, 2026
- 69.80Jun 1, 2026
- 91.00Jun 21, 2026
- 73.10Oct 7, 2026
- 73.10Sep 28, 2026
- 54.00Sep 28, 2026
- DeepSWE v1.1 (Resolved)69.80Sep 10, 2026
- 84.50Sep 28, 2026
- 63.10Sep 28, 2026
- 77.20Sep 28, 2026
- 52.50Sep 28, 2026
- 64.40Oct 7, 2026
- Finance Agent51.50Jun 1, 2026Finance Agent v2
- Finance Agent71.5May 10, 2026Finance Agent evaluation (overall score across six modules)
- Finance Agent evaluation0.81Oct 3, 2026
- Finance Agent v1.164.40Sep 28, 2026
- 51.51Oct 7, 2026
- Finance Agent v251.50Jul 6, 2026Finance Agent v2 Financial analysis and decision-making
- frontiermath_tier_4_v122.92Aug 29, 2026frontiermath_tier_4
- GDPval-AA (Elo)1753.00Jun 1, 2026GDPval-AA
- GDPval-AA (Elo)1753.00Jul 6, 2026GDPval-AA Economically valuable knowledge work
- GraphWalks — Parents task, 256K-token context subset93.60Jun 1, 2026GraphWalks Parents 256K
- 93.57May 3, 2026
- 76.90Jun 1, 2026
- 76.91May 3, 2026
- 58.60May 3, 2026
- 75.10May 3, 2026
- HLE (with tools)54.70Jun 21, 2026Humanity's Last Exam (With tools)
- 46.90Jul 7, 2026
- 7.10Oct 7, 2026
- 76.91Oct 7, 2026
- MathVerse (Vision-Only)77.40Sep 28, 2026
- 91.50Oct 7, 2026
- MMSIBench (circular)17.40Sep 28, 2026
- Office QA Pro [Multimodal]76.50Sep 28, 2026
- 86.30Jun 21, 2026
- 80.60Jun 21, 2026
- 76.50Sep 28, 2026
- 75.60Sep 28, 2026
- 79.50Jun 21, 2026
- 87.60Jun 21, 2026
- 56.50Sep 28, 2026
- 34.50Jun 21, 2026
- SWE-Pro Bench64.30Sep 28, 2026
- 69.40Oct 7, 2026
- Toolathlon66.70Jun 30, 2026Toolathlon Pass@3
- Toolathlon59.30Jun 30, 2026Toolathlon Pass@1
- 52.80Sep 28, 2026
- 25.90Jun 30, 2026
- 52.80Jul 7, 2026
- vectara_answer_rate98.00Aug 29, 2026Answer Rate
- vectara_avg_summary_length149.10Aug 29, 2026Average Summary Length (Words)
- vectara_factual_consistency88.00Aug 29, 2026Factual Consistency Rate
- 32.60Sep 28, 2026
- 35.90Sep 28, 2026
- 98.50Sep 9, 2026
- ZeroBench (main)8.00Sep 28, 2026
- ZeroBench (sub)37.10Sep 28, 2026
- τ²-Bench Telecom (AA run)88.60Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)73.98Oct 8, 2026aa_tau2
Claude Opus 4.7: common questions
Who makes Claude Opus 4.7?
Claude Opus 4.7 is made by Anthropic.
When was Claude Opus 4.7 released?
Claude Opus 4.7 was released on Apr 16, 2026, according to Artificial Analysis.
What is Claude Opus 4.7 good at?
Claude Opus 4.7 is capable in coding, multimodal tasks, reasoning, factuality, long context, and agentic tasks; and behind the leaders in math and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does Claude Opus 4.7 cost?
Claude Opus 4.7 costs $5.00 per million input tokens and $25.00 per million output tokens, according to Anthropic's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 93% of the 330 priced models we track.
How many benchmarks has Claude Opus 4.7 been tested on?
We track 168 results for Claude Opus 4.7 on 106 benchmarks from 33 sources, 39 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Claude Opus 4.7 support?
OpenRouter lists tool calling, structured outputs, and reasoning for Claude Opus 4.7.
About this record
Where Claude Opus 4.7's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Sep 28, 2026
Where the results come from
Verification: 168 scores · 39 independently verified · 35 aggregator-attributed · 43 vendor cross-reference · 51 vendor-reported. How these tiers are assigned
From 33 sources on 17 sites. Artificial Analysis supplies 37 of them; the 39 independently verified results come from 10 sites. Bars are coloured by trust tier.
- artificialanalysis.ai37
- lf3-static.bytednsdoc.com29
- api.llm-stats.com15
- cdn.sanity.io15
- www-cdn.anthropic.com15
- deepmind.google12
- arcprize.org8
- livebench.ai7
- anthropic.com6
- datasets-server.huggingface.co5
- epoch.ai4
- matharena.ai4
- raw.githubusercontent.com4
- huggingface.co2
- labs.scale.com2
- lmarena.ai2
- simple-bench.com1
Also known as
How our sources name Claude Opus 4.7 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | claude 4.7 (low) | — | — |
| medium | claude 4.7 (medium) | — | — |
| high | Opus 4.7 (High) claude 4.7 (high) claude opus 4.7 (non-reasoning, high) | claude opus 4.7 (non-reasoning, high effort) | claude-opus-4-7-high claude-opus-4.7 (high) |
| xhigh | claude opus 4.7 (xhigh) | — | — |
| max | claude 4.7 (max) claude opus 4.7 (max) | claude opus 4.7 (adaptive reasoning, max effort) | claude-opus-4-7 (max) |