Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.7 vs Seed2.1

31 SHARED BENCHMARKS

Across 31 shared benchmarks, Claude Opus 4.7 scores higher on 11 and Seed2.1 on 20. The widest gap is BabyVision, where Seed2.1 scores 62.9 against 22.2.

ANTHROPICVSX31 SHARED11–20 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicx
Released—
Price per 1M tokens input / output$5.00 / $25.00—
Cost of 1M in + 1M out$30.00—
Head-to-head of 31 shared benchmarks11 wins20 wins
Scores tracked independently verified167 39 ◆50 0 ◆

Claude Opus 4.7's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Opus 4.7. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 4.7WINSSeed2.1
Reasoning10Claude Opus 4.7 leads 1 of 1 · widest: ARC-AGI-2 75.8 vs 61.3
Coding11Even, 1–1 of 2
Agentic12Seed2.1 leads 2 of 3 · widest: GDPVal 41.9 vs 82.7
Multimodal04Seed2.1 leads 4 of 4 · widest: BLINK 70.4 vs 79.4

Biggest gaps

Claude Opus 4.7 pulls furthest ahead on

  1. ARC-AGI-275.8 vs 61.3
  2. Terminal-Bench 2.183.1 vs 67.6
  3. AA ApexAgents33.9 vs 29.2

Seed2.1 pulls furthest ahead on

  1. GDPVal82.7 vs 41.9
  2. BLINK79.4 vs 70.4
  3. OCRBenchv262.8 vs 56.9

Every shared benchmark31 · grouped by area

Reasoning 1

BenchmarkClaude Opus 4.7MARGINSeed2.1
ARC-AGI-275.861.3

Coding 2

Agentic 3

BenchmarkClaude Opus 4.7MARGINSeed2.1
GDPVal41.982.7
MCP Atlas77.380.3

Multimodal 4

BenchmarkClaude Opus 4.7MARGINSeed2.1
BLINK70.479.4
MathVista84.490.5
MMMU-Pro78.880.1
OCRBenchv256.962.8

Other shared benchmarks 21

BenchmarkClaude Opus 4.7MARGINSeed2.1
BabyVision22.262.9
ChartQAPro65.570.9
CyberGym73.167
DUDE84.583.1
DynaMath63.168.1
ERQA52.571.3
Finance Agent v1.164.456
MathVerse (Vision-Only)77.489.2
MMSIBench (circular)17.431.4
Office QA Pro [Multimodal]76.571.1
OfficeQA Pro80.662.8
RealWorldQA75.686.3
SimpleVQA56.571.1
SWE-Pro Bench64.357
Toolathlon59.349.1
VisuLogic32.652.9
WorldVQA35.948.6
ZeroBench (main)811
ZeroBench (sub)37.149.1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Claude Opus 4.7 withALL PAIRINGS →

Compare Seed2.1 withALL PAIRINGS →