Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.7 vs Seed 2.1 Pro Preview

35 SHARED BENCHMARKS

Across 35 shared benchmarks, Claude Opus 4.7 scores higher on 13 and Seed 2.1 Pro Preview on 22. The widest gap is BabyVision, where Seed 2.1 Pro Preview scores 73.7 against 22.2.

ANTHROPICVSBYTEDANCE35 SHARED13–22 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicByteDance
Released—
Price per 1M tokens input / output$5.00 / $25.00—
Cost of 1M in + 1M out$30.00—
Head-to-head of 35 shared benchmarks13 wins22 wins
Scores tracked independently verified167 39 ◆66 1 ◆

Claude Opus 4.7's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Opus 4.7. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 4.7WINSSeed 2.1 Pro Preview
Reasoning11Even, 1–1 of 2
Coding31Claude Opus 4.7 leads 3 of 4 · widest: Terminal-Bench 2.1 83.1 vs 71
Agentic14Seed 2.1 Pro Preview leads 4 of 5 · widest: GDPVal 41.9 vs 87.9
Multimodal14Seed 2.1 Pro Preview leads 4 of 5 · widest: BLINK 70.4 vs 81.4

Biggest gaps

Claude Opus 4.7 pulls furthest ahead on

  1. ARC-AGI-275.8 vs 62.5
  2. Terminal-Bench 2.183.1 vs 71
  3. SWE-bench Pro64.3 vs 57.5

Seed 2.1 Pro Preview pulls furthest ahead on

  1. GDPVal87.9 vs 41.9
  2. Humanity's Last Exam55.7 vs 42.3
  3. BLINK81.4 vs 70.4

Every shared benchmark35 · grouped by area

Coding 4

Agentic 5

Multimodal 5

BLINK70.481.4
MathVista84.490.7
MMMU-Pro78.882.7
OCRBenchv256.963.2

Other shared benchmarks 19

BabyVision22.273.7
ChartQAPro65.570.9
CyberGym73.170.2
DeepSWE5432.7
DUDE84.582.8
DynaMath63.173.1
ERQA52.572
Finance Agent v1.164.460.7
MathVerse (Vision-Only)77.489.7
MMSIBench (circular)17.435.9
Office QA Pro [Multimodal]76.572.2
OfficeQA Pro80.672.2
RealWorldQA75.686.7
SimpleVQA56.574.5
Toolathlon59.350.6
WorldVQA35.953
ZeroBench (main)818
ZeroBench (sub)37.149.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Claude Opus 4.7 withALL PAIRINGS →

Compare Seed 2.1 Pro Preview withALL PAIRINGS →