Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.5 vs Seed 1.8

20 SHARED BENCHMARKS

Across 20 shared benchmarks, Claude Opus 4.5 scores higher on 11 and Seed 1.8 on 9. The widest gap is ZeroBench, where Seed 1.8 scores 11 against 3. Seed 1.8 is the cheaper of the two on tracked API pricing ($0.25 against $5.00 per million input tokens).

ANTHROPICVSBYTEDANCE20 SHARED11–9 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicByteDance
Released—
Price per 1M tokens input / output$5.00 / $25.00$0.25 / $2.00
Cost of 1M in + 1M out$30.00$2.25 13.3× less
Head-to-head of 20 shared benchmarks11 wins9 wins
Scores tracked independently verified204 44 ◆40 0 ◆

Claude Opus 4.5's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Opus 4.5; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 4.5WINSSeed 1.8
Coding20Claude Opus 4.5 leads 2 of 2 · widest: SWE-bench Verified 80.9 vs 72.9
Agentic11Even, 1–1 of 2
Instruction Following01Seed 1.8 leads 1 of 1 · widest: Multi-Challenge 59 vs 66.7
Multimodal13Seed 1.8 leads 3 of 4 · widest: MathVista 80.2 vs 87.7

Biggest gaps

Claude Opus 4.5 pulls furthest ahead on

  1. SWE-bench Verified80.9 vs 72.9
  2. OSWorld-Verified66.3 vs 61.9
  3. LiveCodeBench v684.8 vs 79.5

Seed 1.8 pulls furthest ahead on

  1. BrowseComp67.6 vs 37
  2. Multi-Challenge66.7 vs 59
  3. MathVista87.7 vs 80.2

Every shared benchmark20 · grouped by area

Coding 2

Agentic 2

Instruction Following 1

Multimodal 4

CharXiv (RQ)67.271.4
MathVista80.287.7
MMMU80.783.4
MMMU-Pro7473.2

Other shared benchmarks 11

AIME 202592.894.3
MATH-Vision77.181.3
MMLU-Pro89.584.9
MMVU77.373.1
MotionBench60.370.6
SimpleVQA69.765.4
SuperGPQA70.664.8
VideoMMMU84.482.7

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, deepinfra), otherwise the lowest tracked offer.

Compare Claude Opus 4.5 withALL PAIRINGS →