Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.5 vs Seed 2.1 Pro Preview

40 SHARED BENCHMARKS

Across 40 shared benchmarks, GPT-5.5 scores higher on 15 and Seed 2.1 Pro Preview on 25. The widest gap is DeepSWE, where GPT-5.5 scores 70 against 32.7.

OPENAIVSBYTEDANCE40 SHARED15–25 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIByteDance
Released—
Price per 1M tokens input / output$5.00 / $30.00—
Cost of 1M in + 1M out$35.00—
Head-to-head of 40 shared benchmarks15 wins25 wins
Scores tracked independently verified185 29 ◆66 1 ◆

GPT-5.5's release date per Artificial Analysis. Prices: Artificial Analysis for GPT-5.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.5WINSSeed 2.1 Pro Preview
Reasoning11Even, 1–1 of 2
Coding22Even, 2–2 of 4
Agentic14Seed 2.1 Pro Preview leads 4 of 5 · widest: GDPVal 41.8 vs 87.9
Multimodal04Seed 2.1 Pro Preview leads 4 of 4 · widest: MathVista 84.2 vs 90.7

Biggest gaps

GPT-5.5 pulls furthest ahead on

  1. ARC-AGI-285 vs 62.5
  2. Terminal-Bench 2.184.3 vs 71
  3. AA ApexAgents37.7 vs 33.8

Seed 2.1 Pro Preview pulls furthest ahead on

  1. GDPVal87.9 vs 41.8
  2. Humanity's Last Exam55.7 vs 45.8
  3. MCP Atlas83.8 vs 75.3

Every shared benchmark40 · grouped by area

Reasoning 2

Coding 4

Agentic 5

BrowseComp84.486.2
GDPVal41.887.9
MCP Atlas75.383.8

Multimodal 4

BLINK78.381.4
MathVista84.290.7
MMMU-Pro79.982.7
OCRBenchv261.163.2

Other shared benchmarks 25

BabyVision55.973.7
ChartQAPro69.470.9
CyberGym81.870.2
DeepSWE7032.7
DUDE81.782.8
DynaMath75.973.1
ERQA64.572
Finance Agent v1.165.360.7
HLE-Verified (no tool)50.442.9
MathVerse (Vision-Only)84.689.7
MMSIBench (circular)3635.9
NL2Repo50.747
Office QA Pro [Multimodal]69.572.2
OfficeQA Pro54.172.2
RealWorldQA82.286.7
SimpleVQA58.674.5
SuperGPQA72.770.8
Toolathlon55.650.6
WorldVQA34.653
ZeroBench (main)1318
ZeroBench (sub)4149.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare GPT-5.5 withALL PAIRINGS →

Compare Seed 2.1 Pro Preview withALL PAIRINGS →