Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.5 vs Seed2.1

33 SHARED BENCHMARKS

Across 33 shared benchmarks, GPT-5.5 scores higher on 13 and Seed2.1 on 20. The widest gap is DeepSWE, where GPT-5.5 scores 70 against 23.

OPENAIVSX33 SHARED13–20 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIx
Released—
Price per 1M tokens input / output$5.00 / $30.00—
Cost of 1M in + 1M out$35.00—
Head-to-head of 33 shared benchmarks13 wins20 wins
Scores tracked independently verified185 29 ◆50 0 ◆

GPT-5.5's release date per Artificial Analysis. Prices: Artificial Analysis for GPT-5.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.5WINSSeed2.1
Reasoning10GPT-5.5 leads 1 of 1 · widest: ARC-AGI-2 85 vs 61.3
Coding11Even, 1–1 of 2
Agentic12Seed2.1 leads 2 of 3 · widest: GDPVal 41.8 vs 82.7
Multimodal04Seed2.1 leads 4 of 4 · widest: MathVista 84.2 vs 90.5

Biggest gaps

GPT-5.5 pulls furthest ahead on

  1. ARC-AGI-285 vs 61.3
  2. AA ApexAgents37.7 vs 29.2
  3. Terminal-Bench 2.184.3 vs 67.6

Seed2.1 pulls furthest ahead on

  1. GDPVal82.7 vs 41.8
  2. MathVista90.5 vs 84.2
  3. MCP Atlas80.3 vs 75.3

Every shared benchmark33 · grouped by area

Reasoning 1

BenchmarkGPT-5.5MARGINSeed2.1
ARC-AGI-28561.3

Coding 2

BenchmarkGPT-5.5MARGINSeed2.1
SciCode55.857.8

Agentic 3

BenchmarkGPT-5.5MARGINSeed2.1
GDPVal41.882.7
MCP Atlas75.380.3

Multimodal 4

BenchmarkGPT-5.5MARGINSeed2.1
BLINK78.379.4
MathVista84.290.5
MMMU-Pro79.980.1
OCRBenchv261.162.8

Other shared benchmarks 23

BenchmarkGPT-5.5MARGINSeed2.1
BabyVision55.962.9
ChartQAPro69.470.9
CyberGym81.867
DUDE81.783.1
DynaMath75.968.1
ERQA64.571.3
Finance Agent v1.165.356
HLE-Verified (no tool)50.442.4
MathVerse (Vision-Only)84.689.2
MMSIBench (circular)3631.4
Office QA Pro [Multimodal]69.571.1
OfficeQA Pro54.162.8
RealWorldQA82.286.3
SimpleVQA58.671.1
SuperGPQA72.767.4
SWE-Pro Bench58.657
Toolathlon55.649.1
VisuLogic43.552.9
WorldVQA34.648.6
ZeroBench (main)1311
ZeroBench (sub)4149.1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare GPT-5.5 withALL PAIRINGS →

Compare Seed2.1 withALL PAIRINGS →