GPT-5.5 vs Seed2.1
33 SHARED BENCHMARKSAcross 33 shared benchmarks, GPT-5.5 scores higher on 13 and Seed2.1 on 20. The widest gap is DeepSWE, where GPT-5.5 scores 70 against 23.
OPENAIVSX33 SHARED13–20 HEAD-TO-HEADNEWEST SCORE ADDED
At a glance
MakerOpenAIx
Released—
Price per 1M tokens input / output$5.00 / $30.00—
Cost of 1M in + 1M out$35.00—
Head-to-head of 33 shared benchmarks13 wins20 wins
Scores tracked independently verified185 29 ◆50 0 ◆
GPT-5.5's release date per Artificial Analysis. Prices: Artificial Analysis for GPT-5.5. ◆ = independently verified score.
Where each leadsby capability area · benchmark wins
AreaGPT-5.5WINSSeed2.1
Biggest gaps
GPT-5.5 pulls furthest ahead on
- ARC-AGI-285 vs 61.3
- AA ApexAgents37.7 vs 29.2
- Terminal-Bench 2.184.3 vs 67.6
Every shared benchmark33 · grouped by area
Coding 2
Agentic 3
Multimodal 4
Other shared benchmarks 23
Finance Agent v1.165.356
HLE-Verified (no tool)50.442.4
MathVerse (Vision-Only)84.689.2
MMSIBench (circular)3631.4
Office QA Pro [Multimodal]69.571.1
SWE-Pro Bench58.657
ZeroBench (main)1311
ZeroBench (sub)4149.1
Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.