Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Seed 2.1 Pro Preview vs Seed2.1

47 SHARED BENCHMARKS

Across 47 shared benchmarks, Seed 2.1 Pro Preview scores higher on 42 and Seed2.1 on 3, with 2 level. The widest gap is ZeroBench (main), where Seed 2.1 Pro Preview scores 18 against 11.

BYTEDANCEVSX47 SHARED42–3 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerByteDancex
Released——
Price per 1M tokens input / output——
Cost of 1M in + 1M out——
Head-to-head of 47 shared benchmarks42 wins3 wins
Scores tracked independently verified66 1 ◆50 0 ◆

◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaSeed 2.1 Pro PreviewWINSSeed2.1
Reasoning10Seed 2.1 Pro Preview leads 1 of 1
Coding20Seed 2.1 Pro Preview leads 2 of 2 · widest: Terminal-Bench 2.1 71 vs 67.6
Agentic30Seed 2.1 Pro Preview leads 3 of 3 · widest: AA ApexAgents 33.8 vs 29.2
Multimodal50Seed 2.1 Pro Preview leads 5 of 5

Biggest gaps

Seed 2.1 Pro Preview pulls furthest ahead on

  1. AA ApexAgents33.8 vs 29.2
  2. GDPVal87.9 vs 82.7
  3. Terminal-Bench 2.171 vs 67.6

Seed2.1 pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark47 · grouped by area

Reasoning 1

Coding 2

Agentic 3

GDPVal87.982.7
MCP Atlas83.880.3

Multimodal 5

BLINK81.479.4
MathVista90.790.5
MMMU-Pro82.780.1
OCRBenchv263.262.8
Video-MME89.289

Other shared benchmarks 36

BabyVision73.762.9
BrowseComp (with Search)86.284.9
ChartQAPro70.970.9
CyberGym70.267
DeepSWE32.723
DUDE82.883.1
DynaMath73.168.1
ERQA7271.3
Finance Agent v1.160.756
HLE-Verified (no tool)42.942.4
LVBench7876.8
MathVerse (Vision-Only)89.789.2
Minerva70.765.9
MMSIBench (circular)35.931.4
MotionBench74.974.8
MSQA50.242
Office QA Pro [Multimodal]72.271.1
OfficeQA Pro72.262.8
OVBench7069.7
RealWorldQA86.786.3
SimpleVQA74.571.1
SuperGPQA70.867.4
TOMATO79.556.8
Toolathlon50.649.1
TVBench80.577.2
VideoHolmes68.267.6
WorldVQA5348.6
ZeroBench (main)1811
ZeroBench (sub)49.449.1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Seed 2.1 Pro Preview withALL PAIRINGS →

Compare Seed2.1 withALL PAIRINGS →