Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Gemini 3.1 Pro vs Seed 2.1 Pro Preview

50 SHARED BENCHMARKS

Across 50 shared benchmarks, Gemini 3.1 Pro scores higher on 6 and Seed 2.1 Pro Preview on 44. The widest gap is GDPVal, where Seed 2.1 Pro Preview scores 87.9 against 13.8.

GOOGLEVSBYTEDANCE50 SHARED6–44 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerGoogleByteDance
Released—
Price per 1M tokens input / output$2.00 / $12.00—
Cost of 1M in + 1M out$14.00—
Head-to-head of 50 shared benchmarks6 wins44 wins
Scores tracked independently verified215 30 ◆66 1 ◆

Gemini 3.1 Pro's release date per Artificial Analysis. Prices: Artificial Analysis for Gemini 3.1 Pro. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGemini 3.1 ProWINSSeed 2.1 Pro Preview
Reasoning11Even, 1–1 of 2
Coding13Seed 2.1 Pro Preview leads 3 of 4 · widest: SWE-bench Pro 54.2 vs 57.5
Agentic05Seed 2.1 Pro Preview leads 5 of 5 · widest: GDPVal 13.8 vs 87.9
Multimodal06Seed 2.1 Pro Preview leads 6 of 6 · widest: CharXiv (RQ) 83.3 vs 86.4

Biggest gaps

Gemini 3.1 Pro pulls furthest ahead on

  1. ARC-AGI-277.1 vs 62.5

Seed 2.1 Pro Preview pulls furthest ahead on

  1. GDPVal87.9 vs 13.8
  2. MCP Atlas83.8 vs 69.2
  3. Humanity's Last Exam55.7 vs 47

Every shared benchmark50 · grouped by area

Coding 4

Agentic 5

Multimodal 6

BLINK79.181.4
CharXiv (RQ)83.386.4
MathVista90.290.7
MMMU-Pro82.482.7
OCRBenchv262.863.2
Video-MME86.789.2

Other shared benchmarks 33

BabyVision54.473.7
ChartQAPro70.270.9
CyberGym38.870.2
DeepSWE1032.7
DUDE82.182.8
DynaMath72.173.1
ERQA70.872
Finance Agent v1.159.760.7
LVBench66.278
MathVerse (Vision-Only)87.789.7
MATH-Vision89.894.5
Minerva63.570.7
MMSIBench (circular)27.435.9
MotionBench69.974.9
NL2Repo33.447
Office QA Pro [Multimodal]72.572.2
OfficeQA Pro72.572.2
OVBench58.870
RealWorldQA85.486.7
SimpleVQA69.974.5
TOMATO60.479.5
Toolathlon48.850.6
TVBench7180.5
VideoHolmes65.968.2
WorldVQA44.353
ZeroBench (main)1218
ZeroBench (sub)41.949.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Gemini 3.1 Pro withALL PAIRINGS →

Compare Seed 2.1 Pro Preview withALL PAIRINGS →