Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Gemini 3.1 Pro vs Seed 1.8

24 SHARED BENCHMARKS

Across 24 shared benchmarks, Gemini 3.1 Pro scores higher on 18 and Seed 1.8 on 6. The widest gap is Terminal-Bench 2.0, where Gemini 3.1 Pro scores 68.5 against 45.2. Seed 1.8 is the cheaper of the two on tracked API pricing ($0.25 against $2.00 per million input tokens).

GOOGLEVSBYTEDANCE24 SHARED18–6 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerGoogleByteDance
Released—
Price per 1M tokens input / output$2.00 / $12.00$0.25 / $2.00
Cost of 1M in + 1M out$14.00$2.25 6.2× less
Head-to-head of 24 shared benchmarks18 wins6 wins
Scores tracked independently verified215 30 ◆40 0 ◆

Gemini 3.1 Pro's release date per Artificial Analysis. Prices: Artificial Analysis for Gemini 3.1 Pro; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGemini 3.1 ProWINSSeed 1.8
Coding20Gemini 3.1 Pro leads 2 of 2 · widest: LiveCodeBench v6 91.7 vs 79.5
Agentic20Gemini 3.1 Pro leads 2 of 2 · widest: BrowseComp 85.9 vs 67.6
Instruction Following10Gemini 3.1 Pro leads 1 of 1 · widest: Multi-Challenge 71.4 vs 66.7
Multimodal40Gemini 3.1 Pro leads 4 of 4 · widest: CharXiv (RQ) 83.3 vs 71.4

Biggest gaps

Gemini 3.1 Pro pulls furthest ahead on

  1. BrowseComp85.9 vs 67.6
  2. OSWorld-Verified76.2 vs 61.9
  3. CharXiv (RQ)83.3 vs 71.4

Seed 1.8 pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark24 · grouped by area

Coding 2

Agentic 2

Instruction Following 1

BenchmarkGemini 3.1 ProMARGINSeed 1.8
Multi-Challenge71.4 ◆66.7

Multimodal 4

BenchmarkGemini 3.1 ProMARGINSeed 1.8
BLINK79.174.3
CharXiv (RQ)83.371.4
MathVista90.287.7
MMMU-Pro82.473.2

Other shared benchmarks 15

BenchmarkGemini 3.1 ProMARGINSeed 1.8
DUDE82.169.4
DynaMath72.161.5
ERQA70.858.8
LVBench66.273
MATH-Vision89.881.3
Minerva63.562.4
MMLU-Pro9184.9
MotionBench69.970.6
OVBench58.865.1
SimpleVQA69.965.4
TOMATO60.460.8
TVBench7171.5

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, deepinfra), otherwise the lowest tracked offer.

Compare Gemini 3.1 Pro withALL PAIRINGS →