Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3.7 Plus Preview vs Seed2.1

16 SHARED BENCHMARKS

Across 16 shared benchmarks, Qwen3.7 Plus Preview scores higher on 8 and Seed2.1 on 8. The widest gap is GDPVal, where Seed2.1 scores 82.7 against 12.8.

ALIBABAVSX16 SHARED8–8 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAlibabax
Released—
Price per 1M tokens input / output$0.40 / $1.60—
Cost of 1M in + 1M out$2.00—
Head-to-head of 16 shared benchmarks8 wins8 wins
Scores tracked independently verified85 4 ◆50 0 ◆

Qwen3.7 Plus Preview's release date per Artificial Analysis. Prices: Alibaba's own price page for Qwen3.7 Plus Preview. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaQwen3.7 Plus PreviewWINSSeed2.1
Coding02Seed2.1 leads 2 of 2 · widest: SciCode 46.1 vs 57.8
Agentic03Seed2.1 leads 3 of 3 · widest: GDPVal 12.8 vs 82.7
Multimodal21Qwen3.7 Plus Preview leads 2 of 3 · widest: OCRBenchv2 67.1 vs 62.8

Biggest gaps

Qwen3.7 Plus Preview pulls furthest ahead on

  1. OCRBenchv267.1 vs 62.8

Seed2.1 pulls furthest ahead on

  1. GDPVal82.7 vs 12.8
  2. AA ApexAgents29.2 vs 22.4
  3. SciCode57.8 vs 46.1

Every shared benchmark16 · grouped by area

Coding 2

Agentic 3

GDPVal12.882.7
MCP Atlas73.280.3

Multimodal 3

Other shared benchmarks 8

BabyVision70.462.9
ERQA69.871.3
LVBench76.276.8
RealWorldQA86.986.3
SimpleVQA81.771.1
SuperGPQA71.467.4
TVBench78.277.2
WorldVQA61.148.6

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Seed2.1 withALL PAIRINGS →