Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3.5 27B vs Seed 1.8

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Qwen3.5 27B scores higher on 16 and Seed 1.8 on 11. The widest gap is DynaMath, where Qwen3.5 27B scores 87.7 against 61.5. Seed 1.8 is the cheaper of the two on tracked API pricing ($0.25 against $0.30 per million input tokens).

ALIBABAVSBYTEDANCE27 SHARED16–11 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAlibabaByteDance
Released—
Price per 1M tokens input / output$0.30 / $2.40$0.25 / $2.00
Cost of 1M in + 1M out$2.70$2.25 1.2× less
Head-to-head of 27 shared benchmarks16 wins11 wins
Scores tracked independently verified121 11 ◆40 0 ◆

Qwen3.5 27B's release date per Artificial Analysis. Prices: Alibaba's own price page for Qwen3.5 27B; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaQwen3.5 27BWINSSeed 1.8
Coding11Even, 1–1 of 2
Agentic02Seed 1.8 leads 2 of 2 · widest: BrowseComp 61 vs 67.6
Instruction Following01Seed 1.8 leads 1 of 1 · widest: Multi-Challenge 60.8 vs 66.7
Multimodal41Qwen3.5 27B leads 4 of 5 · widest: CharXiv (RQ) 79.5 vs 71.4

Biggest gaps

Qwen3.5 27B pulls furthest ahead on

  1. CharXiv (RQ)79.5 vs 71.4

Seed 1.8 pulls furthest ahead on

  1. BrowseComp67.6 vs 61
  2. OSWorld-Verified61.9 vs 56.2
  3. Multi-Challenge66.7 vs 60.8

Every shared benchmark27 · grouped by area

Coding 2

Agentic 2

Instruction Following 1

BenchmarkQwen3.5 27BMARGINSeed 1.8

Multimodal 5

BenchmarkQwen3.5 27BMARGINSeed 1.8
CharXiv (RQ)79.571.4
MathVista87.887.7
MMMU82.383.4
MMMU-Pro7573.2
MMStar8179.9

Other shared benchmarks 17

BenchmarkQwen3.5 27BMARGINSeed 1.8
AI2D92.989.1
BFCLv468.557.2
DynaMath87.761.5
ERQA60.558.8
LVBench73.673
MMLU-Pro86.184.9
MMVU73.373.1
SimpleVQA5665.4
SuperGPQA65.664.8
VideoMMMU82.382.7
WideSearch61.163.8

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (alibaba-official, deepinfra), otherwise the lowest tracked offer.

Compare Qwen3.5 27B withALL PAIRINGS →