Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3.5 35B A3B vs Seed 1.8

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Qwen3.5 35B A3B scores higher on 10 and Seed 1.8 on 17. The widest gap is DynaMath, where Qwen3.5 35B A3B scores 85 against 61.5.

ALIBABAVSBYTEDANCE27 SHARED10–17 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAlibabaByteDance
Released—
Price per 1M tokens input / output$0.25 / $2.00$0.25 / $2.00
Cost of 1M in + 1M out$2.25$2.25
Head-to-head of 27 shared benchmarks10 wins17 wins
Scores tracked independently verified148 12 ◆40 0 ◆

Qwen3.5 35B A3B's release date per Artificial Analysis. Prices: Alibaba's own price page for Qwen3.5 35B A3B; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaQwen3.5 35B A3BWINSSeed 1.8
Coding02Seed 1.8 leads 2 of 2 · widest: LiveCodeBench v6 74.6 vs 79.5
Agentic02Seed 1.8 leads 2 of 2 · widest: OSWorld-Verified 54.5 vs 61.9
Instruction Following01Seed 1.8 leads 1 of 1 · widest: Multi-Challenge 60 vs 66.7
Multimodal23Seed 1.8 leads 3 of 5

Biggest gaps

Qwen3.5 35B A3B pulls furthest ahead on

  1. CharXiv (RQ)77.5 vs 71.4

Seed 1.8 pulls furthest ahead on

  1. OSWorld-Verified61.9 vs 54.5
  2. Multi-Challenge66.7 vs 60
  3. BrowseComp67.6 vs 61

Every shared benchmark27 · grouped by area

Coding 2

Agentic 2

Instruction Following 1

Multimodal 5

CharXiv (RQ)77.571.4
MathVista86.287.7
MMMU81.483.4
MMMU-Pro72.773.2
MMStar81.979.9

Other shared benchmarks 17

AI2D92.689.1
BFCLv467.357.2
DynaMath8561.5
ERQA64.858.8
LVBench71.473
MATH-Vision83.981.3
MMLU-Pro85.384.9
MMVU72.373.1
SimpleVQA58.365.4
SuperGPQA63.464.8
VideoMMMU80.482.7
WideSearch57.163.8

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Qwen3.5 35B A3B withALL PAIRINGS →