Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3.5 122B A10B vs Seed 1.8

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Qwen3.5 122B A10B scores higher on 16 and Seed 1.8 on 11. The widest gap is DynaMath, where Qwen3.5 122B A10B scores 85.9 against 61.5. Seed 1.8 is the cheaper of the two on tracked API pricing ($0.25 against $0.40 per million input tokens).

ALIBABAVSBYTEDANCE27 SHARED16–11 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAlibabaByteDance
Released—
Price per 1M tokens input / output$0.40 / $3.20$0.25 / $2.00
Cost of 1M in + 1M out$3.60$2.25 1.6× less
Head-to-head of 27 shared benchmarks16 wins11 wins
Scores tracked independently verified98 7 ◆40 0 ◆

Qwen3.5 122B A10B's release date per Artificial Analysis. Prices: Alibaba's own price page for Qwen3.5 122B A10B; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaQwen3.5 122B A10BWINSSeed 1.8
Coding02Seed 1.8 leads 2 of 2
Agentic02Seed 1.8 leads 2 of 2 · widest: OSWorld-Verified 58 vs 61.9
Instruction Following01Seed 1.8 leads 1 of 1 · widest: Multi-Challenge 61.5 vs 66.7
Multimodal41Qwen3.5 122B A10B leads 4 of 5 · widest: CharXiv (RQ) 77.2 vs 71.4

Biggest gaps

Qwen3.5 122B A10B pulls furthest ahead on

  1. CharXiv (RQ)77.2 vs 71.4
  2. MMStar82.9 vs 79.9

Seed 1.8 pulls furthest ahead on

  1. Multi-Challenge66.7 vs 61.5
  2. OSWorld-Verified61.9 vs 58
  3. BrowseComp67.6 vs 63.8

Every shared benchmark27 · grouped by area

Agentic 2

Instruction Following 1

Multimodal 5

CharXiv (RQ)77.271.4
MathVista87.487.7
MMMU83.983.4
MMMU-Pro7573.2
MMStar82.979.9

Other shared benchmarks 17

AI2D93.389.1
BFCLv472.257.2
DynaMath85.961.5
ERQA6258.8
LVBench74.473
MATH-Vision86.281.3
MMLU-Pro86.784.9
MMVU74.773.1
SimpleVQA61.765.4
SuperGPQA67.164.8
VideoMMMU8282.7
WideSearch60.563.8

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (alibaba-official, deepinfra), otherwise the lowest tracked offer.