Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3 VL 235B A22B Reasoning vs Seed 1.8

22 SHARED BENCHMARKS

Across 22 shared benchmarks, Qwen3 VL 235B A22B Reasoning scores higher on 3 and Seed 1.8 on 19. The widest gap is ZeroBench, where Seed 1.8 scores 11 against 4. Seed 1.8 is the cheaper of the two on tracked API pricing ($0.25 against $0.40 per million input tokens).

ALIBABAVSBYTEDANCE22 SHARED3–19 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAlibabaByteDance
Released—
Price per 1M tokens input / output$0.40 / $4.00$0.25 / $2.00
Cost of 1M in + 1M out$4.40$2.25 2.0× less
Head-to-head of 22 shared benchmarks3 wins19 wins
Scores tracked independently verified91 2 ◆40 0 ◆

Qwen3 VL 235B A22B Reasoning's release date per Artificial Analysis. Prices: Alibaba's own price page for Qwen3 VL 235B A22B Reasoning; deepinfra for Seed 1.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaQwen3 VL 235B A22B ReasoningWINSSeed 1.8
Coding01Seed 1.8 leads 1 of 1 · widest: LiveCodeBench v6 70.1 vs 79.5
Agentic01Seed 1.8 leads 1 of 1 · widest: OSWorld-Verified 38.1 vs 61.9
Multimodal06Seed 1.8 leads 6 of 6 · widest: BLINK 67.1 vs 74.3

Biggest gaps

Qwen3 VL 235B A22B Reasoning pulls furthest ahead on

No ratified-area lead of 3 points or more.

Seed 1.8 pulls furthest ahead on

  1. OSWorld-Verified61.9 vs 38.1
  2. LiveCodeBench v679.5 vs 70.1
  3. BLINK74.3 vs 67.1

Every shared benchmark22 · grouped by area

Multimodal 6

BLINK67.174.3
CharXiv (RQ)66.171.4
MathVista85.887.7
MMMU78.783.4
MMMU-Pro68.773.2
MMStar78.779.9

Other shared benchmarks 14

AI2D89.289.1
AIME 202589.794.3
ERQA52.558.8
LVBench63.673
MATH-Vision74.681.3
MMLU-Pro83.884.9
MMVU71.173.1
SimpleVQA61.365.4
SuperGPQA64.364.8
VideoMMMU8082.7

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (alibaba-official, deepinfra), otherwise the lowest tracked offer.