VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude Opus 4.8 or Qwen3.5 122B A10B?
Across 31 shared benchmarks, Claude Opus 4.8 scores higher on 30 and Qwen3.5 122B A10B on 1. The widest gap is OmniScience Non-Hallucination, where Claude Opus 4.8 scores 60.7 against 12.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $5.00 per million input tokens).

Claude Opus 4.8 vs Qwen3.5 122B A10B

Across 31 shared benchmarks, Claude Opus 4.8 scores higher on 30 and Qwen3.5 122B A10B on 1. The widest gap is OmniScience Non-Hallucination, where Claude Opus 4.8 scores 60.7 against 12.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $5.00 per million input tokens).

AnthropicvsAlibaba31 shared benchmarks301 head-to-head
BenchmarkClaude Opus 4.8Qwen3.5 122B A10B
AA Agentic Index49.421.3
AA Intelligence57.332.8
AA-LCR7370.3
AA-Omniscience28.8-41.5
arena_vision12871246
Artificial Analysis Coding Index74.345.7
browsecomp88.563.8
coding_arena_elo15391358
critpt20.90.9
gdpval54.124.3
GPQA Diamond93.686.6
HLE57.947.5
IFBench62.276.1
include87.682.8
longbench_v269.160.2
mathvision86.786.2
MMMU-Pro78.976.9
MMVU79.274.7
OmniScience Accuracy48.824.4
OmniScience Non-Hallucination60.712.9
OSWorld-Verified83.458
scicode53.542
screenspot_pro_no_tools87.970.4
SWE-bench Verified88.672
TauBench V3 - Banking34.215.3
Terminal-Bench 2.074.649.4
Terminal-Bench 2.18547.6
Terminal-Bench Hard58.331.1
WideSearch72.960.5
τ²-Bench Telecom (AA run)94.493.6
τ³-Bench27.613.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.