VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude 4 Opus or Qwen3.5 122B A10B?
Across 26 shared benchmarks, Claude 4 Opus scores higher on 4 and Qwen3.5 122B A10B on 21, with 1 level. The widest gap is HMMT 2025, where Qwen3.5 122B A10B scores 90.3 against 15.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $15.00 per million input tokens).

Claude 4 Opus vs Qwen3.5 122B A10B

Across 26 shared benchmarks, Claude 4 Opus scores higher on 4 and Qwen3.5 122B A10B on 21, with 1 level. The widest gap is HMMT 2025, where Qwen3.5 122B A10B scores 90.3 against 15.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $15.00 per million input tokens).

AnthropicvsAlibaba26 shared benchmarks421 head-to-head
BenchmarkClaude 4 OpusQwen3.5 122B A10B
AA Intelligence31.732.8
AA-LCR4070.3
arena_vision12071246
Artificial Analysis Coding Index3445.7
GPQA Diamond79.686.6
HLE12.347.5
HMMT 202515.990.3
IFBench53.776.1
ifeval87.493.4
LiveCodeBench v647.478.9
longbench_v255.660.2
mmlu_redux94.294
MMLU-Pro86.686.7
mmmlu88.886.7
MMMU73.783.9
multichallenge58.661.5
OJBench19.639.5
scicode40.942
supergpqa56.567.1
SWE-bench Verified79.472
Terminal-Bench Hard31.131.1
vectara_answer_rate9199.8
vectara_avg_summary_length123.286.4
vectara_factual_consistency8888.8
vectara_hallucination_rate1211.2
τ²-Bench Telecom (AA run)73.493.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.