VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude Opus 4.7 or Qwen3.5 122B A10B?
Across 31 shared benchmarks, Claude Opus 4.7 scores higher on 26 and Qwen3.5 122B A10B on 5. The widest gap is OmniScience Non-Hallucination, where Claude Opus 4.7 scores 57.7 against 12.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $5.00 per million input tokens).

Claude Opus 4.7 vs Qwen3.5 122B A10B

Across 31 shared benchmarks, Claude Opus 4.7 scores higher on 26 and Qwen3.5 122B A10B on 5. The widest gap is OmniScience Non-Hallucination, where Claude Opus 4.7 scores 57.7 against 12.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $5.00 per million input tokens).

AnthropicvsAlibaba31 shared benchmarks265 head-to-head
BenchmarkClaude Opus 4.7Qwen3.5 122B A10B
AA Agentic Index64.621.3
AA Intelligence5532.8
AA-LCR75.370.3
AA-Omniscience27.3-41.5
arena_vision13161246
Artificial Analysis Coding Index73.645.7
browsecomp79.863.8
coding_arena_elo15571358
critpt120.9
gdpval58.624.3
GPQA Diamond94.286.6
HLE54.747.5
IFBench58.676.1
mmmlu91.586.7
MMMU-Pro78.876.9
OmniScience Accuracy48.924.4
OmniScience Non-Hallucination57.712.9
OSWorld-Verified82.858
scicode54.542
screenspot_pro_no_tools79.570.4
SWE-bench Verified87.672
TauBench V3 - Banking34.615.3
Terminal-Bench 2.069.449.4
Terminal-Bench 2.183.147.6
Terminal-Bench Hard54.531.1
vectara_answer_rate9899.8
vectara_avg_summary_length149.186.4
vectara_factual_consistency8888.8
vectara_hallucination_rate1211.2
τ²-Bench Telecom (AA run)88.693.6
τ³-Bench28.913.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.