VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs Qwen3.5 397B A17B

AnthropicvsAlibaba62 shared benchmarks3725 head-to-head
BenchmarkClaude Opus 4.5Qwen3.5 397B A17B
AA Agentic Index59.653.3
AA Intelligence41.934.3
AA-LCR7672.7
AA-Omniscience14-30.8
aime_202695.194.2
arena_elo14731442
arena_text_factuality14581445
Artificial Analysis Coding Index47.848.2
browsecomp3769
BrowseComp (context management)59.278.6
browsecomp_zh62.470.3
C-Eval92.293
CC-OCR76.982
charxiv_rq68.580.8
Claw Eval (pass@3)59.648.1
Claw-Eval Avg76.670.7
coding_arena_elo14951399
critpt4.61.7
DynaMath79.786.3
ERQA46.867.5
gdpval47.335.8
GPQA Diamond8789.3
HLE30.829
HLE (with tools)43.448.3
hmmt_feb_202592.994.8
hmmt_feb_202685.387.9
hmmt_nov_202593.392.7
IFBench5878.8
imo_answer_bench8480.9
LiveCodeBench v684.883.6
longbench_v264.463.2
MLVU81.786.7
mmlu_redux95.694.9
MMLU-Pro9087.8
mmmlu90.888.5
MMMU80.785
MMMU-Pro7479
MMStar73.283.8
multichallenge5967.6
MV-Bench67.277.6
nl2repo43.232.2
OmniScience Accuracy46.630.8
OmniScience Non-Hallucination3917.3
QwenClawBench52.351.8
QwenWebBench15361186
RealWorldQA7783.9
scicode5042
Seal-047.746.9
simpleqa_verified41.826
SimpleVQA69.767.1
SkillsBench Avg545.330
supergpqa70.670.4
SWE-bench Multilingual77.569.3
SWE-bench Pro57.150.9
SWE-bench Verified81.576.4
Terminal-Bench 2.059.352.5
Terminal-Bench Hard4740.9
toolathlon43.538.3
V*6795.8
VideoMME (w sub.)77.787.5
VideoMMMU84.484.7
τ²-Bench Telecom (AA run)89.595.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.