VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs Qwen3.7 Plus Preview

AnthropicvsAlibaba43 shared benchmarks1527 head-to-head
BenchmarkClaude Opus 4.5Qwen3.7 Plus Preview
AA Agentic Index59.620.8
AA Intelligence41.939.4
AA-LCR7669
AA-Omniscience142.4
arena_elo14731458
arena_text_factuality14581454
Artificial Analysis Coding Index47.855.9
charxiv_rq68.585.9
Claw Eval (pass@3)59.662.7
critpt4.69.1
ERQA46.869.8
gdpval47.322.2
GPQA Diamond8790.3
HLE30.835.6
hmmt_feb_202685.392.9
IFBench5879.1
imo_answer_bench8486
LiveCodeBench v684.889.6
mathvision77.190.3
MCP Atlas69.873.2
MLVU81.787.4
mmlu_redux95.694.5
MMLU-Pro9088.5
mmmlu90.889
MMMU-Pro7480.5
nl2repo43.241.1
OmniDocBench 1.587.791.4
OmniScience Accuracy46.622.5
OmniScience Non-Hallucination3974.5
OSWorld-Verified66.373.3
QwenClawBench52.361.8
RealWorldQA7786.9
scicode5051.3
SimpleVQA69.70.8
supergpqa70.671.4
SWE-bench Multilingual77.575.8
SWE-bench Pro57.157.6
SWE-bench Verified81.577.7
Terminal-Bench 2.059.370.3
Terminal-Bench Hard4747
VideoMMMU84.485.4
WorldVQA36.861.1
τ²-Bench Telecom (AA run)89.593

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.