VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs Qwen3.6 Plus

AnthropicvsAlibaba60 shared benchmarks3525 head-to-head
BenchmarkClaude Opus 4.5Qwen3.6 Plus
AA Agentic Index59.629
AA Intelligence41.940.5
AA-LCR7672.3
AA-Omniscience142.6
aime_202695.195.3
arena_elo14731444
arena_text_factuality14581438
Artificial Analysis Coding Index47.854.5
C-Eval92.293.3
Claw Eval (pass@3)59.658.7
coding_arena_elo14951460
critpt4.62.9
ERQA46.865.7
frontiermath_tier_44.28.3
gdpval47.332
GPQA Diamond8790.4
HLE30.828.8
HLE (with tools)43.450.6
hmmt_feb_202685.387.8
hmmt_nov_202593.394.6
IFBench5875.2
imo_answer_bench8483.8
livebench7670.8
livebench_agentic_coding39.741.4
livebench_coding79.778.2
livebench_data_analysis74.469.9
livebench_instruction_following62.558.3
livebench_language81.375
livebench_math90.483.7
livebench_reasoning80.175.8
LiveCodeBench v684.887.1
longbench_v264.462
mathvision77.188
MCP Atlas69.874.1
MLVU81.786.7
mmlu_redux95.694.5
MMLU-Pro9088.5
mmmlu90.889.5
MMMU80.786
MMMU-Pro7478.8
MMStar73.283.3
nl2repo43.237.9
OmniDocBench 1.587.791.2
OmniScience Accuracy46.626.4
OmniScience Non-Hallucination3968
OSWorld-Verified66.362.5
RealWorldQA7785.4
scicode5040.7
simpleqa_verified41.849.1
SimpleVQA69.70.7
supergpqa70.671.6
SWE-bench Multilingual77.573.8
SWE-bench Pro57.156.6
SWE-bench Verified81.578.8
Terminal-Bench 2.059.361.6
Terminal-Bench Hard4743.9
tool_decathlon43.539.8
toolathlon43.539.8
VideoMMMU84.484
τ²-Bench Telecom (AA run)89.597.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.