VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.8 vs Qwen3.6 Plus

AnthropicvsAlibaba61 shared benchmarks5110 head-to-head
BenchmarkClaude Opus 4.8Qwen3.6 Plus
AA Agentic Index49.429
AA Intelligence57.340.5
AA-LCR7372.3
AA-Omniscience28.82.6
AgentWorldBench - Android61.557.6
AgentWorldBench - MCP54.955.3
AgentWorldBench - OS66.660.3
AgentWorldBench - Overall56.650.8
AgentWorldBench - Search35.121.9
AgentWorldBench - SWE64.159.1
AgentWorldBench - Terminal59.250.6
AgentWorldBench - Web54.750.8
aime_202610095.3
arena_elo14821444
arena_text_factuality14641438
Artificial Analysis Coding Index74.354.5
coding_arena_elo15391460
critpt20.92.9
Finance Agent v253.940.9
frontiermath_tier_431.38.3
gdpval54.232
GPQA Diamond93.690.4
HLE57.928.8
HLE (with tools)57.950.6
hmmt_feb_202696.787.8
hmmt_nov_202596.594.6
IFBench62.275.2
imo_answer_bench83.583.8
include87.685.1
livebench77.270.8
livebench_agentic_coding50.541.4
livebench_coding81.878.2
livebench_data_analysis6669.9
livebench_instruction_following7258.3
livebench_language79.775
livebench_math94.383.7
livebench_reasoning89.275.8
longbench_v269.162
mathvision86.788
MCP Atlas83.674.1
MMMU-Pro78.978.8
nl2repo69.737.9
OmniScience Accuracy48.826.4
OmniScience Non-Hallucination60.768
OSWorld-Verified83.462.5
scicode53.540.7
screenspot_pro_no_tools87.968.2
simpleqa_verified39.549.1
SWE-bench Multilingual84.473.8
SWE-bench Pro69.256.6
SWE-bench Verified88.678.8
TauBench V3 - Banking34.220.8
Terminal-Bench 2.074.661.6
Terminal-Bench 2.18561.4
Terminal-Bench Hard58.343.9
tool_decathlon59.939.8
toolathlon59.939.8
Video-MME8684.2
WideSearch72.974.3
τ²-Bench Telecom (AA run)94.497.7
τ³-Bench27.670.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.