VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Qwen3.5 397B A17B

AnthropicvsAlibaba53 shared benchmarks4310 head-to-head
BenchmarkClaude Opus 4.6Qwen3.5 397B A17B
AA Agentic Index67.653.3
AA Intelligence44.934.3
AA-LCR74.372.7
AA-Omniscience13.7-30.8
AgentWorldBench - Android61.754.9
AgentWorldBench - MCP69.968.3
AgentWorldBench - OS70.260.9
AgentWorldBench - Overall57.854.7
AgentWorldBench - Search29.330.8
AgentWorldBench - SWE64.564.4
AgentWorldBench - Terminal57.555.3
AgentWorldBench - Web51.448.5
aime_202696.794.2
apexAgents3315.3
arena_elo14971442
arena_text_factuality14871445
arena_vision13001249
Artificial Analysis Coding Index48.148.2
browsecomp86.869
charxiv_rq69.180.8
coding_arena_elo15451399
critpt12.61.7
ERQA40.867.5
gdpval55.935.8
GPQA Diamond91.389.3
GSM8K0.666.7
HLE53.129
HLE (with tools)62.748.3
hmmt_feb_202696.287.9
hmmt_nov_202596.392.7
IFBench62.578.8
imo_answer_bench75.380.9
LiveCodeBench v688.883.6
mcpmark56.746.1
MMLU-Pro89.187.8
mmmlu91.188.5
MMMU-Pro77.379
multichallenge5667.6
nl2repo49.832.2
OmniScience Accuracy4730.8
OmniScience Non-Hallucination37.217.3
RealWorldQA73.983.9
scicode5242
simpleqa_verified46.526
SWE-bench Multilingual77.869.3
SWE-bench Pro57.350.9
SWE-bench Verified80.876.4
Terminal-Bench 2.065.452.5
Terminal-Bench Hard65.440.9
toolathlon56.838.3
usamo_202666.236.3
τ²-Bench Telecom (AA run)99.395.6
τ³-Bench72.413.4

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.