VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-5.6 Sol vs Qwen3.7 Max

OpenAIvsAlibaba45 shared benchmarks396 head-to-head
BenchmarkGPT-5.6 SolQwen3.7 Max
AA Agentic Index57.830.9
AA Intelligence6146.7
AA-LCR77.774.7
AA-Omniscience2213.5
Agents' Last Exam (Pass / Score)53.631.1
aime_202699.997
AndroidBench7456.5
arena_elo14821474
arena_text_factuality14741479
Artificial Analysis Coding Index78.366
coding_arena_elo16191517
critpt32.313.4
DeepSWE7318
DeepSWE 1.17321.6
Finance Agent v253.848.4
gdpval61.138.5
GPQA Diamond94.692.4
HealthBench5754.5
HLE49.541.4
HLE (with tools)5853.5
IFBench72.780.5
itbenchSre56.242.5
JobBench45.431.3
livebench_agentic_coding56.243.6
livebench_coding83.974.2
livebench_data_analysis79.871.8
livebench_instruction_following71.874
livebench_language87.779.7
livebench_math96.285.3
livebench_reasoning91.783.3
longbench_v267.165.3
MCP Atlas83.676.4
MLS Bench Lite46.231.7
MRCR v2 256K (8-needle)93.886.7
OmniScience Accuracy59.431.1
OmniScience Non-Hallucination10.674.4
PaperBench90.564.8
scicode56.953.5
simplebench64.870.4
simpleqa_verified71.658.5
SWE-bench Pro64.660.6
TauBench V3 - Banking44.311.8
Terminal-Bench 2.189.575
Terminal-Bench Hard65.950.8
τ²-Bench Telecom (AA run)85.194.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.