VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GLM-5.2 Full Open Source vs GPT-5.4

Z.aivsOpenAI46 shared benchmarks1431 head-to-head
BenchmarkGLM-5.2 Full Open SourceGPT-5.4
AA Agentic Index45.758.2
AA Intelligence52.653.1
AA-LCR76.777.7
AA-Omniscience4.45.8
aime_202699.299.2
apexAgents33.733.3
arc_agi_17793.7
ARC-AGI-222.874
arena_elo14711477
arena_text_factuality14601476
Artificial Analysis Coding Index68.871.1
coding_arena_elo15821457
critpt20.923.4
DeepSWE 1.14452
gdpval50.350.1
GPQA Diamond91.293
HiL-Bench43.79.7
HLE54.743.7
HLE (with tools)54.752.1
hmmt_feb_202692.597.7
hmmt_nov_202594.495.8
IFBench73.373.9
imo_answer_bench9191.4
itbenchSre42.734.5
livebench_agentic_coding51.853.8
livebench_coding79.777.5
livebench_data_analysis73.779.3
livebench_instruction_following62.370.2
livebench_language76.282.6
livebench_math89.894.2
livebench_reasoning78.688.1
MCP Atlas82.670.6
nl2repo48.941.3
officeqa_pro41.451.1
OmniScience Accuracy24.350.9
OmniScience Non-Hallucination73.717.4
scicode50.556.6
simpleqa_verified38.145.3
SWE-bench Pro62.159.1
TauBench V3 - Banking34.639.6
Terminal-Bench 2.182.778.3
Terminal-Bench Hard50.857.6
tool_decathlon48.254.6
toolathlon48.254.6
τ²-Bench Telecom (AA run)99.187.1
τ³-Bench26.872.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.