VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GLM-5.1 vs GLM-5.2 Full Open Source

Z.aivsZ.ai39 shared benchmarks434 head-to-head
BenchmarkGLM-5.1GLM-5.2 Full Open Source
AA Agentic Index6645.7
AA Intelligence4152.6
AA-LCR6876.7
AA-Omniscience0.84.4
aime_202695.899.2
arena_elo14681471
arena_text_factuality14581460
Artificial Analysis Coding Index55.868.8
coding_arena_elo15101582
critpt4.620.9
DeepSWE1846.2
gdpval49.550.3
GPQA Diamond86.891.2
HLE52.354.7
HLE (with tools)52.354.7
HMMT 20259494.4
hmmt_feb_202689.492.5
hmmt_nov_20259494.4
IFBench76.373.3
imo_answer_bench83.891
itbenchSre40.342.7
MCP Atlas75.682.6
nl2repo42.748.9
OmniScience Accuracy25.224.3
OmniScience Non-Hallucination70.173.7
PostTrainBench20.134.3
Program Bench50.963.7
scicode43.850.5
simplebench55.158.8
simpleqa_verified38.138.1
SWE-bench Pro58.462.1
SWE-Marathon113
TauBench V3 - Banking13.634.6
Terminal-Bench 2.163.582.7
Terminal-Bench Hard43.250.8
tool_decathlon40.748.2
toolathlon40.748.2
τ²-Bench Telecom (AA run)97.799.1
τ³-Bench70.626.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.