VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GLM-5.1 vs GPT-5.4

Z.aivsOpenAI54 shared benchmarks1341 head-to-head
BenchmarkGLM-5.1GPT-5.4
AA Agentic Index6658.2
AA Intelligence4153.1
AA-LCR6877.7
AA-Omniscience0.85.8
AgentWorldBench - Android59.160
AgentWorldBench - MCP67.670.1
AgentWorldBench - OS59.168.6
AgentWorldBench - Overall51.358.3
AgentWorldBench - Search22.537.3
AgentWorldBench - SWE52.166.3
AgentWorldBench - Terminal47.353.7
AgentWorldBench - Web51.551.8
aime_202695.899.2
Apex (Pass@1)11.554.1
Apex Shortlist (Pass@1)72.478.1
arena_elo14681477
arena_text_factuality14581476
Artificial Analysis Coding Index55.871.1
browsecomp79.382.7
browsecomp_with_context_manager79.382.7
coding_arena_elo15101457
critpt4.623.4
cybergym68.766.3
frontiermath_tier_412.527.1
gdpval49.550.1
GDPval-AA (Elo)15351674
GPQA Diamond86.893
HLE52.343.7
HLE (with tools)52.352.1
hmmt_feb_202689.497.7
hmmt_nov_20259495.8
IFBench76.373.9
imo_answer_bench83.891.4
itbenchSre40.334.5
livebench70.280.3
MCP Atlas75.670.6
MCPAtlas Public (Pass@1)71.867.2
MMLU-Pro8687.5
nl2repo42.741.3
OmniScience Accuracy25.250.9
OmniScience Non-Hallucination70.117.4
scicode43.856.6
simpleqa_verified38.145.3
SWE-bench Multilingual73.371.7
SWE-bench Pro58.459.1
SWE-QA72.781.3
TauBench V3 - Banking13.639.6
Terminal-Bench 2.06975.1
Terminal-Bench 2.163.578.3
Terminal-Bench Hard43.257.6
tool_decathlon40.754.6
toolathlon40.754.6
τ²-Bench Telecom (AA run)97.787.1
τ³-Bench70.672.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.