VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.8 vs GLM-5.2 Full Open Source

AnthropicvsZ.ai63 shared benchmarks5210 head-to-head
BenchmarkClaude Opus 4.8GLM-5.2 Full Open Source
AA Agentic Index49.445.7
AA Intelligence57.352.6
AA-LCR7376.7
AA-Omniscience28.84.4
Agents' Last Exam2723.8
aime_202610099.2
apex_agents39.435.6
arc_agi_192.577
ARC-AGI-272.122.8
arena_elo14821471
arena_text_factuality14641460
Artificial Analysis Coding Index74.368.8
AutomationBench (Public)27.212.9
AutomationBench Public27.212.9
coding_arena_elo15391582
critpt20.920.9
DeepSWE5946.2
DeepSWE 1.15944
DSBench-FullStack71.661.8
DSBench-Hard71.754.5
gdpval54.250.3
GDPval-AA v215931514
GDPval-AA v2 (Elo)15931510
GPQA Diamond93.691.2
HiL-Bench35.343.7
HLE57.954.7
HLE (with tools)57.954.7
HLE (wo / w tools)57.954.7
hmmt_feb_202696.792.5
hmmt_nov_202596.594.4
IFBench62.273.3
imo_answer_bench83.591
JobBench48.443.4
livebench_agentic_coding50.551.8
livebench_coding81.879.7
livebench_data_analysis6673.7
livebench_instruction_following7262.3
livebench_language79.776.2
livebench_math94.389.8
livebench_reasoning89.278.6
MCP Atlas83.682.6
MLS Bench Lite42.840.4
nl2repo69.748.9
officeqa_pro66.241.4
OmniScience Accuracy48.824.3
OmniScience Non-Hallucination60.773.7
PostTrainBench37.234.3
Program Bench71.963.7
ResearchRubrics73.571.1
scicode53.550.5
simplebench64.858.8
simpleqa_verified39.538.1
SpreadsheetBench 231.628.1
SWE-bench Pro69.262.1
SWE-Marathon4013
Tau 3 Banking27.626.8
TauBench V3 - Banking34.234.6
Terminal-Bench 2.18582.7
Terminal-Bench Hard58.350.8
tool_decathlon59.948.2
toolathlon59.948.2
τ²-Bench Telecom (AA run)94.499.1
τ³-Bench27.626.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.