VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.7 vs GLM-5.1

AnthropicvsZ.ai38 shared benchmarks325 head-to-head
BenchmarkClaude Opus 4.7GLM-5.1
AA Agentic Index64.666
AA Intelligence5541
AA-LCR75.368
AA-Omniscience27.30.8
aime_202695.895.8
arena_elo14941468
arena_text_factuality14741458
Artificial Analysis Coding Index73.655.8
browsecomp79.879.3
coding_arena_elo15571510
critpt124.6
cybergym73.868.7
Finance Agent v251.544.8
frontiermath_tier_422.912.5
gdpval58.649.5
GDPval-AA (Elo)17531535
GPQA Diamond94.286.8
HLE54.752.3
HLE (with tools)54.752.3
hmmt_feb_202693.989.4
IFBench58.676.3
itbenchSre46.740.3
livebench76.970.2
MCP Atlas79.175.6
OmniScience Accuracy48.925.2
OmniScience Non-Hallucination57.770.1
scicode54.543.8
simplebench61.755.1
simpleqa_verified50.638.1
SWE-bench Multilingual80.573.3
SWE-bench Pro64.358.4
TauBench V3 - Banking34.613.6
Terminal-Bench 2.069.469
Terminal-Bench 2.183.163.5
Terminal-Bench Hard54.543.2
toolathlon59.340.7
τ²-Bench Telecom (AA run)88.697.7
τ³-Bench28.970.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.