VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4.6 vs GLM-5.1

AnthropicvsZ.ai41 shared benchmarks2516 head-to-head
BenchmarkClaude Sonnet 4.6GLM-5.1
AA Agentic Index61.666
AA Intelligence48.441
AA-LCR7468
AA-Omniscience12.20.8
AgentWorldBench - Android5859.1
AgentWorldBench - MCP7067.6
AgentWorldBench - OS63.259.1
AgentWorldBench - Overall5651.3
AgentWorldBench - Search28.822.5
AgentWorldBench - SWE64.552.1
AgentWorldBench - Terminal5747.3
AgentWorldBench - Web50.851.5
arena_elo14721468
arena_text_factuality14601458
Artificial Analysis Coding Index6355.8
browsecomp76.279.3
coding_arena_elo15231510
critpt3.14.6
Finance Agent v25144.8
frontiermath_tier_48.312.5
gdpval54.849.5
GDPval-AA (Elo)16761535
GPQA Diamond89.986.8
HLE4952.3
HLE (with tools)46.852.3
IFBench56.676.3
itbenchSre39.840.3
livebench75.570.2
MCP Atlas69.575.6
OmniScience Accuracy40.925.2
OmniScience Non-Hallucination51.670.1
scicode4743.8
simpleqa_verified2938.1
SWE-bench Pro58.158.4
TauBench V3 - Banking34.413.6
Terminal-Bench 2.059.169
Terminal-Bench 2.171.263.5
Terminal-Bench Hard59.143.2
toolathlon49.440.7
τ²-Bench Telecom (AA run)97.997.7
τ³-Bench30.570.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.