GLM-5.1
GLM-5.1 is capable in instruction following, long context, and coding; and behind the leaders in reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math, Multimodal or Multilingual.
Price
$1.38input$4.40outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
133results on85benchmarks
- 11 independently verified
- 34 aggregator
- 31 vendor-reported
- 57 cross-referenced
From 20 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
10 papers reference GLM-5.1GLM-5.1 benchmark results
133 results on 85 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
17.6% behind the leader2 of 3 ranked benchmarks measured
- IFBench76.26Oct 8, 2026aa_ifbench
- 63.00Jun 10, 2026
Show 1 more instruction following resultHide 1 instruction following result
- IFBench51.97Oct 8, 2026aa_ifbench
20.8% behind the leader1 of 3 ranked benchmarks measured
- 73.67Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 53.33Oct 8, 2026
21.0% behind the leader8 of 10 ranked benchmarks measured
- LiveCodeBench v685.70Jun 6, 2026LiveCodeBench (v6)
- 74.80Jun 6, 2026
- 76.20Jun 6, 2026
- 1508.21May 22, 2026
- Terminal-Bench 2.161.80Oct 8, 2026terminalbenchV21
- SciCode44.79Oct 8, 2026aa_scicode
- Terminal-Bench Hard43.18Oct 8, 2026aa_terminalbench_hard
- 58.40Oct 7, 2026
Show 9 more coding resultsHide 9 coding results
- SciCode36.11Sep 4, 2026aa_scicode
- 72.30Jun 15, 2026
- SWE-bench Multilingual73.30Jun 6, 2026SWE Multilingual (Resolved)
- 17.50Jun 15, 2026
- SWE-bench Pro58.40Jun 6, 2026SWE Pro (Resolved)
- Terminal-Bench 2.169.00Jul 13, 2026Terminal Bench 2.1 (Best Reported Harness)
- Terminal-Bench 2.163.50Jul 13, 2026Terminal Bench 2.1 (Terminus-2)
- 59.30Jun 6, 2026
- Terminal-Bench Hard35.61Oct 8, 2026aa_terminalbench_hard
26.4% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond86.77Oct 8, 2026gpqa
- 55.10Jul 10, 2026
- Humanity's Last Exam30.07Oct 8, 2026aa_hle
- 4.57Oct 8, 2026
Show 11 more reasoning resultsHide 11 reasoning results
- 0.00Oct 8, 2026
- 4.60Jul 13, 2026
- GPQA Diamond83.94Oct 8, 2026gpqa
- GPQA Diamond86.20Oct 7, 2026GPQA
- GPQA Diamond86.10Aug 24, 2026GPQA (no tools)
- GPQA Diamond86.20Jun 6, 2026GPQA Diamond (Pass@1)
- Humanity's Last Exam27.94Oct 8, 2026aa_hle
- 52.30Oct 7, 2026
- 31.00Jul 13, 2026
- Humanity's Last Exam27.20Jun 6, 2026HLE (no tools)
- Humanity's Last Exam34.70Jun 6, 2026HLE (Pass@1)
27.4% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination70.05Oct 8, 2026omniscienceNonHallucination
- 34.00May 20, 2026
- AA-Omniscience · Accuracy23.70Oct 8, 2026omniscienceAccuracy
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy25.18Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination36.36Oct 8, 2026omniscienceNonHallucination
- SimpleQA Verified38.10Jun 27, 2026SimpleQA-Verified (Pass@1)
32.4% behind the leader5 of 7 ranked benchmarks measured
- 79.30Oct 7, 2026
- 71.80Oct 7, 2026
- 30.95Oct 8, 2026
- τ-Bench V3 · Banking13.61Oct 8, 2026tauBanking
- Terminal-Bench 4.02.02Oct 8, 2026
Show 8 more agentic resultsHide 8 agentic results
- 40.25Oct 8, 2026
- 68.00May 18, 2026
- 59.40Jun 6, 2026
- BrowseComp79.30Jun 6, 2026BrowseComp (Pass@1)
- 49.49Jun 15, 2026
- GDPval (win rate)54.70Jun 6, 2026GDPVal
- 75.60Oct 8, 2026
- 12.80Jun 6, 2026
0 of 5 ranked benchmarks measured
- 36.84Sep 9, 2026
- 89.39Sep 2, 2026
- 95.83Sep 2, 2026
Show 7 more math resultsHide 7 math results
- 95.83May 2, 2026
- 95.30Oct 7, 2026
- 89.39May 10, 2026
- HMMT Feb 202682.60Oct 7, 2026HMMT Feb 26
- HMMT Feb 202689.40Jun 6, 2026HMMT 2026 Feb (Pass@1)
- 83.80Oct 7, 2026
- IMO-AnswerBench83.80Jun 6, 2026IMOAnswerBench (Pass@1)
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1465Sep 13, 2026
- frontiermath_tier_4_v112.50May 20, 2026frontiermath_tier_4
- -22.43Oct 8, 2026
- AA Intelligence24.24Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)97.08Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)97.66Oct 8, 2026aa_tau2
Show 61 more resultsHide 61 results
- 25.23Sep 9, 2026
- 66.04Jun 18, 2026
- AA Intelligence26.06Oct 8, 2026aa_intelligence_index
- 0.85Oct 8, 2026
- 59.10Jun 24, 2026
- 67.60Jun 24, 2026
- 59.13Jun 24, 2026
- 51.31Jun 24, 2026
- 22.46Jun 24, 2026
- 52.07Jun 24, 2026
- 47.32Jun 24, 2026
- 51.50Jun 24, 2026
- 11.50Jun 6, 2026
- 72.40Jun 6, 2026
- 71.10Jun 6, 2026
- 79.00Jun 6, 2026
- Artificial Analysis Coding Index55.78Sep 9, 2026aa_coding_index
- 35.77Jun 18, 2026
- browsecomp_with_context_manager79.30May 18, 2026BrowseComp (w/ Context Manage)
- Chinese SimpleQA (C-SimpleQA)75.00Jun 6, 2026Chinese-SimpleQA (Pass@1)
- 3.70Jun 6, 2026
- 68.70Oct 7, 2026
- 18.00Jul 13, 2026
- 44.79Oct 7, 2026
- 1535.00Jun 6, 2026
- GPQA (unspecified)86.10Jun 6, 2026GPQA (no tools)
- HLE (with tools)52.30Jul 13, 2026HLE (w/ Tools)
- 50.40Jun 6, 2026
- 94.00Oct 7, 2026
- 94.00Jul 13, 2026
- 76.60Jun 10, 2026
- 91.10Jun 6, 2026
- 456.50Jun 6, 2026
- 70.18Oct 7, 2026
- 71.80Jun 6, 2026
- MMLU-Pro86.00Jun 27, 2026MMLU-Pro (EM)
- 85.90Jun 6, 2026
- 85.80Jun 10, 2026
- 42.70Oct 7, 2026
- 81.20Jun 6, 2026
- 20.10Jul 13, 2026
- 46.00Jun 6, 2026
- 50.90Jul 13, 2026
- 47.70Jun 6, 2026
- 1.00Jul 13, 2026
- 72.70Jun 15, 2026
- 85.00Jun 6, 2026
- 69.70Jun 6, 2026
- 84.10Jun 6, 2026
- TauBench V3 - Telecom96.90Aug 24, 2026Telecom
- 69.00Oct 7, 2026
- Terminal-Bench 2.063.50May 18, 2026Terminal-Bench 2.0 (Terminus-2)
- Terminal-Bench 2.063.50Jun 6, 2026Terminal Bench 2.0 (Acc)
- 40.70Jul 13, 2026
- 40.70Oct 7, 2026
- Toolathlon40.70Jun 6, 2026Toolathlon (Pass@1)
- 60.70Jun 6, 2026
- 60.20Jun 6, 2026
- 563441.00Oct 7, 2026
- 84.40Jun 10, 2026
- τ³-Bench70.60Oct 7, 2026TAU3-Bench
GLM-5.1: common questions
Who makes GLM-5.1?
GLM-5.1 is made by Z.ai.
When was GLM-5.1 released?
GLM-5.1 was released on Apr 7, 2026, according to Artificial Analysis.
What is GLM-5.1 good at?
GLM-5.1 is capable in instruction following, long context, and coding; and behind the leaders in reasoning, factuality, and agentic tasks. Too few results yet to rate safety, math, multimodal tasks, or multilingual tasks.
How much does GLM-5.1 cost?
GLM-5.1 costs $1.38 per million input tokens and $4.40 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it costs more than 73% of the 330 priced models we track.
How many benchmarks has GLM-5.1 been tested on?
We track 133 results for GLM-5.1 on 85 benchmarks from 20 sources, 11 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GLM-5.1 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GLM-5.1.
About this record
Where GLM-5.1's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 133 scores · 11 independently verified · 34 aggregator-attributed · 57 vendor cross-reference · 31 vendor-reported. How these tiers are assigned
From 20 sources on 9 sites. Hugging Face supplies 71 of them; the 11 independently verified results come from 6 sites. Bars are coloured by trust tier.
- huggingface.co71
- artificialanalysis.ai34
- api.llm-stats.com17
- matharena.ai4
- epoch.ai3
- datasets-server.huggingface.co1
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1