GLM-4.6V
GLM-4.6V is behind the leaders in factuality, multimodal tasks, and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multilingual or Instruction Following.
Price
$0.30input$0.90outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
33results on17benchmarks
- 2 independently verified
- 31 aggregator
From 3 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingJSON modeReasoning
As listed by OpenRouter
Research
2 papers reference GLM-4.6VGLM-4.6V benchmark results
33 results on 17 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
35.0% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination48.54Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy16.18Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy17.40Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination33.37Oct 8, 2026omniscienceNonHallucination
40.1% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro48.55Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- 1163.18Aug 25, 2026
- 1163May 1, 2026
- MMMU-Pro42.20Oct 8, 2026aa_mmmu_pro
43.0% behind the leader1 of 7 ranked benchmarks measured
- 5.46Jun 15, 2026
Show 1 more agentic resultHide 1 agentic result
- 9.58Jun 15, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond71.92Oct 8, 2026gpqa
- Humanity's Last Exam9.64Oct 8, 2026aa_hle
Show 2 more reasoning resultsHide 2 reasoning results
- GPQA Diamond56.57Oct 8, 2026gpqa
- Humanity's Last Exam3.66Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard14.39Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard3.03Oct 8, 2026aa_terminalbench_hard
- SciCode30.44Sep 4, 2026aa_scicode
Show 1 more coding resultHide 1 coding result
- SciCode27.20Sep 4, 2026aa_scicode
0 of 3 ranked benchmarks measured
- 48.67Oct 8, 2026
- 17.00Oct 8, 2026
0 of 3 ranked benchmarks measured
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -26.95Oct 8, 2026
- AA Intelligence11.22Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)31.58Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)30.70Oct 8, 2026aa_tau2
- AA Intelligence8.35Oct 8, 2026aa_intelligence_index
- -37.63Oct 8, 2026
Show 4 more resultsHide 4 results
- 17.51Jun 18, 2026
- 18.69Jun 18, 2026
- 19.74Jun 18, 2026
- 11.09Jun 18, 2026
GLM-4.6V: common questions
Who makes GLM-4.6V?
GLM-4.6V is made by Z.ai.
When was GLM-4.6V released?
GLM-4.6V was released on Dec 8, 2025, according to Artificial Analysis.
What is GLM-4.6V good at?
GLM-4.6V is behind the leaders in factuality, multimodal tasks, and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multilingual tasks, or instruction following.
How much does GLM-4.6V cost?
GLM-4.6V costs $0.30 per million input tokens and $0.90 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 62% of the 330 priced models we track.
How many benchmarks has GLM-4.6V been tested on?
We track 33 results for GLM-4.6V on 17 benchmarks from 3 sources, 2 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GLM-4.6V support?
OpenRouter lists tool calling, json mode, and reasoning for GLM-4.6V.
About this record
Where GLM-4.6V's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- May 2, 2026
Where the results come from
Verification: 33 scores · 2 independently verified · 31 aggregator-attributed. How these tiers are assigned
From 3 sources on 3 sites. Artificial Analysis supplies 31 of them; the 2 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai31
- datasets-server.huggingface.co1
- lmarena.ai1