Qwen3 VL 235B A22B Reasoning
Qwen3 VL 235B A22B Reasoning is capable in multimodal tasks; and behind the leaders in long context, instruction following, and factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Math or Multilingual.
Price
$0.40input$4.00outputper million tokens
From Alibaba's own price page · 4 providers tracked · All prices
Evidence
91results on73benchmarks
- 2 independently verified
- 13 aggregator
- 46 vendor-reported
- 30 cross-referenced
From 8 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Qwen3 VL 235B A22B Reasoning benchmark results
91 results on 73 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
23.6% behind the leader6 of 6 ranked benchmarks measured
- 87.50Oct 7, 2026
- MathVista85.80Oct 7, 2026MathVista-Mini
- 78.70Jun 5, 2026
- 79.00Jun 15, 2026
- MMMU-Pro68.73Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)66.10Oct 7, 2026CharXiv-R
Show 12 more multimodal resultsHide 12 multimodal results
- 67.10Oct 7, 2026
- CharXiv (reasoning)66.10Jun 15, 2026CharXiv (RQ)
- 1206.22Sep 21, 2026
- 1189May 1, 2026
- MathVista85.80Jun 15, 2026MathVista (mini)
- 85.10Jun 5, 2026
- 69.30Oct 7, 2026
- 69.30Jun 15, 2026
- 78.70Oct 7, 2026
- 76.80Jun 5, 2026
- 87.50Jun 15, 2026
- 87.30Jun 5, 2026
30.1% behind the leader1 of 3 ranked benchmarks measured
- 63.67Oct 8, 2026
35.9% behind the leader1 of 3 ranked benchmarks measured
- IFBench56.46Oct 8, 2026aa_ifbench
38.5% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy20.85Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination14.85Oct 8, 2026omniscienceNonHallucination
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond77.17Oct 8, 2026gpqa
- Humanity's Last Exam11.91Oct 8, 2026aa_hle
Show 1 more reasoning resultHide 1 reasoning result
- 13.60Oct 7, 2026
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard11.36Oct 8, 2026aa_terminalbench_hard
- SciCode39.93Sep 4, 2026aa_scicode
- 70.10Oct 7, 2026
0 of 7 ranked benchmarks measured
- OSWorld-Verified38.10Oct 7, 2026OSWorld
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -46.55Oct 8, 2026
- τ²-Bench Telecom (AA run)54.09Oct 8, 2026aa_tau2
- AA Intelligence13.44Oct 8, 2026aa_intelligence_index
- 73.90Oct 7, 2026
- 69.90Oct 7, 2026
- 84.30Oct 7, 2026
Show 55 more resultsHide 55 results
- 89.20Oct 7, 2026
- 89.70Oct 7, 2026
- 83.59Jun 12, 2026
- 71.90Oct 7, 2026
- 81.50Oct 7, 2026
- 95.10Sep 2, 2026
- 63.50Oct 7, 2026
- 96.50Oct 7, 2026
- 52.50Oct 7, 2026
- 66.70Oct 7, 2026
- HMMT 202577.40Oct 7, 2026HMMT25
- 67.71Jun 12, 2026
- 88.20Aug 31, 2026
- 80.00Oct 7, 2026
- 89.50Jun 15, 2026
- 89.50Oct 7, 2026
- Key Information Extraction Overall84.20Sep 2, 2026
- 69.45Jun 12, 2026
- 65.60Jun 15, 2026
- 63.60Oct 7, 2026
- 63.60Jun 15, 2026
- 74.60Oct 7, 2026
- 74.60Jun 15, 2026
- 72.10Jun 5, 2026
- 83.80Oct 7, 2026
- 8.50Oct 7, 2026
- 92.70Jun 5, 2026
- 56.20Oct 7, 2026
- 83.80Oct 7, 2026
- 80.60Oct 7, 2026
- 93.70Oct 7, 2026
- MMMU (val) (Pass@1)80.60Aug 27, 2026MMMUval
- 71.10Jun 15, 2026
- OCRBench KIE94.00Sep 2, 2026
- OCRBench v2 (Chinese)63.50Oct 7, 2026OCRBench-V2 (zh)
- OCRBench v2_en66.80Oct 7, 2026OCRBench-V2 (en)
- OCRBenchv2 KIE (en)85.60Sep 2, 2026
- OCRBenchv2 KIE (zh)62.90Sep 2, 2026
- 82.00Jun 15, 2026
- 68.30Oct 7, 2026
- 81.30Oct 7, 2026
- 95.40Oct 7, 2026
- ScreenSpot-Pro (No tools)61.80Oct 7, 2026ScreenSpot Pro
- 44.40Oct 7, 2026
- 61.30Oct 7, 2026
- 56.80Jun 15, 2026
- 64.30Oct 7, 2026
- 80.00Oct 7, 2026
- 80.00Jun 15, 2026
- 0.34Sep 28, 2026
- 23.50Jun 15, 2026
- 97.30Oct 7, 2026
- 4.00Oct 7, 2026
- 4.00Jun 15, 2026
- 3.00Jun 15, 2026
Qwen3 VL 235B A22B Reasoning: common questions
Who makes Qwen3 VL 235B A22B Reasoning?
Qwen3 VL 235B A22B Reasoning is made by Alibaba.
When was Qwen3 VL 235B A22B Reasoning released?
Qwen3 VL 235B A22B Reasoning was released on Sep 23, 2025, according to Artificial Analysis.
What is Qwen3 VL 235B A22B Reasoning good at?
Qwen3 VL 235B A22B Reasoning is capable in multimodal tasks; and behind the leaders in long context, instruction following, and factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, or multilingual tasks.
How much does Qwen3 VL 235B A22B Reasoning cost?
Qwen3 VL 235B A22B Reasoning costs $0.40 per million input tokens and $4.00 per million output tokens, according to Alibaba's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 63% of the 330 priced models we track.
How many benchmarks has Qwen3 VL 235B A22B Reasoning been tested on?
We track 91 results for Qwen3 VL 235B A22B Reasoning on 73 benchmarks from 8 sources, 2 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 VL 235B A22B Reasoning support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3 VL 235B A22B Reasoning.
About this record
Where Qwen3 VL 235B A22B Reasoning's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 91 scores · 2 independently verified · 13 aggregator-attributed · 30 vendor cross-reference · 46 vendor-reported. How these tiers are assigned
From 8 sources on 5 sites. api.llm-stats.com supplies 46 of them; the 2 independently verified results come from 2 sites. Bars are coloured by trust tier.
- api.llm-stats.com46
- huggingface.co30
- artificialanalysis.ai13
- datasets-server.huggingface.co1
- lmarena.ai1