Qwen3 235B A22B
Qwen3 235B A22B is behind the leaders in factuality and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.70input$2.80outputper million tokens
From Alibaba's own price page · 4 providers tracked · All prices
Evidence
104results on69benchmarks
- 15 independently verified
- 26 aggregator
- 17 vendor-reported
- 46 cross-referenced
From 14 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingJSON modeReasoning
As listed by OpenRouter
Qwen3 235B A22B benchmark results
104 results on 69 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
37.4% behind the leader4 of 4 ranked benchmarks measured
- 9.30May 2, 2026
- 40.44May 20, 2026
- AA-Omniscience · Accuracy18.53Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination22.38Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy18.57Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination12.87Oct 8, 2026omniscienceNonHallucination
41.5% behind the leader1 of 7 ranked benchmarks measured
- 11.82Jun 15, 2026
0 of 6 ranked benchmarks measured
- 31.00May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond70.00Oct 8, 2026gpqa
Show 8 more reasoning resultsHide 8 reasoning results
- GPQA Diamond61.31Oct 8, 2026gpqa
- GPQA Diamond47.47Oct 7, 2026GPQA
- 62.90Jun 15, 2026
- 71.10Jun 5, 2026
- Humanity's Last Exam10.98Oct 8, 2026aa_hle
- Humanity's Last Exam4.19Oct 8, 2026aa_hle
- Humanity's Last Exam5.70Jun 15, 2026Humanity's Last Exam (Text Only)
- Humanity's Last Exam7.60Jun 5, 2026HLE (no tools)
0 of 10 ranked benchmarks measured
- 21.41Oct 8, 2026
- Terminal-Bench Hard6.06Oct 8, 2026aa_terminalbench_hard
- SciCode39.93Sep 4, 2026aa_scicode
Show 5 more coding resultsHide 5 coding results
- LiveCodeBench v637.00Jun 15, 2026LiveCodeBench v6 (Aug 24 - May 25)
- SciCode29.86Sep 4, 2026aa_scicode
- SWE-bench Multilingual20.90Jun 15, 2026SWE-bench Multilingual (Agentic Coding)
- SWE-bench Verified34.40Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified39.40Jun 15, 2026SWE-bench Verified (Agentless Coding)
0 of 3 ranked benchmarks measured
- 0.00Oct 8, 2026
- 50.10Jun 5, 2026
0 of 3 ranked benchmarks measured
Show 2 more instruction following resultsHide 2 instruction following results
- 34.00Jun 15, 2026
- 40.00Jun 5, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 80.83Jul 3, 2026
- 62.50Jun 25, 2026
- 54.00May 3, 2026
- 88.77May 3, 2026
- 99.07May 3, 2026
- 80.38May 3, 2026
Show 65 more resultsHide 65 results
- 18.45Jun 18, 2026
- 19.23Jun 18, 2026
- AA Intelligence9.50Oct 8, 2026aa_intelligence_index
- AA Intelligence8.29Oct 8, 2026aa_intelligence_index
- -44.70Oct 8, 2026
- -52.38Oct 8, 2026
- 70.50Jun 15, 2026
- 61.80Jun 15, 2026
- 40.10Jun 15, 2026
- 85.70Jun 5, 2026
- 81.50Oct 7, 2026
- 24.70Jun 15, 2026
- 81.50Jun 5, 2026
- 95.60Sep 8, 2026
- Artificial Analysis Coding Index17.35Jun 18, 2026aa_coding_index
- 13.99Jun 18, 2026
- 83.30Jun 15, 2026
- 88.87Oct 7, 2026
- BFCLv470.80Oct 7, 2026BFCL
- 48.60Jun 15, 2026
- 77.60Oct 7, 2026
- 62.90Jun 5, 2026
- 94.39Oct 7, 2026
- 11.90Jun 15, 2026
- HMMT Feb. 202562.50Jun 4, 2026HMMT Feb 25
- 83.20Jun 15, 2026
- 73.46Oct 7, 2026
- 77.10Oct 7, 2026
- 67.60Jun 15, 2026
- 70.70Aug 23, 2026
- LiveCodeBench66.50Jun 4, 2026LiveCodeBench (2408-2505)
- 65.90Jun 5, 2026
- MATH-500 (EM)91.20Jun 15, 2026MATH-500
- MATH-500 (EM)96.20Jun 5, 2026MATH-500
- 81.40Oct 7, 2026
- 83.53Oct 7, 2026
- 87.00Jun 15, 2026
- 68.18Oct 7, 2026
- 77.30Jun 15, 2026
- 83.00Jun 5, 2026
- 66.70May 26, 2025
- 87.40Oct 7, 2026
- 89.20Jun 15, 2026
- 86.70Oct 7, 2026
- 65.94Oct 7, 2026
- 78.20Jun 15, 2026
- 11.30Jun 15, 2026
- 27.70Jun 5, 2026
- 51.90Jun 15, 2026
- 13.20Jun 15, 2026
- 11.00Jun 5, 2026
- 44.06Oct 7, 2026
- 50.20Jun 15, 2026
- 34.70Jun 5, 2026
- 58.60Jun 5, 2026
- 26.50Jun 15, 2026
- vectara_answer_rate94.90May 2, 2026Answer Rate
- vectara_avg_summary_length105.60May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency90.70May 2, 2026Factual Consistency Rate
- 37.70Jun 15, 2026
- 80.30Jun 5, 2026
- τ²-Bench22.10Jun 15, 2026Tau2 telecom
- τ²-Bench (Retail)57.00Jun 15, 2026Tau2 retail
- τ²-Bench Telecom (AA run)23.98Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)27.19Oct 8, 2026aa_tau2
Qwen3 235B A22B: common questions
Who makes Qwen3 235B A22B?
Qwen3 235B A22B is made by Alibaba.
When was Qwen3 235B A22B released?
Qwen3 235B A22B was released on Apr 28, 2025, according to Artificial Analysis.
What is Qwen3 235B A22B good at?
Qwen3 235B A22B is behind the leaders in factuality and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does Qwen3 235B A22B cost?
Qwen3 235B A22B costs $0.70 per million input tokens and $2.80 per million output tokens, according to Alibaba's own price page. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 63% of the 330 priced models we track.
How many benchmarks has Qwen3 235B A22B been tested on?
We track 104 results for Qwen3 235B A22B on 69 benchmarks from 14 sources, 15 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 235B A22B support?
OpenRouter lists tool calling, json mode, and reasoning for Qwen3 235B A22B.
About this record
Where Qwen3 235B A22B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Jun 24, 2026
Where the results come from
Verification: 104 scores · 15 independently verified · 26 aggregator-attributed · 46 vendor cross-reference · 17 vendor-reported. How these tiers are assigned
From 14 sources on 10 sites. Hugging Face supplies 46 of them; the 15 independently verified results come from 7 sites. Bars are coloured by trust tier.
- huggingface.co46
- artificialanalysis.ai26
- api.llm-stats.com17
- livecodebench.github.io4
- raw.githubusercontent.com4
- labs.scale.com2
- matharena.ai2
- arxiv.org1
- epoch.ai1
- simple-bench.com1