Qwen3 235B A22B Instruct 2507
Qwen3 235B A22B Instruct 2507 is behind the leaders in factuality, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multimodal or Multilingual.
Price
$0.23input$0.92outputper million tokens
From Alibaba's own price page · 6 providers tracked · All prices
Evidence
41results on40benchmarks
- 8 independently verified
- 15 aggregator
- 18 vendor-reported
From 10 sources · latest Oct 8, 2026 · How verification works
Qwen3 235B A22B Instruct 2507 benchmark results
41 results on 40 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
36.7% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy18.73Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination22.83Oct 8, 2026omniscienceNonHallucination
40.5% behind the leader1 of 3 ranked benchmarks measured
- IFBench46.05Oct 8, 2026aa_ifbench
40.7% behind the leader1 of 7 ranked benchmarks measured
- 14.12Jun 15, 2026
0 of 6 ranked benchmarks measured
- 1.25May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond75.25Oct 8, 2026gpqa
Show 2 more reasoning resultsHide 2 reasoning results
- GPQA Diamond77.50Oct 7, 2026GPQA
- Humanity's Last Exam11.12Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard15.15Oct 8, 2026aa_terminalbench_hard
- SciCode36.00Sep 4, 2026aa_scicode
- 51.80Oct 7, 2026
0 of 3 ranked benchmarks measured
- 33.90Oct 8, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 100.00Sep 7, 2026
- 98.56Sep 7, 2026
- 99.86Sep 7, 2026
- 79.63Sep 7, 2026
- 96.20Sep 7, 2026
- 78.97May 19, 2026
Show 22 more resultsHide 22 results
- 22.83Jun 18, 2026
- AA Intelligence12.00Oct 8, 2026aa_intelligence_index
- -43.98Oct 8, 2026
- 57.30Oct 7, 2026
- 70.30Oct 7, 2026
- 11.00May 10, 2026
- 79.20Oct 7, 2026
- 22.10Jun 18, 2026
- 70.90Oct 7, 2026
- HMMT 202555.40Oct 7, 2026HMMT25
- 88.70Aug 31, 2026
- 79.50Oct 7, 2026
- 83.00Oct 7, 2026
- 79.40Oct 7, 2026
- 93.10Oct 7, 2026
- 87.90Oct 7, 2026
- 54.30Oct 7, 2026
- 62.60Oct 7, 2026
- 44.00Oct 7, 2026
- 95.00Oct 7, 2026
- τ²-Bench (Retail)71.30Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)33.33Oct 8, 2026aa_tau2
Qwen3 235B A22B Instruct 2507: common questions
Who makes Qwen3 235B A22B Instruct 2507?
Qwen3 235B A22B Instruct 2507 is made by Alibaba.
When was Qwen3 235B A22B Instruct 2507 released?
Qwen3 235B A22B Instruct 2507 was released on Jul 21, 2025, according to Artificial Analysis.
What is Qwen3 235B A22B Instruct 2507 good at?
Qwen3 235B A22B Instruct 2507 is behind the leaders in factuality, instruction following, and agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, or multilingual tasks.
How much does Qwen3 235B A22B Instruct 2507 cost?
Qwen3 235B A22B Instruct 2507 costs $0.23 per million input tokens and $0.92 per million output tokens, according to Alibaba's own price page. We track its price at 6 providers. At a mix of three input tokens to one output token, it is cheaper than 65% of the 330 priced models we track.
How many benchmarks has Qwen3 235B A22B Instruct 2507 been tested on?
We track 41 results for Qwen3 235B A22B Instruct 2507 on 40 benchmarks from 10 sources, 8 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Qwen3 235B A22B Instruct 2507's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- May 19, 2026
Where the results come from
Verification: 41 scores · 8 independently verified · 15 aggregator-attributed · 18 vendor-reported. How these tiers are assigned
From 10 sources on 4 sites. api.llm-stats.com supplies 18 of them; the 8 independently verified results come from 2 sites. Bars are coloured by trust tier.
- api.llm-stats.com18
- artificialanalysis.ai15
- storage.googleapis.com6
- arcprize.org2