Qwen3 30B A3B 2507 Thinking
Qwen3 30B A3B 2507 Thinking is behind the leaders in long context and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Math, Multimodal, Multilingual or Factuality.
Price
$0.20input$2.40outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
57results on43benchmarks
- 6 independently verified
- 15 aggregator
- 36 cross-referenced
From 8 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingJSON modeReasoning
As listed by OpenRouter
Qwen3 30B A3B 2507 Thinking benchmark results
57 results on 43 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
31.0% behind the leader1 of 3 ranked benchmarks measured
- 61.33Oct 8, 2026
38.6% behind the leader1 of 3 ranked benchmarks measured
- IFBench50.68Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- 51.11Aug 9, 2026
0 of 6 ranked benchmarks measured
- 0.29Oct 8, 2026
- GPQA Diamond70.71Oct 8, 2026gpqa
- Humanity's Last Exam10.29Oct 8, 2026aa_hle
Show 4 more reasoning resultsHide 4 reasoning results
- ARC-AGI-20.87Jun 15, 2026ArcAGI V2
- GPQA Diamond71.40Jun 15, 2026GPQA-D
- 8.70Jun 15, 2026
- 9.80May 16, 2026
0 of 10 ranked benchmarks measured
- SciCode32.99Oct 8, 2026aa_scicode
- Terminal-Bench Hard5.30Oct 8, 2026aa_terminalbench_hard
- SWE-bench Verified33.50Jun 15, 2026SWE-Bench Verified (AgentLess 4*10)
Show 4 more coding resultsHide 4 coding results
- LiveCodeBench v660.30Jun 15, 2026LiveCodeBench v6 (02/2025-05/2025)
- LiveCodeBench v666.00Jun 13, 2026LCB v6
- SWE-bench Verified31.00Jun 15, 2026SWE-Bench Verified (OpenHands)
- 22.00May 16, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking5.36Oct 8, 2026tauBanking
Show 1 more agentic resultHide 1 agentic result
- 2.29May 16, 2026
0 of 5 ranked benchmarks measured
- 78.79Sep 2, 2026
- 88.33Sep 2, 2026
- 87.50Jun 22, 2026
Show 3 more math resultsHide 3 math results
- 88.33May 2, 2026
- AIME 202666.67Aug 9, 2026AIME26
- 78.79May 10, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy16.30Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination13.46Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence10.00Jun 21, 2026Artificial Analysis Intelligence Index
- τ²-Bench Telecom (AA run)28.07Oct 8, 2026aa_tau2
- -56.13Oct 8, 2026
- AA Intelligence9.81Oct 8, 2026aa_intelligence_index
- τ²-Bench (Retail)56.14Aug 9, 2026Tau² Retail
- 21.93Aug 9, 2026
Show 22 more resultsHide 22 results
- AIME 202487.70Jun 15, 2026AIME24
- AIME 202581.30Jun 15, 2026AIME25
- AIME 202585.00May 16, 2026AIME 25
- AIME25 no tools71.67Aug 9, 2026AIME25
- 56.00Jun 15, 2026
- 73.39Aug 9, 2026
- 50.53Aug 9, 2026
- 73.40May 16, 2026
- 90.82Aug 9, 2026
- 88.00Jun 15, 2026
- 70.20Jun 15, 2026
- MATH500 (Pass@1)86.48Aug 9, 2026MATH500
- 86.90Jun 15, 2026
- 81.90Jun 15, 2026
- 79.00Jun 15, 2026
- 9.50Jun 15, 2026
- RULER94.50Jun 15, 2026RULER (128K)
- 23.60Jun 15, 2026
- 57.30Jun 15, 2026
- TAU-bench (airline)47.00Jun 15, 2026TAU1-Airline
- TAU-bench (retail)58.70Jun 15, 2026TAU1-Retail
- 49.00Jun 13, 2026
Qwen3 30B A3B 2507 Thinking: common questions
Who makes Qwen3 30B A3B 2507 Thinking?
Qwen3 30B A3B 2507 Thinking is made by Alibaba.
When was Qwen3 30B A3B 2507 Thinking released?
Qwen3 30B A3B 2507 Thinking was released on Jul 29, 2025, according to Artificial Analysis.
What is Qwen3 30B A3B 2507 Thinking good at?
Qwen3 30B A3B 2507 Thinking is behind the leaders in long context and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
How much does Qwen3 30B A3B 2507 Thinking cost?
Qwen3 30B A3B 2507 Thinking costs $0.20 per million input tokens and $2.40 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 51% of the 331 priced models we track.
How many benchmarks has Qwen3 30B A3B 2507 Thinking been tested on?
We track 57 results for Qwen3 30B A3B 2507 Thinking on 43 benchmarks from 8 sources, 6 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 30B A3B 2507 Thinking support?
OpenRouter lists tool calling, json mode, and reasoning for Qwen3 30B A3B 2507 Thinking.
About this record
Where Qwen3 30B A3B 2507 Thinking's numbers come from, and every name it appears under.
- Tracked since
- Sep 24, 2026
- Newest source mention
- Sep 24, 2026
Where the results come from
Verification: 57 scores · 6 independently verified · 15 aggregator-attributed · 36 vendor cross-reference. How these tiers are assigned
From 8 sources on 3 sites. Hugging Face supplies 36 of them; the 6 independently verified results come from 2 sites. Bars are coloured by trust tier.
- huggingface.co36
- artificialanalysis.ai16
- matharena.ai5