Qwen3 4B 2507 Instruct
Qwen3 4B 2507 Instruct is behind the leaders in agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Safety, Long Context, Math, Multimodal, Multilingual, Instruction Following or Factuality.
Price
$0.01input$0.03outputper million tokens
From nscale · All prices
Evidence
29results on27benchmarks
- 15 aggregator
- 14 cross-referenced
From 3 sources · latest Oct 7, 2026 · How verification works
Qwen3 4B 2507 Instruct benchmark results
29 results on 27 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
47.0% behind the leader1 of 7 ranked benchmarks measured
- 0.00Jun 15, 2026
0 of 6 ranked benchmarks measured
- 0.00Oct 7, 2026
- GPQA Diamond51.72Oct 7, 2026gpqa
- Humanity's Last Exam4.49Oct 7, 2026aa_hle
Show 1 more reasoning resultHide 1 reasoning result
- GPQA Diamond34.85Oct 7, 2026GPQA
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard4.55Oct 7, 2026aa_terminalbench_hard
- SciCode18.06Sep 4, 2026aa_scicode
- LiveCodeBench v648.72Oct 7, 2026LCB v6
0 of 3 ranked benchmarks measured
- 11.33Oct 7, 2026
0 of 3 ranked benchmarks measured
- IFBench33.54Oct 7, 2026aa_ifbench
- 30.28Oct 7, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy10.90Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination18.26Oct 7, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- -61.93Oct 7, 2026
- AA Intelligence6.74Oct 7, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)26.61Oct 7, 2026aa_tau2
- 8.87Jun 18, 2026
- 9.05Jun 18, 2026
- Creative Writing v351.71Oct 7, 2026
Show 10 more resultsHide 10 results
- 68.46Aug 9, 2026
- 82.32Oct 7, 2026
- 85.62Oct 7, 2026
- LCB v550.80Oct 7, 2026
- math lvl 573.62Oct 7, 2026
- MATH-500 (EM)85.60Oct 7, 2026MATH 500
- 81.76Oct 7, 2026
- 72.25Oct 7, 2026
- 52.31Oct 7, 2026
- 60.67Oct 7, 2026
Qwen3 4B 2507 Instruct: common questions
Who makes Qwen3 4B 2507 Instruct?
Qwen3 4B 2507 Instruct is made by Alibaba.
When was Qwen3 4B 2507 Instruct released?
Qwen3 4B 2507 Instruct was released on Aug 6, 2025, according to Artificial Analysis.
What is Qwen3 4B 2507 Instruct good at?
Qwen3 4B 2507 Instruct is behind the leaders in agentic tasks. Too few results yet to rate reasoning, coding, safety, long context, math, multimodal tasks, multilingual tasks, instruction following, or factuality.
How much does Qwen3 4B 2507 Instruct cost?
Qwen3 4B 2507 Instruct costs $0.01 per million input tokens and $0.03 per million output tokens, according to nscale. At a mix of three input tokens to one output token, it is cheaper than 99% of the 329 priced models we track.
How many benchmarks has Qwen3 4B 2507 Instruct been tested on?
We track 29 results for Qwen3 4B 2507 Instruct on 27 benchmarks from 3 sources. The latest was recorded on Oct 7, 2026.
About this record
Where Qwen3 4B 2507 Instruct's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 29 scores · 0 independently verified · 15 aggregator-attributed · 14 vendor cross-reference. How these tiers are assigned
From 3 sources on 2 sites. Artificial Analysis supplies 15 of them. Bars are coloured by trust tier.
- artificialanalysis.ai15
- huggingface.co14