Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Qwen3 4B 2507 Thinking

Score basis

Qwen3 4B 2507 Thinking is behind the leaders in instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited1 capability
  1. Instruction Following−39.3%1 of 3
Not rated9 capabilities

Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Factuality.

Price

$0.01input$0.03outputper million tokens

From nscale · All prices

Cheaper than 99% of 328 priced models · 3:1 input-to-output blend, log scale

Evidence

14results on13benchmarks

  • 3 independently verified
  • 11 aggregator

From 3 sources · latest Oct 7, 2026 · How verification works

Qwen3 4B 2507 Thinking benchmark results

14 results on 13 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Instruction Following

LimitedFull instruction following ranking

39.3% behind the leader1 of 3 ranked benchmarks measured

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Coding

Not enough dataFull coding ranking

0 of 10 ranked benchmarks measured

Long Context

Not enough dataFull long context ranking

0 of 3 ranked benchmarks measured

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

Factuality

Not enough dataFull factuality ranking

0 of 4 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Qwen3 4B 2507 Thinking: common questions

Who makes Qwen3 4B 2507 Thinking?

Qwen3 4B 2507 Thinking is made by Alibaba.

When was Qwen3 4B 2507 Thinking released?

Qwen3 4B 2507 Thinking was released on Aug 6, 2025, according to Artificial Analysis.

What is Qwen3 4B 2507 Thinking good at?

Qwen3 4B 2507 Thinking is behind the leaders in instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.

How much does Qwen3 4B 2507 Thinking cost?

Qwen3 4B 2507 Thinking costs $0.01 per million input tokens and $0.03 per million output tokens, according to nscale. At a mix of three input tokens to one output token, it is cheaper than 99% of the 328 priced models we track.

How many benchmarks has Qwen3 4B 2507 Thinking been tested on?

We track 14 results for Qwen3 4B 2507 Thinking on 13 benchmarks from 3 sources, 3 of them independently verified. The latest was recorded on Oct 7, 2026.

About this record

Where Qwen3 4B 2507 Thinking's numbers come from, and every name it appears under.

Tracked since
Sep 23, 2026
Newest source mention
Sep 23, 2026

Where the results come from

Verification: 14 scores · 3 independently verified · 11 aggregator-attributed. How these tiers are assigned

From 3 sources on 2 sites. Artificial Analysis supplies 11 of them; the 3 independently verified results come from 1 site. Bars are coloured by trust tier.

  • artificialanalysis.ai11
  • matharena.ai3

Also known as

qwen3-4b-thinking-2507Qwen3 4B 2507 (Reasoning)Qwen3 4B 2507 Think