Qwen3 14B
Qwen3 14B is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.35input$1.40outputper million tokens
From Alibaba's own price page · 5 providers tracked · All prices
Evidence
50results on36benchmarks
- 4 independently verified
- 30 aggregator
- 9 vendor-reported
- 7 cross-referenced
From 4 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
38 papers reference Qwen3 14BQwen3 14B benchmark results
50 results on 36 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
34.4% behind the leader3 of 4 ranked benchmarks measured
- 5.40May 2, 2026
- AA-Omniscience · Non-hallucination24.18Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy15.28Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy13.33Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.69Oct 8, 2026omniscienceNonHallucination
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond46.97Oct 8, 2026gpqa
- Humanity's Last Exam4.09Oct 8, 2026aa_hle
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond60.40Oct 8, 2026gpqa
- 66.30Jun 5, 2026
- Humanity's Last Exam4.53Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- Terminal-Bench Hard5.30Oct 8, 2026aa_terminalbench_hard
- SciCode30.67Oct 8, 2026aa_scicode
- Terminal-Bench 2.14.87Oct 8, 2026terminalbenchV21
Show 2 more coding resultsHide 2 coding results
- SciCode26.50Sep 4, 2026aa_scicode
- Terminal-Bench Hard3.79Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking5.57Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 0.00Oct 8, 2026
0 of 3 ranked benchmarks measured
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- vectara_avg_summary_length111.10May 2, 2026Average Summary Length (Words)
- vectara_answer_rate99.90May 2, 2026Answer Rate
- vectara_factual_consistency94.60May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)32.16Oct 8, 2026aa_tau2
- AA Intelligence6.70Oct 8, 2026aa_intelligence_index
- -66.67Oct 8, 2026
Show 22 more resultsHide 22 results
- 0.93Sep 9, 2026
- 13.60Jun 18, 2026
- AA Intelligence8.15Oct 8, 2026aa_intelligence_index
- -48.95Oct 8, 2026
- AIME 202483.70Jun 5, 2026AIME24
- 70.40Oct 7, 2026
- AIME 202573.70Jun 5, 2026AIME25
- 85.60Oct 7, 2026
- 42.70Jun 5, 2026
- Artificial Analysis Coding Index13.78Sep 9, 2026aa_coding_index
- 12.37Jun 18, 2026
- 89.20Oct 7, 2026
- 70.40Oct 7, 2026
- 86.20Oct 7, 2026
- 59.30Jun 5, 2026
- 87.00Jun 5, 2026
- MATH-500 (EM)96.80Oct 7, 2026MATH-500
- 88.60Oct 7, 2026
- 90.10Oct 7, 2026
- 65.10Jun 5, 2026
- 88.50Oct 7, 2026
- τ²-Bench Telecom (AA run)34.50Oct 8, 2026aa_tau2
Qwen3 14B: common questions
Who makes Qwen3 14B?
Qwen3 14B is made by Alibaba.
When was Qwen3 14B released?
Qwen3 14B was released on Apr 28, 2025, according to Artificial Analysis.
What is Qwen3 14B good at?
Qwen3 14B is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does Qwen3 14B cost?
Qwen3 14B costs $0.35 per million input tokens and $1.40 per million output tokens, according to Alibaba's own price page. We track its price at 5 providers. At a mix of three input tokens to one output token, it is cheaper than 53% of the 330 priced models we track.
How many benchmarks has Qwen3 14B been tested on?
We track 50 results for Qwen3 14B on 36 benchmarks from 4 sources, 4 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Qwen3 14B support?
OpenRouter lists tool calling, structured outputs, and reasoning for Qwen3 14B.
About this record
Where Qwen3 14B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 31, 2026
Where the results come from
Verification: 50 scores · 4 independently verified · 30 aggregator-attributed · 7 vendor cross-reference · 9 vendor-reported. How these tiers are assigned
From 4 sources on 4 sites. Artificial Analysis supplies 30 of them; the 4 independently verified results come from 1 site. Bars are coloured by trust tier.
- artificialanalysis.ai30
- api.llm-stats.com9
- huggingface.co7
- raw.githubusercontent.com4