Phi 4
Phi 4 is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.13input$0.50outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
36results on33benchmarks
- 10 independently verified
- 14 aggregator
- 12 vendor-reported
From 8 sources · latest Oct 8, 2026 · How verification works
API features
Structured outputs
As listed by OpenRouter
Research
30 papers reference Phi 4Phi 4 benchmark results
36 results on 33 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
35.9% behind the leader3 of 4 ranked benchmarks measured
- 3.70May 2, 2026
- AA-Omniscience · Non-hallucination18.81Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy14.07Oct 8, 2026omniscienceAccuracy
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond57.47Oct 8, 2026gpqa
- Humanity's Last Exam3.76Oct 8, 2026aa_hle
Show 1 more reasoning resultHide 1 reasoning result
- GPQA Diamond56.10Oct 7, 2026GPQA
0 of 10 ranked benchmarks measured
- LiveBench · Coding31.25Aug 23, 2026livebench_coding@2025-04-07
- Terminal-Bench Hard3.79Oct 8, 2026aa_terminalbench_hard
- SciCode26.04Sep 4, 2026aa_scicode
0 of 3 ranked benchmarks measured
- 0.00Oct 8, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following54.80Aug 23, 2026livebench_instruction_following@2025-04-07
- IFBench23.54Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language29.33Aug 23, 2026livebench_language@2025-04-07
- AA Intelligence5.00Jul 26, 2026Artificial Analysis Intelligence Index
- AA Intelligence6.00Jun 12, 2026Artificial Analysis Intelligence Index
- vectara_avg_summary_length120.90May 2, 2026Average Summary Length (Words)
- vectara_answer_rate80.70May 2, 2026Answer Rate
- vectara_factual_consistency96.30May 2, 2026Factual Consistency Rate
Show 17 more resultsHide 17 results
- 0.00Jun 18, 2026
- AA Intelligence5.92Oct 8, 2026aa_intelligence_index
- -55.70Oct 8, 2026
- 75.40Sep 8, 2026
- 11.21Jun 18, 2026
- 75.50Oct 7, 2026
- 82.60Oct 7, 2026
- 82.80Oct 7, 2026
- 63.00Aug 31, 2026
- 47.60Oct 7, 2026
- 80.40May 30, 2026
- 80.60Oct 7, 2026
- 84.80May 30, 2026
- 70.40Oct 7, 2026
- 49.90May 26, 2025
- 3.00Oct 7, 2026
- τ²-Bench Telecom (AA run)0.00Oct 8, 2026aa_tau2
Phi 4: common questions
Who makes Phi 4?
Phi 4 is made by Microsoft.
When was Phi 4 released?
Phi 4 was released on Dec 12, 2024, according to Artificial Analysis.
What is Phi 4 good at?
Phi 4 is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does Phi 4 cost?
Phi 4 costs $0.13 per million input tokens and $0.50 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 80% of the 330 priced models we track.
How many benchmarks has Phi 4 been tested on?
We track 36 results for Phi 4 on 33 benchmarks from 8 sources, 10 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Phi 4 support?
OpenRouter lists structured outputs for Phi 4.
About this record
Where Phi 4's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- May 26, 2026
Where the results come from
Verification: 36 scores · 10 independently verified · 14 aggregator-attributed · 12 vendor-reported. How these tiers are assigned
From 8 sources on 5 sites. Artificial Analysis supplies 16 of them; the 10 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai16
- api.llm-stats.com10
- huggingface.co5
- raw.githubusercontent.com4
- arxiv.org1