Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Phi 3 Medium 4K Instruct vs Qwen 2.5 14B

Microsoft · released

Alibaba · released

Scores updated · 7 tests both models report · How we compare

Every test, side by side

All 7 tests both models report. The winning score is in its model's colour; marks a score checked independently.

Other results7 tests
  • MGSMQwen 2.5 14B by 26.153.579.6+26.1
  • DROPQwen 2.5 14B by 17.268.385.5+17.2
  • MATHQwen 2.5 14B by 14.561.175.6+14.5
  • MMLUQwen 2.5 14B by 12.467.579.9+12.4
  • GPQA (unspecified)Qwen 2.5 14B by 11.731.242.9+11.7
  • HumanEvalQwen 2.5 14B by 4.367.872.1+4.3
  • SimpleQAPhi 3 Medium 4K Instruct by 2.27.65.4+2.2

Questions people ask

How do you compare the two?

We use the 7 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.