Hunyuan Large Instruct is the stronger all-rounder.
Scores updated · 10 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Hunyuan Large Instruct pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+7.1points ahead
Where DeepSeek-V2 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 10 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningHunyuan Large Instruct
- GPQA DiamondHunyuan Large Instruct by 7.135.342.4+7.1
Other results9 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HumanEvalHunyuan Large Instruct by 46.743.390+46.7
- MATHHunyuan Large Instruct by 3443.477.4+34
- IFEvalHunyuan Large Instruct by 27.357.785+27.3
- MMLUHunyuan Large Instruct by 14.375.689.9+14.3
- BBHHunyuan Large Instruct by 10.778.889.5+10.7
- C-EvalHunyuan Large Instruct by 7.281.488.6+7.2
- CMMLUHunyuan Large Instruct by 6.48490.4+6.4
- ARC-ChallengeHunyuan Large Instruct by 2.492.294.6+2.4
- HellaSwagHunyuan Large Instruct by 1.487.188.5+1.4
Questions people ask
Which is better, DeepSeek-V2 or Hunyuan Large Instruct?
Hunyuan Large Instruct wins the one area where both have results: reasoning. DeepSeek-V2 wins none.
How do you compare the two?
We use the 10 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 9 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.