Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Hunyuan Large vs Llama 3.1 Instruct 405B

Tencent

Meta · released

Scores updated · 17 tests both models report · How we compare

Every test, side by side

All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.

Other results17 tests
  • C-EvalHunyuan Large by 19.491.972.5+19.4
  • HumanEvalLlama 3.1 Instruct 405B by 17.671.489+17.6
  • CMMLUHunyuan Large by 16.590.273.7+16.5
  • MMLU-ProLlama 3.1 Instruct 405B by 13.160.273.3+13.1
  • NaturalQuestionsHunyuan Large by 11.352.841.5+11.3
  • CommonsenseQAHunyuan Large by 7.192.985.8+7.1
  • MBPPHunyuan Large by 4.272.668.4+4.2
  • DROPHunyuan Large by 4.188.984.8+4.1
  • GSM8KLlama 3.1 Instruct 405B by 492.896.8+4
  • MATHLlama 3.1 Instruct 405B by 469.873.8+4
  • WinoGrandeHunyuan Large by 3.588.785.2+3.5
  • BBHHunyuan Large by 3.486.382.9+3.4
  • C3Hunyuan Large by 2.682.379.7+2.6
  • HellaSwagLlama 3.1 Instruct 405B by 2.486.889.2+2.4
  • PIQAHunyuan Large by 2.488.385.9+2.4
  • ARC-ChallengeLlama 3.1 Instruct 405B by 1.99596.9+1.9
  • MMLUtie88.488.6tie

Questions people ask

How do you compare the two?

We use the 17 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.