Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude 3.5 Haiku vs Nanbeige4.1-3B

Anthropic · released

Wins 5 of 6 areas

Coding · Agents · Facts · Long documents · Following instructions

Nanbeige · released

Wins 1 of 6 areas

Reasoning

Claude 3.5 Haiku is the stronger all-rounder.Nanbeige4.1-3B is better at reasoning.

Scores updated · 16 tests both models report · How we compare

The biggest differences

The tests each model wins by the widest margin, up to three each. Scores are out of 100.

Where Claude 3.5 Haiku pulls ahead

  • Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+17points aheadClaude 3.5 Haiku58.9Nanbeige4.1-3B41.9
  • Follows unfamiliar, precisely checkable instructionsIFBench+7.4points aheadClaude 3.5 Haiku42.8Nanbeige4.1-3B35.4
  • Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+3.3points aheadClaude 3.5 Haiku13.2Nanbeige4.1-3B9.9

Where Nanbeige4.1-3B pulls ahead

  • Graduate-level biology, physics and chemistry questionsGPQA Diamond+44.1points aheadNanbeige4.1-3B84.9Claude 3.5 Haiku40.8
  • Very hard expert questions across many subjectsHumanity's Last Exam+7.3points aheadNanbeige4.1-3B10.9Claude 3.5 Haiku3.6

Every test, side by side

All 16 tests both models report. The winning score is in its model's colour; marks a score checked independently.

CodingClaude 3.5 Haiku

Full coding ranking

AgentsClaude 3.5 Haiku

Full agents ranking

ReasoningNanbeige4.1-3B

Full reasoning ranking

FactsClaude 3.5 Haiku

Full facts ranking

Long documentsClaude 3.5 Haiku
  • AA-LCRClaude 3.5 Haiku by 27.327.30+27.3

Full long documents ranking

Following instructionsClaude 3.5 Haiku
  • IFBenchClaude 3.5 Haiku by 7.442.835.4+7.4

Full following instructions ranking

Other results5 tests, not counted

Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.

Questions people ask

Which is better, Claude 3.5 Haiku or Nanbeige4.1-3B?

Claude 3.5 Haiku wins five of the six areas where both have results: coding, agents, facts, long documents and following instructions. Nanbeige4.1-3B wins reasoning.

Which is better for coding?

Claude 3.5 Haiku. It wins 1 of the 2 coding tests both models report; Nanbeige4.1-3B wins none, and 1 is a tie.

How do you compare the two?

We use the 16 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 5 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.