Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude 3.5 Haiku vs JT-MINI

Anthropic · released

Wins 2 of 6 areas

Long documents · Following instructions

China Mobile · released

Wins 3 of 6 areas

Coding · Agents · Reasoning

JT-MINI wins more areas, narrowly.Claude 3.5 Haiku is better at long documents.

Scores updated · 15 tests both models report · How we compare

Where each one wins

Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.

Agents, long documents and following instructions rest on a single test each.

The biggest differences

The three tests each model wins by the widest margin. Scores are out of 100.

Where JT-MINI pulls ahead

  • Graduate-level biology, physics and chemistry questionsGPQA Diamond+26.8points aheadJT-MINI67.6Claude 3.5 Haiku40.8
  • Hard command-line tasks in a real terminalTerminal-Bench Hard+15.9points aheadJT-MINI18.2Claude 3.5 Haiku2.3
  • Very hard expert questions across many subjectsHumanity's Last Exam+2.8points aheadJT-MINI6.4Claude 3.5 Haiku3.6

Where Claude 3.5 Haiku pulls ahead

  • Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+50.2points aheadClaude 3.5 Haiku58.9JT-MINI8.7
  • Reasons across sets of long documentsAA-LCR+14.3points aheadClaude 3.5 Haiku27.3JT-MINI13
  • Follows unfamiliar, precisely checkable instructionsIFBench+6.1points aheadClaude 3.5 Haiku42.8JT-MINI36.7

Every test, side by side

All 15 tests both models report. The winning score is in its model's colour; marks a score checked independently.

CodingJT-MINI

Full coding ranking

AgentsJT-MINI
  • GDPValJT-MINI by 16.6016.6+16.6

Full agents ranking

ReasoningJT-MINI

Full reasoning ranking

FactsEven

Full facts ranking

Long documentsClaude 3.5 Haiku
  • AA-LCRClaude 3.5 Haiku by 14.327.313+14.3

Full long documents ranking

Following instructionsClaude 3.5 Haiku
  • IFBenchClaude 3.5 Haiku by 6.142.836.7+6.1

Full following instructions ranking

Other results5 tests, not counted

Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.

Questions people ask

Which is better, Claude 3.5 Haiku or JT-MINI?

JT-MINI wins three of the six areas where both have results: coding, agents and reasoning. Claude 3.5 Haiku wins long documents and following instructions. They are level on facts.

Which is better for coding?

JT-MINI. It wins 1 of the 2 coding tests both models report; Claude 3.5 Haiku wins none, and 1 is a tie.

How do you compare the two?

We use the 15 benchmark tests both models have published scores on. The verdict counts the 10 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 5 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.