Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Fable 5 vs MiMo V2.5

29 SHARED BENCHMARKS

Across 29 shared benchmarks, Claude Fable 5 scores higher on 26 and MiMo V2.5 on 3. The widest gap is CritPt, where Claude Fable 5 scores 28.6 against 3.7. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $10.00 per million input tokens).

ANTHROPICVSXIAOMI29 SHARED26–3 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicXiaomi
Released
Price per 1M tokens input / output$10.00 / $50.00$0.14 / $0.28
Cost of 1M in + 1M out$60.00$0.42 142.9× less
Head-to-head of 29 shared benchmarks26 wins3 wins
Scores tracked independently verified153 30 ◆51 3 ◆

Release dates per Artificial Analysis. Prices: Anthropic's own price page for Claude Fable 5; Artificial Analysis for MiMo V2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Fable 5WINSMiMo V2.5
Reasoning30Claude Fable 5 leads 3 of 3 · widest: CritPt 28.6 vs 3.7
Coding60Claude Fable 5 leads 6 of 6 · widest: Terminal-Bench Hard 62.9 vs 41.7
Agentic30Claude Fable 5 leads 3 of 3 · widest: τ-Bench V3 · Banking 38.1 vs 8.7
Factuality21Claude Fable 5 leads 2 of 3 · widest: SimpleQA Verified 70.7 vs 16.1
Instruction Following01MiMo V2.5 leads 1 of 1 · widest: IFBench 63.5 vs 67.1
Long Context10Claude Fable 5 leads 1 of 1 · widest: AA-LCR 82.3 vs 73
Multimodal20Claude Fable 5 leads 2 of 2 · widest: MMMU-Pro 84.2 vs 75.4

Biggest gaps

Claude Fable 5 pulls furthest ahead on

  1. CritPt28.6 vs 3.7
  2. τ-Bench V3 · Banking38.1 vs 8.7
  3. SimpleQA Verified70.7 vs 16.1

MiMo V2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination68.1 vs 36.4
  2. IFBench67.1 vs 63.5

Every shared benchmark29 · grouped by area

Reasoning 3

Agentic 3

GDPVal54.824.3
Terminal-Bench 4.042.40

Instruction Following 1

IFBench63.567.1

Long Context 1

AA-LCR82.373

Multimodal 2

LMArena · Vision1326 ◆◆ 1247
MMMU-Pro84.275.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.

Compare Claude Fable 5 withALL PAIRINGS →