Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Sol vs MiMo V2.5

29 SHARED BENCHMARKS

Across 29 shared benchmarks, GPT-5.6 Sol scores higher on 26 and MiMo V2.5 on 3. The widest gap is AA-Omniscience · Non-hallucination, where MiMo V2.5 scores 68.1 against 7.8. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $4.00 per million input tokens).

OPENAIVSXIAOMI29 SHARED26–3 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIXiaomi
Released
Price per 1M tokens input / output$4.00 / $20.00$0.14 / $0.28
Cost of 1M in + 1M out$24.00$0.42 57.1× less
Head-to-head of 29 shared benchmarks26 wins3 wins
Scores tracked independently verified219 31 ◆51 3 ◆

Release dates per Artificial Analysis. Prices: OpenAI's own price page for GPT-5.6 Sol; Artificial Analysis for MiMo V2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 SolWINSMiMo V2.5
Reasoning30GPT-5.6 Sol leads 3 of 3 · widest: CritPt 32.3 vs 3.7
Coding50GPT-5.6 Sol leads 5 of 5 · widest: Terminal-Bench Hard 65.9 vs 41.7
Agentic30GPT-5.6 Sol leads 3 of 3 · widest: τ-Bench V3 · Banking 44.3 vs 8.7
Factuality21GPT-5.6 Sol leads 2 of 3 · widest: SimpleQA Verified 69.7 vs 16.1
Instruction Following10GPT-5.6 Sol leads 1 of 1 · widest: IFBench 72.7 vs 67.1
Long Context10GPT-5.6 Sol leads 1 of 1 · widest: AA-LCR 84 vs 73
Math10GPT-5.6 Sol leads 1 of 1 · widest: AIME 2026 99.9 vs 93.6
Multimodal30GPT-5.6 Sol leads 3 of 3 · widest: MMMU-Pro 83.4 vs 75.4

Biggest gaps

GPT-5.6 Sol pulls furthest ahead on

  1. CritPt32.3 vs 3.7
  2. τ-Bench V3 · Banking44.3 vs 8.7
  3. SimpleQA Verified69.7 vs 16.1

MiMo V2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination68.1 vs 7.8

Every shared benchmark29 · grouped by area

Reasoning 3

Coding 5

Agentic 3

BenchmarkGPT-5.6 SolMARGINMiMo V2.5
GDPVal54.424.3
Terminal-Bench 4.039.90

Instruction Following 1

BenchmarkGPT-5.6 SolMARGINMiMo V2.5
IFBench72.767.1

Long Context 1

BenchmarkGPT-5.6 SolMARGINMiMo V2.5
AA-LCR8473

Math 1

BenchmarkGPT-5.6 SolMARGINMiMo V2.5
AIME 202699.993.6

Multimodal 3

BenchmarkGPT-5.6 SolMARGINMiMo V2.5
LMArena · Vision1280 ◆◆ 1247
MMMU-Pro83.475.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (openai-official, direct), otherwise the lowest tracked offer.

Compare GPT-5.6 Sol withALL PAIRINGS →