Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

AI Model Tracker — LLM benchmarks
& leaderboard.

GPT-6 Astra currently leads 2 of 10 capability axes. The five leading models on each, on one honest scale — composite scores are weighted averages over each axis's audited benchmark basket.

10 AXES·480 MODELS·44 BENCHMARKSUPDATED OCT 9 · 06:19 UTCScore basis
Tier 1 — Established axesHigh coverage, audited harnesses, narrow confidence.

Reasoning

480 MODELS SCOREDDEAD HEAT
Composite /100 · Most authoritative basis
2AClaude Fable 5.192.2−0.9
3AClaude Opus 590.9−2.2
4OGPT-5.6 Sol90.7−2.4
5AClaude Fable 589.9−3.1
LeadingFieldWINDOW 89.9–93.1 / 100FULL STANDINGS →

Coding

259 MODELS SCOREDCONTESTED
Composite /100 · Most authoritative basis
2AClaude Fable 5.179.3−3.7
3AClaude Opus 578.2−4.9
4OGPT-5.6 Sol77.9−5.1
5AClaude Opus 5.575.4−7.6
LeadingFieldWINDOW 75.4–83.0 / 100FULL STANDINGS →

Agentic

325 MODELS SCOREDCONTESTED
Composite /100 · Most authoritative basis
2AClaude Fable 582.1−3.2
3OGPT-5.6 Sol81.6−3.6
4KKimi K380.5−4.8
5AClaude Fable 5.179.0−6.2
LeadingFieldWINDOW 79.0–85.2 / 100FULL STANDINGS →

Safety

DATA-THIN
Composite /100 · Most authoritative basis
Scale's SEAL safety boards (MASK honesty, Fortress adversarial risk, PropensityBench misuse propensity) now flow daily — 46 models measured on two or more. The cross-vendor safety ranking is parked pending calibration research; board scores appear on each model's page, labeled, with direction stated (Fortress and PropensityBench count risk — lower is better).
Tier 2 — Emerging axesSufficient signal to rank; bands wider where coverage thins.

Long Context

372 MODELS SCOREDTHIN BASKETCONTESTED
Composite /100 · Most authoritative basis
2GGemini 3.7 Flash75.4−1.0
4KKimi K373.5−3.0
5SStep 5 Preview73.4−3.1
LeadingFieldWINDOW 73.4–76.4 / 100FULL STANDINGS →

Math

80 MODELS SCOREDCONTESTED
Composite /100 · Most authoritative basis
2OGPT-6 Astra79.3−1.2
3AClaude Opus 5.577.8−2.7
4OGPT-6 Sol76.4−4.1
5AClaude Fable 5.175.6−4.9
LeadingFieldWINDOW 75.6–80.5 / 100FULL STANDINGS →

Multimodal

190 MODELS SCOREDCONTESTED
Composite /100 · Most authoritative basis
3KKimi K2.570.1−4.9
4GGemini 3.1 Pro68.5−6.5
5QQwen3.5 122B A10B68.0−7.0
LeadingFieldWINDOW 68.0–75.0 / 100FULL STANDINGS →

Multilingual

DATA-THIN
Composite /100 · Most authoritative basis
Scale's MultiNRC multilingual board now flows (36 models, 8 vendors — French, Spanish and Chinese). A single-board three-language ranking is parked pending research; multilingual scores appear on model pages. Re-opens with a broader source or the research pass.

Factuality

357 MODELS SCOREDDEAD HEAT
Composite /100 · Most authoritative basis
2OGPT-6.1 Sol78.7−0.8
3AClaude Opus 5.577.5−2.0
4GGemini 3.1 Pro77.2−2.2
5GGemini 3.8 Flash76.2−3.3
LeadingFieldWINDOW 76.2–79.5 / 100FULL STANDINGS →
Tier 3 — Provisional axesSmall basket or volatile field; treat as directional.

Instruction Following

355 MODELS SCOREDTHIN BASKETCLEAR LEADER
Composite /100 · Most authoritative basis
3QQwen3.8 Flash Next79.4−6.2
5GGemini 3.5 Flash76.1−9.5
LeadingFieldWINDOW 76.1–85.6 / 100FULL STANDINGS →
Composite — absolute weighted average over each axis basket. Methodology & sources →

Head-to-head model comparisons → curated pairs scored on shared benchmarks