AI Model Tracker — LLM benchmarks
& leaderboard.
GPT-6 Astra currently leads 2 of 10 capability axes. The five leading models on each, on one honest scale — composite scores are weighted averages over each axis's audited benchmark basket.
10 AXES·480 MODELS·44 BENCHMARKSUPDATED OCT 9 · 06:19 UTCScore basis
Tier 1 — Established axesHigh coverage, audited harnesses, narrow confidence.
Reasoning
480 MODELS SCOREDDEAD HEATComposite /100 · Most authoritative basis
LeadingFieldWINDOW 89.9–93.1 / 100FULL STANDINGS →
Coding
259 MODELS SCOREDCONTESTEDComposite /100 · Most authoritative basis
LeadingFieldWINDOW 75.4–83.0 / 100FULL STANDINGS →
Agentic
325 MODELS SCOREDCONTESTEDComposite /100 · Most authoritative basis
LeadingFieldWINDOW 79.0–85.2 / 100FULL STANDINGS →
Safety
DATA-THINComposite /100 · Most authoritative basis
Scale's SEAL safety boards (MASK honesty, Fortress adversarial risk, PropensityBench misuse propensity) now flow daily — 46 models measured on two or more. The cross-vendor safety ranking is parked pending calibration research; board scores appear on each model's page, labeled, with direction stated (Fortress and PropensityBench count risk — lower is better).
Tier 2 — Emerging axesSufficient signal to rank; bands wider where coverage thins.
Long Context
372 MODELS SCOREDTHIN BASKETCONTESTEDComposite /100 · Most authoritative basis
LeadingFieldWINDOW 73.4–76.4 / 100FULL STANDINGS →
Math
80 MODELS SCOREDCONTESTEDComposite /100 · Most authoritative basis
LeadingFieldWINDOW 75.6–80.5 / 100FULL STANDINGS →
Multimodal
190 MODELS SCOREDCONTESTEDComposite /100 · Most authoritative basis
LeadingFieldWINDOW 68.0–75.0 / 100FULL STANDINGS →
Multilingual
DATA-THINComposite /100 · Most authoritative basis
Scale's MultiNRC multilingual board now flows (36 models, 8 vendors — French, Spanish and Chinese). A single-board three-language ranking is parked pending research; multilingual scores appear on model pages. Re-opens with a broader source or the research pass.
Tier 3 — Provisional axesSmall basket or volatile field; treat as directional.
Composite — absolute weighted average over each axis basket. Methodology & sources →
Head-to-head model comparisons → curated pairs scored on shared benchmarks