Best LLMs for Reasoning
GPT-6 Astra leads Claude Fable 5.1 by just 0.9 at 93.1 — inside the margin, with no model clear of the pack. The median of the 480-model field trails the leader by 46.4 points.
- 1GPT-6 AstraOpenAI93.1Reasoning score, out of 1006 of 6 benchmarks measured4 independent sources0.9 ahead of Claude Fable 5.146.4 above the field medianCompare the top two
- 2Claude Fable 5.1AnthropicWithin 1% of the leader−0.992.26 of 6
- 3Claude Opus 5Anthropic−2.290.96 of 6
- 4GPT-5.6 SolOpenAI−2.490.76 of 6
- 5Claude Fable 5Anthropic−3.189.96 of 6
- 6Claude Opus 5.5Anthropic−4.888.35 of 6
- 7Gemini 3.8 FlashGoogle−5.287.96 of 6
- 8GPT-6.1 SolOpenAI−5.487.65 of 6
- 9GPT-5.5OpenAI−5.687.56 of 6
- 10Grok 4.6SpaceXAI−7.585.66 of 6
- 11Claude Opus 4.8Anthropic−8.085.16 of 6
- 12GPT-5.6 TerraOpenAI−8.284.96 of 6
Show all 200 ranked modelsShow the top 12 only
- 13Muse Spark 1.3Meta−8.584.65 of 6
- 14Gemini 3.1 ProGoogle−8.584.56 of 6
- 15Kimi K3Moonshot−8.684.46 of 6
- 16GPT-6 SolOpenAI−9.283.95 of 6
- 17GPT-5.5 ProOpenAI−9.983.24 of 6
- 18GPT-5.4OpenAI−10.282.95 of 6
- 19Gemini 3.7 FlashGoogle−10.382.85 of 6
- 20Muse Spark 1.2Meta−11.481.75 of 6
- 21Claude Opus 4.7Anthropic−12.580.66 of 6
- 22Gemini 3.5 FlashGoogle−12.680.56 of 6
- 23Claude Opus 4.6Anthropic−12.880.36 of 6
- 24Grok4.5SpaceXAI−12.880.36 of 6
- 25DeepSeek-V4.1-FlashDeepSeek−12.980.26 of 6
- 26Qwen3.8 2.4T A95BAlibaba−13.679.54 of 6
- 27GLM-5.3Z.ai−13.979.25 of 6
- 28Qwen3.8 Max PreviewAlibaba−14.478.74 of 6
- 29Claude Sonnet 5Anthropic−14.478.75 of 6
- 30GPT-5.6 LunaOpenAI−15.178.06 of 6
- 31Muse Spark 1.1Meta−15.677.44 of 6
- 32Qwen3.7 MaxAlibaba−16.077.15 of 6
- 33DeepSeek-V4-FlashDeepSeek−16.176.96 of 6
- 34GPT-5.3 CodexOpenAI−16.376.83 of 6
- 35Gemini 3.6 FlashGoogle−16.576.65 of 6
- 36Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−16.576.62 of 6
- 37Gemini 3 ProGoogle−16.776.45 of 6
- 38GLM-5.2Z.ai−16.876.36 of 6
- 39GLM 5.3 FlashZ.ai−16.976.25 of 6
- 40Gemini 4 ArgonGoogle−17.076.12 of 6
- 41Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−17.275.92 of 6
- 42MiMo V2.6 ProXiaomi−17.775.42 of 6
- 43Agnes 3.0 FlashSapiens AI−17.975.23 of 6
- 44Muse SparkMeta−18.075.14 of 6
- 45JT-4.1 Flash 236B A21BChina Mobile−18.374.83 of 6
- 46Agnes 2.5 Pro BetaSapiens AI−18.374.83 of 6
- 47Qwen3.8 Flash NextAlibaba−18.474.74 of 6
- 48Step 5 PreviewStepFun−18.774.42 of 6
- 49Grok 4.7SpaceXAI−18.974.24 of 6
- 50Gemini 3 Flash PreviewGoogle−19.773.45 of 6
- 51GPT-5.2OpenAI−19.773.46 of 6
- 52Haiku 5.5Anthropic−19.773.43 of 6
- 53DeepSeek-V4-ProDeepSeek−20.672.55 of 6
- 54GPT-6 LunaOpenAI−20.772.44 of 6
- 55Grok 4.20SpaceXAI−20.872.34 of 6
- 56Qwen3.7 Plus PreviewAlibaba−21.072.13 of 6
- 57Motif-3-BetaMotif Technologies−21.072.03 of 6
- 58Kimi K2.7 CodeMoonshot−21.172.05 of 6
- 59GPT-5.4 ProOpenAI−21.172.03 of 6
- 60Ling-3.1-flashInclusionAI−21.172.02 of 6
- 61Agnes 2.5 Pro AlphaSapiens AI−21.571.63 of 6
- 62Claude 5.5Anthropic−21.571.62 of 6
- 63Inkling-SmallThinking Machines−21.771.44 of 6
- 64Nex-N2-ProNex AGI−22.071.13 of 6
- 65Kimi K2.6Moonshot−22.470.74 of 6
- 66Qwen3.8 27BAlibaba−22.570.66 of 6
- 67Motif-3Motif Technologies−22.970.13 of 6
- 68GPT-5.2 CodexOpenAI−23.070.14 of 6
- 69Qwen3.6 Max PreviewAlibaba−23.469.74 of 6
- 70Hy3-previewTencent−23.469.73 of 6
- 71Claude Sonnet 4.6Anthropic−23.569.65 of 6
- 72A.X-K2SK Telecom−23.669.53 of 6
- 73Grok 4.3SpaceXAI−23.769.44 of 6
- 74MiMo V2.5 ProXiaomi−23.869.33 of 6
- 75Apodex 1.1Apodex−24.069.13 of 6
- 76Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−24.169.02 of 6
- 77MiMo V2.6 Flash RLXiaomi−24.169.02 of 6
- 78DeepSeek-V3.2-SpecialeDeepSeek−24.169.04 of 6
- 79Solar Pro4Upstage−24.268.93 of 6
- 80Claude Mythos PreviewAnthropic−24.468.72 of 6
- 81K2 Horizon 375B A23BMBZUAI−24.468.73 of 6
- 82GLM-5.1Z.ai−24.668.54 of 6
- 83Sonnet 5.5Anthropic−24.868.33 of 6
- 84Claude Mythos 5.1Anthropic−24.868.32 of 6
- 85MiniMax M3MiniMax−24.868.25 of 6
- 86Solar Open2 250BUpstage−24.968.23 of 6
- 87Mistral Large 4Mistral−25.267.93 of 6
- 88Claude Mythos 5Anthropic−25.367.82 of 6
- 89GPT-5.1OpenAI−25.367.85 of 6
- 90GPT-5OpenAI−25.467.75 of 6
- 91InklingThinking Machines−25.667.56 of 6
- 92GPT-5.1 CodexOpenAI−25.767.43 of 6
- 93Claude Opus 4.5Anthropic−25.767.45 of 6
- 94Kimi K2.5Moonshot−25.767.44 of 6
- 95GPT-5.4 miniOpenAI−25.867.35 of 6
- 96GPT-5 CodexOpenAI−26.266.93 of 6
- 97Gemini 3 Deep ThinkGoogle−26.566.62 of 6
- 98MiMo V2.5Xiaomi−26.866.33 of 6
- 99GPT-5.2 ProOpenAI−27.066.04 of 6
- 100LongCat-2.0Meituan−27.265.93 of 6
- 101Qwen3.5 397B A17BAlibaba−27.265.93 of 6
- 102Hy4Tencent−27.365.82 of 6
- 103GPT-5.4 nanoOpenAI−27.465.75 of 6
- 104MiMo V2 FlashXiaomi−27.665.53 of 6
- 105Seed 2.1 Pro PreviewByteDance−27.765.42 of 6
- 106Qwen3 Max (Reasoning)Alibaba−28.165.03 of 6
- 107Grok 4SpaceXAI−28.264.94 of 6
- 108Qwen3.6 PlusAlibaba−28.764.34 of 6
- 109Grok 4.1 FastSpaceXAI−28.864.34 of 6
- 110GLM-4.7Z.ai−29.164.04 of 6
- 111Ling-3.0-flashInclusionAI−29.164.03 of 6
- 112GPT-5.5 InstantOpenAI−29.164.03 of 6
- 113Ling-3.0-flash-VLInclusionAI−29.263.83 of 6
- 114Muse GlimmerMeta−29.363.83 of 6
- 115K2 Horizon 36B-A4BMBZUAI−29.563.63 of 6
- 116Nemotron 3 Super 120B A12BNVIDIA−29.663.53 of 6
- 117Gemma 4 31BGoogle−29.663.53 of 6
- 118Nemotron 3 Ultra 550B A55BNVIDIA−29.763.45 of 6
- 119MiniMax M2.7MiniMax−29.763.43 of 6
- 120Step 3.5 FlashStepFun−29.963.13 of 6
- 121GLM-5Z.ai−30.063.15 of 6
- 122DeepSeek-V3.2DeepSeek−30.362.84 of 6
- 123Step 3.7 FlashStepFun−30.462.73 of 6
- 124Kimi K2 ThinkingMoonshot−30.962.24 of 6
- 125Grok 4 FastSpaceXAI−31.062.14 of 6
- 126Qwen3.5 27BAlibaba−31.062.13 of 6
- 127MiMo V2 ProXiaomi−31.461.73 of 6
- 128MiMo V2 OmniXiaomi−31.461.73 of 6
- 129Qwen3.5 122B A10BAlibaba−31.561.63 of 6
- 130Ling-3.0-flash-FinInclusionAI−31.761.42 of 6
- 131Qwen3.5 35B A3BAlibaba−32.260.93 of 6
- 132Solar Mini4Upstage−32.360.82 of 6
- 133Gemini 2.5 ProGoogle−32.460.75 of 6
- 134DeepSeek-V3.1-TerminusDeepSeek−32.760.43 of 6
- 135GLM-5-TurboZ.ai−33.060.13 of 6
- 136Claude Sonnet 4.5Anthropic−33.060.15 of 6
- 137Gemini 3.1 Flash Lite PreviewGoogle−33.160.03 of 6
- 138O3OpenAI−33.160.05 of 6
- 139K-EXAONE 2.0LG AI−33.259.93 of 6
- 140MiniMax M2.5MiniMax−33.359.84 of 6
- 141Qwen3.6 27BAlibaba−33.559.64 of 6
- 142DeepSeek-V3.2-ExpDeepSeek−33.559.53 of 6
- 143Qwen3.6 35B A3BAlibaba−34.358.83 of 6
- 144ERNIE 5.0 Thinking PreviewBaidu−34.658.53 of 6
- 145GLM-4.6Z.ai−34.658.53 of 6
- 146DeepSeek-V3.1DeepSeek−34.758.44 of 6
- 147MiMo V2.5 Pro FP4 DFlashXiaomi−34.758.41 of 6
- 148GLM 5V TurboZ.ai−34.958.23 of 6
- 149K-EXAONELG AI−35.158.03 of 6
- 150Mercury 2Inception−35.157.93 of 6
- 151Trinity Large ThinkingArcee AI−35.657.53 of 6
- 152GPT Oss 120bOpenAI−35.757.44 of 6
- 153Qwen3.5 Omni PlusAlibaba−35.857.33 of 6
- 154MiniMax M2MiniMax−35.857.23 of 6
- 155MiniMax M2.1MiniMax−35.957.24 of 6
- 156MAI-Code-1-FlashMicrosoft−35.957.22 of 6
- 157G9v3-39A5BAI9Stars−36.556.63 of 6
- 158O4 MiniOpenAI−36.656.44 of 6
- 159Apriel-v1.5-15B-ThinkerServiceNow−36.956.23 of 6
- 160K2 Horizon 7BMBZUAI−37.255.93 of 6
- 161Qwen3.5 9BAlibaba−37.355.83 of 6
- 162NVIDIA Nemotron 3 Nano 30B A3BNVIDIA−37.455.73 of 6
- 163Nemotron Cascade 2 30B A3BNVIDIA−37.755.43 of 6
- 164GPT Oss 20bOpenAI−37.755.43 of 6
- 165Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−37.955.23 of 6
- 166Ring-1TInclusionAI−38.055.13 of 6
- 167EXAONE 4.5 33BLG AI−38.254.83 of 6
- 168Doubao Seed CodeByteDance−38.354.83 of 6
- 169Qwen3 235B ThinkingAlibaba−38.754.41 of 6
- 170Nova 2.0 LiteAmazon−38.854.33 of 6
- 171INTELLECT-3Prime Intellect−38.854.33 of 6
- 172Claude Opus 4Anthropic−38.954.24 of 6
- 173Gemini 2.5 FlashGoogle−39.253.95 of 6
- 174Seed 2.0ByteDance−39.353.81 of 6
- 175NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16NVIDIA−39.353.81 of 6
- 176GPT-5.1 InstantOpenAI−39.453.71 of 6
- 177PaCoRe-8BStepFun−39.453.71 of 6
- 178Command A+Cohere−39.453.73 of 6
- 179North-Mini-Code-1.0Cohere−40.252.93 of 6
- 180MAI-Thinking-1Microsoft−40.352.81 of 6
- 181Apriel-v1.6-15B-ThinkerServiceNow−41.251.93 of 6
- 182O1 ProOpenAI−41.251.91 of 6
- 183Nemotron 3 SuperNVIDIA−41.251.81 of 6
- 184Mistral Small 4Mistral−41.351.83 of 6
- 185Magistral Medium 1.2Mistral−41.551.63 of 6
- 186Granite 4.2 30BIBM−41.851.33 of 6
- 187Grok 3 BetaSpaceXAI−41.951.21 of 6
- 188Falcon H1R 7BTII−42.151.03 of 6
- 189Grok 3 Mini ReasoningSpaceXAI−42.151.04 of 6
- 190O3 ProOpenAI−42.250.92 of 6
- 191Nemotron 3 NanoNVIDIA−42.250.91 of 6
- 192ERNIE 4.5Baidu−42.250.91 of 6
- 193DeepSeek-R1-ZeroDeepSeek−42.450.61 of 6
- 194Ling-1TInclusionAI−42.650.53 of 6
- 195Ling 2.6 1tInclusionAI−42.650.53 of 6
- 196Claude Sonnet 4Anthropic−42.750.45 of 6
- 197Phi 4 Reasoning PlusMicrosoft−43.150.01 of 6
- 198MiniCPM5-2BOpenBMB−43.249.93 of 6
- 199MiMo V2.5 Pro BaseXiaomi−43.449.71 of 6
- 200Grok 3 Mini BetaSpaceXAI−43.549.61 of 6
The benchmarks behind the reasoning ranking
Six public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.
- 28%CritPt
Research-level physics problems
Best on this testGPT-5.6 Sol32.3
- 27%Humanity's Last Exam
Very hard expert questions across many subjects
Best on this testClaude Mythos Preview64.7
- 15%ARC-AGI-2
Abstract visual puzzles that people can solve
Best on this testGPT-6 Astra95.0
- 10%GPQA Diamond
Graduate-level biology, physics and chemistry questions
Best on this testGPT-6 Astra96.1
- 10%SimpleBench
Common-sense trick questions
Best on this testClaude Opus 5.588.4
- 10%LiveBench Reasoning
Fresh logic puzzles that change regularly
Best on this testGPT-6 Astra92.7
- Tracked, not ranked
ARC-AGI-3 — solves new puzzle games it has never seen. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.
Why these weights
Ranked on critpt, HLE, ARC-AGI-2/3, GPQA Diamond, SimpleBench and LiveBench reasoning (2026-Q3 v2, ratified 2026-08-23). Weights follow measured frontier separation, construct centrality and field recognition; no saturated benchmark ranks.
- CritPtresearch-physics; strongest measured frontier separation (D .43), AA day-0
- Humanity's Last Exambroad frontier exam; D .15, field standard (82 card sources), AA day-0
- ARC-AGI-2abstract reasoning; ARC Prize board, ~9d
- GPQA Diamondscience QA; compressing at the top (D .02) — operator floor .10, anchor for presence
- SimpleBenchrobustness facet, single site source, ~9d — operator floor .10
- LiveBench Reasoningrolling contamination-free reasoning (43 models/12 labs, continuous) — probation, operator floor .10
How the reasoning score is calculated
- 1
Rank on each benchmark
Every model gets a percentile on each benchmark it has been measured on.
- 2
Steady the thin fields
Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.
- 3
Weigh and average
The percentiles are averaged with the weights above into one score out of 100.
- 4
Qualify
A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.
Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.
Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.
Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.
The gap. Points behind the leader on the reasoning score.
480 models from 53 vendors have a reasoning score, each measured on at least 3 of the 6 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.
Questions about the reasoning ranking
Which LLM is best at reasoning right now?
GPT-6 Astra, with a reasoning score of 93.1 out of 100. Claude Fable 5.1 is 0.9 points behind at 92.2 — inside the margin, so the top of the ranking is a dead heat.
Which benchmarks make up the reasoning score?
Six public benchmarks: CritPt (28%), Humanity's Last Exam (27%), ARC-AGI-2 (15%), and GPQA Diamond, SimpleBench and LiveBench Reasoning (10% each).
How many models are ranked?
480 models from 53 vendors have a reasoning score. Each needs results on at least 3 of the 6 benchmarks to be ranked.
How current is the ranking?
It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.