Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best LLMs for Reasoning

Tier 1 · establishedDead heat at the top

GPT-6 Astra leads Claude Fable 5.1 by just 0.9 at 93.1 — inside the margin, with no model clear of the pack. The median of the 480-model field trails the leader by 46.4 points.

Updated 9 Oct 2026Score basis
  1. 1GPT-6 AstraOpenAI93.1Reasoning score, out of 100

    CritPt28%

    31.7Aggregator

    Top 32.3 · GPT-5.6 Sol

    Humanity's Last Exam27%

    54.7Aggregator

    Top 64.7 · Claude Mythos Preview

    ARC-AGI-215%

    95.0Vendor card

    Best score on this test

    GPQA Diamond10%

    96.1Aggregator

    Best score on this test

    SimpleBench10%

    83.6Independent

    Top 88.4 · Claude Opus 5.5

    LiveBench Reasoning10%

    92.7Independent

    Best score on this test

    6 of 6 benchmarks measured4 independent sources0.9 ahead of Claude Fable 5.146.4 above the field medianCompare the top two
188 more ranked, from 84.6 down to 49.6280 further models score below 49.6Field median 46.7, across all 480 scored
Show all 200 ranked modelsShow the top 12 only
  1. 13Muse Spark 1.3Meta−8.584.65 of 6
  2. 14Gemini 3.1 ProGoogle−8.584.56 of 6
  3. 15Kimi K3Moonshot−8.684.46 of 6
  4. 16GPT-6 SolOpenAI−9.283.95 of 6
  5. 17GPT-5.5 ProOpenAI−9.983.24 of 6
  6. 18GPT-5.4OpenAI−10.282.95 of 6
  7. 19Gemini 3.7 FlashGoogle−10.382.85 of 6
  8. 20Muse Spark 1.2Meta−11.481.75 of 6
  9. 21Claude Opus 4.7Anthropic−12.580.66 of 6
  10. 22Gemini 3.5 FlashGoogle−12.680.56 of 6
  11. 23Claude Opus 4.6Anthropic−12.880.36 of 6
  12. 24Grok4.5SpaceXAI−12.880.36 of 6
  13. 25DeepSeek-V4.1-FlashDeepSeek−12.980.26 of 6
  14. 26Qwen3.8 2.4T A95BAlibaba−13.679.54 of 6
  15. 27GLM-5.3Z.ai−13.979.25 of 6
  16. 28Qwen3.8 Max PreviewAlibaba−14.478.74 of 6
  17. 29Claude Sonnet 5Anthropic−14.478.75 of 6
  18. 30GPT-5.6 LunaOpenAI−15.178.06 of 6
  19. 31Muse Spark 1.1Meta−15.677.44 of 6
  20. 32Qwen3.7 MaxAlibaba−16.077.15 of 6
  21. 33DeepSeek-V4-FlashDeepSeek−16.176.96 of 6
  22. 34GPT-5.3 CodexOpenAI−16.376.83 of 6
  23. 35Gemini 3.6 FlashGoogle−16.576.65 of 6
  24. 36Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−16.576.62 of 6
  25. 37Gemini 3 ProGoogle−16.776.45 of 6
  26. 38GLM-5.2Z.ai−16.876.36 of 6
  27. 39GLM 5.3 FlashZ.ai−16.976.25 of 6
  28. 40Gemini 4 ArgonGoogle−17.076.12 of 6
  29. 41Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−17.275.92 of 6
  30. 42MiMo V2.6 ProXiaomi−17.775.42 of 6
  31. 43Agnes 3.0 FlashSapiens AI−17.975.23 of 6
  32. 44Muse SparkMeta−18.075.14 of 6
  33. 45JT-4.1 Flash 236B A21BChina Mobile−18.374.83 of 6
  34. 46Agnes 2.5 Pro BetaSapiens AI−18.374.83 of 6
  35. 47Qwen3.8 Flash NextAlibaba−18.474.74 of 6
  36. 48Step 5 PreviewStepFun−18.774.42 of 6
  37. 49Grok 4.7SpaceXAI−18.974.24 of 6
  38. 50Gemini 3 Flash PreviewGoogle−19.773.45 of 6
  39. 51GPT-5.2OpenAI−19.773.46 of 6
  40. 52Haiku 5.5Anthropic−19.773.43 of 6
  41. 53DeepSeek-V4-ProDeepSeek−20.672.55 of 6
  42. 54GPT-6 LunaOpenAI−20.772.44 of 6
  43. 55Grok 4.20SpaceXAI−20.872.34 of 6
  44. 56Qwen3.7 Plus PreviewAlibaba−21.072.13 of 6
  45. 57Motif-3-BetaMotif Technologies−21.072.03 of 6
  46. 58Kimi K2.7 CodeMoonshot−21.172.05 of 6
  47. 59GPT-5.4 ProOpenAI−21.172.03 of 6
  48. 60Ling-3.1-flashInclusionAI−21.172.02 of 6
  49. 61Agnes 2.5 Pro AlphaSapiens AI−21.571.63 of 6
  50. 62Claude 5.5Anthropic−21.571.62 of 6
  51. 63Inkling-SmallThinking Machines−21.771.44 of 6
  52. 64Nex-N2-ProNex AGI−22.071.13 of 6
  53. 65Kimi K2.6Moonshot−22.470.74 of 6
  54. 66Qwen3.8 27BAlibaba−22.570.66 of 6
  55. 67Motif-3Motif Technologies−22.970.13 of 6
  56. 68GPT-5.2 CodexOpenAI−23.070.14 of 6
  57. 69Qwen3.6 Max PreviewAlibaba−23.469.74 of 6
  58. 70Hy3-previewTencent−23.469.73 of 6
  59. 71Claude Sonnet 4.6Anthropic−23.569.65 of 6
  60. 72A.X-K2SK Telecom−23.669.53 of 6
  61. 73Grok 4.3SpaceXAI−23.769.44 of 6
  62. 74MiMo V2.5 ProXiaomi−23.869.33 of 6
  63. 75Apodex 1.1Apodex−24.069.13 of 6
  64. 76Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−24.169.02 of 6
  65. 77MiMo V2.6 Flash RLXiaomi−24.169.02 of 6
  66. 78DeepSeek-V3.2-SpecialeDeepSeek−24.169.04 of 6
  67. 79Solar Pro4Upstage−24.268.93 of 6
  68. 80Claude Mythos PreviewAnthropic−24.468.72 of 6
  69. 81K2 Horizon 375B A23BMBZUAI−24.468.73 of 6
  70. 82GLM-5.1Z.ai−24.668.54 of 6
  71. 83Sonnet 5.5Anthropic−24.868.33 of 6
  72. 84Claude Mythos 5.1Anthropic−24.868.32 of 6
  73. 85MiniMax M3MiniMax−24.868.25 of 6
  74. 86Solar Open2 250BUpstage−24.968.23 of 6
  75. 87Mistral Large 4Mistral−25.267.93 of 6
  76. 88Claude Mythos 5Anthropic−25.367.82 of 6
  77. 89GPT-5.1OpenAI−25.367.85 of 6
  78. 90GPT-5OpenAI−25.467.75 of 6
  79. 91InklingThinking Machines−25.667.56 of 6
  80. 92GPT-5.1 CodexOpenAI−25.767.43 of 6
  81. 93Claude Opus 4.5Anthropic−25.767.45 of 6
  82. 94Kimi K2.5Moonshot−25.767.44 of 6
  83. 95GPT-5.4 miniOpenAI−25.867.35 of 6
  84. 96GPT-5 CodexOpenAI−26.266.93 of 6
  85. 97Gemini 3 Deep ThinkGoogle−26.566.62 of 6
  86. 98MiMo V2.5Xiaomi−26.866.33 of 6
  87. 99GPT-5.2 ProOpenAI−27.066.04 of 6
  88. 100LongCat-2.0Meituan−27.265.93 of 6
  89. 101Qwen3.5 397B A17BAlibaba−27.265.93 of 6
  90. 102Hy4Tencent−27.365.82 of 6
  91. 103GPT-5.4 nanoOpenAI−27.465.75 of 6
  92. 104MiMo V2 FlashXiaomi−27.665.53 of 6
  93. 105Seed 2.1 Pro PreviewByteDance−27.765.42 of 6
  94. 106Qwen3 Max (Reasoning)Alibaba−28.165.03 of 6
  95. 107Grok 4SpaceXAI−28.264.94 of 6
  96. 108Qwen3.6 PlusAlibaba−28.764.34 of 6
  97. 109Grok 4.1 FastSpaceXAI−28.864.34 of 6
  98. 110GLM-4.7Z.ai−29.164.04 of 6
  99. 111Ling-3.0-flashInclusionAI−29.164.03 of 6
  100. 112GPT-5.5 InstantOpenAI−29.164.03 of 6
  101. 113Ling-3.0-flash-VLInclusionAI−29.263.83 of 6
  102. 114Muse GlimmerMeta−29.363.83 of 6
  103. 115K2 Horizon 36B-A4BMBZUAI−29.563.63 of 6
  104. 116Nemotron 3 Super 120B A12BNVIDIA−29.663.53 of 6
  105. 117Gemma 4 31BGoogle−29.663.53 of 6
  106. 118Nemotron 3 Ultra 550B A55BNVIDIA−29.763.45 of 6
  107. 119MiniMax M2.7MiniMax−29.763.43 of 6
  108. 120Step 3.5 FlashStepFun−29.963.13 of 6
  109. 121GLM-5Z.ai−30.063.15 of 6
  110. 122DeepSeek-V3.2DeepSeek−30.362.84 of 6
  111. 123Step 3.7 FlashStepFun−30.462.73 of 6
  112. 124Kimi K2 ThinkingMoonshot−30.962.24 of 6
  113. 125Grok 4 FastSpaceXAI−31.062.14 of 6
  114. 126Qwen3.5 27BAlibaba−31.062.13 of 6
  115. 127MiMo V2 ProXiaomi−31.461.73 of 6
  116. 128MiMo V2 OmniXiaomi−31.461.73 of 6
  117. 129Qwen3.5 122B A10BAlibaba−31.561.63 of 6
  118. 130Ling-3.0-flash-FinInclusionAI−31.761.42 of 6
  119. 131Qwen3.5 35B A3BAlibaba−32.260.93 of 6
  120. 132Solar Mini4Upstage−32.360.82 of 6
  121. 133Gemini 2.5 ProGoogle−32.460.75 of 6
  122. 134DeepSeek-V3.1-TerminusDeepSeek−32.760.43 of 6
  123. 135GLM-5-TurboZ.ai−33.060.13 of 6
  124. 136Claude Sonnet 4.5Anthropic−33.060.15 of 6
  125. 137Gemini 3.1 Flash Lite PreviewGoogle−33.160.03 of 6
  126. 138O3OpenAI−33.160.05 of 6
  127. 139K-EXAONE 2.0LG AI−33.259.93 of 6
  128. 140MiniMax M2.5MiniMax−33.359.84 of 6
  129. 141Qwen3.6 27BAlibaba−33.559.64 of 6
  130. 142DeepSeek-V3.2-ExpDeepSeek−33.559.53 of 6
  131. 143Qwen3.6 35B A3BAlibaba−34.358.83 of 6
  132. 144ERNIE 5.0 Thinking PreviewBaidu−34.658.53 of 6
  133. 145GLM-4.6Z.ai−34.658.53 of 6
  134. 146DeepSeek-V3.1DeepSeek−34.758.44 of 6
  135. 147MiMo V2.5 Pro FP4 DFlashXiaomi−34.758.41 of 6
  136. 148GLM 5V TurboZ.ai−34.958.23 of 6
  137. 149K-EXAONELG AI−35.158.03 of 6
  138. 150Mercury 2Inception−35.157.93 of 6
  139. 151Trinity Large ThinkingArcee AI−35.657.53 of 6
  140. 152GPT Oss 120bOpenAI−35.757.44 of 6
  141. 153Qwen3.5 Omni PlusAlibaba−35.857.33 of 6
  142. 154MiniMax M2MiniMax−35.857.23 of 6
  143. 155MiniMax M2.1MiniMax−35.957.24 of 6
  144. 156MAI-Code-1-FlashMicrosoft−35.957.22 of 6
  145. 157G9v3-39A5BAI9Stars−36.556.63 of 6
  146. 158O4 MiniOpenAI−36.656.44 of 6
  147. 159Apriel-v1.5-15B-ThinkerServiceNow−36.956.23 of 6
  148. 160K2 Horizon 7BMBZUAI−37.255.93 of 6
  149. 161Qwen3.5 9BAlibaba−37.355.83 of 6
  150. 162NVIDIA Nemotron 3 Nano 30B A3BNVIDIA−37.455.73 of 6
  151. 163Nemotron Cascade 2 30B A3BNVIDIA−37.755.43 of 6
  152. 164GPT Oss 20bOpenAI−37.755.43 of 6
  153. 165Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−37.955.23 of 6
  154. 166Ring-1TInclusionAI−38.055.13 of 6
  155. 167EXAONE 4.5 33BLG AI−38.254.83 of 6
  156. 168Doubao Seed CodeByteDance−38.354.83 of 6
  157. 169Qwen3 235B ThinkingAlibaba−38.754.41 of 6
  158. 170Nova 2.0 LiteAmazon−38.854.33 of 6
  159. 171INTELLECT-3Prime Intellect−38.854.33 of 6
  160. 172Claude Opus 4Anthropic−38.954.24 of 6
  161. 173Gemini 2.5 FlashGoogle−39.253.95 of 6
  162. 174Seed 2.0ByteDance−39.353.81 of 6
  163. 175NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16NVIDIA−39.353.81 of 6
  164. 176GPT-5.1 InstantOpenAI−39.453.71 of 6
  165. 177PaCoRe-8BStepFun−39.453.71 of 6
  166. 178Command A+Cohere−39.453.73 of 6
  167. 179North-Mini-Code-1.0Cohere−40.252.93 of 6
  168. 180MAI-Thinking-1Microsoft−40.352.81 of 6
  169. 181Apriel-v1.6-15B-ThinkerServiceNow−41.251.93 of 6
  170. 182O1 ProOpenAI−41.251.91 of 6
  171. 183Nemotron 3 SuperNVIDIA−41.251.81 of 6
  172. 184Mistral Small 4Mistral−41.351.83 of 6
  173. 185Magistral Medium 1.2Mistral−41.551.63 of 6
  174. 186Granite 4.2 30BIBM−41.851.33 of 6
  175. 187Grok 3 BetaSpaceXAI−41.951.21 of 6
  176. 188Falcon H1R 7BTII−42.151.03 of 6
  177. 189Grok 3 Mini ReasoningSpaceXAI−42.151.04 of 6
  178. 190O3 ProOpenAI−42.250.92 of 6
  179. 191Nemotron 3 NanoNVIDIA−42.250.91 of 6
  180. 192ERNIE 4.5Baidu−42.250.91 of 6
  181. 193DeepSeek-R1-ZeroDeepSeek−42.450.61 of 6
  182. 194Ling-1TInclusionAI−42.650.53 of 6
  183. 195Ling 2.6 1tInclusionAI−42.650.53 of 6
  184. 196Claude Sonnet 4Anthropic−42.750.45 of 6
  185. 197Phi 4 Reasoning PlusMicrosoft−43.150.01 of 6
  186. 198MiniCPM5-2BOpenBMB−43.249.93 of 6
  187. 199MiMo V2.5 Pro BaseXiaomi−43.449.71 of 6
  188. 200Grok 3 Mini BetaSpaceXAI−43.549.61 of 6
Back to the top of the ranking

The benchmarks behind the reasoning ranking

Six public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

Why these weights

Ranked on critpt, HLE, ARC-AGI-2/3, GPQA Diamond, SimpleBench and LiveBench reasoning (2026-Q3 v2, ratified 2026-08-23). Weights follow measured frontier separation, construct centrality and field recognition; no saturated benchmark ranks.

  • CritPtresearch-physics; strongest measured frontier separation (D .43), AA day-0
  • Humanity's Last Exambroad frontier exam; D .15, field standard (82 card sources), AA day-0
  • ARC-AGI-2abstract reasoning; ARC Prize board, ~9d
  • GPQA Diamondscience QA; compressing at the top (D .02) — operator floor .10, anchor for presence
  • SimpleBenchrobustness facet, single site source, ~9d — operator floor .10
  • LiveBench Reasoningrolling contamination-free reasoning (43 models/12 labs, continuous) — probation, operator floor .10

How the reasoning score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.

The gap. Points behind the leader on the reasoning score.

480 models from 53 vendors have a reasoning score, each measured on at least 3 of the 6 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the reasoning ranking

Which LLM is best at reasoning right now?

GPT-6 Astra, with a reasoning score of 93.1 out of 100. Claude Fable 5.1 is 0.9 points behind at 92.2 — inside the margin, so the top of the ranking is a dead heat.

Which benchmarks make up the reasoning score?

Six public benchmarks: CritPt (28%), Humanity's Last Exam (27%), ARC-AGI-2 (15%), and GPQA Diamond, SimpleBench and LiveBench Reasoning (10% each).

How many models are ranked?

480 models from 53 vendors have a reasoning score. Each needs results on at least 3 of the 6 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.

Full methodology and sources