Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best Long-Context LLMs

Tier 2 · emergingContested

GPT-5.6 Terra holds the line at 76.4 — Gemini 3.7 Flash runs +1.0 back. The median of the 372-model field trails the leader by 26.8 points.

Updated 9 Oct 2026Score basis
  1. 1GPT-5.6 TerraOpenAI76.4Long Context score, out of 100

    MRCR v2 (8-needle, 128K)40%

    93.5Cross-reference

    Top 97.0 · Gemini 3.7 Flash

    AA-LCR35%

    83.0Aggregator

    Top 88.7 · Kimi K3

    LongBench v225%

    —Not reported by the vendor

    Top 66.3 · Qwen3.8 Max Preview

    2 of 3 benchmarks measured1 not reported by the vendor2 independent sources1.0 ahead of Gemini 3.7 Flash26.8 above the field medianCompare the top two
188 more ranked, from 72.6 down to 48.0172 further models score below 48.0Field median 49.7, across all 372 scored
Show all 200 ranked modelsShow the top 12 only
  1. 13DeepSeek-V4.1-FlashDeepSeek−3.972.61 of 3
  2. 14GPT-5.6 SolOpenAI−3.972.61 of 3
  3. 15GPT-6 SolOpenAI−4.172.31 of 3
  4. 16Claude Sonnet 5Anthropic−4.372.22 of 3
  5. 17Claude Sonnet 4.6Anthropic−4.472.02 of 3
  6. 18GPT-6 LunaOpenAI−4.571.91 of 3
  7. 19GPT-5.3 CodexOpenAI−4.571.91 of 3
  8. 20Solar Mini4Upstage−4.571.91 of 3
  9. 21Muse GlimmerMeta−4.571.91 of 3
  10. 22GPT-5.2OpenAI−4.671.83 of 3
  11. 23Nemotron 3 Ultra 550B A55BNVIDIA−4.971.52 of 3
  12. 24Agnes 2.5 Pro BetaSapiens AI−5.171.31 of 3
  13. 25Muse Spark 1.3Meta−5.171.31 of 3
  14. 26GPT-6.1 SolOpenAI−5.171.31 of 3
  15. 27MiniMax M3MiniMax−5.171.31 of 3
  16. 28Ling-3.1-flashInclusionAI−5.171.31 of 3
  17. 29Haiku 5.5Anthropic−5.770.71 of 3
  18. 30Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−5.770.71 of 3
  19. 31Qwen3.6 PlusAlibaba−5.770.72 of 3
  20. 32Claude Fable 5Anthropic−6.070.41 of 3
  21. 33GPT-5.2 CodexOpenAI−6.070.41 of 3
  22. 34Qwen3.8 27BAlibaba−6.470.11 of 3
  23. 35GPT-5.4OpenAI−6.470.11 of 3
  24. 36Claude Opus 4.5Anthropic−6.569.92 of 3
  25. 37Gemini 3 ProGoogle−6.669.93 of 3
  26. 38Nex-N2-ProNex AGI−6.869.71 of 3
  27. 39Qwen3.5 397B A17BAlibaba−6.969.52 of 3
  28. 40Mistral Large 4Mistral−7.069.41 of 3
  29. 41Gemini 3.8 FlashGoogle−7.069.41 of 3
  30. 42Agnes 3.0 FlashSapiens AI−7.369.11 of 3
  31. 43Kimi K2.6Moonshot−7.369.11 of 3
  32. 44Grok 4.6SpaceXAI−7.369.11 of 3
  33. 45Qwen3.6 Max PreviewAlibaba−7.668.81 of 3
  34. 46GPT-6 AstraOpenAI−7.668.81 of 3
  35. 47Qwen3.8 2.4T A95BAlibaba−7.968.61 of 3
  36. 48Claude Opus 4.6Anthropic−8.168.42 of 3
  37. 49Grok4.5SpaceXAI−8.168.42 of 3
  38. 50Kimi K2.5Moonshot−8.268.22 of 3
  39. 51K2 Horizon 375B A23BMBZUAI−8.368.11 of 3
  40. 52GPT-5.1OpenAI−8.368.11 of 3
  41. 53GLM 5.3 FlashZ.ai−8.368.11 of 3
  42. 54DeepSeek-V4-FlashDeepSeek−8.967.51 of 3
  43. 55MiMo V2.5 ProXiaomi−8.967.51 of 3
  44. 56GLM-5.3Z.ai−8.967.51 of 3
  45. 57Gemini 4 ArgonGoogle−8.967.51 of 3
  46. 58Qwen3.8 Flash NextAlibaba−8.967.51 of 3
  47. 59Apodex 1.1Apodex−9.666.91 of 3
  48. 60Kimi K2.7 CodeMoonshot−9.666.91 of 3
  49. 61Claude Opus 5Anthropic−9.666.91 of 3
  50. 62Qwen3.5 27BAlibaba−9.866.72 of 3
  51. 63Qwen3.7 MaxAlibaba−10.166.41 of 3
  52. 64Hy3-previewTencent−10.166.41 of 3
  53. 65Muse Spark 1.2Meta−10.166.41 of 3
  54. 66Gemini 3.1 ProGoogle−10.366.12 of 3
  55. 67MiniMax M2.7MiniMax−10.765.71 of 3
  56. 68Ling-3.0-flash-VLInclusionAI−10.765.71 of 3
  57. 69JT-4.1 Flash 236B A21BChina Mobile−10.765.71 of 3
  58. 70GLM-5.2Z.ai−10.765.71 of 3
  59. 71GPT-5OpenAI−11.165.41 of 3
  60. 72Muse SparkMeta−11.465.11 of 3
  61. 73Gemini 3 Flash PreviewGoogle−11.764.82 of 3
  62. 74Qwen3.5 122B A10BAlibaba−11.864.72 of 3
  63. 75Claude Opus 4.8Anthropic−11.864.61 of 3
  64. 76Muse Spark 1.1Meta−11.864.61 of 3
  65. 77Claude Opus 4.7Anthropic−11.864.62 of 3
  66. 78Qwen3.6 27BAlibaba−12.364.21 of 3
  67. 79InklingThinking Machines−12.364.21 of 3
  68. 80GPT-5.4 miniOpenAI−12.663.91 of 3
  69. 81Grok 4.7SpaceXAI−12.863.71 of 3
  70. 82GPT-5.4 nanoOpenAI−12.863.71 of 3
  71. 83Claude Sonnet 4.5Anthropic−12.963.52 of 3
  72. 84Claude 5.5Anthropic−13.163.41 of 3
  73. 85Qwen3 Max (Reasoning)Alibaba−13.163.32 of 3
  74. 86Gemini 3.5 Flash LiteGoogle−13.662.91 of 3
  75. 87Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−13.662.91 of 3
  76. 88Claude Opus 4.1Anthropic−13.662.91 of 3
  77. 89Motif-3Motif Technologies−13.662.91 of 3
  78. 90GLM-5Z.ai−14.162.41 of 3
  79. 91Inkling-SmallThinking Machines−14.162.41 of 3
  80. 92Agnes 2.5 Pro AlphaSapiens AI−14.162.41 of 3
  81. 93O3OpenAI−14.362.22 of 3
  82. 94MiMo V2 OmniXiaomi−14.362.11 of 3
  83. 95Motif-3-BetaMotif Technologies−14.661.81 of 3
  84. 96DeepSeek-V4-ProDeepSeek−14.661.81 of 3
  85. 97Gemini 3.1 Flash Lite PreviewGoogle−15.161.31 of 3
  86. 98Claude Haiku 4.5Anthropic−15.161.31 of 3
  87. 99MiMo V2.6 Flash RLXiaomi−15.161.31 of 3
  88. 100Solar Pro4Upstage−15.560.91 of 3
  89. 101Grok 4.1 FastSpaceXAI−15.560.91 of 3
  90. 102Step 3.7 FlashStepFun−15.960.61 of 3
  91. 103Grok 4 FastSpaceXAI−15.960.61 of 3
  92. 104GLM-5.1Z.ai−15.960.61 of 3
  93. 105Ling-3.0-flash-FinInclusionAI−15.960.61 of 3
  94. 106MiniMax M2.5MiniMax−16.360.11 of 3
  95. 107DeepSeek-V3.2DeepSeek−16.559.92 of 3
  96. 108Qwen3.5 35B A3BAlibaba−16.659.92 of 3
  97. 109Qwen3.7 Plus PreviewAlibaba−16.859.71 of 3
  98. 110Grok 4.3SpaceXAI−16.859.71 of 3
  99. 111MiMo V2.5Xiaomi−16.859.71 of 3
  100. 112Ling-3.0-flashInclusionAI−16.859.71 of 3
  101. 113MiMo V2 FlashXiaomi−16.859.72 of 3
  102. 114GPT-5 miniOpenAI−17.259.21 of 3
  103. 115DeepSeek-V3.2-ExpDeepSeek−17.259.21 of 3
  104. 116Qwen3.6 35B A3BAlibaba−18.158.41 of 3
  105. 117GPT-5 CodexOpenAI−18.158.41 of 3
  106. 118GLM-5-TurboZ.ai−18.158.41 of 3
  107. 119Solar Open2 250BUpstage−18.158.41 of 3
  108. 120Mercury 2.5 PreviewInception−18.158.41 of 3
  109. 121K2 Horizon 36B-A4BMBZUAI−18.158.41 of 3
  110. 122Doubao Seed CodeByteDance−18.158.41 of 3
  111. 123Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−18.657.81 of 3
  112. 124GLM-4.7Z.ai−18.657.81 of 3
  113. 125Claude Sonnet 4Anthropic−19.057.41 of 3
  114. 126GLM 5V TurboZ.ai−19.057.41 of 3
  115. 127A.X-K2SK Telecom−19.357.11 of 3
  116. 128DeepSeek-V3.2-SpecialeDeepSeek−19.357.11 of 3
  117. 129K2 Horizon 7BMBZUAI−19.656.81 of 3
  118. 130Gemini 3.5 FlashGoogle−19.656.82 of 3
  119. 131GPT-5.1 CodexOpenAI−20.056.41 of 3
  120. 132Mistral Medium 3.5 128BMistral−20.056.41 of 3
  121. 133DeepSeek-V3.1-TerminusDeepSeek−20.056.41 of 3
  122. 134Gemini 2.5 ProGoogle−20.056.42 of 3
  123. 135Gemma 4 31BGoogle−20.555.92 of 3
  124. 136Qwen3.5 9BAlibaba−20.655.82 of 3
  125. 137GPT-4.1OpenAI−20.655.81 of 3
  126. 138MiMo V2 ProXiaomi−20.655.81 of 3
  127. 139Grok 4SpaceXAI−20.855.61 of 3
  128. 140Claude Opus 4Anthropic−20.955.52 of 3
  129. 141Grok 4.20SpaceXAI−21.055.41 of 3
  130. 142MiniMax M2.1MiniMax−21.055.41 of 3
  131. 143MiniMax M1 80KMiniMax−21.155.32 of 3
  132. 144GPT-5.1 Codex miniOpenAI−21.255.21 of 3
  133. 145GPT-5.5 InstantOpenAI−21.355.11 of 3
  134. 146Kimi K2 ThinkingMoonshot−21.455.02 of 3
  135. 147Nemotron 3 Super 120B A12BNVIDIA−21.654.91 of 3
  136. 148JT-35B-FlashChina Mobile−21.654.91 of 3
  137. 149Gemini 2.5 FlashGoogle−21.954.61 of 3
  138. 150G9v3-39A5BAI9Stars−21.954.61 of 3
  139. 151GPT-5 (ChatGPT)OpenAI−22.254.21 of 3
  140. 152LongCat-2.0Meituan−22.254.21 of 3
  141. 153O1OpenAI−22.254.21 of 3
  142. 154Gemini 2.5 Flash Lite Preview 09-2025Google−22.454.01 of 3
  143. 155MiniMax M1 40KMiniMax−22.553.92 of 3
  144. 156MiniMax M2MiniMax−22.653.91 of 3
  145. 157Nova 2.0 Pro PreviewAmazon−22.753.71 of 3
  146. 158Qwen3 VL 235B A22B ReasoningAlibaba−23.053.41 of 3
  147. 159Qwen3 Next 80B A3BAlibaba−23.053.41 of 3
  148. 160Qwen3.5 Omni PlusAlibaba−23.053.41 of 3
  149. 161K2 Horizon 3.7BMBZUAI−23.453.01 of 3
  150. 162Gemma 4 26B A4BGoogle−23.752.82 of 3
  151. 163Qwen3 30B A3B 2507 ThinkingAlibaba−23.752.71 of 3
  152. 164Seed Oss 36B InstructByteDance−23.752.71 of 3
  153. 165K-EXAONELG AI−23.752.71 of 3
  154. 166O4 MiniOpenAI−23.952.51 of 3
  155. 167Ling 3.0 TinyInclusionAI−24.252.21 of 3
  156. 168Nemotron 3.5 LightningNVIDIA−24.252.21 of 3
  157. 169Nova 2.0 LiteAmazon−24.252.21 of 3
  158. 170K-EXAONE 2.0LG AI−24.452.01 of 3
  159. 171Nova 2.0 Omni Reasoning MediumAmazon−24.651.91 of 3
  160. 172MiniCPM5-2BOpenBMB−24.751.71 of 3
  161. 173Grok 3SpaceXAI−24.851.61 of 3
  162. 174K2 Think V2MBZUAI−25.251.21 of 3
  163. 175DeepSeek-V3.1DeepSeek−25.351.11 of 3
  164. 176DeepSeek-R1DeepSeek−25.650.92 of 3
  165. 177Gemini 2.5 Flash LiteGoogle−25.750.71 of 3
  166. 178Gemma 4 12BGoogle−25.750.72 of 3
  167. 179Qwen3 VL 32B ReasoningAlibaba−26.050.41 of 3
  168. 180Grok 3 Mini ReasoningSpaceXAI−26.050.41 of 3
  169. 181Qwen3.5 4BAlibaba−26.150.32 of 3
  170. 182EXAONE 4.5 33BLG AI−26.350.21 of 3
  171. 183Ring-1TInclusionAI−26.350.21 of 3
  172. 184GLM-4.6Z.ai−26.450.01 of 3
  173. 185Kimi K2 InstructMoonshot−26.849.71 of 3
  174. 186Grok Code Fast 1SpaceXAI−26.849.71 of 3
  175. 187Kimi K2 (Non-Reasoning)Moonshot−26.849.71 of 3
  176. 188Magistral Medium 1.2Mistral−26.849.71 of 3
  177. 189GLM-4.5Z.ai−27.249.31 of 3
  178. 190Qwen3 Next 80B A3B InstructAlibaba−27.249.31 of 3
  179. 191Command A+Cohere−27.249.31 of 3
  180. 192GPT Oss 120bOpenAI−27.548.91 of 3
  181. 193Qwen3.5 Omni FlashAlibaba−27.548.91 of 3
  182. 194Claude 3.7 SonnetAnthropic−27.748.81 of 3
  183. 195Apriel-v1.6-15B-ThinkerServiceNow−27.848.61 of 3
  184. 196Qwen3 Max (Non-Reasoning)Alibaba−28.148.41 of 3
  185. 197Step 3.5 FlashStepFun−28.148.41 of 3
  186. 198Llama 4 MaverickMeta−28.148.41 of 3
  187. 199Mistral Small 4Mistral−28.348.11 of 3
  188. 200Granite 4.2 30BIBM−28.448.01 of 3
Back to the top of the ranking

The benchmarks behind the long context ranking

Three public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

Why these weights

Ranked on MRCR v2, AA-LCR and LongBench v2 (2026-Q3 v2.1, coverage pass 2026-08-23). No living universal long-context board exists beyond AA's; card depth carries the rest.

  • MRCR v2 (8-needle, 128K)needle recall; vendor cards (n=20 post-merge, D .102 — the axis's discriminator)
  • AA-LCRlong-context reasoning, universal day-0 (AA); compressed at the top (D .0196)
  • LongBench v2long-context understanding; vendor cards (n=24, span 6.2) — slow-fill cap

How the long context score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 180 days count.

The gap. Points behind the leader on the long context score.

372 models from 51 vendors have a long context score, each measured on at least 2 of the 3 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the long context ranking

Which LLM is best at long context right now?

GPT-5.6 Terra, with a long context score of 76.4 out of 100. Gemini 3.7 Flash is second at 75.4, 1.0 points behind.

Which benchmarks make up the long context score?

Three public benchmarks: MRCR v2 (8-needle, 128K) (40%), AA-LCR (35%), and LongBench v2 (25%).

How many models are ranked?

372 models from 51 vendors have a long context score. Each needs results on at least 2 of the 3 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 180 days count.

Full methodology and sources