Best Long-Context LLMs
GPT-5.6 Terra holds the line at 76.4 — Gemini 3.7 Flash runs +1.0 back. The median of the 372-model field trails the leader by 26.8 points.
- 1GPT-5.6 TerraOpenAI76.4Long Context score, out of 1002 of 3 benchmarks measured1 not reported by the vendor2 independent sources1.0 ahead of Gemini 3.7 Flash26.8 above the field medianCompare the top two
- 2Gemini 3.7 FlashGoogle−1.075.42 of 3
- 3Qwen3.8 Max PreviewAlibaba−1.475.02 of 3
- 4Kimi K3Moonshot−3.073.51 of 3
- 5Step 5 PreviewStepFun−3.173.41 of 3
- 6MiMo V2.6 ProXiaomi−3.273.21 of 3
- 7Claude Fable 5.1Anthropic−3.373.11 of 3
- 8Claude Opus 5.5Anthropic−3.573.01 of 3
- 9GPT-5.5OpenAI−3.672.81 of 3
- 10Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−3.672.81 of 3
- 11Gemini 3.6 FlashGoogle−3.872.62 of 3
- 12GPT-5.6 LunaOpenAI−3.872.62 of 3
Show all 200 ranked modelsShow the top 12 only
- 13DeepSeek-V4.1-FlashDeepSeek−3.972.61 of 3
- 14GPT-5.6 SolOpenAI−3.972.61 of 3
- 15GPT-6 SolOpenAI−4.172.31 of 3
- 16Claude Sonnet 5Anthropic−4.372.22 of 3
- 17Claude Sonnet 4.6Anthropic−4.472.02 of 3
- 18GPT-6 LunaOpenAI−4.571.91 of 3
- 19GPT-5.3 CodexOpenAI−4.571.91 of 3
- 20Solar Mini4Upstage−4.571.91 of 3
- 21Muse GlimmerMeta−4.571.91 of 3
- 22GPT-5.2OpenAI−4.671.83 of 3
- 23Nemotron 3 Ultra 550B A55BNVIDIA−4.971.52 of 3
- 24Agnes 2.5 Pro BetaSapiens AI−5.171.31 of 3
- 25Muse Spark 1.3Meta−5.171.31 of 3
- 26GPT-6.1 SolOpenAI−5.171.31 of 3
- 27MiniMax M3MiniMax−5.171.31 of 3
- 28Ling-3.1-flashInclusionAI−5.171.31 of 3
- 29Haiku 5.5Anthropic−5.770.71 of 3
- 30Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−5.770.71 of 3
- 31Qwen3.6 PlusAlibaba−5.770.72 of 3
- 32Claude Fable 5Anthropic−6.070.41 of 3
- 33GPT-5.2 CodexOpenAI−6.070.41 of 3
- 34Qwen3.8 27BAlibaba−6.470.11 of 3
- 35GPT-5.4OpenAI−6.470.11 of 3
- 36Claude Opus 4.5Anthropic−6.569.92 of 3
- 37Gemini 3 ProGoogle−6.669.93 of 3
- 38Nex-N2-ProNex AGI−6.869.71 of 3
- 39Qwen3.5 397B A17BAlibaba−6.969.52 of 3
- 40Mistral Large 4Mistral−7.069.41 of 3
- 41Gemini 3.8 FlashGoogle−7.069.41 of 3
- 42Agnes 3.0 FlashSapiens AI−7.369.11 of 3
- 43Kimi K2.6Moonshot−7.369.11 of 3
- 44Grok 4.6SpaceXAI−7.369.11 of 3
- 45Qwen3.6 Max PreviewAlibaba−7.668.81 of 3
- 46GPT-6 AstraOpenAI−7.668.81 of 3
- 47Qwen3.8 2.4T A95BAlibaba−7.968.61 of 3
- 48Claude Opus 4.6Anthropic−8.168.42 of 3
- 49Grok4.5SpaceXAI−8.168.42 of 3
- 50Kimi K2.5Moonshot−8.268.22 of 3
- 51K2 Horizon 375B A23BMBZUAI−8.368.11 of 3
- 52GPT-5.1OpenAI−8.368.11 of 3
- 53GLM 5.3 FlashZ.ai−8.368.11 of 3
- 54DeepSeek-V4-FlashDeepSeek−8.967.51 of 3
- 55MiMo V2.5 ProXiaomi−8.967.51 of 3
- 56GLM-5.3Z.ai−8.967.51 of 3
- 57Gemini 4 ArgonGoogle−8.967.51 of 3
- 58Qwen3.8 Flash NextAlibaba−8.967.51 of 3
- 59Apodex 1.1Apodex−9.666.91 of 3
- 60Kimi K2.7 CodeMoonshot−9.666.91 of 3
- 61Claude Opus 5Anthropic−9.666.91 of 3
- 62Qwen3.5 27BAlibaba−9.866.72 of 3
- 63Qwen3.7 MaxAlibaba−10.166.41 of 3
- 64Hy3-previewTencent−10.166.41 of 3
- 65Muse Spark 1.2Meta−10.166.41 of 3
- 66Gemini 3.1 ProGoogle−10.366.12 of 3
- 67MiniMax M2.7MiniMax−10.765.71 of 3
- 68Ling-3.0-flash-VLInclusionAI−10.765.71 of 3
- 69JT-4.1 Flash 236B A21BChina Mobile−10.765.71 of 3
- 70GLM-5.2Z.ai−10.765.71 of 3
- 71GPT-5OpenAI−11.165.41 of 3
- 72Muse SparkMeta−11.465.11 of 3
- 73Gemini 3 Flash PreviewGoogle−11.764.82 of 3
- 74Qwen3.5 122B A10BAlibaba−11.864.72 of 3
- 75Claude Opus 4.8Anthropic−11.864.61 of 3
- 76Muse Spark 1.1Meta−11.864.61 of 3
- 77Claude Opus 4.7Anthropic−11.864.62 of 3
- 78Qwen3.6 27BAlibaba−12.364.21 of 3
- 79InklingThinking Machines−12.364.21 of 3
- 80GPT-5.4 miniOpenAI−12.663.91 of 3
- 81Grok 4.7SpaceXAI−12.863.71 of 3
- 82GPT-5.4 nanoOpenAI−12.863.71 of 3
- 83Claude Sonnet 4.5Anthropic−12.963.52 of 3
- 84Claude 5.5Anthropic−13.163.41 of 3
- 85Qwen3 Max (Reasoning)Alibaba−13.163.32 of 3
- 86Gemini 3.5 Flash LiteGoogle−13.662.91 of 3
- 87Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−13.662.91 of 3
- 88Claude Opus 4.1Anthropic−13.662.91 of 3
- 89Motif-3Motif Technologies−13.662.91 of 3
- 90GLM-5Z.ai−14.162.41 of 3
- 91Inkling-SmallThinking Machines−14.162.41 of 3
- 92Agnes 2.5 Pro AlphaSapiens AI−14.162.41 of 3
- 93O3OpenAI−14.362.22 of 3
- 94MiMo V2 OmniXiaomi−14.362.11 of 3
- 95Motif-3-BetaMotif Technologies−14.661.81 of 3
- 96DeepSeek-V4-ProDeepSeek−14.661.81 of 3
- 97Gemini 3.1 Flash Lite PreviewGoogle−15.161.31 of 3
- 98Claude Haiku 4.5Anthropic−15.161.31 of 3
- 99MiMo V2.6 Flash RLXiaomi−15.161.31 of 3
- 100Solar Pro4Upstage−15.560.91 of 3
- 101Grok 4.1 FastSpaceXAI−15.560.91 of 3
- 102Step 3.7 FlashStepFun−15.960.61 of 3
- 103Grok 4 FastSpaceXAI−15.960.61 of 3
- 104GLM-5.1Z.ai−15.960.61 of 3
- 105Ling-3.0-flash-FinInclusionAI−15.960.61 of 3
- 106MiniMax M2.5MiniMax−16.360.11 of 3
- 107DeepSeek-V3.2DeepSeek−16.559.92 of 3
- 108Qwen3.5 35B A3BAlibaba−16.659.92 of 3
- 109Qwen3.7 Plus PreviewAlibaba−16.859.71 of 3
- 110Grok 4.3SpaceXAI−16.859.71 of 3
- 111MiMo V2.5Xiaomi−16.859.71 of 3
- 112Ling-3.0-flashInclusionAI−16.859.71 of 3
- 113MiMo V2 FlashXiaomi−16.859.72 of 3
- 114GPT-5 miniOpenAI−17.259.21 of 3
- 115DeepSeek-V3.2-ExpDeepSeek−17.259.21 of 3
- 116Qwen3.6 35B A3BAlibaba−18.158.41 of 3
- 117GPT-5 CodexOpenAI−18.158.41 of 3
- 118GLM-5-TurboZ.ai−18.158.41 of 3
- 119Solar Open2 250BUpstage−18.158.41 of 3
- 120Mercury 2.5 PreviewInception−18.158.41 of 3
- 121K2 Horizon 36B-A4BMBZUAI−18.158.41 of 3
- 122Doubao Seed CodeByteDance−18.158.41 of 3
- 123Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−18.657.81 of 3
- 124GLM-4.7Z.ai−18.657.81 of 3
- 125Claude Sonnet 4Anthropic−19.057.41 of 3
- 126GLM 5V TurboZ.ai−19.057.41 of 3
- 127A.X-K2SK Telecom−19.357.11 of 3
- 128DeepSeek-V3.2-SpecialeDeepSeek−19.357.11 of 3
- 129K2 Horizon 7BMBZUAI−19.656.81 of 3
- 130Gemini 3.5 FlashGoogle−19.656.82 of 3
- 131GPT-5.1 CodexOpenAI−20.056.41 of 3
- 132Mistral Medium 3.5 128BMistral−20.056.41 of 3
- 133DeepSeek-V3.1-TerminusDeepSeek−20.056.41 of 3
- 134Gemini 2.5 ProGoogle−20.056.42 of 3
- 135Gemma 4 31BGoogle−20.555.92 of 3
- 136Qwen3.5 9BAlibaba−20.655.82 of 3
- 137GPT-4.1OpenAI−20.655.81 of 3
- 138MiMo V2 ProXiaomi−20.655.81 of 3
- 139Grok 4SpaceXAI−20.855.61 of 3
- 140Claude Opus 4Anthropic−20.955.52 of 3
- 141Grok 4.20SpaceXAI−21.055.41 of 3
- 142MiniMax M2.1MiniMax−21.055.41 of 3
- 143MiniMax M1 80KMiniMax−21.155.32 of 3
- 144GPT-5.1 Codex miniOpenAI−21.255.21 of 3
- 145GPT-5.5 InstantOpenAI−21.355.11 of 3
- 146Kimi K2 ThinkingMoonshot−21.455.02 of 3
- 147Nemotron 3 Super 120B A12BNVIDIA−21.654.91 of 3
- 148JT-35B-FlashChina Mobile−21.654.91 of 3
- 149Gemini 2.5 FlashGoogle−21.954.61 of 3
- 150G9v3-39A5BAI9Stars−21.954.61 of 3
- 151GPT-5 (ChatGPT)OpenAI−22.254.21 of 3
- 152LongCat-2.0Meituan−22.254.21 of 3
- 153O1OpenAI−22.254.21 of 3
- 154Gemini 2.5 Flash Lite Preview 09-2025Google−22.454.01 of 3
- 155MiniMax M1 40KMiniMax−22.553.92 of 3
- 156MiniMax M2MiniMax−22.653.91 of 3
- 157Nova 2.0 Pro PreviewAmazon−22.753.71 of 3
- 158Qwen3 VL 235B A22B ReasoningAlibaba−23.053.41 of 3
- 159Qwen3 Next 80B A3BAlibaba−23.053.41 of 3
- 160Qwen3.5 Omni PlusAlibaba−23.053.41 of 3
- 161K2 Horizon 3.7BMBZUAI−23.453.01 of 3
- 162Gemma 4 26B A4BGoogle−23.752.82 of 3
- 163Qwen3 30B A3B 2507 ThinkingAlibaba−23.752.71 of 3
- 164Seed Oss 36B InstructByteDance−23.752.71 of 3
- 165K-EXAONELG AI−23.752.71 of 3
- 166O4 MiniOpenAI−23.952.51 of 3
- 167Ling 3.0 TinyInclusionAI−24.252.21 of 3
- 168Nemotron 3.5 LightningNVIDIA−24.252.21 of 3
- 169Nova 2.0 LiteAmazon−24.252.21 of 3
- 170K-EXAONE 2.0LG AI−24.452.01 of 3
- 171Nova 2.0 Omni Reasoning MediumAmazon−24.651.91 of 3
- 172MiniCPM5-2BOpenBMB−24.751.71 of 3
- 173Grok 3SpaceXAI−24.851.61 of 3
- 174K2 Think V2MBZUAI−25.251.21 of 3
- 175DeepSeek-V3.1DeepSeek−25.351.11 of 3
- 176DeepSeek-R1DeepSeek−25.650.92 of 3
- 177Gemini 2.5 Flash LiteGoogle−25.750.71 of 3
- 178Gemma 4 12BGoogle−25.750.72 of 3
- 179Qwen3 VL 32B ReasoningAlibaba−26.050.41 of 3
- 180Grok 3 Mini ReasoningSpaceXAI−26.050.41 of 3
- 181Qwen3.5 4BAlibaba−26.150.32 of 3
- 182EXAONE 4.5 33BLG AI−26.350.21 of 3
- 183Ring-1TInclusionAI−26.350.21 of 3
- 184GLM-4.6Z.ai−26.450.01 of 3
- 185Kimi K2 InstructMoonshot−26.849.71 of 3
- 186Grok Code Fast 1SpaceXAI−26.849.71 of 3
- 187Kimi K2 (Non-Reasoning)Moonshot−26.849.71 of 3
- 188Magistral Medium 1.2Mistral−26.849.71 of 3
- 189GLM-4.5Z.ai−27.249.31 of 3
- 190Qwen3 Next 80B A3B InstructAlibaba−27.249.31 of 3
- 191Command A+Cohere−27.249.31 of 3
- 192GPT Oss 120bOpenAI−27.548.91 of 3
- 193Qwen3.5 Omni FlashAlibaba−27.548.91 of 3
- 194Claude 3.7 SonnetAnthropic−27.748.81 of 3
- 195Apriel-v1.6-15B-ThinkerServiceNow−27.848.61 of 3
- 196Qwen3 Max (Non-Reasoning)Alibaba−28.148.41 of 3
- 197Step 3.5 FlashStepFun−28.148.41 of 3
- 198Llama 4 MaverickMeta−28.148.41 of 3
- 199Mistral Small 4Mistral−28.348.11 of 3
- 200Granite 4.2 30BIBM−28.448.01 of 3
The benchmarks behind the long context ranking
Three public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.
- 40%MRCR v2 (8-needle, 128K)
Finds one of eight look-alike replies in a long chat
Best on this testGemini 3.7 Flash97.0
- 35%AA-LCR
Reasons across sets of long documents
Best on this testKimi K388.7
- 25%LongBench v2
Questions about very long texts
Best on this testQwen3.8 Max Preview66.3
Why these weights
Ranked on MRCR v2, AA-LCR and LongBench v2 (2026-Q3 v2.1, coverage pass 2026-08-23). No living universal long-context board exists beyond AA's; card depth carries the rest.
- MRCR v2 (8-needle, 128K)needle recall; vendor cards (n=20 post-merge, D .102 — the axis's discriminator)
- AA-LCRlong-context reasoning, universal day-0 (AA); compressed at the top (D .0196)
- LongBench v2long-context understanding; vendor cards (n=24, span 6.2) — slow-fill cap
How the long context score is calculated
- 1
Rank on each benchmark
Every model gets a percentile on each benchmark it has been measured on.
- 2
Steady the thin fields
Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.
- 3
Weigh and average
The percentiles are averaged with the weights above into one score out of 100.
- 4
Qualify
A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.
Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.
Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.
Printed values. Every score is the value its source printed. Only scores first seen in the last 180 days count.
The gap. Points behind the leader on the long context score.
372 models from 51 vendors have a long context score, each measured on at least 2 of the 3 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.
Questions about the long context ranking
Which LLM is best at long context right now?
GPT-5.6 Terra, with a long context score of 76.4 out of 100. Gemini 3.7 Flash is second at 75.4, 1.0 points behind.
Which benchmarks make up the long context score?
Three public benchmarks: MRCR v2 (8-needle, 128K) (40%), AA-LCR (35%), and LongBench v2 (25%).
How many models are ranked?
372 models from 51 vendors have a long context score. Each needs results on at least 2 of the 3 benchmarks to be ranked.
How current is the ranking?
It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 180 days count.