Best LLMs for Coding
Claude Fable 5 holds the line at 83.0 — Claude Fable 5.1 runs +3.7 back. The median of the 259-model field trails the leader by 31.0 points.
- 1Claude Fable 5Anthropic83.0Coding score, out of 1009 of 10 benchmarks measured1 not reported by the vendor5 independent sources3.7 ahead of Claude Fable 5.131.0 above the field medianCompare the top two
- 2Claude Fable 5.1Anthropic−3.779.37 of 10
- 3Claude Opus 5Anthropic−4.978.28 of 10
- 4GPT-5.6 SolOpenAI−5.177.97 of 10
- 5Claude Opus 5.5Anthropic−7.675.46 of 10
- 6Claude Opus 4.8Anthropic−8.174.99 of 10
- 7Claude Opus 4.7Anthropic−9.173.99 of 10
- 8GPT-5.6 TerraOpenAI−10.073.07 of 10
- 9GPT-5.5OpenAI−10.872.37 of 10
- 10Qwen3.8 Flash NextAlibaba−11.971.18 of 10
- 11Claude Sonnet 5Anthropic−11.971.18 of 10
- 12Qwen3.8 Max PreviewAlibaba−12.970.16 of 10
Show all 200 ranked modelsShow the top 12 only
- 13GLM-5.2Z.ai−13.469.67 of 10
- 14Kimi K2.6Moonshot−15.167.910 of 10
- 15Qwen3.8 27BAlibaba−15.467.68 of 10
- 16Kimi K3Moonshot−15.967.15 of 10
- 17DeepSeek-V4.1-FlashDeepSeek−15.967.15 of 10
- 18Muse Spark 1.3Meta−16.067.05 of 10
- 19Qwen3.7 MaxAlibaba−16.266.810 of 10
- 20Qwen3.7 Plus PreviewAlibaba−16.466.67 of 10
- 21Muse Spark 1.1Meta−16.866.26 of 10
- 22Gemini 3.1 ProGoogle−16.966.110 of 10
- 23GPT-5.6 LunaOpenAI−17.265.87 of 10
- 24GPT-5.4OpenAI−17.265.88 of 10
- 25GPT-6 AstraOpenAI−17.365.75 of 10
- 26GLM-5.1Z.ai−17.465.68 of 10
- 27Claude Opus 4.6Anthropic−17.565.59 of 10
- 28Sonnet 5.5Anthropic−18.065.05 of 10
- 29GLM-5.3Z.ai−18.364.75 of 10
- 30Claude Sonnet 4.6Anthropic−18.764.38 of 10
- 31Gemini 3.7 FlashGoogle−19.263.85 of 10
- 32Muse SparkMeta−19.363.76 of 10
- 33Hy3-previewTencent−19.463.67 of 10
- 34MiMo V2.5 ProXiaomi−19.463.67 of 10
- 35GPT-5.1OpenAI−19.563.56 of 10
- 36MiniMax M2.7MiniMax−19.563.58 of 10
- 37Haiku 5.5Anthropic−20.063.05 of 10
- 38Grok4.5SpaceXAI−20.162.96 of 10
- 39Gemini 3.5 FlashGoogle−20.862.27 of 10
- 40GPT-5.3 CodexOpenAI−21.161.94 of 10
- 41GLM 5.3 FlashZ.ai−21.161.95 of 10
- 42DeepSeek-V4-ProDeepSeek−21.261.810 of 10
- 43Hy4Tencent−21.261.84 of 10
- 44MiMo V2.6 ProXiaomi−21.361.73 of 10
- 45DeepSeek-V4-FlashDeepSeek−21.561.510 of 10
- 46Grok 4.6SpaceXAI−21.661.45 of 10
- 47Muse Spark 1.2Meta−21.661.45 of 10
- 48Gemini 3 ProGoogle−21.961.17 of 10
- 49Seed 2.1 Pro PreviewByteDance−22.061.14 of 10
- 50Gemini 3 Flash PreviewGoogle−22.260.86 of 10
- 51GLM-5Z.ai−22.460.66 of 10
- 52Claude Mythos 5Anthropic−22.660.43 of 10
- 53MiMo V2 ProXiaomi−22.760.35 of 10
- 54Qwen3.6 PlusAlibaba−22.860.210 of 10
- 55Kimi K2.5Moonshot−22.860.28 of 10
- 56MiMo V2.6 Flash RLXiaomi−22.960.13 of 10
- 57MiMo V2.5Xiaomi−23.459.66 of 10
- 58GPT-5.2 CodexOpenAI−23.559.57 of 10
- 59GPT-6.1 SolOpenAI−23.559.54 of 10
- 60GLM-4.7Z.ai−23.659.47 of 10
- 61GPT-6 SolOpenAI−23.659.44 of 10
- 62Claude Mythos PreviewAnthropic−23.759.33 of 10
- 63Step 3.5 FlashStepFun−23.759.34 of 10
- 64Qwen3.5 397B A17BAlibaba−23.959.18 of 10
- 65Gemini 3.8 FlashGoogle−24.059.05 of 10
- 66GPT-5 CodexOpenAI−24.159.03 of 10
- 67MiniMax M2.5MiniMax−24.458.65 of 10
- 68DeepSeek-V3.2-SpecialeDeepSeek−24.558.53 of 10
- 69Ember-1fireworks−24.858.22 of 10
- 70GPT-5.5 InstantOpenAI−24.958.23 of 10
- 71MiniMax M3MiniMax−24.958.18 of 10
- 72Step 3.7 FlashStepFun−25.257.84 of 10
- 73Qwen3.8 2.4T A95BAlibaba−25.457.72 of 10
- 74Inkling-SmallThinking Machines−25.557.55 of 10
- 75Agnes 3.0 FlashSapiens AI−25.757.32 of 10
- 76Grok 4.7SpaceXAI−26.057.04 of 10
- 77Seed2.1ByteDance−26.057.02 of 10
- 78O3OpenAI−26.156.93 of 10
- 79Claude Opus 4.1Anthropic−26.156.94 of 10
- 80Claude Opus 4.5Anthropic−26.256.89 of 10
- 81MiMo V2 FlashXiaomi−26.356.87 of 10
- 82Mistral Medium 3.5 128BMistral−26.356.75 of 10
- 83MiMo V2 OmniXiaomi−26.456.63 of 10
- 84Gemini 3.6 FlashGoogle−26.456.66 of 10
- 85Claude Sonnet 4.5Anthropic−26.456.68 of 10
- 86GPT-6 LunaOpenAI−26.856.24 of 10
- 87GPT-5.2OpenAI−26.856.28 of 10
- 88Laguna S 2.1poolside−26.856.23 of 10
- 89Ling-3.0-flashInclusionAI−26.856.25 of 10
- 90Kimi K2.7 CodeMoonshot−27.155.96 of 10
- 91Agnes 2.5 Pro BetaSapiens AI−27.355.72 of 10
- 92Qwen3 Max (Reasoning)Alibaba−27.355.74 of 10
- 93Qwen3.5 122B A10BAlibaba−27.455.66 of 10
- 94Kimi K2 ThinkingMoonshot−27.555.55 of 10
- 95GPT-5.1 CodexOpenAI−27.555.54 of 10
- 96Atria Dawnshanghai-ai-lab−27.655.42 of 10
- 97JT-4.1 Flash 236B A21BChina Mobile−27.755.32 of 10
- 98North-Mini-Code-1.0Cohere−27.755.36 of 10
- 99Apodex 1.1Apodex−27.855.22 of 10
- 100Seed 2.0ByteDance−27.855.22 of 10
- 101DeepSeek-V4-Flash-Vision-ExpDeepSeek−27.955.23 of 10
- 102Motif-3-BetaMotif Technologies−27.955.12 of 10
- 103K2 Horizon 375B A23BMBZUAI−28.154.92 of 10
- 104Motif-3Motif Technologies−28.254.82 of 10
- 105Claude Opus 4Anthropic−28.354.84 of 10
- 106GPT-5OpenAI−28.354.78 of 10
- 107Nex-N2-ProNex AGI−28.454.62 of 10
- 108Ling-3.0-flash-VLInclusionAI−28.554.52 of 10
- 109Agnes 2.5 Pro AlphaSapiens AI−28.654.42 of 10
- 110Qwen3.6 35B A3BAlibaba−28.754.37 of 10
- 111O4 MiniOpenAI−28.854.23 of 10
- 112Claude 3.7 SonnetAnthropic−29.153.93 of 10
- 113MiMo V2.5 Pro FP4 DFlashXiaomi−29.153.92 of 10
- 114GPT-5.4 nanoOpenAI−29.153.96 of 10
- 115DeepSeek-V3.1-TerminusDeepSeek−29.153.93 of 10
- 116Ling-3.0-flash-FinInclusionAI−29.253.82 of 10
- 117Qwen3.5 27BAlibaba−29.453.67 of 10
- 118Nova 2.0 Pro PreviewAmazon−29.553.53 of 10
- 119Qwen3.6 27BAlibaba−29.553.59 of 10
- 120GPT-5.4 miniOpenAI−29.553.57 of 10
- 121Solar Open2 250BUpstage−29.653.42 of 10
- 122DeepSeek-V3.2DeepSeek−29.853.28 of 10
- 123Solar Pro4Upstage−30.053.04 of 10
- 124GPT-5.1 Codex MaxOpenAI−30.152.91 of 10
- 125MAI-Thinking-1Microsoft−30.352.73 of 10
- 126K2 Horizon 36B-A4BMBZUAI−30.452.62 of 10
- 127Muse GlimmerMeta−30.652.45 of 10
- 128InklingThinking Machines−30.752.36 of 10
- 129Seed 1.8ByteDance−31.052.02 of 10
- 130GPT-5.1 InstantOpenAI−31.052.01 of 10
- 131K-EXAONE 2.0LG AI−31.052.02 of 10
- 132Command A+Cohere−31.251.83 of 10
- 133Gemma 4 12BGoogle−31.351.74 of 10
- 134Gemini 3.1 Flash Lite PreviewGoogle−31.351.74 of 10
- 135A.X-K2SK Telecom−31.351.72 of 10
- 136DeepSeek-V3.1DeepSeek−31.551.54 of 10
- 137Gemini 2.5 ProGoogle−31.851.25 of 10
- 138Nemotron 3 Ultra 550B A55BNVIDIA−31.851.28 of 10
- 139Qwen3 Coder NextAlibaba−32.051.03 of 10
- 140K-EXAONELG AI−32.051.03 of 10
- 141Claude Haiku 4.5Anthropic−32.051.07 of 10
- 142MiniMax M2.1MiniMax−32.250.86 of 10
- 143LongCat-2.0Meituan−32.250.82 of 10
- 144Mistral Small 4Mistral−32.450.63 of 10
- 145Gemma 4 31BGoogle−32.650.48 of 10
- 146GLM-4.7-FlashZ.ai−32.850.23 of 10
- 147Qwen3.5 9BAlibaba−33.050.04 of 10
- 148DeepSeek-V3.2-ExpDeepSeek−33.050.05 of 10
- 149GPT-5 miniOpenAI−33.149.94 of 10
- 150G9v3-39A5BAI9Stars−33.249.82 of 10
- 151Nemotron Cascade 2 30B A3BNVIDIA−33.249.83 of 10
- 152Claude Sonnet 4Anthropic−33.349.77 of 10
- 153Trinity Large ThinkingArcee AI−33.449.64 of 10
- 154Nova 2.0 LiteAmazon−33.449.63 of 10
- 155EXAONE 4.5 33BLG AI−33.449.64 of 10
- 156Magistral Medium 1.2Mistral−33.549.53 of 10
- 157Mercury 2Inception−33.749.34 of 10
- 158Gemini 3.5 Flash LiteGoogle−33.849.27 of 10
- 159MiniMax M2MiniMax−33.949.15 of 10
- 160GLM-4.5-AirZ.ai−34.049.03 of 10
- 161GLM-4.5Z.ai−34.148.94 of 10
- 162Qwen3 Next 80B A3BAlibaba−34.448.64 of 10
- 163Nemotron 3 Super 120B A12BNVIDIA−34.548.55 of 10
- 164Qwen3.5 35B A3BAlibaba−34.648.48 of 10
- 165Grok Code Fast 1SpaceXAI−34.648.44 of 10
- 166GPT-4.1OpenAI−34.748.34 of 10
- 167MAI-Code-1-FlashMicrosoft−34.848.23 of 10
- 168Ling 2.6 FlashInclusionAI−34.948.13 of 10
- 169Qwen3 Coder PlusAlibaba−34.948.12 of 10
- 170K2 Horizon 7BMBZUAI−34.948.12 of 10
- 171GLM-4.6Z.ai−35.347.78 of 10
- 172Grok 4.3SpaceXAI−35.347.76 of 10
- 173O1OpenAI−35.447.63 of 10
- 174GPT-5 nanoOpenAI−35.547.53 of 10
- 175Gemini 2.5 FlashGoogle−35.547.53 of 10
- 176Laguna-XS-2.1poolside−35.647.43 of 10
- 177Devstral Small 2Mistral−35.747.35 of 10
- 178Mistral Medium 3.1 (Non-Reasoning)Mistral−35.747.33 of 10
- 179GPT-4.5 PreviewOpenAI−36.346.71 of 10
- 180GPT Oss 120bOpenAI−36.446.64 of 10
- 181Mistral Large 3Mistral−36.446.64 of 10
- 182K2 Think V2MBZUAI−36.546.63 of 10
- 183Devstral 2Mistral−36.646.46 of 10
- 184Ling 3.0 TinyInclusionAI−36.746.32 of 10
- 185Nemotron 3 SuperNVIDIA−36.746.34 of 10
- 186Mistral Small 3.1Mistral−36.946.13 of 10
- 187DeepSeek-V3DeepSeek−37.046.05 of 10
- 188Granite 4.2 3BIBM−37.145.93 of 10
- 189K2 Horizon 3.7BMBZUAI−37.245.82 of 10
- 190MiniMax M1 80KMiniMax−37.245.83 of 10
- 191Qwen3.5 4BAlibaba−37.445.64 of 10
- 192DeepSeek-V2.5DeepSeek−37.445.61 of 10
- 193Gemini 2.0 Flash (Reasoning)Google−37.545.51 of 10
- 194MiniMax M1 40KMiniMax−37.645.43 of 10
- 195Qwen3 Coder 480B A35B InstructAlibaba−37.645.46 of 10
- 196MiniCPM5-2BOpenBMB−37.845.32 of 10
- 197O1 MiniOpenAI−37.845.23 of 10
- 198O3 MiniOpenAI−38.045.04 of 10
- 199Nemotron 3 Nano OmniNVIDIA−38.144.93 of 10
- 200Magistral Small 1.2Mistral−38.144.93 of 10
The benchmarks behind the coding ranking
Ten public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.
- 18%LiveBench Agentic Coding
Fixes real GitHub issues, working as an agent
Best on this testDeepSeek-V4.1-Flash77.3
- 11%Terminal-Bench Hard
Hard command-line tasks in a real terminal
Best on this testGPT-5.6 Sol65.9
- 10%SWE-bench Pro
Long, multi-file coding tasks in real codebases
Best on this testClaude Opus 5.589.9
- 10%LiveBench Coding
Programming problems, refreshed regularly
Best on this testSonnet 5.591.4
- 10%SciCode
Code for real scientific research problems
Best on this testClaude Opus 5.566.9
- 9%SWE-bench Verified
Fixes real GitHub issues in Python projects
Best on this testClaude Opus 596.0
- 9%Terminal-Bench 2.1
Command-line tasks in a real terminal
Best on this testClaude Fable 5.191.4
- 8%SWE-bench Multilingual
Fixes real GitHub issues in many programming languages
Best on this testClaude Opus 5.593.9
- 8%LiveCodeBench v6
Recent programming contest problems
Best on this testDeepSeek-V4-Pro92.5
- 7%LMArena WebDev
People's blind votes on coding answers
Best on this testClaude Opus 5.51814
Why these weights
Ranked on the SWE-bench family, LiveBench agentic coding + coding, Terminal-Bench Hard + 2.1, scicode, LiveCodeBench v6 and the webdev arena (2026-Q3 v3, AA Coding Index removed 2026-09-19 after AA withdrew it; replaced 1:1 by Terminal-Bench Hard).
- LiveBench Agentic Codingrolling agentic coding (javascript/typescript/python), D .21 — probation cap
- Terminal-Bench Hardterminal agentic coding (hard tier), AA day-0 (mig 845/AA resumed it, fresh since); REPLACES the withdrawn AA Coding Index 1:1 by weight (AA dropped codingIndex 2026-09-09) — impartial coverage OpenAI 28 / Anthropic 15 / Chinese 98, mirrors the removed line (2026-09-19, operator-ratified)
- SWE-bench Proreal-repo, SEAL/cards; SWE family line .27 (Pro .10 / Verified .09 / Multilingual .08)
- LiveBench Codingrolling code generation/completion, D .16 — probation
- SciCodescientific-coding facet, AA day-0 — operator .10
- SWE-bench Verifiedreal-repo; 69 card sources, day-0 when the vendor reports
- Terminal-Bench 2.1terminal agentic coding, AA day-0
- SWE-bench Multilingualreal-repo across languages; 62 models/11 labs incl. DeepSeek/Moonshot/MiniMax/Xiaomi/Z.ai
- LiveCodeBench v6contest-coding facet; vendor cards (official board dead since 2025-08)
- LMArena WebDevLMArena webdev preference proxy (≤ .10 by rule)
How the coding score is calculated
- 1
Rank on each benchmark
Every model gets a percentile on each benchmark it has been measured on.
- 2
Steady the thin fields
Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.
- 3
Weigh and average
The percentiles are averaged with the weights above into one score out of 100.
- 4
Qualify
A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.
Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.
Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.
Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.
The gap. Points behind the leader on the coding score.
259 models from 43 vendors have a coding score, each measured on at least 4 of the 10 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.
Questions about the coding ranking
Which LLM is best at coding right now?
Claude Fable 5, with a coding score of 83.0 out of 100. Claude Fable 5.1 is second at 79.3, 3.7 points behind.
Which benchmarks make up the coding score?
Ten public benchmarks: LiveBench Agentic Coding (18%), Terminal-Bench Hard (11%), SWE-bench Pro, LiveBench Coding and SciCode (10% each), SWE-bench Verified and Terminal-Bench 2.1 (9% each), SWE-bench Multilingual and LiveCodeBench v6 (8% each), and LMArena WebDev (7%).
How many models are ranked?
259 models from 43 vendors have a coding score. Each needs results on at least 4 of the 10 benchmarks to be ranked.
How current is the ranking?
It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.