Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best LLMs for Instruction Following

Tier 3 · provisionalClear leader

Gemini 3.1 Pro holds the line at 85.6 — Nemotron 3 Ultra 550B A55B runs +4.7 back. The median of the 355-model field trails the leader by 35.7 points.

Updated 9 Oct 2026Score basis
  1. 1Gemini 3.1 ProGoogle85.6Instruction Following score, out of 100

    LiveBench Instruction Following40%

    79.1Independent

    Top 81.4 · Gemini 3.8 Flash

    IFBench35%

    77.1Aggregator

    Top 82.9 · Grok 4.20

    MultiChallenge25%

    71.4Independent

    Top 82.0 · GPT-6.1 Sol

    3 of 3 benchmarks measured3 independent sources4.7 ahead of Nemotron 3 Ultra 550B A55B35.7 above the field medianCompare the top two
188 more ranked, from 69.8 down to 47.1155 further models score below 47.1Field median 49.9, across all 355 scored
Show all 200 ranked modelsShow the top 12 only
  1. 13Gemini 3.1 Flash Lite PreviewGoogle−15.869.82 of 3
  2. 14Qwen3.5 122B A10BAlibaba−16.569.12 of 3
  3. 15Grok 4.20SpaceXAI−16.569.11 of 3
  4. 16Inkling-SmallThinking Machines−16.968.71 of 3
  5. 17GPT-5.1OpenAI−17.168.52 of 3
  6. 18GPT-5.5OpenAI−17.268.42 of 3
  7. 19Qwen3.5 27BAlibaba−17.268.42 of 3
  8. 20Claude Fable 5Anthropic−17.468.22 of 3
  9. 21Nemotron Cascade 2 30B A3BNVIDIA−17.668.01 of 3
  10. 22MiMo V2.5 ProXiaomi−17.767.91 of 3
  11. 23GPT-5.6 SolOpenAI−18.067.62 of 3
  12. 24Granite 4.2 8BIBM−18.067.61 of 3
  13. 25Nova 2.0 Pro PreviewAmazon−18.367.31 of 3
  14. 26Gemini 3 ProGoogle−18.567.12 of 3
  15. 27Qwen3.7 Plus PreviewAlibaba−18.567.01 of 3
  16. 28Gemini 3 Flash PreviewGoogle−18.567.01 of 3
  17. 29Kimi K2.6Moonshot−18.766.93 of 3
  18. 30GPT-5 miniOpenAI−18.766.92 of 3
  19. 31Granite 4.2 30BIBM−19.066.61 of 3
  20. 32Qwen3 Max (Reasoning)Alibaba−19.066.52 of 3
  21. 33GPT-5.4OpenAI−19.266.42 of 3
  22. 34Muse GlimmerMeta−19.266.41 of 3
  23. 35Qwen3.6 Max PreviewAlibaba−19.366.31 of 3
  24. 36GPT-6.1 SolOpenAI−19.566.12 of 3
  25. 37GPT-5.2 CodexOpenAI−20.065.62 of 3
  26. 38DeepSeek-V4-FlashDeepSeek−20.265.42 of 3
  27. 39GPT-5.4 nanoOpenAI−20.465.22 of 3
  28. 40Qwen3.5 35B A3BAlibaba−20.465.22 of 3
  29. 41Gemma 4 31BGoogle−20.565.11 of 3
  30. 42O3OpenAI−20.864.82 of 3
  31. 43GPT-5.3 CodexOpenAI−20.964.71 of 3
  32. 44Kimi K2.5Moonshot−21.064.62 of 3
  33. 45MAI-Code-1-FlashMicrosoft−21.264.41 of 3
  34. 46Ling-3.0-flashInclusionAI−21.364.31 of 3
  35. 47Granite 4.2 3BIBM−21.464.21 of 3
  36. 48GPT-5 CodexOpenAI−21.564.11 of 3
  37. 49Command A+Cohere−21.763.91 of 3
  38. 50Gemma 4 12BGoogle−21.963.71 of 3
  39. 51GLM-5-TurboZ.ai−22.263.41 of 3
  40. 52Gemma 4 26B A4BGoogle−22.862.81 of 3
  41. 53Claude Opus 4.8Anthropic−22.862.82 of 3
  42. 54MiniMax M2MiniMax−23.062.61 of 3
  43. 55GLM-5Z.ai−23.062.61 of 3
  44. 56Grok 4.3SpaceXAI−23.162.52 of 3
  45. 57Gemini 3.8 FlashGoogle−23.162.51 of 3
  46. 58Nemotron 3.5 LightningNVIDIA−23.162.51 of 3
  47. 59Nemotron 3 Super 120B A12BNVIDIA−23.262.42 of 3
  48. 60MiniMax M2.5MiniMax−23.362.31 of 3
  49. 61GPT-5.5 InstantOpenAI−23.462.21 of 3
  50. 62Gemini 3.7 FlashGoogle−23.562.11 of 3
  51. 63Solar Pro 3Upstage−23.861.81 of 3
  52. 64Nova 2.0 LiteAmazon−24.161.51 of 3
  53. 65Muse Spark 1.3Meta−24.361.31 of 3
  54. 66O1OpenAI−24.461.21 of 3
  55. 67GPT-5.1 CodexOpenAI−24.760.91 of 3
  56. 68MiniMax M2.1MiniMax−24.860.81 of 3
  57. 69MiniMax M2.7MiniMax−24.960.72 of 3
  58. 70Mercury 2Inception−24.960.71 of 3
  59. 71Kimi K2 ThinkingMoonshot−24.960.72 of 3
  60. 72Apriel-v1.6-15B-ThinkerServiceNow−25.060.61 of 3
  61. 73MiMo V2 ProXiaomi−25.460.21 of 3
  62. 74Mistral Medium 3.5 128BMistral−25.560.11 of 3
  63. 75GLM-4.7Z.ai−25.959.71 of 3
  64. 76GPT-5.1 Codex miniOpenAI−25.959.71 of 3
  65. 77GPT-6 AstraOpenAI−26.059.61 of 3
  66. 78DeepSeek-V4-ProDeepSeek−26.159.52 of 3
  67. 79GPT-5 nanoOpenAI−26.159.51 of 3
  68. 80Step 3.7 FlashStepFun−26.359.31 of 3
  69. 81MAI-Thinking-1Microsoft−26.459.22 of 3
  70. 82Gemini 3.6 FlashGoogle−26.459.21 of 3
  71. 83MiMo V2.5Xiaomi−26.559.11 of 3
  72. 84MiniMax M3MiniMax−26.758.92 of 3
  73. 85Muse Spark 1.1Meta−26.758.92 of 3
  74. 86Grok 4.7SpaceXAI−26.858.81 of 3
  75. 87Qwen3.5 9BAlibaba−26.958.72 of 3
  76. 88Nex-N2-ProNex AGI−26.958.71 of 3
  77. 89Nova 2.0 Omni Reasoning MediumAmazon−26.958.71 of 3
  78. 90GPT-5.6 TerraOpenAI−27.158.52 of 3
  79. 91GPT Oss 20bOpenAI−27.158.51 of 3
  80. 92K-EXAONELG AI−27.258.41 of 3
  81. 93Muse Spark 1.2Meta−27.258.41 of 3
  82. 94Step 3.5 FlashStepFun−27.358.31 of 3
  83. 95Qwen3.6 35B A3BAlibaba−27.458.11 of 3
  84. 96MiMo V2 FlashXiaomi−27.658.01 of 3
  85. 97DeepSeek-V3.2-SpecialeDeepSeek−27.757.91 of 3
  86. 98Nemotron 3 Nano OmniNVIDIA−27.957.71 of 3
  87. 99K2 Think V2MBZUAI−28.157.41 of 3
  88. 100GPT-5.2OpenAI−28.357.32 of 3
  89. 101GPT Oss 120bOpenAI−28.357.32 of 3
  90. 102Apriel-v1.5-15B-ThinkerServiceNow−28.457.21 of 3
  91. 103GLM 5V TurboZ.ai−28.557.11 of 3
  92. 104GLM-4.7-FlashZ.ai−28.657.01 of 3
  93. 105Qwen3 Next 80B A3BAlibaba−28.856.81 of 3
  94. 106DeepSeek-V3.2DeepSeek−28.856.81 of 3
  95. 107K2-V2 (high)MBZUAI−29.056.61 of 3
  96. 108Step3 VL 10BStepFun−29.156.52 of 3
  97. 109Qwen3 VL 32B ReasoningAlibaba−29.156.51 of 3
  98. 110GLM-5.2Z.ai−29.156.52 of 3
  99. 111LFM2.5-2.6BLiquid AI−29.256.41 of 3
  100. 112Claude Fable 5.1Anthropic−29.356.31 of 3
  101. 113NVIDIA Nemotron 3 Nano 4BNVIDIA−29.456.21 of 3
  102. 114EXAONE 4.5 33BLG AI−29.656.01 of 3
  103. 115Claude Sonnet 4.5Anthropic−29.656.02 of 3
  104. 116Solar Open 100B (Reasoning)Upstage−29.855.81 of 3
  105. 117Claude Opus 4.1Anthropic−29.855.82 of 3
  106. 118NVIDIA Nemotron 3 Nano 30B A3BNVIDIA−29.855.82 of 3
  107. 119North-Mini-Code-1.0Cohere−29.955.71 of 3
  108. 120O4 MiniOpenAI−29.955.72 of 3
  109. 121Ling 2.6 FlashInclusionAI−30.055.61 of 3
  110. 122Claude Opus 4.7Anthropic−30.155.42 of 3
  111. 123Claude Sonnet 4Anthropic−30.255.42 of 3
  112. 124Claude Opus 4Anthropic−30.255.42 of 3
  113. 125Motif-2-12.7B-ReasoningMotif Technologies−30.355.31 of 3
  114. 126Ling 2.6 1tInclusionAI−30.555.11 of 3
  115. 127Grok 4.6SpaceXAI−30.655.01 of 3
  116. 128Qwen3.6 PlusAlibaba−30.754.92 of 3
  117. 129Qwen3 VL 235B A22B ReasoningAlibaba−30.754.91 of 3
  118. 130DeepSeek-V3.1-TerminusDeepSeek−30.854.82 of 3
  119. 131Trinity Large ThinkingArcee AI−30.854.81 of 3
  120. 132GPT-5.4 miniOpenAI−30.954.72 of 3
  121. 133LFM2.5-8B-A1BLiquid AI−30.954.71 of 3
  122. 134Tri-21B-ThinkTrillion Labs−31.354.31 of 3
  123. 135Falcon H1R 7BTII−31.454.21 of 3
  124. 136Grok4.5SpaceXAI−31.454.21 of 3
  125. 137DeepSeek-V3.2-ExpDeepSeek−31.654.01 of 3
  126. 138Kimi K3Moonshot−31.853.81 of 3
  127. 139Grok 4SpaceXAI−31.953.71 of 3
  128. 140MiMo V2 OmniXiaomi−32.053.61 of 3
  129. 141O3 MiniOpenAI−32.053.52 of 3
  130. 142Grok 4.1 FastSpaceXAI−32.253.41 of 3
  131. 143DeepSeek-V4-Flash-Vision-ExpDeepSeek−32.353.31 of 3
  132. 144Gemini 2.5 Flash Lite Preview 09-2025Google−32.353.31 of 3
  133. 145Jamba Reasoning 3BAI21−32.553.11 of 3
  134. 146Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−32.653.01 of 3
  135. 147Doubao Seed CodeByteDance−32.852.81 of 3
  136. 148Qwen3.5 Omni PlusAlibaba−32.952.71 of 3
  137. 149Qwen3 30B A3B 2507 ThinkingAlibaba−33.052.61 of 3
  138. 150Grok 4 FastSpaceXAI−33.152.41 of 3
  139. 151Gemini 2.5 FlashGoogle−33.352.31 of 3
  140. 152Gemini 2.5 Flash LiteGoogle−33.552.11 of 3
  141. 153Claude Haiku 4.5Anthropic−33.652.02 of 3
  142. 154Qwen3 4B 2507 ThinkingAlibaba−33.652.01 of 3
  143. 155Mi:dm K 2.5 ProKorea Telecom−33.851.81 of 3
  144. 156MiniCPM5-1BOpenBMB−33.851.81 of 3
  145. 157DeepSeek-V4.1-FlashDeepSeek−33.951.71 of 3
  146. 158Claude Opus 4.5Anthropic−34.151.53 of 3
  147. 159Mistral Small 4Mistral−34.251.41 of 3
  148. 160Hy3-previewTencent−34.351.31 of 3
  149. 161Llama 3.3 70B InstructMeta−34.451.21 of 3
  150. 162Grok 3SpaceXAI−34.551.01 of 3
  151. 163Qwen3 235B A22B Instruct 2507Alibaba−34.750.91 of 3
  152. 164GLM-5.3Z.ai−34.850.81 of 3
  153. 165LFM2-24B-A2BLiquid AI−34.850.81 of 3
  154. 166Grok 3 Mini ReasoningSpaceXAI−34.850.81 of 3
  155. 167Qwen3.5 4BAlibaba−35.050.62 of 3
  156. 168Gemini 2.5 ProGoogle−35.050.62 of 3
  157. 169Qwen3 VL 30B A3B ReasoningAlibaba−35.050.61 of 3
  158. 170GPT-5 (ChatGPT)OpenAI−35.150.51 of 3
  159. 171GPT-6 SolOpenAI−35.250.41 of 3
  160. 172Ring-1TInclusionAI−35.250.31 of 3
  161. 173Magistral Small 1.2Mistral−35.450.21 of 3
  162. 174Granite 4.1 30BIBM−35.450.21 of 3
  163. 175Gemini 3.5 Flash LiteGoogle−35.650.01 of 3
  164. 176Gemma 4 E4BGoogle−35.650.01 of 3
  165. 177Claude 3.7 SonnetAnthropic−35.650.02 of 3
  166. 178Qwen3 Max (Non-Reasoning)Alibaba−35.749.91 of 3
  167. 179GLM-4.5Z.ai−35.849.81 of 3
  168. 180LFM2.5-1.2B-InstructLiquid AI−35.949.71 of 3
  169. 181Claude Sonnet 4.6Anthropic−36.049.62 of 3
  170. 182GLM-4.6Z.ai−36.149.52 of 3
  171. 183Qwen3 Omni 30B A3BAlibaba−36.149.51 of 3
  172. 184Ring-flash-2.0InclusionAI−36.349.31 of 3
  173. 185LongCat Flash LiteMeituan−36.449.21 of 3
  174. 186Llama 4 MaverickMeta−36.649.01 of 3
  175. 187Magistral Medium 1.2Mistral−36.649.01 of 3
  176. 188Claude 3.5 HaikuAnthropic−36.948.71 of 3
  177. 189Qwen3 VL 235B A22B InstructAlibaba−37.048.61 of 3
  178. 190JT-35B-FlashChina Mobile−37.148.51 of 3
  179. 191Seed Oss 36B InstructByteDance−37.248.41 of 3
  180. 192Claude Opus 5.5Anthropic−37.348.31 of 3
  181. 193LFM2.5-1.2B-ThinkingLiquid AI−37.348.31 of 3
  182. 194Kimi K2 (Non-Reasoning)Moonshot−37.847.81 of 3
  183. 195Qwen3 30B A3BAlibaba−37.847.81 of 3
  184. 196ERNIE 5.0 Thinking PreviewBaidu−38.147.51 of 3
  185. 197Grok Code Fast 1SpaceXAI−38.147.51 of 3
  186. 198Qwen3.6 27BAlibaba−38.247.42 of 3
  187. 199Kimi K2 InstructMoonshot−38.347.32 of 3
  188. 200LFM2.5-350MLiquid AI−38.547.11 of 3
Back to the top of the ranking

The benchmarks behind the instruction following ranking

Three public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

  • 40%
    LiveBench Instruction Following

    Follows detailed instructions exactly

    Best on this testGemini 3.8 Flash81.4

  • 35%
    IFBench

    Follows unfamiliar, precisely checkable instructions

    Best on this testGrok 4.2082.9

  • 25%
    MultiChallenge

    Keeps track of context across a multi-turn chat

    Best on this testGPT-6.1 Sol82.0

  • Tracked, not ranked

    Multi-IF — follows instructions over several turns and languages. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

Why these weights

Ranked on LiveBench instruction following, MultiChallenge, IFBench and Multi-IF (2026-Q3 v2.1, coverage pass 2026-08-23).

  • LiveBench Instruction Followingrolling IF, 43+ models/12 labs, continuous — probation waived: the liveliest IF signal
  • IFBenchIF core, AA (decaying: AA stopped running it on new models — watched)
  • MultiChallengemulti-turn conversational IF; SEAL board + cards (n=45, span 8.4) — slow-fill cap

How the instruction following score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.

The gap. Points behind the leader on the instruction following score.

355 models from 48 vendors have a instruction following score, each measured on at least 2 of the 3 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the instruction following ranking

Which LLM is best at instruction following right now?

Gemini 3.1 Pro, with a instruction following score of 85.6 out of 100. Nemotron 3 Ultra 550B A55B is second at 80.9, 4.7 points behind.

Which benchmarks make up the instruction following score?

Three public benchmarks: LiveBench Instruction Following (40%), IFBench (35%), and MultiChallenge (25%).

How many models are ranked?

355 models from 48 vendors have a instruction following score. Each needs results on at least 2 of the 3 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.

Full methodology and sources