Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best LLMs for Agentic Tasks

Tier 1 · establishedContested

Claude Opus 5 holds the line at 85.2 — Claude Fable 5 runs +3.2 back. The median of the 325-model field trails the leader by 36.7 points.

Updated 9 Oct 2026Score basis
  1. 1Claude Opus 5Anthropic85.2Agentic score, out of 100

    τ-Bench V3 Banking19%

    42.1Aggregator

    Top 50.5 · Muse Spark 1.3

    OSWorld-Verified16%

    83.4Cross-reference

    Top 86.1 · Qwen3.8 Max Preview

    MCP Atlas16%

    85.8Vendor card

    Top 88.1 · Muse Spark 1.1

    Terminal-Bench 4.015%

    49.0Aggregator

    Top 63.6 · Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)

    GDPval (win rate)14%

    61.2Aggregator

    Top 87.9 · Seed 2.1 Pro Preview

    BrowseComp10%

    90.8Vendor card

    Top 92.5 · Atria Dawn

    AA AnalystAgent10%

    53.8Aggregator

    Top 60.0 · Gemini 3.7 Flash

    7 of 7 benchmarks measured4 independent sources3.2 ahead of Claude Fable 536.7 above the field medianCompare the top two
188 more ranked, from 72.0 down to 45.1125 further models score below 45.1Field median 48.6, across all 325 scored
Show all 200 ranked modelsShow the top 12 only
  1. 13GPT-5.6 TerraOpenAI−13.372.04 of 7
  2. 14Muse Spark 1.3Meta−13.571.83 of 7
  3. 15Qwen3.8 27BAlibaba−14.071.34 of 7
  4. 16GLM 5.3 FlashZ.ai−14.670.73 of 7
  5. 17Muse Spark 1.1Meta−14.670.65 of 7
  6. 18Qwen3.8 Flash NextAlibaba−15.569.83 of 7
  7. 19Gemini 3.5 FlashGoogle−15.769.56 of 7
  8. 20Grok 4.6SpaceXAI−15.969.34 of 7
  9. 21Claude Opus 4.7Anthropic−17.068.26 of 7
  10. 22Gemini 3.8 FlashGoogle−17.368.03 of 7
  11. 23Agnes 3.0 FlashSapiens AI−17.467.83 of 7
  12. 24GPT-5.6 LunaOpenAI−17.967.35 of 7
  13. 25Gemini 3.7 FlashGoogle−18.167.14 of 7
  14. 26Seed 2.1 Pro PreviewByteDance−18.367.04 of 7
  15. 27Gemini 3.6 FlashGoogle−18.666.64 of 7
  16. 28MiMo V2.6 ProXiaomi−18.866.43 of 7
  17. 29DeepSeek-V4-FlashDeepSeek−18.866.46 of 7
  18. 30Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−18.966.43 of 7
  19. 31Claude Opus 5.5Anthropic−19.166.13 of 7
  20. 32Grok4.5SpaceXAI−19.266.04 of 7
  21. 33Muse Spark 1.2Meta−19.965.43 of 7
  22. 34GPT-5.4OpenAI−20.365.05 of 7
  23. 35GLM-5.2Z.ai−20.664.74 of 7
  24. 36DeepSeek-V4-ProDeepSeek−20.664.66 of 7
  25. 37GPT-6.1 SolOpenAI−20.964.33 of 7
  26. 38MiMo V2.6 Flash RLXiaomi−21.064.23 of 7
  27. 39Gemini 4 ArgonGoogle−21.763.52 of 7
  28. 40Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−22.362.92 of 7
  29. 41K2 Horizon 375B A23BMBZUAI−22.562.83 of 7
  30. 42GPT-6 SolOpenAI−22.962.32 of 7
  31. 43Ling-3.1-flashInclusionAI−23.062.22 of 7
  32. 44Grok 4.7SpaceXAI−23.361.92 of 7
  33. 45Haiku 5.5Anthropic−23.361.92 of 7
  34. 46JT-4.1 Flash 236B A21BChina Mobile−23.361.93 of 7
  35. 47Agnes 2.5 Pro BetaSapiens AI−23.561.82 of 7
  36. 48Step 5 PreviewStepFun−23.561.73 of 7
  37. 49Claude Sonnet 4.6Anthropic−23.761.57 of 7
  38. 50DeepSeek-V4.1-FlashDeepSeek−23.961.32 of 7
  39. 51Hy3-previewTencent−24.161.25 of 7
  40. 52Gemini 3.1 ProGoogle−24.261.17 of 7
  41. 53Inkling-SmallThinking Machines−24.660.66 of 7
  42. 54Motif-3Motif Technologies−24.760.52 of 7
  43. 55Mistral Large 4Mistral−24.960.32 of 7
  44. 56Seed2.1ByteDance−24.960.32 of 7
  45. 57InklingThinking Machines−24.960.35 of 7
  46. 58Kimi K2.6Moonshot−25.260.16 of 7
  47. 59Claude 5.5Anthropic−25.459.92 of 7
  48. 60MiniMax M3MiniMax−25.859.57 of 7
  49. 61GPT-6 LunaOpenAI−26.259.12 of 7
  50. 62Motif-3-BetaMotif Technologies−26.658.62 of 7
  51. 63Claude Opus 4.6Anthropic−26.858.44 of 7
  52. 64Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−26.958.32 of 7
  53. 65K2 Horizon 7BMBZUAI−27.158.23 of 7
  54. 66Muse SparkMeta−27.457.93 of 7
  55. 67Kimi K2.7 CodeMoonshot−27.457.94 of 7
  56. 68GLM-5.1Z.ai−27.657.75 of 7
  57. 69G9v3-39A5BAI9Stars−28.057.32 of 7
  58. 70Gemini 3.5 Flash LiteGoogle−28.656.64 of 7
  59. 71Ling-3.0-flash-VLInclusionAI−28.856.53 of 7
  60. 72Solar Open2 250BUpstage−28.856.52 of 7
  61. 73Qwen3.6 Max PreviewAlibaba−29.056.31 of 7
  62. 74K2 Horizon 36B-A4BMBZUAI−29.056.33 of 7
  63. 75Solar Pro4Upstage−29.056.24 of 7
  64. 76GLM-5-TurboZ.ai−29.056.21 of 7
  65. 77Muse GlimmerMeta−29.355.95 of 7
  66. 78GPT-5.3 CodexOpenAI−29.355.92 of 7
  67. 79GLM-5Z.ai−29.555.73 of 7
  68. 80Ling-3.0-flash-FinInclusionAI−29.555.74 of 7
  69. 81Nex-N2-ProNex AGI−29.655.62 of 7
  70. 82MiMo V2 ProXiaomi−29.755.61 of 7
  71. 83Qwen3.7 MaxAlibaba−29.855.55 of 7
  72. 84Qwen3.7 Plus PreviewAlibaba−29.955.35 of 7
  73. 85MiniMax M2.5MiniMax−30.055.32 of 7
  74. 86GPT-5.4 miniOpenAI−30.155.16 of 7
  75. 87Gemini 3 Deep ThinkGoogle−30.155.11 of 7
  76. 88Agnes 2.5 Pro AlphaSapiens AI−30.155.13 of 7
  77. 89Qwen3.6 PlusAlibaba−30.255.14 of 7
  78. 90MiMo V2 OmniXiaomi−30.255.01 of 7
  79. 91GPT-5OpenAI−30.355.03 of 7
  80. 92Apodex 1.1Apodex−30.355.03 of 7
  81. 93GPT-5.2 CodexOpenAI−30.354.91 of 7
  82. 94GPT-5 CodexOpenAI−30.654.71 of 7
  83. 95Gemini 3 Flash PreviewGoogle−30.854.44 of 7
  84. 96GPT-5.1 CodexOpenAI−30.854.41 of 7
  85. 97Qwen3.5 Omni PlusAlibaba−30.954.41 of 7
  86. 98GLM 5V TurboZ.ai−31.054.32 of 7
  87. 99Solar Mini4Upstage−31.054.32 of 7
  88. 100MiniMax M2.1MiniMax−31.753.62 of 7
  89. 101Step 3.5 FlashStepFun−31.753.52 of 7
  90. 102JT-35B-FlashChina Mobile−31.953.31 of 7
  91. 103Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−32.053.21 of 7
  92. 104GPT-5.1 Codex miniOpenAI−32.253.11 of 7
  93. 105Claude 3.7 SonnetAnthropic−32.353.01 of 7
  94. 106Ling 2.6 1tInclusionAI−32.352.91 of 7
  95. 107Grok 4.20SpaceXAI−32.452.91 of 7
  96. 108Qwen3 Max (Non-Reasoning)Alibaba−32.552.71 of 7
  97. 109GPT-5.2OpenAI−32.652.74 of 7
  98. 110Gemini 3 ProGoogle−32.952.33 of 7
  99. 111MiniMax M1 80KMiniMax−33.152.21 of 7
  100. 112Grok 4SpaceXAI−33.252.11 of 7
  101. 113Qwen3.5 27BAlibaba−33.351.94 of 7
  102. 114Ling-3.0-flashInclusionAI−33.351.95 of 7
  103. 115Doubao Seed CodeByteDance−33.451.91 of 7
  104. 116Claude Opus 4.5Anthropic−33.451.84 of 7
  105. 117O4 MiniOpenAI−33.651.72 of 7
  106. 118Qwen3.5 Omni FlashAlibaba−33.851.51 of 7
  107. 119Gemini 2.5 Flash Lite Preview 09-2025Google−34.051.21 of 7
  108. 120LongCat Flash LiteMeituan−34.251.11 of 7
  109. 121JT-MINIChina Mobile−34.251.01 of 7
  110. 122MiniMax M2MiniMax−34.351.02 of 7
  111. 123Grok 4 FastSpaceXAI−34.350.92 of 7
  112. 124Devstral SmallMistral−34.350.91 of 7
  113. 125Kimi K2.5Moonshot−34.350.95 of 7
  114. 126Step 3.7 FlashStepFun−34.550.72 of 7
  115. 127K-EXAONE 2.0LG AI−34.650.72 of 7
  116. 128ERNIE 5.0 Thinking PreviewBaidu−34.650.71 of 7
  117. 129Nova 2.0 Omni Reasoning MediumAmazon−34.650.61 of 7
  118. 130Qwen3 235B A22B Instruct 2507Alibaba−34.750.61 of 7
  119. 131Grok 4.1 FastSpaceXAI−34.750.51 of 7
  120. 132GPT-4.1OpenAI−34.850.51 of 7
  121. 133DeepSeek-V3.2-ExpDeepSeek−35.050.22 of 7
  122. 134Grok Code Fast 1SpaceXAI−35.050.21 of 7
  123. 135Kimi K2 ThinkingMoonshot−35.150.22 of 7
  124. 136Seed Oss 36B InstructByteDance−35.150.11 of 7
  125. 137K2 Horizon 3.7BMBZUAI−35.150.13 of 7
  126. 138GPT-5 nanoOpenAI−35.250.01 of 7
  127. 139INTELLECT-3Prime Intellect−35.250.01 of 7
  128. 140Qwen3 235B A22BAlibaba−35.349.91 of 7
  129. 141GPT-5.4 nanoOpenAI−35.449.95 of 7
  130. 142O1OpenAI−35.449.81 of 7
  131. 143Qwen3 Coder 30B A3B InstructAlibaba−35.549.71 of 7
  132. 144A.X-K2SK Telecom−35.649.63 of 7
  133. 145Gemini 2.5 FlashGoogle−35.649.61 of 7
  134. 146Qwen3.6 27BAlibaba−35.749.54 of 7
  135. 147Devstral MediumMistral−35.849.51 of 7
  136. 148Granite 4.2 30BIBM−35.849.42 of 7
  137. 149Ring-1TInclusionAI−35.849.41 of 7
  138. 150GLM-4.7-FlashZ.ai−35.949.42 of 7
  139. 151Qwen3 VL 235B A22B InstructAlibaba−35.949.32 of 7
  140. 152Grok 3SpaceXAI−36.049.21 of 7
  141. 153O3OpenAI−36.149.22 of 7
  142. 154DeepSeek-V3.1-TerminusDeepSeek−36.149.23 of 7
  143. 155Magistral Medium 1Mistral−36.149.21 of 7
  144. 156Solar Open 100B (Reasoning)Upstage−36.149.11 of 7
  145. 157GLM-4.6Z.ai−36.249.13 of 7
  146. 158Mi:dm K 2.5 ProKorea Telecom−36.249.01 of 7
  147. 159Sarvam 105BSarvam AI−36.348.92 of 7
  148. 160DeepSeek-V3.2DeepSeek−36.448.84 of 7
  149. 161Qwen3 Next 80B A3B InstructAlibaba−36.448.81 of 7
  150. 162MiniCPM5-2BOpenBMB−36.448.83 of 7
  151. 163GLM-4.6VZ.ai−36.748.61 of 7
  152. 164Gemini 2.0 Flash (Non-Reasoning)Google−36.848.51 of 7
  153. 165GLM-4.5VZ.ai−37.148.21 of 7
  154. 166Kimi K2 InstructMoonshot−37.148.12 of 7
  155. 167Nova PremierAmazon−37.248.11 of 7
  156. 168Qwen3 Coder 480B A35B InstructAlibaba−37.248.01 of 7
  157. 169EXAONE 4.0 32BLG AI−37.348.01 of 7
  158. 170GLM-4.7Z.ai−37.348.04 of 7
  159. 171GPT-5 miniOpenAI−37.347.93 of 7
  160. 172Nemotron 3 Ultra 550B A55BNVIDIA−37.447.96 of 7
  161. 173GPT-5.1OpenAI−37.447.84 of 7
  162. 174Qwen3.5 397B A17BAlibaba−37.547.74 of 7
  163. 175GPT-5.5 InstantOpenAI−37.647.73 of 7
  164. 176MiMo V2.5 ProXiaomi−37.647.74 of 7
  165. 177LongCat-2.0Meituan−37.747.63 of 7
  166. 178Ling 3.0 TinyInclusionAI−37.747.53 of 7
  167. 179Qwen3.5 122B A10BAlibaba−38.247.15 of 7
  168. 180Qwen3 VL 4B (Reasoning)Alibaba−38.546.82 of 7
  169. 181Qwen3 VL 32B ReasoningAlibaba−38.846.42 of 7
  170. 182DeepSeek-V3.1DeepSeek−39.046.22 of 7
  171. 183Qwen3 VL 8B InstructAlibaba−39.146.22 of 7
  172. 184G9v3-3BAI9Stars−39.246.12 of 7
  173. 185Qwen3 VL Thinking (8B)Alibaba−39.246.12 of 7
  174. 186MiMo V2.5Xiaomi−39.445.93 of 7
  175. 187K-EXAONELG AI−39.545.82 of 7
  176. 188Grok 4.3SpaceXAI−39.545.74 of 7
  177. 189Claude Sonnet 4Anthropic−39.745.54 of 7
  178. 190Qwen3 VL 30B A3B ReasoningAlibaba−39.945.42 of 7
  179. 191GLM-4.5-AirZ.ai−40.045.32 of 7
  180. 192Ling-mini-2.0InclusionAI−40.245.11 of 7
  181. 193Qwen3 4B 2507 InstructAlibaba−40.245.11 of 7
  182. 194NVIDIA Nemotron Nano 9B V2NVIDIA−40.245.11 of 7
  183. 195Gemini 2.5 Flash LiteGoogle−40.245.11 of 7
  184. 196Grok 3 Mini ReasoningSpaceXAI−40.245.11 of 7
  185. 197Llama 3.1 Nemotron Instruct 70BNVIDIA−40.245.11 of 7
  186. 198Granite 4.0 MicroIBM−40.245.11 of 7
  187. 199Llama 3.1 70B InstructMeta−40.245.11 of 7
  188. 200LFM2-2.6BLiquid AI−40.245.11 of 7
Back to the top of the ranking

The benchmarks behind the agentic ranking

Seven public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

Why these weights

Ranked on TauBench V3, OSWorld-Verified, MCP Atlas, GDPval, Terminal-Bench 4.0, BrowseComp and analystAgent (2026-Q3 v4). apexAgents/itbenchSre moved to DEPTH — too thin (33 models) to rank the newest launches; Terminal-Bench 4.0 (115 models incl. GPT-6 Astra, 27 vendors) added as the broad agentic-execution leg — operator Option A, 2026-09-20.

  • τ-Bench V3 Bankingtool-use core, AA day-0, D .23 — probation cap
  • OSWorld-Verifiedcomputer-use facet; official board + cards
  • MCP AtlasMCP tool-chain facet; SEAL board (registered 2026-08-23) + cards
  • Terminal-Bench 4.0terminal-agent task completion (AA Terminal-Bench v4.0); broad — 115 models incl. GPT-6 Astra + all 14 recent frontier launches, 27 vendors (Chinese 36 / OpenAI 12 / Google 11 / Anthropic 8 = impartial). REPLACES the thin apexAgents/itbenchSre (33 each, missed the newest launches) as the agentic ranked leg — operator-ratified 2026-09-20 (Option A). Raw AA field-key canonical per this axis's convention (cf. gdpval/browsecomp); ⚠️ fragmented from card-sourced 'Terminal-Bench 4.0' (18 rows) — flagged for a benchmark-identity pass, NOT folded here (surgical).
  • GDPval (win rate)real-world task completion (AA win-rate), day-0 — the broad AA agentic successor (OpenAI 31 / Chinese 98), weight held per operator 2026-09-19
  • BrowseCompresearch-agent facet; vendor cards (35 sources)
  • AA AnalystAgentanalyst facet; AA new suite — probation, operator .10

How the agentic score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.

The gap. Points behind the leader on the agentic score.

325 models from 45 vendors have a agentic score, each measured on at least 3 of the 7 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the agentic ranking

Which LLM is best at agentic right now?

Claude Opus 5, with a agentic score of 85.2 out of 100. Claude Fable 5 is second at 82.1, 3.2 points behind.

Which benchmarks make up the agentic score?

Seven public benchmarks: τ-Bench V3 Banking (19%), OSWorld-Verified and MCP Atlas (16% each), Terminal-Bench 4.0 (15%), GDPval (win rate) (14%), and BrowseComp and AA AnalystAgent (10% each).

How many models are ranked?

325 models from 45 vendors have a agentic score. Each needs results on at least 3 of the 7 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.

Full methodology and sources