Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best Multimodal LLMs

Tier 2 · emergingContested

Gemini 3 Pro holds the line at 75.0 — Seed 2.1 Pro Preview runs +2.1 back. The median of the 190-model field trails the leader by 24.3 points.

Updated 9 Oct 2026Score basis
  1. 1Gemini 3 ProGoogle75.0Multimodal score, out of 100

    Video-MME21%

    88.4Cross-reference

    Top 89.2 · Seed 2.1 Pro Preview

    OCRBench21%

    90.3Cross-reference

    Top 92.3 · Kimi K2.5

    MMMU17%

    —Not reported by the vendor

    Top 86.0 · Qwen3.6 Plus

    CharXiv (reasoning)15%

    81.4Vendor card

    Top 93.2 · Claude Mythos Preview

    MMMU-Pro13%

    80.2Aggregator

    Top 87.7 · Claude Opus 5.5

    MathVista13%

    89.8Cross-reference

    Top 90.7 · Seed 2.1 Pro Preview

    5 of 6 benchmarks measured1 not reported by the vendor3 independent sources2.1 ahead of Seed 2.1 Pro Preview24.3 above the field medianCompare the top two
178 more ranked, from 63.6 down to 28.2Field median 50.7, across all 190 scored
Show all 190 ranked modelsShow the top 12 only
  1. 13Gemini 3.6 FlashGoogle−11.463.62 of 6
  2. 14Gemini 3.5 FlashGoogle−11.563.53 of 6
  3. 15Gemini 3.7 FlashGoogle−11.763.32 of 6
  4. 16Qwen3.8 Flash NextAlibaba−12.262.82 of 6
  5. 17Claude Fable 5Anthropic−12.362.72 of 6
  6. 18Qwen3.6 PlusAlibaba−12.762.34 of 6
  7. 19Gemini 3.8 FlashGoogle−12.862.22 of 6
  8. 20Claude Opus 4.8Anthropic−13.161.92 of 6
  9. 21GPT-5.6 SolOpenAI−13.761.32 of 6
  10. 22Muse SparkMeta−13.761.32 of 6
  11. 23Claude Opus 5Anthropic−13.961.12 of 6
  12. 24Qwen3.6 35B A3BAlibaba−14.060.94 of 6
  13. 25GPT-5.6 TerraOpenAI−14.260.82 of 6
  14. 26Gemini 2.5 ProGoogle−14.360.75 of 6
  15. 27Kimi K2.6Moonshot−14.360.72 of 6
  16. 28Qwen3.8 27BAlibaba−14.460.62 of 6
  17. 29Qwen3.5 397B A17BAlibaba−14.560.53 of 6
  18. 30O3OpenAI−14.860.24 of 6
  19. 31GPT-5.5OpenAI−15.060.02 of 6
  20. 32Claude Sonnet 5Anthropic−15.459.62 of 6
  21. 33Seed 1.8ByteDance−15.659.44 of 6
  22. 34GPT-5.1OpenAI−15.659.42 of 6
  23. 35GPT-5OpenAI−16.258.83 of 6
  24. 36Seed 1.5 VLByteDance−16.358.73 of 6
  25. 37Claude Opus 5.5Anthropic−16.358.71 of 6
  26. 38Grok4.5SpaceXAI−16.358.72 of 6
  27. 39GPT-6 AstraOpenAI−16.458.61 of 6
  28. 40Step3 VL 10BStepFun−16.458.64 of 6
  29. 41GPT-6.1 SolOpenAI−16.558.51 of 6
  30. 42Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−16.658.41 of 6
  31. 43Claude Opus 4.5Anthropic−17.057.95 of 6
  32. 44GPT-5.6 LunaOpenAI−17.157.92 of 6
  33. 45GPT-5.4OpenAI−17.257.82 of 6
  34. 46MiMo V2.5Xiaomi−17.257.83 of 6
  35. 47GPT-6 SolOpenAI−17.357.71 of 6
  36. 48Qwen3.8 Max PreviewAlibaba−17.457.61 of 6
  37. 49GLM-4.6V (106B-A12B)Z.ai−17.557.53 of 6
  38. 50Qwen3 VL 235B A22B InstructAlibaba−17.557.54 of 6
  39. 51Muse Spark 1.3Meta−17.657.31 of 6
  40. 52Qwen3 VL 235B A22B ReasoningAlibaba−17.757.36 of 6
  41. 53Gemini 3 Flash PreviewGoogle−17.957.12 of 6
  42. 54MiniMax M3MiniMax−18.256.82 of 6
  43. 55GPT-6 LunaOpenAI−18.856.21 of 6
  44. 56O4 MiniOpenAI−18.856.24 of 6
  45. 57Apodex 1.1Apodex−19.056.01 of 6
  46. 58GPT-5.5 InstantOpenAI−19.056.02 of 6
  47. 59EXAONE 4.5 33BLG AI−19.056.03 of 6
  48. 60Ling-3.0-flash-VLInclusionAI−19.155.81 of 6
  49. 61Gemma 4 31BGoogle−19.255.85 of 6
  50. 62GPT-5.2OpenAI−19.355.75 of 6
  51. 63GPT-5.3 CodexOpenAI−19.655.41 of 6
  52. 64Grok 4.7SpaceXAI−19.855.21 of 6
  53. 65Grok 4.3SpaceXAI−19.955.11 of 6
  54. 66DeepSeek-V4.1-FlashDeepSeek−20.354.71 of 6
  55. 67Mistral Large 4Mistral−20.554.51 of 6
  56. 68Step 5 PreviewStepFun−20.654.41 of 6
  57. 69Qwen3 VL 32B InstructAlibaba−20.754.34 of 6
  58. 70GPT-5.2 CodexOpenAI−20.754.31 of 6
  59. 71Gemini 3.5 Flash LiteGoogle−20.854.22 of 6
  60. 72Gemini 2.5 FlashGoogle−21.054.02 of 6
  61. 73Agnes 2.5 Pro BetaSapiens AI−21.253.81 of 6
  62. 74Step 3.7 FlashStepFun−21.553.51 of 6
  63. 75GLM-4.6V-Flash (9B)Z.ai−21.853.23 of 6
  64. 76Muse GlimmerMeta−22.152.92 of 6
  65. 77Claude Opus 4.6Anthropic−22.552.52 of 6
  66. 78GPT-5 CodexOpenAI−22.652.41 of 6
  67. 79Gemini 3.1 Flash Lite PreviewGoogle−22.952.12 of 6
  68. 80GPT-5.4 miniOpenAI−23.052.01 of 6
  69. 81InklingThinking Machines−23.151.92 of 6
  70. 82Claude Sonnet 4.5Anthropic−23.251.84 of 6
  71. 83Grok 4.20SpaceXAI−23.251.81 of 6
  72. 84Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−23.351.71 of 6
  73. 85Gemma 4 26B A4BGoogle−23.451.64 of 6
  74. 86MiMo V2.6 Flash RLXiaomi−23.451.61 of 6
  75. 87GLM 5V TurboZ.ai−23.551.51 of 6
  76. 88GPT-5.1 CodexOpenAI−23.751.31 of 6
  77. 89Qwen3.5 Omni PlusAlibaba−23.851.21 of 6
  78. 90Inkling-SmallThinking Machines−23.851.22 of 6
  79. 91GPT-5 miniOpenAI−23.951.11 of 6
  80. 92Qwen3 VL 32B ReasoningAlibaba−24.051.04 of 6
  81. 93MiMo V2 OmniXiaomi−24.050.91 of 6
  82. 94Gemma 4 12BGoogle−24.150.81 of 6
  83. 95MiMo VL RL 2508 (7B)Xiaomi−24.250.83 of 6
  84. 96Qwen3.5 9BAlibaba−24.350.71 of 6
  85. 97GPT-5.1 Codex miniOpenAI−24.650.41 of 6
  86. 98Grok 4SpaceXAI−24.750.31 of 6
  87. 99Qwen3.5 4BAlibaba−25.050.04 of 6
  88. 100Doubao Seed CodeByteDance−25.050.01 of 6
  89. 101Claude Opus 4.1Anthropic−25.149.91 of 6
  90. 102Qwen3 VL 30B A3B InstructAlibaba−25.249.75 of 6
  91. 103Claude Sonnet 4.6Anthropic−25.349.72 of 6
  92. 104GPT-5.4 nanoOpenAI−25.449.61 of 6
  93. 105Gemini 2.5 Flash Lite Preview 09-2025Google−25.649.41 of 6
  94. 106Mistral Medium 3.5 128BMistral−25.649.31 of 6
  95. 107Qwen3.5 Omni FlashAlibaba−25.749.21 of 6
  96. 108ERNIE 5.0 Thinking PreviewBaidu−25.849.21 of 6
  97. 109Nova 2.0 Pro PreviewAmazon−25.949.11 of 6
  98. 110JT-4.1 Flash 236B A21BChina Mobile−26.148.91 of 6
  99. 111Claude 3.7 SonnetAnthropic−26.148.92 of 6
  100. 112Claude Sonnet 4Anthropic−26.248.82 of 6
  101. 113Qwen2.5 VL 72BAlibaba−26.248.84 of 6
  102. 114Nova 2.0 LiteAmazon−26.348.71 of 6
  103. 115Grok 4.1 FastSpaceXAI−26.548.51 of 6
  104. 116Llama 4 MaverickMeta−26.648.43 of 6
  105. 117InternVL-3.5 (8B)OpenGVLab−26.848.23 of 6
  106. 118Nova 2.0 Omni Reasoning MediumAmazon−26.948.11 of 6
  107. 119Grok 4 FastSpaceXAI−27.147.91 of 6
  108. 120GPT-5 nanoOpenAI−27.347.61 of 6
  109. 121Qwen3 Omni 30B A3BAlibaba−27.447.61 of 6
  110. 122Magistral Medium 1.2Mistral−27.647.41 of 6
  111. 123Claude Haiku 4.5Anthropic−27.847.21 of 6
  112. 124Gemini 2.5 Flash LiteGoogle−28.047.02 of 6
  113. 125Apriel-v1.5-15B-ThinkerServiceNow−28.047.01 of 6
  114. 126Command A+Cohere−28.246.84 of 6
  115. 127Mistral Small 4Mistral−28.246.81 of 6
  116. 128Mistral Large 3Mistral−28.546.51 of 6
  117. 129Magistral Small 1.2Mistral−28.646.41 of 6
  118. 130Qwen3 Omni 30B A3B InstructAlibaba−28.646.41 of 6
  119. 131Mistral Medium 3.1 (Non-Reasoning)Mistral−28.946.01 of 6
  120. 132GPT-4.5 PreviewOpenAI−29.145.83 of 6
  121. 133Gemini 2.0 Flash ExpGoogle−29.445.54 of 6
  122. 134GLM-4.5VZ.ai−29.845.21 of 6
  123. 135Ministral 3 14BMistral−29.945.11 of 6
  124. 136MiniMax VL 01MiniMax−29.945.04 of 6
  125. 137GLM-4.6VZ.ai−30.144.91 of 6
  126. 138Gemini 1.5 Flash (001)Google−30.244.81 of 6
  127. 139Mistral Small 3.2Mistral−30.444.61 of 6
  128. 140Ministral 3 8BMistral−30.744.31 of 6
  129. 141GPT-4.1OpenAI−30.844.24 of 6
  130. 142Claude 3.5 HaikuAnthropic−30.844.21 of 6
  131. 143Devstral Small 2Mistral−30.944.11 of 6
  132. 144Qwen2.5 VL 32B InstructAlibaba−31.543.54 of 6
  133. 145Qwen2 VL 72B InstructAlibaba−31.543.55 of 6
  134. 146Llama 4 ScoutMeta−31.743.33 of 6
  135. 147MiniCPM-V 4.6 1.3BOpenBMB−31.943.11 of 6
  136. 148Qwen3 VL 8B InstructAlibaba−32.043.05 of 6
  137. 149Qwen3 VL 30B A3B ReasoningAlibaba−32.142.95 of 6
  138. 150Molmo2-8BAllenAI−32.142.81 of 6
  139. 151GPT-4.1 miniOpenAI−32.242.84 of 6
  140. 152Nemotron 3 Nano OmniNVIDIA−32.442.64 of 6
  141. 153Claude 3 HaikuAnthropic−32.542.51 of 6
  142. 154Llama 3.2 Instruct 11B (Vision)Meta−32.842.21 of 6
  143. 155Gemma 3n E4B InstructGoogle−33.141.91 of 6
  144. 156Qwen3.5 0.8BAlibaba−33.241.81 of 6
  145. 157Molmo 7B-DAllenAI−33.341.71 of 6
  146. 158Qwen3 VL 4B InstructAlibaba−33.641.44 of 6
  147. 159LFM2.5-VL-3B-DSparkLiquid AI−33.741.31 of 6
  148. 160Pixtral LargeMistral−34.041.03 of 6
  149. 161InternVL2.5-78BOpenGVLab−34.041.04 of 6
  150. 162Qwen3 VL Thinking (8B)Alibaba−34.340.76 of 6
  151. 163Nova ProAmazon−34.640.32 of 6
  152. 164Qwen3.5 2BAlibaba−34.840.23 of 6
  153. 165Kimi VL A3B ThinkingMoonshot−34.840.22 of 6
  154. 166LFM2.5-VL-3BLiquid AI−35.139.93 of 6
  155. 167NVIDIA Nemotron Nano 12B v2 VLNVIDIA−35.439.63 of 6
  156. 168Gemma 3 27BGoogle−35.439.63 of 6
  157. 169Gemma 4 E4BGoogle−35.839.22 of 6
  158. 170Nova LiteAmazon−36.438.62 of 6
  159. 171InternVL3_5-4BOpenGVLab−36.538.52 of 6
  160. 172LFM2.5-VL-1.6BLiquid AI−37.537.52 of 6
  161. 173Gemini 1.5 ProGoogle−37.637.45 of 6
  162. 174Claude 3.5 SonnetAnthropic−37.837.24 of 6
  163. 175Kimi VL A3B InstructMoonshot−38.037.03 of 6
  164. 176Gemma 3 4BGoogle−38.436.62 of 6
  165. 177Ministral 3 3BMistral−38.436.62 of 6
  166. 178Qwen2.5 Omni 7BAlibaba−38.436.53 of 6
  167. 179Gemma 3 12BGoogle−38.636.43 of 6
  168. 180GPT-4oOpenAI−38.936.15 of 6
  169. 181LFM2-VL-3BLiquid AI−39.036.03 of 6
  170. 182GPT-4o miniOpenAI−39.235.83 of 6
  171. 183InternVL3_5-2BOpenGVLab−39.335.73 of 6
  172. 184Qwen3 VL 4B (Reasoning)Alibaba−41.133.95 of 6
  173. 185Llama 3.2 Instruct 90B (Vision)Meta−41.833.24 of 6
  174. 186Gemini 1.5 Flash 8BGoogle−43.831.24 of 6
  175. 187Gemma 4 E2BGoogle−44.630.43 of 6
  176. 188Phi 4 Multimodal InstructMicrosoft−44.930.15 of 6
  177. 189GPT-4.1 nanoOpenAI−45.829.24 of 6
  178. 190Phi 3.5 Vision InstructMicrosoft−46.828.23 of 6
Back to the top of the ranking

The benchmarks behind the multimodal ranking

Six public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

  • 21%
    Video-MME

    Questions about short and long videos

    Best on this testSeed 2.1 Pro Preview89.2

  • 21%
    OCRBench

    Reads text in images

    Best on this testKimi K2.592.3

  • 17%
    MMMU

    College exam questions with charts, maps and diagrams

    Best on this testQwen3.6 Plus86.0

  • 15%
    CharXiv (reasoning)

    Reasoning about charts from research papers

    Best on this testClaude Mythos Preview93.2

  • 13%
    MMMU-Pro

    Harder college exam questions with images

    Best on this testClaude Opus 5.587.7

  • 13%
    MathVista

    Math problems shown in pictures and charts

    Best on this testSeed 2.1 Pro Preview90.7

  • Tracked, not ranked

    BLINK — visual perception tasks people solve at a glance. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MMStar — questions that truly need the image to answer. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    OCRBench v2 — reads and reasons about text in images. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    LMArena Vision — people's blind votes on answers about images. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

Why these weights

Ranked on MMMU-Pro, OCRBench, Video-MME, MMMU, CharXiv (reasoning) and MathVista (2026-Q3 v2.3, ratified 2026-09-08). MMMU and MMMU-Pro are ONE suite — they agree on 97.6% of ranked pairs, so the pair is capped at 0.30 of the axis rather than counted twice. OCRBench is admitted as the most independent entry measured (text-in-image). BLINK, MMStar, OCRBench v2 and the vision arena stay tracked as depth, visible on model pages but not ranked.

  • Video-MMEvideo facet; the only video-construct entry (VideoMMMU rejected: ρ 0.930 with MMMU-Pro). D .1637 — the set's widest top-12 span
  • OCRBenchOCR / text-in-image facet — the most independent entry measured (ρ .57 vs MMMU-Pro); 57 models, 14 sources — v2.3
  • MMMUmultimodal understanding; ratified C=1.0 core. Suite line with MMMU-Pro (ρ 0.976) — v2.3
  • CharXiv (reasoning)chart-reasoning facet; cards (+Gemini eval-PDF vision rows 2026-08-24); ρ .736 vs MMMU-Pro
  • MMMU-Promultimodal understanding standard, AA day-0 — the universal anchor. Suite line with MMMU — v2.3
  • MathVistavisual math; ratified C=0.6 facet

How the multimodal score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 180 days count.

The gap. Points behind the leader on the multimodal score.

190 models from 30 vendors have a multimodal score, each measured on at least 3 of the 6 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the multimodal ranking

Which LLM is best at multimodal right now?

Gemini 3 Pro, with a multimodal score of 75.0 out of 100. Seed 2.1 Pro Preview is second at 72.9, 2.1 points behind.

Which benchmarks make up the multimodal score?

Six public benchmarks: Video-MME and OCRBench (21% each), MMMU (17%), CharXiv (reasoning) (15%), and MMMU-Pro and MathVista (13% each).

How many models are ranked?

190 models from 30 vendors have a multimodal score. Each needs results on at least 3 of the 6 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 180 days count.

Full methodology and sources