Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Best LLMs for Math

Tier 2 · emergingContested

GPT-6.1 Sol holds the line at 80.5 — GPT-6 Astra runs +1.2 back. The median of the 80-model field trails the leader by 32.3 points.

Updated 9 Oct 2026Score basis
  1. 1GPT-6.1 SolOpenAI80.5Math score, out of 100

    FrontierMath Tier 430%

    100.0Vendor card

    Best score on this test

    MathArena BrokenArXiv (Jun 2026)29%

    100.0Independent

    Best score on this test

    MathArena ArxivMath (Jun 2026)19%

    94.4Independent

    Best score on this test

    FrontierMath Tiers 1-3 (v2)12%

    93.7Independent

    Best score on this test

    LiveBench Mathematics10%

    96.8Independent

    Top 97.1 · Claude Opus 5.5

    5 of 5 benchmarks measured4 independent sources1.2 ahead of GPT-6 Astra32.3 above the field medianCompare the top two
68 more ranked, from 63.1 down to 30.3Field median 48.2, across all 80 scored
Show all 80 ranked modelsShow the top 12 only
  1. 13GPT-5.5 ProOpenAI−17.463.12 of 5
  2. 14Claude Opus 4.8Anthropic−19.461.15 of 5
  3. 15GPT-5.4 ProOpenAI−19.860.82 of 5
  4. 16GPT-5.4OpenAI−19.960.63 of 5
  5. 17Qwen3.8 Max PreviewAlibaba−21.059.65 of 5
  6. 18GPT-6 LunaOpenAI−21.658.93 of 5
  7. 19GPT-5.6 LunaOpenAI−21.658.93 of 5
  8. 20GPT-5.2 ProOpenAI−23.557.02 of 5
  9. 21Claude Opus 4.7Anthropic−24.655.93 of 5
  10. 22Kimi K3Moonshot−24.655.95 of 5
  11. 23GPT-5.2OpenAI−25.355.23 of 5
  12. 24Grok 4.6SpaceXAI−25.455.23 of 5
  13. 25Claude Sonnet 5Anthropic−26.753.83 of 5
  14. 26Union AlphaStealth−28.152.41 of 5
  15. 27GLM-5.3Z.ai−28.452.13 of 5
  16. 28Qwen3.7 MaxAlibaba−28.552.03 of 5
  17. 29Mistral Large 4Mistral−28.751.91 of 5
  18. 30Gemini 3.7 FlashGoogle−29.251.35 of 5
  19. 31Claude Opus 4.6Anthropic−29.351.23 of 5
  20. 32Muse Spark 1.2Meta−29.850.71 of 5
  21. 33Grok 4.7SpaceXAI−30.849.75 of 5
  22. 34GPT-5.2 CodexOpenAI−31.049.51 of 5
  23. 35Nemotron 3 Ultra 550B A55BNVIDIA−31.249.31 of 5
  24. 36DeepSeek-V4.1-FlashDeepSeek−31.449.13 of 5
  25. 37DeepSeek-V4-Flash-Vision-ExpDeepSeek−31.748.81 of 5
  26. 38Claude Sonnet 4.6Anthropic−32.148.41 of 5
  27. 39Gemini 3.1 ProGoogle−32.248.35 of 5
  28. 40Gemini 3.8 FlashGoogle−32.348.25 of 5
  29. 41Haiku 5.5Anthropic−32.448.11 of 5
  30. 42Qwen3.8 27BAlibaba−32.647.91 of 5
  31. 43Qwen3.8 Flash NextAlibaba−32.847.71 of 5
  32. 44Muse Spark 1.1Meta−33.047.53 of 5
  33. 45Grok 4.3SpaceXAI−33.247.41 of 5
  34. 46GPT-5OpenAI−33.247.32 of 5
  35. 47Grok4.5SpaceXAI−33.646.95 of 5
  36. 48GPT-5 ProOpenAI−33.846.72 of 5
  37. 49Kimi K2.6Moonshot−33.946.63 of 5
  38. 50Ox Alpha MaxStealth−34.246.31 of 5
  39. 51MiniMax M3MiniMax−34.446.21 of 5
  40. 52GLM-5.2Z.ai−35.245.35 of 5
  41. 53Gemini 3 Flash PreviewGoogle−35.744.82 of 5
  42. 54Inkling-SmallThinking Machines−36.044.52 of 5
  43. 55Grok 4.20SpaceXAI−36.444.12 of 5
  44. 56Gemini 3.5 FlashGoogle−37.143.55 of 5
  45. 57Step 3.7 FlashStepFun−37.143.42 of 5
  46. 58Qwen3.6 PlusAlibaba−37.443.12 of 5
  47. 59Grok 4.3 BetaSpaceXAI−37.842.72 of 5
  48. 60GPT-5 miniOpenAI−37.942.62 of 5
  49. 61GPT-5.4 nanoOpenAI−38.042.53 of 5
  50. 62GLM 5.3 FlashZ.ai−38.042.53 of 5
  51. 63Qwen3.6 27BAlibaba−38.442.12 of 5
  52. 64DeepSeek-V4-FlashDeepSeek−39.041.55 of 5
  53. 65Qwen3.5 2BAlibaba−40.540.02 of 5
  54. 66O4 MiniOpenAI−40.639.92 of 5
  55. 67Kimi K2.7 CodeMoonshot−40.639.93 of 5
  56. 68Gemini 3.6 FlashGoogle−40.939.65 of 5
  57. 69Claude Opus 4.5Anthropic−41.039.53 of 5
  58. 70DeepSeek-V4-ProDeepSeek−41.239.33 of 5
  59. 71GPT-5.4 miniOpenAI−41.938.63 of 5
  60. 72InklingThinking Machines−42.138.53 of 5
  61. 73GPT-5.5 InstantOpenAI−43.437.22 of 5
  62. 74Claude Sonnet 4.5Anthropic−43.836.72 of 5
  63. 75GPT-5 nanoOpenAI−44.236.32 of 5
  64. 76Claude Opus 4.1Anthropic−45.135.52 of 5
  65. 77Gemini 2.5 ProGoogle−45.235.32 of 5
  66. 78O3 MiniOpenAI−46.234.32 of 5
  67. 79Gemini 3.5 Flash LiteGoogle−49.131.43 of 5
  68. 80Qwen3.6 35B A3BAlibaba−50.230.33 of 5
Back to the top of the ranking

The benchmarks behind the math ranking

Five public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.

  • 30%
    FrontierMath Tier 4

    Extremely hard research-level math problems

    Best on this testGPT-6.1 Sol100.0

  • 29%
    MathArena BrokenArXiv (Jun 2026)

    Spots subtly false math statements it's asked to prove

    Best on this testClaude Opus 5.5100.0

  • 19%
    MathArena ArxivMath (Jun 2026)

    Research-level math questions from recent papers

    Best on this testGPT-6 Astra94.4

  • 12%
    FrontierMath Tiers 1-3 (v2)

    Unpublished advanced math problems

    Best on this testGPT-6 Astra93.7

  • 10%
    LiveBench Mathematics

    Fresh competition math problems

    Best on this testClaude Opus 5.597.1

  • Tracked, not ranked

    MathArena ArxivMath (overall) — research-level math questions from recent papers. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena BrokenArXiv (overall) — spots subtly false math statements it's asked to prove. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena ArxivMath (May 2026) — research-level math questions from recent papers. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena ArxivMath (Apr 2026) — research-level math questions from recent papers. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena BrokenArXiv (May 2026) — spots subtly false math statements it's asked to prove. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena BrokenArXiv (Apr 2026) — spots subtly false math statements it's asked to prove. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena ArxivLean (Jun 2026) — writes computer-checked proofs of new research results. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    IMO-AnswerBench — olympiad math problems with short answers. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    HMMT Feb 2026 — harvard-MIT high-school math contest problems. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    AIME 2026 — US invitational high-school math exam problems. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    MathArena Apex 2025 — very hard math contest problems from 2025. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

  • Tracked, not ranked

    USAMO 2026 — proof problems from the US Math Olympiad. It updates too slowly to score new models within a week, so it sits outside the ranking; its scores appear on each model's page, and it returns to the ranking when it keeps pace.

Why these weights

Ranked on FrontierMath Tier 4 and Tiers 1-3 v2, MathArena BrokenArXiv and ArxivMath (newest live monthly cut) and LiveBench Mathematics (2026-Q3 v2.7, 2026-09-10). Every ranked entry currently scores newly released models. MathArena's cross-month "overall" boards are tracked as depth rather than ranked: they only score a model that has been run on every live monthly cut, and print N/A for four current flagships including GPT-5.6 Sol and Claude Opus 5. The competition boards — AIME 2026, HMMT Feb 2026, IMO-AnswerBench, Apex 2025, USAMO 2026 — are tracked as depth: they are no longer re-run on new launches, and AIME is saturated at the frontier (every measured top-25 model scores 95+). Depth entries stay visible on model pages and re-promote if a board resumes.

  • FrontierMath Tier 4research-grade held-out frontier math (Epoch v2 task); D .331 on 57 models
  • MathArena BrokenArXiv (Jun 2026)perturbed research problems — reasoning over recall of a published solution; newest LIVE cut, 21 models (MATURE), D .334; RANKED in v2.7 in place of the overall, which prints N/A for four current flagships
  • MathArena ArxivMath (Jun 2026)unseen research-paper problems, monthly re-cut — contamination-resistant by construction; newest LIVE cut, 21 models (MATURE), D .156; RANKED in v2.7 in place of the overall
  • FrontierMath Tiers 1-3 (v2)Epoch's graded ladder below Tier 4 — same construct, same runner, 70 gated models; ADDED v2.6 after mig 1009 ingested it
  • LiveBench Mathematicsrolling contamination-free math — the third independent publisher line

How the math score is calculated

  1. 1

    Rank on each benchmark

    Every model gets a percentile on each benchmark it has been measured on.

  2. 2

    Steady the thin fields

    Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.

  3. 3

    Weigh and average

    The percentiles are averaged with the weights above into one score out of 100.

  4. 4

    Qualify

    A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.

Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.

Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.

Printed values. Every score is the value its source printed. Only scores first seen in the last 210 days count.

The gap. Points behind the leader on the math score.

80 models from 15 vendors have a math score, each measured on at least 3 of the 5 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.

Questions about the math ranking

Which LLM is best at math right now?

GPT-6.1 Sol, with a math score of 80.5 out of 100. GPT-6 Astra is second at 79.3, 1.2 points behind.

Which benchmarks make up the math score?

Five public benchmarks: FrontierMath Tier 4 (30%), MathArena BrokenArXiv (Jun 2026) (29%), MathArena ArxivMath (Jun 2026) (19%), FrontierMath Tiers 1-3 (v2) (12%), and LiveBench Mathematics (10%).

How many models are ranked?

80 models from 15 vendors have a math score. Each needs results on at least 3 of the 5 benchmarks to be ranked.

How current is the ranking?

It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 210 days count.

Full methodology and sources