Most Factually Accurate LLMs
GPT-6 Astra leads GPT-6.1 Sol by just 0.8 at 79.5 — inside the margin, with no model clear of the pack. The median of the 357-model field trails the leader by 29.9 points.
- 1GPT-6 AstraOpenAI79.5Factuality score, out of 1003 of 4 benchmarks measured1 awaiting a board run2 independent sources0.8 ahead of GPT-6.1 Sol29.9 above the field medianCompare the top two
- 2GPT-6.1 SolOpenAIWithin 1% of the leader−0.878.73 of 4
- 3Claude Opus 5.5Anthropic−2.077.53 of 4
- 4Gemini 3.1 ProGoogle−2.277.24 of 4
- 5Gemini 3.8 FlashGoogle−3.376.23 of 4
- 6Claude Fable 5Anthropic−3.576.03 of 4
- 7Grok 4.7SpaceXAI−4.475.13 of 4
- 8Claude Fable 5.1Anthropic−4.575.03 of 4
- 9Muse Spark 1.2Meta−4.774.83 of 4
- 10Gemini 3.7 FlashGoogle−5.174.43 of 4
- 11Gemini 3.6 FlashGoogle−5.574.03 of 4
- 12Gemini 3.5 FlashGoogle−6.273.33 of 4
Show all 200 ranked modelsShow the top 12 only
- 13Muse Spark 1.1Meta−6.373.23 of 4
- 14GPT-6 SolOpenAI−6.672.93 of 4
- 15Claude Opus 4.8Anthropic−6.772.83 of 4
- 16Grok 4.6SpaceXAI−6.772.73 of 4
- 17Claude Opus 5Anthropic−6.972.63 of 4
- 18Qwen3.7 MaxAlibaba−7.072.53 of 4
- 19Qwen3.6 Max PreviewAlibaba−9.669.93 of 4
- 20Kimi K3Moonshot−9.769.83 of 4
- 21Gemini 4 ArgonGoogle−10.469.12 of 4
- 22Grok4.5SpaceXAI−10.768.83 of 4
- 23Qwen3.8 Max PreviewAlibaba−11.168.43 of 4
- 24Claude Opus 4.7Anthropic−11.667.94 of 4
- 25Muse SparkMeta−11.767.83 of 4
- 26Muse Spark 1.3Meta−13.765.82 of 4
- 27Grok 4.3SpaceXAI−13.865.72 of 4
- 28Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic−14.065.42 of 4
- 29GPT-5.5OpenAI−14.365.24 of 4
- 30Grok 4SpaceXAI−14.664.93 of 4
- 31GLM-5.3Z.ai−14.764.83 of 4
- 32GPT-5.6 SolOpenAI−14.964.63 of 4
- 33Motif-3Motif Technologies−15.264.32 of 4
- 34Step 5 PreviewStepFun−15.464.12 of 4
- 35Claude 5.5Anthropic−15.663.92 of 4
- 36Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic−15.763.82 of 4
- 37Haiku 5.5Anthropic−15.963.52 of 4
- 38JT-4.1 Flash 236B A21BChina Mobile−16.163.42 of 4
- 39MiMo V2.6 ProXiaomi−16.263.32 of 4
- 40Qwen3.6 PlusAlibaba−16.363.23 of 4
- 41Qwen3.8 2.4T A95BAlibaba−16.662.92 of 4
- 42Gemini 3.5 Flash LiteGoogle−16.662.92 of 4
- 43Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic−16.662.92 of 4
- 44MiMo V2 ProXiaomi−16.962.62 of 4
- 45Ling-3.1-flashInclusionAI−17.162.32 of 4
- 46Gemini 2.5 ProGoogle−17.462.14 of 4
- 47GPT-5.1OpenAI−17.562.04 of 4
- 48InklingThinking Machines−17.661.93 of 4
- 49Claude Opus 4.5Anthropic−17.861.74 of 4
- 50Claude 3.7 SonnetAnthropic−17.861.72 of 4
- 51MiMo V2.5 ProXiaomi−18.261.32 of 4
- 52Claude Opus 4.6Anthropic−18.561.04 of 4
- 53Gemini 3 ProGoogle−18.561.04 of 4
- 54Qwen3.7 Plus PreviewAlibaba−18.660.92 of 4
- 55Claude Sonnet 5Anthropic−18.960.63 of 4
- 56GLM-5Z.ai−19.060.53 of 4
- 57GPT-5.5 InstantOpenAI−19.060.53 of 4
- 58Mistral Large 4Mistral−19.060.42 of 4
- 59GPT-6 LunaOpenAI−19.659.93 of 4
- 60DeepSeek-V3.2-ExpDeepSeek−19.659.93 of 4
- 61Claude Sonnet 4Anthropic−19.859.73 of 4
- 62Agnes 3.0 FlashSapiens AI−19.859.62 of 4
- 63GPT-5.2 CodexOpenAI−19.859.62 of 4
- 64MiMo V2.6 Flash RLXiaomi−20.259.22 of 4
- 65Qwen3.8 Flash NextAlibaba−20.359.22 of 4
- 66GLM-5.2Z.ai−20.359.23 of 4
- 67O1OpenAI−20.359.13 of 4
- 68Solar Pro4Upstage−20.559.02 of 4
- 69DeepSeek-V3.1DeepSeek−20.858.73 of 4
- 70Solar Open2 250BUpstage−20.858.72 of 4
- 71GLM-5-TurboZ.ai−20.858.72 of 4
- 72GPT-5 CodexOpenAI−20.858.72 of 4
- 73Kimi K2.6Moonshot−21.058.54 of 4
- 74GPT-5.1 CodexOpenAI−21.158.42 of 4
- 75O3OpenAI−21.158.43 of 4
- 76GLM 5.3 FlashZ.ai−21.258.33 of 4
- 77Grok 4.20SpaceXAI−21.358.23 of 4
- 78Hy3-previewTencent−21.657.92 of 4
- 79GLM 5V TurboZ.ai−21.657.92 of 4
- 80K2 Horizon 375B A23BMBZUAI−21.757.82 of 4
- 81GLM-5.1Z.ai−21.857.73 of 4
- 82Gemini 3 Flash PreviewGoogle−21.957.64 of 4
- 83GPT-5.1 Codex miniOpenAI−22.057.52 of 4
- 84Gemini 2.5 FlashGoogle−22.057.53 of 4
- 85Llama 3.1 Instruct 405BMeta−22.157.32 of 4
- 86Kimi K2 ThinkingMoonshot−22.357.22 of 4
- 87Claude Sonnet 4.6Anthropic−22.357.24 of 4
- 88GPT-5.4OpenAI−22.457.14 of 4
- 89Gemini 3.1 Flash Lite PreviewGoogle−22.457.13 of 4
- 90LongCat-2.0Meituan−22.557.02 of 4
- 91MiniMax M3MiniMax−22.557.02 of 4
- 92Motif-3-BetaMotif Technologies−22.556.92 of 4
- 93A.X-K2SK Telecom−22.756.82 of 4
- 94Apodex 1.1Apodex−23.056.52 of 4
- 95Gemini 2.5 Flash LiteGoogle−23.056.53 of 4
- 96DeepSeek-V3.1-TerminusDeepSeek−23.156.42 of 4
- 97JT-35B-FlashChina Mobile−23.356.22 of 4
- 98Solar Mini4Upstage−23.356.12 of 4
- 99GPT-5OpenAI−23.456.14 of 4
- 100Nemotron 3 Ultra 550B A55BNVIDIA−23.555.93 of 4
- 101GLM-4.5Z.ai−23.655.92 of 4
- 102Kimi K2 (Non-Reasoning)Moonshot−23.755.82 of 4
- 103MiMo V2 OmniXiaomi−23.855.72 of 4
- 104Qwen3.6 27BAlibaba−23.955.62 of 4
- 105Grok 3SpaceXAI−24.055.53 of 4
- 106G9v3-39A5BAI9Stars−24.255.32 of 4
- 107GPT-5.6 TerraOpenAI−24.255.33 of 4
- 108Magistral Medium 1.2Mistral−24.355.22 of 4
- 109Agnes 2.5 Pro BetaSapiens AI−24.455.12 of 4
- 110Ling-3.0-flash-FinInclusionAI−24.455.12 of 4
- 111Qwen3.5 Omni PlusAlibaba−24.555.02 of 4
- 112Ling-3.0-flashInclusionAI−24.654.92 of 4
- 113Magistral Medium 1Mistral−24.754.82 of 4
- 114Qwen3.6 35B A3BAlibaba−24.854.72 of 4
- 115Mistral Small 4Mistral−24.954.62 of 4
- 116DeepSeek-V3DeepSeek−24.954.63 of 4
- 117Kimi K2.7 CodeMoonshot−24.954.53 of 4
- 118GPT-5.3 CodexOpenAI−25.054.52 of 4
- 119Ling-3.0-flash-VLInclusionAI−25.154.42 of 4
- 120GPT-5.4 ProOpenAI−25.254.32 of 4
- 121Qwen3.8 27BAlibaba−25.254.22 of 4
- 122GPT-5.2OpenAI−25.354.14 of 4
- 123Nova ProAmazon−25.554.03 of 4
- 124GPT-4OpenAI−25.853.72 of 4
- 125Doubao Seed CodeByteDance−25.953.62 of 4
- 126Nova PremierAmazon−25.953.62 of 4
- 127Muse GlimmerMeta−26.053.52 of 4
- 128K2 Horizon 36B-A4BMBZUAI−26.253.32 of 4
- 129K-EXAONE 2.0LG AI−26.253.32 of 4
- 130Devstral MediumMistral−26.453.12 of 4
- 131DeepSeek-V4-ProDeepSeek−26.453.14 of 4
- 132Grok Code Fast 1SpaceXAI−26.453.02 of 4
- 133Mistral Large 2Mistral−26.553.02 of 4
- 134ERNIE 4.5 300B A47BBaidu−26.852.62 of 4
- 135Mistral Medium 3.5 128BMistral−27.052.52 of 4
- 136Trinity Large ThinkingArcee AI−27.252.33 of 4
- 137Qwen3 Coder 480B A35B InstructAlibaba−27.252.32 of 4
- 138Mercury 2.5 PreviewInception−27.252.32 of 4
- 139Qwen3 14BAlibaba−27.352.23 of 4
- 140DeepSeek-V3.2-SpecialeDeepSeek−27.452.12 of 4
- 141Agnes 2.5 Pro AlphaSapiens AI−27.452.12 of 4
- 142Gemma 4 26B A4BGoogle−27.452.13 of 4
- 143Nemotron 3.5 LightningNVIDIA−27.651.92 of 4
- 144Claude Sonnet 4.5Anthropic−27.651.94 of 4
- 145Command A+Cohere−27.651.92 of 4
- 146Step 3.7 FlashStepFun−27.751.82 of 4
- 147Granite 4.2 30BIBM−27.751.82 of 4
- 148Kimi K2.5Moonshot−27.851.74 of 4
- 149GLM-4.6VZ.ai−27.851.72 of 4
- 150GPT-4oOpenAI−28.051.54 of 4
- 151K2 Horizon 7BMBZUAI−28.351.12 of 4
- 152MiniCPM5-2BOpenBMB−28.351.12 of 4
- 153Phi 4Microsoft−28.551.03 of 4
- 154Llama 3.3 70B InstructMeta−28.650.93 of 4
- 155MiniMax M2.5MiniMax−28.650.93 of 4
- 156G9v3-3BAI9Stars−28.650.92 of 4
- 157Granite 4.2 8BIBM−28.650.92 of 4
- 158Qwen3 32BAlibaba−28.650.83 of 4
- 159Granite 4.2 3BIBM−28.750.82 of 4
- 160Llama 3.1 70B InstructMeta−28.850.72 of 4
- 161K2 Think V2MBZUAI−28.850.72 of 4
- 162Inkling-SmallThinking Machines−28.950.63 of 4
- 163Nova LiteAmazon−29.050.53 of 4
- 164MiniMax M2.1MiniMax−29.050.53 of 4
- 165Nova MicroAmazon−29.150.43 of 4
- 166GPT-5.6 LunaOpenAI−29.250.33 of 4
- 167Llama 3.1 Nemotron Instruct 70BNVIDIA−29.250.32 of 4
- 168Qwen3 235B A22B Instruct 2507Alibaba−29.250.32 of 4
- 169GPT-5.4 nanoOpenAI−29.350.24 of 4
- 170LFM2.5-2.6BLiquid AI−29.450.12 of 4
- 171Step 3.5 FlashStepFun−29.649.92 of 4
- 172Magistral Small 1Mistral−29.749.82 of 4
- 173Llama 3.1 Nemotron Ultra 253B V1NVIDIA−29.749.82 of 4
- 174Qwen3 235B A22BAlibaba−29.749.84 of 4
- 175MiniCPM5-1BOpenBMB−29.749.82 of 4
- 176Ling 3.0 TinyInclusionAI−29.849.72 of 4
- 177Gemini 2.5 Flash (Sep) (Non-Reasoning)Google−29.849.72 of 4
- 178Gemma 4 E4BGoogle−29.949.62 of 4
- 179DeepSeek-R1-Distill-Llama-70BDeepSeek−29.949.62 of 4
- 180Kimi K2 InstructMoonshot−29.949.63 of 4
- 181Gemini 2.0 Flash (Non-Reasoning)Google−30.149.42 of 4
- 182Nemotron 3 Super 120B A12BNVIDIA−30.149.42 of 4
- 183GLM-4.5VZ.ai−30.449.12 of 4
- 184Qwen3 VL 235B A22B ReasoningAlibaba−30.648.92 of 4
- 185K2 Horizon 3.7BMBZUAI−30.648.92 of 4
- 186Grok 4.1 FastSpaceXAI−30.748.73 of 4
- 187Llama 4 MaverickMeta−30.848.72 of 4
- 188Devstral 2Mistral−30.848.72 of 4
- 189Gemma 4 E2BGoogle−31.048.52 of 4
- 190Command ACohere−31.148.43 of 4
- 191DeepSeek-V4.1-FlashDeepSeek−31.148.42 of 4
- 192ERNIE 5.0 Thinking PreviewBaidu−31.248.32 of 4
- 193Llama Nemotron Super 49B V1.5NVIDIA−31.348.22 of 4
- 194Grok 4 FastSpaceXAI−31.548.03 of 4
- 195LFM2.5-8B-A1BLiquid AI−31.548.02 of 4
- 196Claude 3 HaikuAnthropic−31.548.02 of 4
- 197Gemma 3 270MGoogle−31.647.92 of 4
- 198Llama 4 ScoutMeta−31.647.93 of 4
- 199North-Mini-Code-1.0Cohere−31.647.92 of 4
- 200Llama 3.1 8B InstructMeta−31.747.82 of 4
The benchmarks behind the factuality ranking
Four public benchmarks decide this ranking. The heavier a benchmark's weight, the more it moves a model's score.
- 39%SimpleQA Verified
Short factual questions, answered correctly
Best on this testGPT-6 Astra75.6
- 21%Vectara HHEM hallucination rate
How often its summaries add facts not in the source
Best on this testFinix_s1_32b1.8%
- 20%AA-Omniscience Accuracy
Answers hard knowledge questions correctly
Best on this testClaude Fable 5.167.2
- 20%AA-Omniscience Non-hallucination
Avoids making up answers it doesn't know
Best on this testG9v3-3B88.3
Why these weights
Ranked on AA-Omniscience (accuracy + non-hallucination), SimpleQA Verified and Vectara HHEM (2026-Q3 v2, ratified 2026-08-23). Third-party day-0 coverage is AA's alone; new launches reach eligibility as Epoch and Vectara publish.
- SimpleQA Verifiedknowledge accuracy; Epoch + cards (~18d)
- Vectara HHEM hallucination ratesummarization faithfulness, inverted (lower is better); Vectara board
- AA-Omniscience Accuracyknowledge accuracy (0–100 component), AA day-0; Omniscience line .40
- AA-Omniscience Non-hallucinationcalibration (0–100 component), AA day-0
How the factuality score is calculated
- 1
Rank on each benchmark
Every model gets a percentile on each benchmark it has been measured on.
- 2
Steady the thin fields
Where few models have taken a benchmark, that percentile is pulled toward the middle of the field.
- 3
Weigh and average
The percentiles are averaged with the weights above into one score out of 100.
- 4
Qualify
A model enters once it is measured on at least half the basket by weight, including one anchor benchmark.
Missing scores. A missing score on a well-covered benchmark counts as the middle of the field, never as zero.
Suites count once. Members of one suite, such as SWE-bench, count together, so a lab that reports one member is not penalised three times.
Printed values. Every score is the value its source printed. Only scores first seen in the last 120 days count.
The gap. Points behind the leader on the factuality score.
357 models from 51 vendors have a factuality score, each measured on at least 2 of the 4 benchmarks. Scores come from AA-graded results, official model cards and third-party evaluations.
Questions about the factuality ranking
Which LLM is best at factuality right now?
GPT-6 Astra, with a factuality score of 79.5 out of 100. GPT-6.1 Sol is 0.8 points behind at 78.7 — inside the margin, so the top of the ranking is a dead heat.
Which benchmarks make up the factuality score?
Four public benchmarks: SimpleQA Verified (39%), Vectara HHEM hallucination rate (21%), and AA-Omniscience Accuracy and AA-Omniscience Non-hallucination (20% each).
How many models are ranked?
357 models from 51 vendors have a factuality score. Each needs results on at least 2 of the 4 benchmarks to be ranked.
How current is the ranking?
It updates as new results are published and was last updated on 9 October 2026. Only scores first seen in the last 120 days count.