Llama 3.1 Instruct 405B
Llama 3.1 Instruct 405B is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
No current price is tracked for this model. See the rate card
Evidence
151results on110benchmarks
- 17 independently verified
- 15 aggregator
- 21 vendor-reported
- 98 cross-referenced
From 24 sources · latest Oct 7, 2026 · How verification works
Llama 3.1 Instruct 405B benchmark results
151 results on 110 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
27.8% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination47.56Oct 7, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy23.18Oct 7, 2026omniscienceAccuracy
0 of 6 ranked benchmarks measured
- 23.00Aug 29, 2026
- 0.00Oct 7, 2026
- GPQA Diamond51.52Oct 7, 2026gpqa
Show 5 more reasoning resultsHide 5 reasoning results
- GPQA Diamond50.70Oct 6, 2026GPQA
- GPQA Diamond49.00Sep 3, 2026GPQA Diamond (CoT)
- GPQA Diamond50.70Jun 6, 2026GPQA (diamond)
- 51.10May 31, 2026
- Humanity's Last Exam3.98Oct 7, 2026aa_hle
0 of 10 ranked benchmarks measured
- 11.18Oct 7, 2026
- LiveBench · Coding41.41Aug 23, 2026livebench_coding@2025-04-07
- Terminal-Bench Hard6.82Oct 7, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- SciCode29.86Sep 4, 2026aa_scicode
- 24.50May 3, 2026
0 of 7 ranked benchmarks measured
- 0.00Jun 15, 2026
0 of 3 ranked benchmarks measured
- 25.33Oct 7, 2026
- 36.10May 3, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following76.83Aug 23, 2026livebench_instruction_following@2025-04-07
- IFBench39.05Oct 7, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 95.89Aug 29, 2026
- 98.75Aug 29, 2026
- 62.69Aug 29, 2026
- 94.50Aug 29, 2026
- 96.46Aug 29, 2026
- 45.60Aug 29, 2026
Show 125 more resultsHide 125 results
- 6.34Jun 18, 2026
- AA Intelligence7.29Oct 7, 2026aa_intelligence_index
- -17.10Oct 7, 2026
- AGIEval60.60Aug 31, 2026AGIEval (Acc.)
- 63.90May 3, 2026
- 5.80May 3, 2026
- 23.30May 3, 2026
- 58.63Aug 29, 2026
- 6.00May 31, 2026
- AlpacaEval 2 LC39.30Sep 10, 2026
- 39.30May 31, 2026
- ARC-Challenge96.90Oct 6, 2026ARC-C
- ARC-Challenge96.90May 31, 2026ARC-C
- ARC-Challenge96.10May 31, 2026ARC-C
- 95.30May 1, 2026
- 95.30Jul 6, 2026
- 98.40May 1, 2026
- 69.30Sep 10, 2026
- 63.50Jun 6, 2026
- 14.50Jun 18, 2026
- 85.90May 31, 2026
- 82.90May 1, 2026
- 82.90Jul 6, 2026
- 81.10Aug 29, 2026
- BFCLv488.50Oct 6, 2026BFCL
- C-Eval72.50Aug 31, 2026C-Eval (Acc.)
- 61.50May 3, 2026
- 79.70May 1, 2026
- 79.70Jul 6, 2026
- CCPM78.60Aug 31, 2026CCPM (Acc.)
- Chinese SimpleQA (C-SimpleQA)54.70Jun 6, 2026C-SimpleQA
- 50.40May 3, 2026
- CLUEWSC83.00Aug 31, 2026CLUEWSC (EM)
- 84.70May 3, 2026
- 77.30May 1, 2026
- CMMLU73.70May 1, 2026chinese_cmmlu_acc
- 73.70Jul 6, 2026
- CMRC76.00Aug 31, 2026CMRC (EM)
- 6.80May 3, 2026
- 25.30May 3, 2026
- 85.80May 31, 2026
- CRUXEval-I (input prediction)58.50May 1, 2026CRUXEval-I (Acc.)
- CRUXEval-O (output prediction)59.90May 1, 2026CRUXEval-O (Acc.)
- 84.80Oct 6, 2026
- 84.80May 31, 2026
- 88.70May 3, 2026
- 92.50Jun 6, 2026
- 86.00May 1, 2026
- 70.00May 3, 2026
- 94.90Jul 2, 2026
- 96.80Oct 6, 2026
- 96.70Jun 6, 2026
- 89.00May 31, 2026
- 83.50May 1, 2026
- 89.20May 1, 2026
- 89.20Jul 6, 2026
- 89.00Oct 6, 2026
- 89.00Jun 6, 2026
- 61.00May 31, 2026
- HumanEval54.90May 1, 2026HumanEval (Pass@1)
- 77.20May 3, 2026
- 88.60Sep 11, 2026
- IFEval86.00May 31, 2026IFEval strict-prompt
- 86.40Jun 6, 2026
- livebench_language50.02Aug 23, 2026livebench_language@2025-04-07
- LiveCodeBench27.70May 1, 2026code_livecodebench_1001202402012025
- LiveCodeBench30.10May 3, 2026LiveCodeBench (Pass@1)
- LiveCodeBench15.50May 1, 2026code_livecodebenchbase_pass1
- 28.40May 3, 2026
- 15.50Jul 6, 2026
- 82.70Aug 29, 2026
- 73.80May 1, 2026
- 73.80Jun 6, 2026
- 53.80May 31, 2026
- 49.00May 1, 2026
- 73.80Jul 6, 2026
- 49.00Jul 6, 2026
- 73.80May 3, 2026
- 73.40May 31, 2026
- MBPP68.40May 1, 2026MBPP (Pass@1)
- 88.60Aug 29, 2026
- 73.00Jun 6, 2026
- 91.60Sep 11, 2026
- 69.90May 1, 2026
- 69.90Jul 6, 2026
- 75.91Aug 29, 2026
- 88.60May 1, 2026
- 88.60Jun 6, 2026
- 87.30May 31, 2026
- 85.20May 31, 2026
- 81.30May 1, 2026
- 84.40Jul 6, 2026
- 88.60Jul 6, 2026
- 73.30Jul 6, 2026
- 73.30Oct 6, 2026
- MMLU-Pro73.40May 1, 2026reasoning_knowledge_mmlu_pro
- MMLU-Pro52.80Jul 3, 2026english_mmlupro_acc
- 73.30Jun 6, 2026
- 61.60May 31, 2026
- 52.80Jul 6, 2026
- 86.20May 3, 2026
- 81.30Jul 6, 2026
- MMMLU73.80Jul 6, 2026Multilingual (MMMLU-non-English (Acc.))
- 73.80Jul 7, 2026
- MT-Bench8.49Sep 10, 2026MT-Bench (GPT-4-Turbo)
- 9.10May 31, 2026
- 75.20Oct 6, 2026
- 65.70Oct 6, 2026
- 74.94Aug 29, 2026
- 41.50May 1, 2026
- 41.50Jul 6, 2026
- 94.00Jul 2, 2026
- Pile-test (BPB)lower is better0.54Jul 6, 2026
- 85.90May 1, 2026
- 85.90Jul 6, 2026
- RACE-High56.80Aug 31, 2026RACE-High (Acc.)
- RACE-Middle74.20Aug 31, 2026RACE-Middle (Acc.)
- 23.20Jun 6, 2026
- 17.10May 3, 2026
- The Pile (Test, BPB)lower is better0.54May 1, 2026
- 82.70May 1, 2026
- 86.70May 31, 2026
- 85.20May 1, 2026
- 85.20Jul 6, 2026
- τ²-Bench Telecom (AA run)19.01Oct 7, 2026aa_tau2
Llama 3.1 Instruct 405B: common questions
Who makes Llama 3.1 Instruct 405B?
Llama 3.1 Instruct 405B is made by Meta.
When was Llama 3.1 Instruct 405B released?
Llama 3.1 Instruct 405B was released on Jul 23, 2024, according to Artificial Analysis.
What is Llama 3.1 Instruct 405B good at?
Llama 3.1 Instruct 405B is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How many benchmarks has Llama 3.1 Instruct 405B been tested on?
We track 151 results for Llama 3.1 Instruct 405B on 110 benchmarks from 24 sources, 17 of them independently verified. The latest was recorded on Oct 7, 2026.
About this record
Where Llama 3.1 Instruct 405B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 10, 2026
Where the results come from
Verification: 151 scores · 17 independently verified · 15 aggregator-attributed · 98 vendor cross-reference · 21 vendor-reported. How these tiers are assigned
From 24 sources on 8 sites. raw.githubusercontent.com supplies 59 of them; the 17 independently verified results come from 4 sites. Bars are coloured by trust tier.
- raw.githubusercontent.com59
- huggingface.co36
- arxiv.org18
- artificialanalysis.ai15
- storage.googleapis.com12
- api.llm-stats.com9
- labs.scale.com1
- simple-bench.com1