Llama 3.1 70B Instruct
Llama 3.1 70B Instruct is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Instruction Following.
Price
$0.56input$0.56outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
67results on53benchmarks
- 15 independently verified
- 15 aggregator
- 18 vendor-reported
- 19 cross-referenced
From 18 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Llama 3.1 70B Instruct benchmark results
67 results on 53 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
36.2% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy19.70Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination21.79Oct 8, 2026omniscienceNonHallucination
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond40.91Oct 8, 2026gpqa
- Humanity's Last Exam4.47Oct 8, 2026aa_hle
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond41.70Oct 7, 2026GPQA
- GPQA Diamond48.00Sep 3, 2026GPQA Diamond (CoT)
- 46.70May 31, 2026
0 of 10 ranked benchmarks measured
- LiveBench · Coding34.38Aug 23, 2026livebench_coding@2025-04-07
- Terminal-Bench Hard3.03Oct 8, 2026aa_terminalbench_hard
- SciCode26.74Sep 4, 2026aa_scicode
0 of 7 ranked benchmarks measured
- 0.00Jun 15, 2026
0 of 3 ranked benchmarks measured
- 8.00Sep 4, 2026
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following88.83Aug 23, 2026livebench_instruction_following@2025-04-07
- IFBench34.42Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language42.57Aug 23, 2026livebench_language@2025-04-07
- 93.80Jul 2, 2026
- 93.80Jul 2, 2026
- 42.50May 19, 2026
- 94.50May 10, 2026
- 92.50May 10, 2026
Show 46 more resultsHide 46 results
- 5.07Jun 18, 2026
- AA Intelligence6.60Oct 8, 2026aa_intelligence_index
- -43.10Oct 8, 2026
- 5.90May 31, 2026
- AlpacaEval 2 LC38.10Sep 10, 2026
- 34.30May 31, 2026
- 93.21May 10, 2026
- ARC-Challenge94.80Oct 7, 2026ARC-C
- ARC-Challenge94.80May 31, 2026ARC-C
- 55.70Sep 10, 2026
- 10.93Jun 18, 2026
- 81.60May 31, 2026
- 95.40May 10, 2026
- 77.50May 1, 2026
- BFCLv484.80Oct 7, 2026BFCL
- 84.10May 31, 2026
- 79.60Oct 7, 2026
- 79.60May 31, 2026
- 83.70May 31, 2026
- 46.88May 10, 2026
- 80.50Oct 7, 2026
- 80.50May 31, 2026
- 87.50Aug 31, 2026
- IFEval83.60May 31, 2026IFEval strict-prompt
- 78.33May 10, 2026
- 68.00May 1, 2026
- 68.00May 31, 2026
- 68.00Jul 6, 2026
- 68.60May 31, 2026
- 86.00May 1, 2026
- 86.90May 1, 2026
- 70.91May 10, 2026
- 86.00May 1, 2026
- 83.60May 31, 2026
- 86.00Jul 6, 2026
- 66.40Jul 6, 2026
- 66.40Oct 7, 2026
- 53.80May 31, 2026
- MT-Bench8.22Sep 10, 2026MT-Bench (GPT-4-Turbo)
- 8.80May 31, 2026
- 65.50Oct 7, 2026
- 62.00Oct 7, 2026
- 77.24May 10, 2026
- 45.17May 10, 2026
- 85.30May 31, 2026
- τ²-Bench Telecom (AA run)15.20Oct 8, 2026aa_tau2
Llama 3.1 70B Instruct: common questions
Who makes Llama 3.1 70B Instruct?
Llama 3.1 70B Instruct is made by Meta.
When was Llama 3.1 70B Instruct released?
Llama 3.1 70B Instruct was released on Jul 23, 2024, according to Artificial Analysis.
What is Llama 3.1 70B Instruct good at?
Llama 3.1 70B Instruct is behind the leaders in factuality. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or instruction following.
How much does Llama 3.1 70B Instruct cost?
Llama 3.1 70B Instruct costs $0.56 per million input tokens and $0.56 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 55% of the 330 priced models we track.
How many benchmarks has Llama 3.1 70B Instruct been tested on?
We track 67 results for Llama 3.1 70B Instruct on 53 benchmarks from 18 sources, 15 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Llama 3.1 70B Instruct support?
OpenRouter lists tool calling and structured outputs for Llama 3.1 70B Instruct.
About this record
Where Llama 3.1 70B Instruct's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 3, 2026
Where the results come from
Verification: 67 scores · 15 independently verified · 15 aggregator-attributed · 19 vendor cross-reference · 18 vendor-reported. How these tiers are assigned
From 18 sources on 5 sites. Hugging Face supplies 22 of them; the 15 independently verified results come from 2 sites. Bars are coloured by trust tier.
- huggingface.co22
- artificialanalysis.ai15
- storage.googleapis.com12
- api.llm-stats.com9
- raw.githubusercontent.com9