Llama 3.3 70B Instruct
Llama 3.3 70B Instruct is behind the leaders in factuality and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multimodal or Multilingual.
Price
$0.71input$0.72outputper million tokens
From Artificial Analysis · 7 providers tracked · All prices
Evidence
52results on44benchmarks
- 14 independently verified
- 18 aggregator
- 13 vendor-reported
- 7 cross-referenced
From 14 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Llama 3.3 70B Instruct benchmark results
52 results on 44 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
36.0% behind the leader3 of 4 ranked benchmarks measured
- 4.10May 2, 2026
- AA-Omniscience · Accuracy18.95Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.79Oct 8, 2026omniscienceNonHallucination
40.2% behind the leader1 of 3 ranked benchmarks measured
- IFBench47.07Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- LiveBench · Instruction Following78.20Aug 23, 2026livebench_instruction_following@2025-04-07
0 of 6 ranked benchmarks measured
- 19.90May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond49.80Oct 8, 2026gpqa
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond50.50Sep 3, 2026GPQA Diamond (CoT)
- GPQA Diamond49.10May 30, 2026GPQA
- Humanity's Last Exam3.56Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- LiveBench · Coding36.72Aug 23, 2026livebench_coding@2025-04-07
- Terminal-Bench 2.14.87Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard3.03Oct 8, 2026aa_terminalbench_hard
Show 1 more coding resultHide 1 coding result
- 26.04Jul 19, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- 0.56Jul 19, 2026
- τ-Bench V3 · Banking1.03Jul 19, 2026tauBanking
0 of 3 ranked benchmarks measured
- 15.67Oct 8, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language42.63Aug 23, 2026livebench_language@2025-04-07
- 92.80Jul 2, 2026
- 94.20Jul 2, 2026
- 43.07May 10, 2026
- 79.13May 10, 2026
- 69.99May 10, 2026
Show 27 more resultsHide 27 results
- 0.34Jul 19, 2026
- AA Intelligence7.66Oct 8, 2026aa_intelligence_index
- -54.17Oct 8, 2026
- Artificial Analysis Coding Index11.93Sep 9, 2026aa_coding_index
- 77.30Oct 7, 2026
- 90.20May 30, 2026
- 88.40Oct 7, 2026
- 78.90May 30, 2026
- 92.10Aug 31, 2026
- LiveCodeBench33.30May 1, 2026code_livecodebench_1001202402012025
- 80.79May 10, 2026
- 77.00May 1, 2026
- 66.30May 30, 2026
- 77.00Jul 6, 2026
- 87.60May 1, 2026
- 91.10Oct 7, 2026
- 89.10May 30, 2026
- 86.00May 1, 2026
- 86.30May 30, 2026
- 86.00Jul 6, 2026
- 68.90Jul 6, 2026
- 68.90Oct 7, 2026
- 20.90May 30, 2026
- vectara_answer_rate99.50May 2, 2026Answer Rate
- vectara_avg_summary_length64.60May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency95.90May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)26.61Oct 8, 2026aa_tau2
Llama 3.3 70B Instruct: common questions
Who makes Llama 3.3 70B Instruct?
Llama 3.3 70B Instruct is made by Meta.
When was Llama 3.3 70B Instruct released?
Llama 3.3 70B Instruct was released on Dec 6, 2024, according to Artificial Analysis.
What is Llama 3.3 70B Instruct good at?
Llama 3.3 70B Instruct is behind the leaders in factuality and instruction following. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multimodal tasks, or multilingual tasks.
How much does Llama 3.3 70B Instruct cost?
Llama 3.3 70B Instruct costs $0.71 per million input tokens and $0.72 per million output tokens, according to Artificial Analysis. We track its price at 7 providers. At a mix of three input tokens to one output token, it costs more than 50% of the 331 priced models we track.
How many benchmarks has Llama 3.3 70B Instruct been tested on?
We track 52 results for Llama 3.3 70B Instruct on 44 benchmarks from 14 sources, 14 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Llama 3.3 70B Instruct support?
OpenRouter lists tool calling and structured outputs for Llama 3.3 70B Instruct.
About this record
Where Llama 3.3 70B Instruct's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 1, 2026
Where the results come from
Verification: 52 scores · 14 independently verified · 18 aggregator-attributed · 7 vendor cross-reference · 13 vendor-reported. How these tiers are assigned
From 14 sources on 6 sites. Artificial Analysis supplies 18 of them; the 14 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai18
- raw.githubusercontent.com12
- huggingface.co10
- storage.googleapis.com6
- api.llm-stats.com5
- simple-bench.com1