DeepSeek-V3 vs Llama 3 Instruct 8B
Wins 5 of 5 areas
Coding · Reasoning · Facts · Long documents · Following instructions
Wins 0 of 5 areas
—
DeepSeek-V3 is the stronger all-rounder.Llama 3 Instruct 8B is cheaper.
Scores updated · 36 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20DeepSeek-V32 of 2 tests
- FactsGetting facts right instead of making them up20DeepSeek-V32 of 2 tests
- ReasoningHard problems that need careful thinking10DeepSeek-V31 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts10DeepSeek-V31 of 1 test
- Following instructionsDoing exactly what it is asked10DeepSeek-V31 of 1 test
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Llama 3 Instruct 8B costs 83% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+35.9points ahead
- Code for real scientific research problemsSciCode+27.1points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+16.4points ahead
Where Llama 3 Instruct 8B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 36 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V3
- SciCodeDeepSeek-V3 by 27.13911.9+27.1
- Terminal-Bench HardDeepSeek-V3 by 15.215.20+15.2
ReasoningDeepSeek-V3
- GPQA DiamondDeepSeek-V3 by 35.965.529.6+35.9
- Humanity's Last Examtie4.75.1tie
- CritPttie00tie
FactsDeepSeek-V3
- AA-Omniscience · AccuracyDeepSeek-V3 by 15.125.410.4+15.1
- AA-Omniscience · Non-hallucinationDeepSeek-V3 by 3.914.110.2+3.9
Following instructionsDeepSeek-V3
- IFBenchDeepSeek-V3 by 16.44124.6+16.4
Other results27 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MultiPL-EDeepSeek-V3 by 60.583.122.6+60.5
- MATHDeepSeek-V3 by 51.190.239.1+51.1
- τ²-Bench Telecom (AA run)DeepSeek-V3 by 47.147.10+47.1
- GSM8KDeepSeek-V3 by 44.19449.9+44.1
- MMLU-ProDeepSeek-V3 by 40.575.935.4+40.5
- CMMLUDeepSeek-V3 by 3888.850.8+38
- C-EvalDeepSeek-V3 by 3786.549.5+37
- AIR-Bench 2024Llama 3 Instruct 8B by 30.140.870.9+30.1
- BBHDeepSeek-V3 by 29.887.557.7+29.8
- AA-OmniscienceDeepSeek-V3 by 29.4-40.7-70.1+29.4
- MMLUDeepSeek-V3 by 28.388.560.1+28.3
- HarmBenchLlama 3 Instruct 8B by 2349.772.7+23
- bbqDeepSeek-V3 by 20.296.776.5+20.2
- WinoGrandeDeepSeek-V3 by 19.984.965+19.9
- Artificial Analysis Coding IndexDeepSeek-V3 by 19234+19
- OpenBookQADeepSeek-V3 by 18.895.476.6+18.8
- HellaSwagDeepSeek-V3 by 17.888.971.1+17.8
- ARC-ChallengeDeepSeek-V3 by 12.595.382.8+12.5
- PIQADeepSeek-V3 by 984.775.7+9
- naturalquestions_closedbookDeepSeek-V3 by 8.946.737.8+8.9
- MBPPDeepSeek-V3 by 7.775.467.7+7.7
- AA IntelligenceDeepSeek-V3 by 4.99.74.8+4.9
- HumanEvalDeepSeek-V3 by 4.865.260.4+4.8
- NarrativeQADeepSeek-V3 by 4.279.675.4+4.2
- simple_safety_testsLlama 3 Instruct 8B by 495.399.3+4
- anthropic_red_teamLlama 3 Instruct 8B by 1.797.198.8+1.7
- XSTestDeepSeek-V3 by 1.597.195.6+1.5
Questions people ask
Which is better, DeepSeek-V3 or Llama 3 Instruct 8B?
DeepSeek-V3 wins all five areas where both have results: coding, reasoning, facts, long documents and following instructions. Llama 3 Instruct 8B wins none, but costs 83% less.
Which is better for coding?
DeepSeek-V3. It wins 2 of the 2 coding tests both models report; Llama 3 Instruct 8B wins none.
Which is cheaper?
DeepSeek-V3 costs $0.24 per million input tokens and $0.90 per million output tokens; Llama 3 Instruct 8B costs $0.04 and $0.14. That makes Llama 3 Instruct 8B about 83% cheaper for the same work.
How do you compare the two?
We use the 36 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 27 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.