VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs Llama 3.1 Instruct 405B

AnthropicvsMeta46 shared benchmarks369 head-to-head
BenchmarkClaude 3.5 SonnetLlama 3.1 Instruct 405B
AA Intelligence108.5
aider_edit_acc84.263.9
aider_polyglot45.35.8
AIME 20241623.3
AIR-Bench 202485.958.6
anthropic_red_team99.896.5
Arena Hard87.669.3
Artificial Analysis Coding Index30.214.5
bbh93.185.9
bbq94.994.5
C-Eval76.772.5
Chinese SimpleQA (C-SimpleQA)56.854.7
CLUEWSC85.484.7
Codeforces71725.3
DROP87.184.8
DROP (3-shot F1)88.388.7
DROP (F1)88.892.5
Fortress1320.6
FRAMES (Acc.)72.570
GPQA Diamond67.251.5
GSM8K96.996.8
harmbench98.162.7
HLE3.94.2
humaneval93.789
HumanEval-Mul (Pass@1)81.777.2
ifeval90.188.6
IFEval (avg)90.186.4
LiveCodeBench32.830.1
livecodebench_pass1cot36.328.4
longbench_v24136.1
MATH81.382.7
math_500_em78.373.8
MBPP+ (EvalPlus-augmented)75.173
mgsm91.691.6
mmlu88.388.6
mmlu_redux88.986.2
MMLU-Pro7873.4
narrativeqa74.674.9
naturalquestions_closedbook50.245.6
OpenBookQA97.294
scicode36.629.9
simple_safety_tests10098.8
simplebench41.423
simpleqa28.423.2
SWE-bench Verified50.824.5
xstest95.695.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.