VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Llama 3.1 Instruct 405B vs o1

MetavsOpenAI37 shared benchmarks433 head-to-head
BenchmarkLlama 3.1 Instruct 405Bo1
AA Agentic Index6.331.1
AA Intelligence8.524
AA-LCR24.763.3
AA-Omniscience-17.1-10.6
aider_polyglot5.861.7
AIME 202423.379.2
AIR-Bench 202458.680
anthropic_red_team96.598.3
Artificial Analysis Coding Index14.539.7
bbq94.597.3
Codeforces25.32061
critpt00.3
DROP (3-shot F1)88.790.2
Fortress20.619.4
gdpval011.5
GPQA Diamond51.575.7
GSM8K96.897.1
harmbench62.796.3
HLE4.27.7
humaneval8988.1
IFBench3970.3
livecodebench_pass1cot28.463.4
MATH82.796.4
math_500_em73.896.4
mgsm91.690.8
mmlu88.691.8
mmmlu73.887.7
OmniScience Accuracy23.234.7
OmniScience Non-Hallucination4930.7
scicode29.935.8
simple_safety_tests98.899
simplebench2341.7
simpleqa23.247
SWE-bench Verified24.548.9
Terminal-Bench Hard6.812.9
xstest95.997
τ²-Bench Telecom (AA run)1962.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.