VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs Qwen2.5 Instruct 72B

AnthropicvsAlibaba57 shared benchmarks4114 head-to-head
BenchmarkClaude 3.5 SonnetQwen2.5 Instruct 72B
AA Intelligence1010
aider_edit_acc84.265.4
aider_polyglot45.37.6
AIME 20241623.3
AIR-Bench 202485.959
anthropic_red_team99.899.6
Arena Hard87.681.2
Artificial Analysis Coding Index30.211.9
bbh93.179.8
bbq94.995.4
C-Eval76.789.2
Chinese SimpleQA (C-SimpleQA)56.852.2
CLUEWSC85.491.4
Codeforces71724.8
DROP87.176.7
DROP (3-shot F1)88.376.7
DROP (F1)88.885
Fortress1356.4
FRAMES (Acc.)72.569.8
GPQA Diamond67.249.1
GSM8K96.995.8
harmbench98.172.8
HLE3.94.2
humaneval93.786.6
HumanEval-Mul (Pass@1)81.777.3
ifeval90.187.2
IFEval (avg)90.187.2
LiveCodeBench32.855.5
livecodebench_pass1cot36.331.1
LongBench v2 easy (w/ CoT)55.247.9
LongBench v2 easy (w/o CoT)46.942.7
LongBench v2 hard (w/ CoT)41.540.8
LongBench v2 hard (w/o CoT)37.341.8
LongBench v2 long (w/ CoT)44.439.8
LongBench v2 long (w/o CoT)3744.4
LongBench v2 medium (w/ CoT)41.940.9
LongBench v2 medium (w/o CoT)38.638.1
LongBench v2 overall (w/ CoT)46.743.5
LongBench v2 overall (w/o CoT)4142.1
LongBench v2 short (w/ CoT)53.948.9
LongBench v2 short (w/o CoT)46.145.6
longbench_v24139.4
MATH81.388.4
math_500_em78.380
MBPP+ (EvalPlus-augmented)75.177
mgsm91.687.3
mmlu88.386.1
mmlu_redux88.986.8
MMLU-Pro7871.6
narrativeqa74.674.5
naturalquestions_closedbook50.235.9
OpenBookQA97.296.2
scicode36.626.7
simple_safety_tests100100
simpleqa28.410.3
SWE-bench Verified50.823.8
xstest95.697.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.