VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs GPT-4o mini

AnthropicvsOpenAI38 shared benchmarks2810 head-to-head
BenchmarkClaude 3.5 SonnetGPT-4o mini
AA Intelligence107
AI2D94.775.2
aider_polyglot45.33.6
AIR-Bench 202485.956.3
anthropic_red_team99.898.3
Artificial Analysis Coding Index30.211.4
bbq94.988.2
BLINK56.551.9
DROP87.179.7
Fortress1348.1
GPQA Diamond67.242.6
GSM8K96.984.3
harmbench98.184.9
HLE3.94.2
humaneval93.787.2
InterGPS (test)45.639.9
LiveCodeBench32.835.5
MATH81.380.2
MathVista67.756.7
mgsm91.687
MMBench (dev-en)82.383.8
mmlu88.381.8
MMMU7259.4
MMMU (val) (Pass@1)68.352.1
MMMU-Pro54.741.5
narrativeqa74.676.8
naturalquestions_closedbook50.238.5
OpenBookQA97.292
POPE (test)76.683.6
scicode36.622.9
ScienceQA (img-test)73.884
simple_safety_tests10097.8
simplebench41.410.7
simpleqa28.49.9
SWE-bench Verified50.88.7
TextVQA (val)70.570.9
Video-MME Overall55.961.2
xstest95.696

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.