VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V3 vs o4-mini

DeepSeekvsOpenAI36 shared benchmarks728 head-to-head
BenchmarkDeepSeek-V3o4-mini
AA Agentic Index1.636.1
AA Intelligence15.426.1
AA-LCR41.360
AA-Omniscience-37.6-35.7
aider_polyglot55.172
AIME 202551.392.7
AIR-Bench 202440.878.5
anthropic_red_team97.198.2
Artificial Analysis Coding Index2325.6
bbq96.794
critpt00.6
gdpval025.4
GPQA Diamond68.481.4
GSM8K96.70
harmbench49.797
HLE5.216.5
IFBench4168.7
LiveCodeBench49.684.5
livecodebench_easy83.398.8
livecodebench_hard14.262.9
livecodebench_medium53.592.2
mmlu_prox70.569.3
multichallenge31.444.9
OmniScience Accuracy25.424.8
OmniScience Non-Hallucination23.319.5
scicode35.846.5
simple_safety_tests95.3100
simplebench27.238.7
SWE-bench Verified4268.1
Terminal-Bench Hard15.215.2
vectara_answer_rate97.599.2
vectara_avg_summary_length81.7130.9
vectara_factual_consistency93.981.4
vectara_hallucination_rate6.118.6
xstest97.197.4
τ²-Bench Telecom (AA run)47.155.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.