VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4 vs DeepSeek-V4-Flash

AnthropicvsDeepSeek28 shared benchmarks325 head-to-head
BenchmarkClaude Sonnet 4DeepSeek-V4-Flash
AA Agentic Index39.261.3
AA Intelligence29.851.8
AA-LCR69.774.3
AA-Omniscience0.2-14.3
arc_agi_163.784
ARC-AGI-213.646
Artificial Analysis Coding Index37.669.1
browsecomp12.273.2
critpt1.116.6
gdpval31.252.9
GPQA Diamond7890.8
HLE10.738.6
HLE (with tools)20.345.1
IFBench5579.2
LiveCodeBench68.591.6
MMLU-Pro8486.4
OmniScience Accuracy22.740.4
OmniScience Non-Hallucination70.915.6
scicode4049.9
simplebench45.546.3
SWE-bench Multilingual56.973.3
SWE-bench Pro42.752.6
SWE-bench Verified80.279
TauBench V3 - Banking16.739.4
Terminal-Bench 2.136.382.7
Terminal-Bench Hard31.138.6
τ²-Bench Telecom (AA run)64.695.6
τ³-Bench13.822.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.