VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs GPT-5.4

DeepSeekvsOpenAI50 shared benchmarks1238 head-to-head
BenchmarkDeepSeek-V4-FlashGPT-5.4
AA Agentic Index61.358.2
AA Intelligence51.853.1
AA-LCR74.377.7
AA-Omniscience-14.35.8
aime_202695.899.2
Apex (Pass@1)3354.1
Apex Shortlist (Pass@1)85.778.1
arc_agi_18493.7
ARC-AGI-24674
arena_elo14391477
arena_text_factuality14341476
Artificial Analysis Coding Index69.171.1
browsecomp73.282.7
Chinese SimpleQA (C-SimpleQA)78.976.8
Codeforces30523168
coding_arena_elo14311457
critpt16.623.4
cybergym76.766.3
gdpval52.950.1
GDPval-AA (Elo)13951674
GPQA Diamond90.893
HLE38.643.7
HLE (with tools)45.152.1
hmmt_feb_202694.897.7
IFBench79.273.9
imo_answer_bench88.491.4
itbenchSre31.534.5
livebench_agentic_coding46.853.8
livebench_coding7577.5
livebench_data_analysis79.379.3
livebench_instruction_following65.570.2
livebench_language79.282.6
livebench_math86.894.2
livebench_reasoning86.688.1
MCP Atlas6970.6
MMLU-Pro86.487.5
nl2repo54.241.3
OmniScience Accuracy40.450.9
OmniScience Non-Hallucination15.617.4
scicode49.956.6
simpleqa_verified4545.3
SWE-bench Multilingual73.371.7
SWE-bench Pro52.659.1
TauBench V3 - Banking39.439.6
Terminal-Bench 2.056.975.1
Terminal-Bench 2.182.778.3
Terminal-Bench Hard38.657.6
toolathlon70.354.6
τ²-Bench Telecom (AA run)95.687.1
τ³-Bench22.972.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.