VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs GPT-5.2

DeepSeekvsOpenAI46 shared benchmarks2323 head-to-head
BenchmarkDeepSeek-V4-FlashGPT-5.2
AA Agentic Index61.360.2
AA Intelligence51.843.3
AA-LCR74.379.3
AA-Omniscience-14.30.3
aime_202695.898.3
arc_agi_18486.2
ARC-AGI-24652.9
arena_elo14391437
arena_text_factuality14341442
Artificial Analysis Coding Index69.148.7
browsecomp73.265.8
BrowseComp (context management)73.270
coding_arena_elo14311418
critpt16.611.6
gdpval52.948.3
GDPval-AA (Elo)13951462
GPQA Diamond90.892.4
HLE38.645.5
HLE (with tools)45.145.5
hmmt_feb_202694.897
IFBench79.275.4
imo_answer_bench88.486.3
livebench_agentic_coding46.850.3
livebench_coding7576.1
livebench_data_analysis79.378.2
livebench_instruction_following65.561.8
livebench_language79.279.8
livebench_math86.893.2
livebench_reasoning86.683.2
LiveCodeBench91.689
MCP Atlas6968
MMLU-Pro86.487
OmniScience Accuracy40.444.3
OmniScience Non-Hallucination15.638.4
scicode49.952.1
simplebench46.345.8
simpleqa_verified4538.9
SWE-bench Multilingual73.372
SWE-bench Pro52.655.6
SWE-bench Verified7980
TauBench V3 - Banking39.411.1
Terminal-Bench 2.056.954
Terminal-Bench Hard38.662.2
toolathlon70.346.3
τ²-Bench Telecom (AA run)95.698.7
τ³-Bench22.912

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.