VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs GLM-5.2 Full Open Source

DeepSeekvsZ.ai52 shared benchmarks2229 head-to-head
BenchmarkDeepSeek-V4-FlashGLM-5.2 Full Open Source
AA Agentic Index61.345.7
AA Intelligence51.852.6
AA-LCR74.376.7
AA-Omniscience-14.34.4
Agents' Last Exam25.223.8
aime_202695.899.2
arc_agi_18477
ARC-AGI-24622.8
arena_elo14391471
arena_text_factuality14341460
Artificial Analysis Coding Index69.168.8
AutomationBench (Public)25.112.9
AutomationBench Public25.112.9
coding_arena_elo14311582
critpt16.620.9
DeepSWE54.446.2
DSBench-FullStack68.761.8
DSBench-Hard59.654.5
gdpval52.950.3
GDPval-AA v211891514
Global-MMLU-Lite88.489.2
GPQA Diamond90.891.2
HLE38.654.7
HLE (with tools)45.154.7
HLE (wo / w tools)51.554.7
hmmt_feb_202694.892.5
IFBench79.273.3
imo_answer_bench88.491
itbenchSre31.542.7
livebench_agentic_coding46.851.8
livebench_coding7579.7
livebench_data_analysis79.373.7
livebench_instruction_following65.562.3
livebench_language79.276.2
livebench_math86.889.8
livebench_reasoning86.678.6
MCP Atlas6982.6
nl2repo54.248.9
OmniScience Accuracy40.424.3
OmniScience Non-Hallucination15.673.7
scicode49.950.5
simplebench46.358.8
simpleqa_verified4538.1
strongreject97.498.5
SWE-bench Pro52.662.1
Tau 3 Banking22.926.8
TauBench V3 - Banking39.434.6
Terminal-Bench 2.182.782.7
Terminal-Bench Hard38.650.8
toolathlon70.348.2
τ²-Bench Telecom (AA run)95.699.1
τ³-Bench22.926.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.