VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs Qwen3.5 397B A17B

DeepSeekvsAlibaba47 shared benchmarks3313 head-to-head
BenchmarkDeepSeek-V4-FlashQwen3.5 397B A17B
AA Agentic Index61.353.3
AA Intelligence51.834.3
AA-LCR74.372.7
AA-Omniscience-14.3-30.8
aime_202695.894.2
arena_elo14391442
arena_text_factuality14341445
Artificial Analysis Coding Index69.148.2
browsecomp73.269
BrowseComp (context management)73.278.6
BrowseComp (with context management)73.278.6
BrowseComp with context management73.278.6
coding_arena_elo14311399
critpt16.61.7
FORTRESS (adversarial)3277.3
FORTRESS (benign)99.295.4
gdpval52.935.8
GDPval-AA v21189962
Global-MMLU-Lite88.490
GPQA Diamond90.889.3
HLE38.629
HLE (with tools)45.148.3
hmmt_feb_202694.887.9
IFBench79.278.8
imo_answer_bench88.480.9
itbenchSre31.534.1
MMLU-Pro86.487.8
nl2repo54.232.2
OmniScience Accuracy40.430.8
OmniScience Non-Hallucination15.617.3
scicode49.942
simpleqa_verified4526
strongreject97.499.4
SWE-bench Multilingual73.369.3
SWE-bench Pro52.650.9
SWE-bench Verified7976.4
SWEBench Pro (public)52.650.9
SWEBench Pro Public52.650.9
Tau 3 Banking22.913.4
TauBench V3 - Banking39.413.4
Terminal Bench 2.1 (best harness)61.851.3
Terminal-Bench 2.056.952.5
Terminal-Bench 2.182.751.3
Terminal-Bench Hard38.640.9
toolathlon70.338.3
τ²-Bench Telecom (AA run)95.695.6
τ³-Bench22.913.4

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.