VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4.5 vs DeepSeek-V3.2

AnthropicvsDeepSeek83 shared benchmarks4537 head-to-head
BenchmarkClaude Sonnet 4.5DeepSeek-V3.2
AA Agentic Index50.639.8
AA Intelligence37.432.8
AA-LCR68.370.7
AA-Omniscience-0.1-22.5
AIME 20258893.1
AIME25 no tools8889.3
AIME25 w/ python10058.1
arc_agi_163.757
ARC-AGI-213.64
arena_text_factuality14481433
Arena-Hard (Creative Writing)76.788.8
Arena-Hard (Hard Prompt)63.353.4
ArtifactsBench61.555.8
Artificial Analysis Coding Index52.144.2
browsecomp24.151.4
BrowseComp (context management)26.167.6
BrowseComp w/ tools24.140.1
browsecomp_zh42.465
BrowseComp-ZH w/ tools42.447.9
coding_arena_elo13921368
critpt1.12.9
FinSearchComp-global60.826.2
FinSearchComp-T34427
FinSearchComp-T3 w/ tools4427
Frames8580.2
Frames w/ tools8580.2
frontiermath_tier_44.22.1
GAIA (text only)71.263.5
gdpval40.418.8
GPQA (unspecified)83.479.9
GPQA Diamond83.484
HealthBench44.246.9
HealthBench no tools44.246.9
HLE19.840.8
HLE (with tools)3240.8
HMMT 202574.690.2
hmmt_feb_202579.292.5
hmmt_nov_202581.790.2
HMMT25 no tools74.683.6
HMMT25 w/ python88.849.5
IFBench57.361
imo_answer_bench65.978.3
LiveCodeBench7186
LiveCodeBench v66483.3
LiveCodeBenchV6 no tools6474.1
longbench_v261.859.8
Longform Writing eval (Kimi K2 Thinking system card)79.872.5
MCP Atlas59.562.2
mmlu_redux95.693.7
MMLU-Pro88.286
mrcr55.455.5
Multi-SWE-Bench44.337.4
OctoCodingbench22.826
OJ-Bench (cpp)30.454.7
OJ-Bench (cpp) no tools30.438.2
OmniScience Accuracy32.933
OmniScience Non-Hallucination50.917.3
scicode4539
Seal-053.449.5
Seal-0 w/ tools53.438.5
simpleqa_verified23.627.5
swe_bench_bash71.460
SWE-bench Multilingual6870.2
SWE-bench Pro43.615.6
SWE-bench Verified8273.1
SWE-Perf30.9
SWT-bench69.562
TauBench V3 - Banking24.518.8
Terminal-Bench5137.7
Terminal-Bench 2.05046.4
Terminal-Bench 2.155.846.8
Terminal-Bench Hard35.635.6
Terminal-Bench w/ simulated tools (JSON)5137.7
theagentcompany4134
toolathlon4135.2
vectara_answer_rate95.692.6
vectara_avg_summary_length127.862
vectara_factual_consistency8893.7
vectara_hallucination_rate126.3
xbench_deepsearch6671
τ²-Bench9891
τ²-Bench Telecom (AA run)78.190.6
τ³-Bench1969.2

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.