VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs DeepSeek-V3.2

AnthropicvsDeepSeek50 shared benchmarks473 head-to-head
BenchmarkClaude Opus 4.6DeepSeek-V3.2
AA Agentic Index67.639.8
AA Intelligence44.932.8
AA-LCR74.370.7
AA-Omniscience13.7-22.5
AIME 202599.893.1
aime_202696.795.1
AIME25 no tools95.689.3
apexAgents3314.5
arc_agi_19357
ARC-AGI-268.84
arena_text_factuality14871433
Artificial Analysis Coding Index48.144.2
browsecomp86.851.4
browsecomp_with_context_manager8467.6
coding_arena_elo15451368
critpt12.62.9
cybergym73.817.3
deepsearchqa_f191.360.9
frontiermath_tier_422.92.1
gdpval55.918.8
GPQA Diamond91.384
HLE53.140.8
HLE (with tools)62.740.8
hmmt_feb_202696.284.1
hmmt_nov_202596.390.2
IFBench62.561
imo_answer_bench75.378.3
LiveCodeBench88.886
LiveCodeBench v688.883.3
MCP Atlas76.862.2
mcpmark56.738
MMLU-Pro89.186
OmniScience Accuracy4733
OmniScience Non-Hallucination37.217.3
scicode5239
simpleqa_verified46.527.5
swe_bench_bash75.660
SWE-bench Multilingual77.870.2
SWE-bench Pro57.315.6
SWE-bench Verified80.873.1
Terminal-Bench 2.065.446.4
Terminal-Bench Hard65.435.6
tool_decathlon47.235.2
toolathlon56.835.2
vectara_answer_rate99.892.6
vectara_avg_summary_length137.662
vectara_factual_consistency87.893.7
vectara_hallucination_rate12.26.3
τ²-Bench Telecom (AA run)99.390.6
τ³-Bench72.469.2

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.