VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, DeepSeek-V3 or Qwen3.5 122B A10B?
Across 36 shared benchmarks, DeepSeek-V3 scores higher on 6 and Qwen3.5 122B A10B on 30. The widest gap is HLE, where Qwen3.5 122B A10B scores 47.5 against 5.2. DeepSeek-V3 is the cheaper of the two on tracked API pricing ($0.27 against $0.40 per million input tokens).

DeepSeek-V3 vs Qwen3.5 122B A10B

Across 36 shared benchmarks, DeepSeek-V3 scores higher on 6 and Qwen3.5 122B A10B on 30. The widest gap is HLE, where Qwen3.5 122B A10B scores 47.5 against 5.2. DeepSeek-V3 is the cheaper of the two on tracked API pricing ($0.27 against $0.40 per million input tokens).

DeepSeekvsAlibaba36 shared benchmarks630 head-to-head
BenchmarkDeepSeek-V3Qwen3.5 122B A10B
AA Agentic Index1.621.3
AA Intelligence15.432.8
AA-LCR41.370.3
AA-Omniscience-37.6-41.5
Artificial Analysis Coding Index2345.7
C-Eval90.191.9
critpt00.9
gdpval024.3
GPQA Diamond68.486.6
GSM8K96.794.5
HLE5.247.5
HMMT 202527.590.3
IFBench4176.1
ifeval87.393.4
LiveCodeBench v646.978.9
longbench_v248.760.2
mmlu_prox70.582.2
mmlu_redux90.594
MMLU-Pro81.286.7
mmmlu79.486.7
multichallenge31.461.5
OJBench2439.5
OmniScience Accuracy25.424.4
OmniScience Non-Hallucination23.312.9
scicode35.842
supergpqa53.767.1
SWE-bench Verified4272
TauBench V3 - Banking4.715.3
Terminal-Bench 2.116.947.6
Terminal-Bench Hard15.231.1
vectara_answer_rate97.599.8
vectara_avg_summary_length81.786.4
vectara_factual_consistency93.988.8
vectara_hallucination_rate6.111.2
τ²-Bench Telecom (AA run)47.193.6
τ³-Bench4.713.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.