VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Kimi K2.5 or Qwen3.5 122B A10B?
Across 47 shared benchmarks, Kimi K2.5 scores higher on 36 and Qwen3.5 122B A10B on 11. The widest gap is τ³-Bench, where Kimi K2.5 scores 66 against 13.6. Tracked API pricing per million tokens: Kimi K2.5 $0.56 in / $2.94 out, Qwen3.5 122B A10B $0.40 in / $3.20 out.

Kimi K2.5 vs Qwen3.5 122B A10B

Across 47 shared benchmarks, Kimi K2.5 scores higher on 36 and Qwen3.5 122B A10B on 11. The widest gap is τ³-Bench, where Kimi K2.5 scores 66 against 13.6. Tracked API pricing per million tokens: Kimi K2.5 $0.56 in / $2.94 out, Qwen3.5 122B A10B $0.40 in / $3.20 out.

MoonshotvsAlibaba47 shared benchmarks3611 head-to-head
BenchmarkKimi K2.5Qwen3.5 122B A10B
AA Agentic Index52.821.3
AA Intelligence3632.8
AA-LCR7370.3
AA-Omniscience-7.3-41.5
arena_vision12671246
Artificial Analysis Coding Index46.845.7
baby_vision36.540.2
browsecomp74.963.8
browsecomp_zh62.369.9
coding_arena_elo14361358
critpt3.10.9
gdpval38.324.3
GPQA Diamond87.986.6
HLE50.247.5
HMMT 202595.490.3
IFBench70.276.1
LiveCodeBench v68578.9
longbench_v26160.2
LVBench75.974.4
mathvision84.286.2
MathVista90.187.4
MMLU-Pro87.186.7
MMMU-Pro78.576.9
MMVU80.474.7
multichallenge61.461.5
OCRBench92.392.1
OmniDocBench 1.588.889.8
OmniScience Accuracy35.224.4
OmniScience Non-Hallucination5012.9
OSWorld-Verified63.358
scicode4942
Seal-057.444.1
SimpleVQA71.20.6
SWE-bench Verified76.872
TauBench V3 - Banking14.215.3
Terminal-Bench 2.050.849.4
Terminal-Bench 2.145.747.6
Terminal-Bench Hard34.831.1
vectara_answer_rate92.299.8
vectara_avg_summary_length11286.4
vectara_factual_consistency85.888.8
vectara_hallucination_rate14.211.2
VideoMMMU86.682
WideSearch7960.5
ZeroBench90.1
τ²-Bench Telecom (AA run)95.993.6
τ³-Bench6613.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (moonshot-official, alibaba-official), otherwise the lowest tracked offer.