VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Kimi K2.5

AnthropicvsMoonshot72 shared benchmarks5715 head-to-head
BenchmarkClaude Opus 4.6Kimi K2.5
AA Agentic Index67.652.8
AA Intelligence44.936
AA-LCR74.373
AA-Omniscience13.7-7.3
AIME 202599.896.1
aime_202696.795.8
apexAgents3311.5
arc_agi_19365.3
ARC-AGI-268.811.8
arena_elo14971450
arena_text_factuality14871445
arena_vision13001247
Artificial Analysis Coding Index48.146.8
baby_vision14.836.5
baby_vision_with_python38.440.5
browsecomp86.874.9
BrowseComp (Agent Swarm)83.778.4
browsecomp_with_context_manager8474.9
charxiv_rq69.178.7
charxiv_rq_with_python84.778.7
coding_arena_elo15451436
critpt12.63.1
cybergym73.841.3
deepsearchqa_accuracy80.677.1
deepsearchqa_f191.389
Fortress20.541.1
frontiermath_tier_422.94.2
gdpval55.938.3
GPQA Diamond91.387.9
HLE53.150.2
HLE (with tools)62.751.8
hmmt_feb_202696.287.1
hmmt_nov_202596.391.1
IFBench62.570.2
imo_answer_bench75.381.8
livebench76.369.1
LiveCodeBench v688.885
matharena_visual_math_overall72.380.6
mathvision71.284.2
mathvision_with_python84.685
MCP Atlas76.864
mcpmark56.729.5
MMLU-Pro89.187.1
mmmu_pro_with_python77.377.7
MMMU-Pro77.378.5
multichallenge5661.4
MultiNRC57.135.2
nl2repo49.832
OJBench (python)60.354.7
OmniDocBench 1.586.688.8
OmniScience Accuracy4735.2
OmniScience Non-Hallucination37.250
OSWorld-Verified72.763.3
scicode5249
SEAL VISTA46.141.9
simplebench67.646.8
simpleqa_verified46.536.9
swe_bench_bash75.670.8
SWE-bench Multilingual77.873
SWE-bench Pro57.353.8
SWE-bench Verified80.876.8
Terminal-Bench 2.065.450.8
Terminal-Bench Hard65.434.8
tool_decathlon47.227.8
toolathlon56.827.8
V* (w/ python)86.486.9
vectara_answer_rate99.892.2
vectara_avg_summary_length137.6112
vectara_factual_consistency87.885.8
vectara_hallucination_rate12.214.2
τ²-Bench Telecom (AA run)99.395.9
τ³-Bench72.466

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.