VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Kimi K2.6

AnthropicvsMoonshot79 shared benchmarks5029 head-to-head
BenchmarkClaude Opus 4.6Kimi K2.6
AA Agentic Index67.658.7
AA Intelligence44.945.1
AA-LCR74.376.7
AA-Omniscience13.76.4
AgentWorldBench - Android61.758.9
AgentWorldBench - MCP69.965.2
AgentWorldBench - OS70.260.8
AgentWorldBench - Overall57.853.4
AgentWorldBench - Search29.327.5
AgentWorldBench - SWE64.558.8
AgentWorldBench - Terminal57.552.5
AgentWorldBench - Web51.450.2
aime_202696.796.4
Apex (Pass@1)34.524
Apex Shortlist (Pass@1)85.975.5
apexAgents3328.5
arena_elo14971461
arena_text_factuality14871455
arena_vision13001264
Artificial Analysis Coding Index48.161.8
baby_vision14.868.5
baby_vision_with_python38.468.5
browsecomp86.886.3
BrowseComp (Agent Swarm)83.786.3
charxiv_rq69.186.7
charxiv_rq_with_python84.786.7
Chinese SimpleQA (C-SimpleQA)76.475.9
coding_arena_elo15451509
critpt12.68
deepsearchqa_accuracy80.683
deepsearchqa_f191.392.5
frontiermath_tier_422.914.6
gdpval55.941.4
GDPval-AA (Elo)16191482
GPQA Diamond91.391.1
HiL-Bench38.318.7
HLE53.137.5
HLE (with tools)62.754
hmmt_feb_202696.294.7
IFBench62.576
imo_answer_bench75.386
livebench76.372.2
livebench_agentic_coding4946.9
livebench_coding78.278.6
livebench_data_analysis69.965.1
livebench_instruction_following63.364.4
livebench_language83.375.1
livebench_math89.384.3
livebench_reasoning88.779.4
LiveCodeBench88.889.6
LiveCodeBench v688.889.6
mathvision71.293.2
mathvision_with_python84.693.2
MCP Atlas76.868.1
MCPAtlas Public (Pass@1)73.866.6
mcpmark56.755.9
MMLU-Pro89.187.1
mmmu_pro_with_python77.380.1
MMMU-Pro77.380.1
OJBench (python)60.360.6
OmniScience Accuracy4732.8
OmniScience Non-Hallucination37.260.7
OSWorld-Verified72.773.1
scicode5253.5
simpleqa_verified46.538.7
SWE-bench Multilingual77.876.7
SWE-bench Pro57.358.6
SWE-bench Verified80.880.2
Terminal-Bench 2.065.466.7
Terminal-Bench Hard65.443.9
toolathlon56.850
usamo_202666.251.2
V* (w/ python)86.496.9
vectara_answer_rate99.899.7
vectara_avg_summary_length137.6116.7
vectara_factual_consistency87.889.2
vectara_hallucination_rate12.210.8
τ²-Bench Telecom (AA run)99.395.9
τ³-Bench72.420.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.