VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Gemini 3.1 Pro vs Kimi K2.6

GooglevsMoonshot94 shared benchmarks6130 head-to-head
BenchmarkGemini 3.1 ProKimi K2.6
AA Agentic Index2358.7
AA Intelligence47.745.1
AA-LCR7976.7
AA-Omniscience32.96.4
AgentWorldBench - Android61.458.9
AgentWorldBench - MCP59.165.2
AgentWorldBench - OS66.960.8
AgentWorldBench - Overall54.653.4
AgentWorldBench - Search30.227.5
AgentWorldBench - SWE59.158.8
AgentWorldBench - Terminal52.552.5
AgentWorldBench - Web52.850.2
aime_202698.396.4
Apex (Pass@1)60.924
Apex Shortlist (Pass@1)89.175.5
apex_agents33.527.9
apexAgents33.528.5
arena_elo14861461
arena_text_factuality14711455
arena_vision12801264
Artificial Analysis Coding Index68.861.8
baby_vision51.668.5
baby_vision_with_python68.368.5
browsecomp85.986.3
BrowseComp (Agent Swarm)85.986.3
BrowseComp (w/ Ctx)85.983.2
charxiv_rq89.986.7
charxiv_rq_with_python89.986.7
Chinese SimpleQA (C-SimpleQA)85.975.9
coding_arena_elo14471509
critpt17.78
deepsearchqa_accuracy60.283
deepsearchqa_f181.992.5
Finance Agent v24344.9
FORTRESS (Adversarial)65.265.6
FORTRESS (Benign)9897.2
frontiermath_tier_416.714.6
gdpval23.341.4
GDPval-AA (Elo)13171482
GDPval-AA v29621190
Global-MMLU-Lite92.788.4
GPQA Diamond94.391.1
HiL-Bench35.318.7
HLE51.437.5
HLE (with tools)51.654
hmmt_feb_202694.794.7
IFBench77.176
imo_answer_bench9186
itbenchSre30.331.2
livebench79.972.2
livebench_agentic_coding44.146.9
livebench_coding76.578.6
livebench_data_analysis78.565.1
livebench_instruction_following79.164.4
livebench_language85.475.1
livebench_math9184.3
livebench_reasoning8479.4
LiveCodeBench91.789.6
LiveCodeBench v691.789.6
mathvision89.893.2
mathvision_with_python95.793.2
MCP Atlas78.268.1
MCPAtlas Public (Pass@1)69.266.6
mcpmark55.955.9
MMLU-Pro9187.1
mmmu_pro_with_python85.380.1
MMMU-Pro8380.1
OJBench (python)70.760.6
OmniScience Accuracy55.332.8
OmniScience Non-Hallucination50.160.7
OSWorld-Verified76.273.1
Prophet Arena0.10.1
scicode5953.5
simpleqa_verified77.338.7
strongreject9899.8
SWE-bench Multilingual76.976.7
SWE-bench Pro54.258.6
SWE-bench Verified80.680.2
SWEBench Pro (Public)54.258.6
Tau 3 Banking16.520.6
TauBench V3 - Banking21.423.3
Terminal Bench 2.1 (Best Harness)73.871.3
Terminal-Bench 2.068.566.7
Terminal-Bench 2.17465.9
Terminal-Bench Hard68.543.9
toolathlon48.850
usamo_202674.451.2
V* (w/ python)96.996.9
vectara_answer_rate99.499.7
vectara_avg_summary_length107.7116.7
vectara_factual_consistency89.689.2
vectara_hallucination_rate10.410.8
τ²-Bench Telecom (AA run)99.395.9
τ³-Bench67.120.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.