VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4.6 vs Kimi K3

AnthropicvsMoonshot39 shared benchmarks534 head-to-head
BenchmarkClaude Sonnet 4.6Kimi K3
AA Agentic Index61.654.3
AA Intelligence48.460
AA-LCR7482.7
AA-Omniscience12.246
apexAgents2841.3
arc_agi_186.565.7
ARC-AGI-260.412.4
arena_elo14721489
arena_text_factuality14601471
Artificial Analysis Coding Index6376.2
browsecomp76.291.2
charxiv_rq72.484.8
coding_arena_elo15231674
critpt3.123.4
DeepSWE 1.13069
Finance Agent v25154.4
gdpval54.859.4
GDPval-AA v213811687
GPQA Diamond89.993.5
HLE4956
itbenchSre39.847.7
livebench_agentic_coding42.662.2
livebench_coding79.381.5
livebench_data_analysis7878.7
livebench_instruction_following63.271.4
livebench_language76.185.5
livebench_math8784.4
livebench_reasoning84.890.7
MCP Atlas69.584.2
MMMU-Pro75.683.4
officeqa_pro53.463.3
OmniScience Accuracy40.947.6
OmniScience Non-Hallucination51.649.1
OSWorld-Verified78.584.8
scicode4758.7
simpleqa_verified2942.7
TauBench V3 - Banking34.446
Terminal-Bench 2.171.288.3
toolathlon49.473.2

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.