VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-5.6 Sol vs Kimi K2.6

OpenAIvsMoonshot52 shared benchmarks475 head-to-head
BenchmarkGPT-5.6 SolKimi K2.6
AA Agentic Index57.858.7
AA Intelligence6145.1
AA-LCR77.776.7
AA-Omniscience226.4
aime_202699.996.4
apex_agents56.727.9
arena_elo14821461
arena_text_factuality14741455
Artificial Analysis Coding Index78.361.8
baby_vision_with_python88.968.5
browsecomp90.486.3
BrowseComp (w/ Ctx)89.483.2
charxiv_rq87.886.7
coding_arena_elo16191509
critpt32.38
Finance Agent v253.844.9
FORTRESS (Adversarial)82.465.6
FORTRESS (Benign)98.197.2
gdpval61.141.4
GDPval-AA v217481190
Global-MMLU-Lite91.888.4
GPQA Diamond94.691.1
HiL-Bench32.318.7
HLE49.537.5
HLE (with tools)5854
IFBench72.776
itbenchSre56.231.2
livebench_agentic_coding56.246.9
livebench_coding83.978.6
livebench_data_analysis79.865.1
livebench_instruction_following71.864.4
livebench_language87.775.1
livebench_math96.284.3
livebench_reasoning91.779.4
mathvision95.893.2
MCP Atlas83.668.1
MMMU-Pro84.680.1
OmniScience Accuracy59.432.8
OmniScience Non-Hallucination10.660.7
OSWorld-Verified8373.1
scicode56.953.5
simpleqa_verified71.638.7
strongreject98.599.8
SWE-bench Pro64.658.6
SWEBench Pro (Public)64.658.6
Tau 3 Banking3320.6
TauBench V3 - Banking44.323.3
Terminal Bench 2.1 (Best Harness)89.571.3
Terminal-Bench 2.189.565.9
Terminal-Bench Hard65.943.9
toolathlon5850
τ²-Bench Telecom (AA run)85.195.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.