VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-5.6 Sol vs Kimi K2.5

OpenAIvsMoonshot50 shared benchmarks455 head-to-head
BenchmarkGPT-5.6 SolKimi K2.5
AA Agentic Index57.852.8
AA Intelligence6136
AA-LCR77.773
AA-Omniscience22-7.3
aime_202699.995.8
arc_agi_197.565.3
ARC-AGI-292.511.8
arena_elo14821450
arena_text_factuality14741445
Artificial Analysis Coding Index78.346.8
baby_vision_with_python88.940.5
browsecomp90.474.9
BrowseComp (w/ Ctx)89.474.9
charxiv_rq87.878.7
coding_arena_elo16191436
critpt32.33.1
cybergym83.641.3
FORTRESS (Adversarial)82.454.1
FORTRESS (Benign)98.198.3
gdpval61.138.3
GDPval-AA v217481009
Global-MMLU-Lite91.884
GPQA Diamond94.687.9
HLE49.550.2
HLE (with tools)5851.8
IFBench72.770.2
longbench_v267.161
mathvision95.884.2
MCP Atlas83.664
MMMU-Pro84.678.5
MMVU81.280.4
OmniScience Accuracy59.435.2
OmniScience Non-Hallucination10.650
OSWorld-Verified8363.3
PaperBench90.563.5
ResearchRubrics73.859.5
scicode56.949
simplebench64.846.8
simpleqa_verified71.636.9
strongreject98.599.5
SWE-bench Pro64.653.8
SWEBench Pro (Public)64.650.7
Tau 3 Banking3314.2
TauBench V3 - Banking44.314.2
Terminal Bench 2.1 (Best Harness)89.551.3
Terminal-Bench 2.189.545.7
Terminal-Bench Hard65.934.8
toolathlon5827.8
Video-MME89.587.4
τ²-Bench Telecom (AA run)85.195.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.