VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-5.4 vs Kimi K2.6

OpenAIvsMoonshot79 shared benchmarks6118 head-to-head
BenchmarkGPT-5.4Kimi K2.6
AA Agentic Index58.258.7
AA Intelligence53.145.1
AA-LCR77.776.7
AA-Omniscience5.86.4
AgentWorldBench - Android6058.9
AgentWorldBench - MCP70.165.2
AgentWorldBench - OS68.660.8
AgentWorldBench - Overall58.353.4
AgentWorldBench - Search37.327.5
AgentWorldBench - SWE66.358.8
AgentWorldBench - Terminal53.752.5
AgentWorldBench - Web51.850.2
aime_202699.296.4
Apex (Pass@1)54.124
Apex Shortlist (Pass@1)78.175.5
apexAgents33.328.5
arena_elo14771461
arena_text_factuality14761455
arena_vision12771264
Artificial Analysis Coding Index71.161.8
baby_vision49.768.5
baby_vision_with_python80.268.5
browsecomp82.786.3
BrowseComp (Agent Swarm)82.786.3
charxiv_rq82.886.7
charxiv_rq_with_python9086.7
Chinese SimpleQA (C-SimpleQA)76.875.9
coding_arena_elo14571509
critpt23.48
deepsearchqa_accuracy63.783
deepsearchqa_f178.692.5
frontiermath_tier_427.114.6
gdpval50.141.4
GDPval-AA (Elo)16741482
GPQA Diamond9391.1
HiL-Bench9.718.7
HLE43.737.5
HLE (with tools)52.154
hmmt_feb_202697.794.7
IFBench73.976
imo_answer_bench91.486
itbenchSre34.531.2
livebench80.372.2
livebench_agentic_coding53.846.9
livebench_coding77.578.6
livebench_data_analysis79.365.1
livebench_instruction_following70.264.4
livebench_language82.675.1
livebench_math94.284.3
livebench_reasoning88.179.4
mathvision9293.2
mathvision_with_python96.193.2
MCP Atlas70.668.1
MCPAtlas Public (Pass@1)67.266.6
mcpmark62.555.9
MMLU-Pro87.587.1
mmmu_pro_with_python82.180.1
MMMU-Pro81.280.1
OmniScience Accuracy50.932.8
OmniScience Non-Hallucination17.460.7
OSWorld-Verified7573.1
scicode56.653.5
simpleqa_verified45.338.7
SWE-bench Multilingual71.776.7
SWE-bench Pro59.158.6
SWE-QA81.371.6
TauBench V3 - Banking39.623.3
Terminal-Bench 2.075.166.7
Terminal-Bench 2.178.365.9
Terminal-Bench Hard57.643.9
toolathlon54.650
usamo_202695.251.2
V* (w/ python)98.496.9
vectara_answer_rate99.999.7
vectara_avg_summary_length81.7116.7
vectara_factual_consistency9389.2
vectara_hallucination_rate710.8
τ²-Bench Telecom (AA run)87.195.9
τ³-Bench72.920.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.