VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-5.5 vs Opus 5

OpenAIvsAnthropic47 shared benchmarks938 head-to-head
BenchmarkGPT-5.5Opus 5
AA Agentic Index47.459.2
AA Intelligence56.363.1
AA-LCR7978.7
AA-Omniscience20.537.1
arc_agi_19597.5
arc_agi_30.430.2
ARC-AGI-28590.4
arena_elo14821487
arena_text_factuality14821481
Artificial Analysis Coding Index74.978
browsecomp84.490.8
coding_arena_elo14571662
critpt27.129.1
DeepSWE 1.16768.8
Finance Agent v251.858.6
GDP (Surge AI)16.785.5
gdpval49.567.2
GDPval-AA (Elo)17691827
GDPval-AA v215091861
GPQA Diamond93.693.7
HealthBench56.567.1
healthbench_length_adjusted56.557.8
healthbench_professional51.859.8
HiL-Bench39.757
HLE52.264.7
HLE (with tools)52.264.7
Legal Agent Benchmark2.123
livebench_agentic_coding5465.2
livebench_coding82.281.5
livebench_data_analysis81.674.5
livebench_instruction_following70.763.8
livebench_language87.488.7
livebench_math95.995.7
livebench_reasoning89.791.2
MCP Atlas82.885.8
MMMU-Pro83.284.7
officeqa_pro60.966.9
OmniScience Accuracy5860.9
OmniScience Non-Hallucination1240.5
OSWorld-Verified78.74
scicode56.155.7
simplebench6980.6
simpleqa_verified63.156.7
SWE-bench Pro58.679.2
SWE-bench Verified80.696
TauBench V3 - Banking3944.7
Terminal-Bench 2.184.389.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.