VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Fable 5 vs GPT-5.5

AnthropicvsOpenAI63 shared benchmarks5011 head-to-head
BenchmarkClaude Fable 5GPT-5.5
AA Agentic Index56.647.4
AA Intelligence62.156.3
AA-LCR76.779
AA-Omniscience43.320.5
Agents' Last Exam25.726.6
apex_agents59.238.5
arc_agi_190.595
ARC-AGI-276.885
arena_elo15081482
arena_text_factuality14861482
arena_vision13071283
Artificial Analysis Coding Index76.574.9
browsecomp8884.4
coding_arena_elo16261457
critpt28.627.1
DeepSWE7070
DeepSWE 1.17067
Finance Agent v256.351.8
GDP (Surge AI)29.816.7
gdpval61.949.5
GDPval-AA (Elo)19321769
GDPval-AA v217601509
GDPval-AA v2 (Elo)17471491
GPQA Diamond92.693.6
healthbench_professional6651.8
HiL-Bench56.339.7
HLE64.552.2
HLE (with tools)64.552.2
IFBench63.575.9
JobBench57.438.3
Legal Agent Benchmark13.32.1
Legal Agent Benchmark (Harvey's Held-Out Set)13.32.1
livebench78.380.7
livebench_agentic_coding62.254
livebench_coding8682.2
livebench_data_analysis80.581.6
livebench_instruction_following75.870.7
livebench_language90.787.4
livebench_math9695.9
livebench_reasoning89.789.7
MCP Atlas84.782.8
MCP Mark Verified87.492.9
MLS Bench Lite49.935.5
MMMU-Pro84.283.2
officeqa_pro69.960.9
OmniScience Accuracy65.358
OmniScience Non-Hallucination36.412
OSWorld-Verified8578.7
PostTrainBench41.428.4
Program Bench76.870.8
scicode60.256.1
simplebench81.969
simpleqa_verified68.363.1
SpreadsheetBench 234.729.1
SWE-bench Pro8058.6
SWE-bench Verified9580.6
SWE-Marathon3514
TauBench V3 - Banking38.139
Terminal-Bench 2.18884.3
Terminal-Bench Hard62.960.6
toolathlon61.755.6
τ²-Bench Telecom (AA run)98.593.9
τ³-Bench26.831.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.