VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Fable 5 vs GPT-5.6 Sol

AnthropicvsOpenAI79 shared benchmarks4831 head-to-head
BenchmarkClaude Fable 5GPT-5.6 Sol
AA Agentic Index56.657.8
AA Intelligence62.161
AA-LCR76.777.7
AA-Omniscience43.322
Agents' Last Exam25.752.7
AndroidBench84.574
apex_agents59.256.7
arc_agi_190.597.5
ARC-AGI-276.892.5
arena_elo15081482
arena_text_factuality14861474
Artificial Analysis Coding Index76.578.3
baby_vision_with_python90.588.9
browsecomp8890.4
BrowseComp (w/ Ctx)8889.4
charxiv_rq89.487.8
coding_arena_elo16261619
CorpFin v271.864.4
critpt28.632.3
CursorBench v3.270.567.2
DeepSWE7073
DeepSWE 1.17073
Finance Agent v256.353.8
FORTRESS (Adversarial)9682.4
FORTRESS (Benign)55.198.1
FrontierCode v1.1 (Extended)63.660.6
GDP (Surge AI)29.830.7
gdpval61.961.1
GDPval-AA v217601748
GDPval-AA v2 (Elo)17471736
Global-MMLU-Lite93.391.8
GPQA Diamond92.694.6
Harvey LAB (Vals)11.32.5
healthbench_professional6660.5
HiL-Bench56.332.3
HLE64.549.5
HLE (with tools)64.558
IFBench63.572.7
JobBench57.445.4
Legal Research Bench49.548.1
livebench_agentic_coding62.256.2
livebench_coding8683.9
livebench_data_analysis80.579.8
livebench_instruction_following75.871.8
livebench_language90.787.7
livebench_math9696.2
livebench_reasoning89.791.7
mathvision94.895.8
MCP Atlas84.783.6
MCP Mark Verified87.492.9
MLS Bench Lite49.946.2
MMMU-Pro84.284.6
officeqa_pro69.963.2
OmniDocBench89.885.8
OmniScience Accuracy65.359.4
OmniScience Non-Hallucination36.410.6
OSWorld 2.066.162.6
OSWorld-Verified8583
PaperBench88.890.5
PerceptionBench57.259.7
PostTrainBench41.434.6
Program Bench76.877.6
scicode60.256.9
simplebench81.964.8
simpleqa_verified68.371.6
SpreadsheetBench 234.732.4
strongreject98.798.5
SWE-bench Pro8064.6
SWE-Marathon3539
SWEBench Pro (Public)8064.6
Tau 3 Banking26.833
TauBench V3 - Banking38.144.3
Terminal Bench 2.1 (Best Harness)84.689.5
Terminal-Bench 2.18889.5
Terminal-bench 3.034.134.6
Terminal-Bench Hard62.965.9
toolathlon61.758
WorldVQA ForceAnswer56.741.8
τ²-Bench Telecom (AA run)98.585.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.