VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs GPT-5.2

AnthropicvsOpenAI87 shared benchmarks3552 head-to-head
BenchmarkClaude Opus 4.5GPT-5.2
AA Agentic Index59.660.2
AA Intelligence41.943.3
AA-LCR7679.3
AA-Omniscience140.3
AIME 202592.8100
aime_202695.198.3
AIME25 no tools9198
arc_agi_18286.2
ARC-AGI-237.652.9
arena_elo14731437
arena_text_factuality14581442
Artificial Analysis Coding Index47.848.7
browsecomp3765.8
BrowseComp (context management)59.270
browsecomp_with_context_manager67.865.8
browsecomp_zh62.476.1
charxiv_rq68.582.1
coding_arena_elo14951418
critpt4.611.6
deepsearchqa_f176.171.3
Fortress13.617.5
frontiermath_tier_44.218.8
gdpval47.348.3
GPQA Diamond8792.4
HLE30.845.5
HLE (with tools)43.445.5
hmmt_feb_202592.9100
hmmt_feb_202685.397
hmmt_nov_202593.399.2
IFBench5875.4
imo_answer_bench8486.3
InfoVQA (val)76.984
livebench7674.8
livebench_agentic_coding39.750.3
livebench_coding79.776.1
livebench_data_analysis74.478.2
livebench_instruction_following62.561.8
livebench_language81.379.8
livebench_math90.493.2
livebench_reasoning80.183.2
LiveCodeBench8789
longbench_v264.454.5
LongVideoBench67.276.5
mathvision77.183
MathVista80.282.8
MCP Atlas69.868
metr_hcast73.477
metr_re_bench4035
metr_swaa99.599.6
MMLU-Pro9087
mmmlu90.889.6
MMMU-Pro7479.5
MMVU77.380.8
MotionBench60.364.8
MultiNRC48.642.2
OCRBench86.580.7
OmniDocBench 1.587.785.7
OmniScience Accuracy46.644.3
OmniScience Non-Hallucination3938.4
PaperBench72.963.7
scicode5052.1
SEAL VISTA46.446.6
Seal-047.745
simplebench6245.8
simpleqa_verified41.838.9
SimpleVQA69.755.8
swe_bench_bash76.872.8
SWE-bench Multilingual77.572
SWE-bench Pro57.155.6
SWE-bench Verified81.580
SWE-Perf4.73.6
SWT-bench80.280.7
Tau2 retail88.982
Terminal-Bench 2.059.354
Terminal-Bench Hard4762.2
tool_decathlon43.546.3
toolathlon43.546.3
vectara_answer_rate98.7100
vectara_avg_summary_length114.5186.3
vectara_factual_consistency89.189.2
vectara_hallucination_rate10.910.8
VideoMMMU84.485.9
WorldVQA36.828
ZeroBench39
ZeroBench (w/ tools)97
τ²-Bench98.285.5
τ²-Bench Telecom (AA run)89.598.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.