VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.8 vs Gemini 3.1 Pro

AnthropicvsGoogle86 shared benchmarks5529 head-to-head
BenchmarkClaude Opus 4.8Gemini 3.1 Pro
AA Agentic Index49.423
AA Intelligence57.347.7
AA-LCR7379
AA-Omniscience28.832.9
AgentWorldBench - Android61.561.4
AgentWorldBench - MCP54.959.1
AgentWorldBench - OS66.666.9
AgentWorldBench - Overall56.654.6
AgentWorldBench - Search35.130.2
AgentWorldBench - SWE64.159.1
AgentWorldBench - Terminal59.252.5
AgentWorldBench - Web54.752.8
aime_202610098.3
AIRS-Bench8483
apex_agents39.433.5
arc_agi_192.598
arc_agi_31.50.4
ARC-AGI-272.177.1
arena_elo14821486
arena_text_factuality14641471
arena_vision12801280
Artificial Analysis Coding Index74.368.8
baby_vision_with_python81.268.3
browsecomp88.585.9
charxiv_reasoning_with_tools89.983.2
charxiv_rq80.589.9
coding_arena_elo15391447
critpt20.917.7
cybergym83.138.8
deepsearchqa_f193.181.9
DeepSWE5910
DeepSWE 1.15912
Finance Agent v253.943
finance_agent53.959.7
Fortress18.229.8
frontiermath_tier_431.316.7
gdpval54.223.3
GDPval-AA (Elo)18901317
GDPval-AA v21593962
GPQA Diamond93.694.3
HiL-Bench35.335.3
HLE57.951.4
HLE (with tools)57.951.6
HLE Calibration26.550.4
hmmt_feb_202696.794.7
hmmt_nov_202596.594.8
IFBench62.277.1
imo_answer_bench83.591
Legal Agent Benchmark100
livebench77.279.9
livebench_agentic_coding50.544.1
livebench_coding81.876.5
livebench_data_analysis6678.5
livebench_instruction_following7279.1
livebench_language79.785.4
livebench_math94.391
livebench_reasoning89.284
matharena_visual_math_overall81.689.4
mathvision86.789.8
MCP Atlas83.678.2
MMMU-Pro78.983
nl2repo69.733.4
officeqa_pro66.242.9
OmniScience Accuracy48.855.3
OmniScience Non-Hallucination60.750.1
OR-Bench (FRR)3.32.5
OSWorld-Verified83.476.2
PostTrainBench37.221.6
Program Bench71.939.5
scicode53.559
simplebench64.879.6
simpleqa_verified39.577.3
SWE-bench Multilingual84.476.9
SWE-bench Pro69.254.2
SWE-bench Verified88.680.6
SWE-Marathon404
Tau 3 Banking27.616.5
TauBench V3 - Banking34.221.4
Terminal-Bench 2.074.668.5
Terminal-Bench 2.18574
Terminal-Bench Hard58.368.5
tool_decathlon59.948.8
toolathlon59.948.8
usamo_202696.774.4
τ²-Bench Telecom (AA run)94.499.3
τ³-Bench27.667.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.