VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.6 vs Gemini 3.1 Pro

AnthropicvsGoogle101 shared benchmarks4357 head-to-head
BenchmarkClaude Opus 4.6Gemini 3.1 Pro
AA Agentic Index67.623
AA Intelligence44.947.7
AA-LCR74.379
AA-Omniscience13.732.9
AgentWorldBench - Android61.761.4
AgentWorldBench - MCP69.959.1
AgentWorldBench - OS70.266.9
AgentWorldBench - Overall57.854.6
AgentWorldBench - Search29.330.2
AgentWorldBench - SWE64.559.1
AgentWorldBench - Terminal57.552.5
AgentWorldBench - Web51.452.8
aime_202696.798.3
Apex (Pass@1)34.560.9
Apex Shortlist (Pass@1)85.989.1
apexAgents3333.5
arc_agi_19398
arc_agi_30.50.4
ARC-AGI-268.877.1
arena_elo14971486
arena_text_factuality14871471
arena_vision13001280
Artificial Analysis Coding Index48.168.8
baby_vision14.851.6
baby_vision_with_python38.468.3
browsecomp86.885.9
BrowseComp (Agent Swarm)83.785.9
browsecomp_with_context_manager8485.9
charxiv_reasoning_with_tools84.783.2
charxiv_rq69.189.9
charxiv_rq_with_python84.789.9
Chinese SimpleQA (C-SimpleQA)76.485.9
coding_arena_elo15451447
CorpusQA 1M (ACC)71.753.8
critpt12.617.7
cybergym73.838.8
deepsearchqa_accuracy80.660.2
deepsearchqa_f191.381.9
finance_agent76.759.7
Fortress20.529.8
frontiermath_tier_422.916.7
gdpval55.923.3
GDPval-AA (Elo)16191317
GPQA Diamond91.394.3
HiL-Bench38.335.3
HLE53.151.4
HLE (with tools)62.751.6
hmmt_feb_202696.294.7
hmmt_nov_202596.394.8
IFBench62.577.1
imo_answer_bench75.391
Legal Agent Benchmark4.20
livebench76.379.9
livebench_agentic_coding4944.1
livebench_coding78.276.5
livebench_data_analysis69.978.5
livebench_instruction_following63.379.1
livebench_language83.385.4
livebench_math89.391
livebench_reasoning88.784
LiveCodeBench88.891.7
LiveCodeBench v688.891.7
matharena_visual_math_overall72.389.4
mathvision71.289.8
mathvision_with_python84.695.7
MCP Atlas76.878.2
MCPAtlas Public (Pass@1)73.869.2
mcpmark56.755.9
MMLU-Pro89.191
mmmlu91.192.6
mmmu_pro_with_python77.385.3
MMMU-Pro77.383
MRCR 1M (MMR)92.976.3
mrcr_v2_8needle_128k_average8484.9
multichallenge5671.4
MultiNRC57.164.7
nl2repo49.833.4
officeqa_pro57.142.9
OJBench (python)60.370.7
OmniScience Accuracy4755.3
OmniScience Non-Hallucination37.250.1
OSWorld-Verified72.776.2
scicode5259
simplebench67.679.6
simpleqa_verified46.577.3
SWE-bench Multilingual77.876.9
SWE-bench Pro57.354.2
SWE-bench Verified80.880.6
Terminal-Bench 2.065.468.5
Terminal-Bench Hard65.468.5
tool_decathlon47.248.8
toolathlon56.848.8
usamo_202666.274.4
V* (w/ python)86.496.9
vectara_answer_rate99.899.4
vectara_avg_summary_length137.6107.7
vectara_factual_consistency87.889.6
vectara_hallucination_rate12.210.4
VisualToolBench27.529
τ²-Bench Telecom (AA run)99.399.3
τ³-Bench72.467.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.