VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Gemini 3.1 Pro vs Gemini 3.5 Flash

GooglevsGoogle62 shared benchmarks3131 head-to-head
BenchmarkGemini 3.1 ProGemini 3.5 Flash
AA Agentic Index2370.4
AA Intelligence47.752
AA-LCR7981
AA-Omniscience32.921.2
aime_202698.395
apexAgents33.547.1
arc_agi_19892.5
ARC-AGI-277.172.1
arena_elo14861477
arena_text_factuality14711466
Artificial Analysis Coding Index68.870.1
Blueprint-Bench 226.533.6
charxiv_reasoning_with_tools83.284.9
charxiv_rq89.984.2
coding_arena_elo14471490
critpt17.713.1
DeepSWE1037
DeepSWE 1.11237
Finance Agent v24357.9
finance_agent59.757.9
frontiermath_tier_416.714.6
GDM-MRCR v2 (8-needle) (1M (pointwise))26.326.6
gdpval23.357.8
GDPval-AA (Elo)13171656
GDPval-AA v29621357
GDPVal-AA v2 (Elo)9651349
GPQA Diamond94.392.2
HiL-Bench35.327.7
HLE51.442.7
hmmt_feb_202694.795.5
Humanity’s Last Exam44.440.2
IFBench77.176.3
itbenchSre30.340.3
Legal Agent Benchmark00.8
livebench79.975
livebench_agentic_coding44.149
livebench_coding76.578.2
livebench_data_analysis78.564.9
livebench_instruction_following79.175.6
livebench_language85.484.6
livebench_math9188.2
livebench_reasoning8482
matharena_visual_math_overall89.489.9
MCP Atlas78.283.6
MLE-Bench42.649.7
MMMU-Pro8384.3
mrcr_v2_8needle_128k_average84.977.3
mrcr_v2_8needle_1m_pointwise26.322.1
OmniScience Accuracy55.351.4
OmniScience Non-Hallucination50.138.2
OSWorld-Verified76.278.4
scicode5953.1
simplebench79.676.7
simpleqa_verified77.368.4
SWE-bench Pro54.255.1
TauBench V3 - Banking21.432.2
Terminal-Bench 2.068.576.2
Terminal-Bench 2.17478.7
Terminal-Bench Hard68.546.2
toolathlon48.856.5
τ²-Bench Telecom (AA run)99.395.6
τ³-Bench67.125.4

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.