VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4.6 vs Gemini 3.5 Flash

AnthropicvsGoogle53 shared benchmarks2033 head-to-head
BenchmarkClaude Sonnet 4.6Gemini 3.5 Flash
AA Agentic Index61.670.4
AA Intelligence48.452
AA-LCR7481
AA-Omniscience12.221.2
apexAgents2847.1
arc_agi_186.592.5
ARC-AGI-260.472.1
arena_elo14721477
arena_text_factuality14601466
Artificial Analysis Coding Index6370.1
Blueprint-Bench 26.733.6
charxiv_reasoning_with_tools85.384.9
charxiv_rq72.484.2
coding_arena_elo15231490
critpt3.113.1
DeepSWE 1.13037
Finance Agent v25157.9
finance_agent63.357.9
frontiermath_tier_48.314.6
gdpval54.857.8
GDPval-AA (Elo)16761656
GDPval-AA v213811357
GPQA Diamond89.992.2
HLE4942.7
Humanity’s Last Exam33.240.2
IFBench56.676.3
itbenchSre39.840.3
Legal Agent Benchmark5.40.8
Legal Agent Benchmark (Harvey's Held-Out Set)5.40.8
livebench75.575
livebench_agentic_coding42.649
livebench_coding79.378.2
livebench_data_analysis7864.9
livebench_instruction_following63.275.6
livebench_language76.184.6
livebench_math8788.2
livebench_reasoning84.882
MCP Atlas69.583.6
MMMU-Pro75.684.3
mrcr_v2_8needle_128k_average84.977.3
OmniScience Accuracy40.951.4
OmniScience Non-Hallucination51.638.2
OSWorld-Verified78.578.4
scicode4753.1
simpleqa_verified2968.4
SWE-bench Pro58.155.1
TauBench V3 - Banking34.432.2
Terminal-Bench 2.059.176.2
Terminal-Bench 2.171.278.7
Terminal-Bench Hard59.146.2
toolathlon49.456.5
τ²-Bench Telecom (AA run)97.995.6
τ³-Bench30.525.4

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.