VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs Gemma 4 26B A4B

AnthropicvsGoogle55 shared benchmarks523 head-to-head
BenchmarkClaude Opus 4.5Gemma 4 26B A4B
AA Agentic Index59.632.1
AA Intelligence41.926.1
AA-LCR7661.7
AA-Omniscience14-50.8
AIME 202592.890
aime_202695.188.3
arena_elo14731438
arena_text_factuality14581445
Artificial Analysis Coding Index47.839.3
browsecomp3726.3
C-Eval92.282.5
CC-OCR76.974.5
charxiv_rq68.569
Claw Eval (pass@3)59.628
Claw-Eval Avg76.658.8
coding_arena_elo14951361
critpt4.60
gdpval47.325.7
GPQA Diamond8782.3
HLE30.819.3
HLE (with tools)43.417.2
hmmt_feb_202592.991.7
hmmt_feb_202685.379
hmmt_nov_202593.387.5
IFBench5872.4
imo_answer_bench8474.3
LiveCodeBench8779.8
LiveCodeBench v684.877.1
mathvision77.182.4
MathVista80.279.4
MCP Atlas69.850
mmlu_redux95.692.7
MMLU-Pro9085.2
mmmlu90.886.3
MMMU80.778.4
MMMU-Pro7473.8
nl2repo43.211.6
OmniDocBench 1.587.774.4
OmniScience Accuracy46.619.1
OmniScience Non-Hallucination3913.6
QwenClawBench52.338.7
QwenWebBench15361178
RealWorldQA7772.2
scicode5040.3
SimpleVQA69.752.2
SkillsBench Avg545.312.3
supergpqa70.661.4
SWE-bench Multilingual77.543.4
SWE-bench Pro57.113.8
SWE-bench Verified81.557.4
Terminal-Bench 2.059.334.2
Terminal-Bench Hard4725
tool_decathlon43.512
τ²-Bench98.268.2
τ²-Bench Telecom (AA run)89.543.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.