Gemini 3 Pro vs MAI-Thinking-1
Wins 4 of 6 areas
Reasoning · Facts · Long documents · Following instructions
Wins 0 of 6 areas
—
Gemini 3 Pro is the stronger all-rounder.
Scores updated · 14 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Following instructionsDoing exactly what it is asked20Gemini 3 Pro2 of 2 tests
- ReasoningHard problems that need careful thinking10Gemini 3 Pro1 of 1 test
- FactsGetting facts right instead of making them up10Gemini 3 Pro1 of 1 test
- Long documentsFinding answers in very long texts10Gemini 3 Pro1 of 1 test
- CodingWriting and fixing software11Even1 each · 1 tie
- MathCompetition and research-level math11Even1 each
Reasoning, facts and long documents rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 3 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+41.9points ahead
- Keeps track of context across a multi-turn chatMulti-Challenge+12.7points ahead
- Questions about very long textsLongBench v2+7.2points ahead
Where MAI-Thinking-1 pulls ahead
- Long, multi-file coding tasks in real codebasesSWE-bench Pro+9.5points ahead
- US invitational high-school math exam problemsAIME 2026+2.8points ahead
Every test, side by side
All 14 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- SWE-bench ProMAI-Thinking-1 by 9.543.352.8+9.5
- SWE-bench VerifiedGemini 3 Pro by 2.776.273.5+2.7
- LiveCodeBench v6tie87.487.7tie
MathEven
- AIME 2026MAI-Thinking-1 by 2.891.794.5+2.8
- HMMT Feb. 2026Gemini 3 Pro by 1.586.484.9+1.5
Following instructionsGemini 3 Pro
- Multi-ChallengeGemini 3 Pro by 12.765.753+12.7
- IFBenchGemini 3 Pro by 1.470.469+1.4
Other results4 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIR-Bench 2024MAI-Thinking-1 by 14.873.288+14.8
- Terminal-Bench 2.0Gemini 3 Pro by 8.254.246+8.2
- MMLU-ProGemini 3 Pro by 5.190.185+5.1
- AIME 2025Gemini 3 Pro by 310097+3
Questions people ask
Which is better, Gemini 3 Pro or MAI-Thinking-1?
Gemini 3 Pro wins four of the six areas where both have results: reasoning, facts, long documents and following instructions. MAI-Thinking-1 wins none. They are level on coding and math.
Which is better for coding?
Neither. They win 1 coding test each of the 3 both models report, and 1 is a tie.
How do you compare the two?
We use the 14 benchmark tests both models have published scores on. The verdict counts the 10 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 4 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.