The two are evenly matched.GPT-4o is better at images and charts; Kimi VL A3B Instruct at agents.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Agents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
Where Kimi VL A3B Instruct pulls ahead
- Math problems shown in pictures and chartsMathVista+7.3points ahead
- Completes tasks by operating a computer desktopOSWorld-Verified+3.2points ahead
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsGPT-4o
Other results21 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ScreenSpot-V2Kimi VL A3B Instruct by 74.718.192.8+74.7
- MMVU-Val (Pass@1)GPT-4o by 15.267.452.2+15.2
- BLINK (Acc)GPT-4o by 10.76857.3+10.7
- MLVU-MCQ (Acc)Kimi VL A3B Instruct by 9.664.674.2+9.6
- MATH-VisionGPT-4o by 930.421.4+9
- MMLongBench-DOC (Acc)GPT-4o by 7.742.835.1+7.7
- RealWorldQA (Acc)GPT-4o by 7.375.468.1+7.3
- EgoSchema (full)Kimi VL A3B Instruct by 6.372.278.5+6.3
- TOMATOGPT-4o by 637.731.7+6
- OCRBench (Acc)Kimi VL A3B Instruct by 52 rating points815867+52 rating
- VideoMME (w sub.)GPT-4o by 4.677.272.6+4.6
- Video-MME (w/o sub.)GPT-4o by 4.171.967.8+4.1
- MMStar (Acc)GPT-4o by 3.464.761.3+3.4
- VSI-BenchKimi VL A3B Instruct by 3.43437.4+3.4
- InfoVQA (Acc)Kimi VL A3B Instruct by 2.580.783.2+2.5
- MMVet (Acc)GPT-4o by 2.469.166.7+2.4
- MMVet (Pass@1)GPT-4o by 2.469.166.7+2.4
- LongVideoBench (val)GPT-4o by 2.266.764.5+2.2
- WindowsAgentArena (Pass@1)Kimi VL A3B Instruct by 19.410.4+1
- AI2D (Acc)tie84.684.9tie
- MMBench-EN-v1.1 (Acc)tie83.183.1tie
Questions people ask
Which is better, GPT-4o or Kimi VL A3B Instruct?
GPT-4o and Kimi VL A3B Instruct each win one of the two areas where both have results. GPT-4o wins images and charts; Kimi VL A3B Instruct wins agents.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 21 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.