Across 2 shared benchmarks, GPT-4o mini scores higher on 1 and LlaVA-Interleave-Qwen-7B on 1. The widest gap is Video-MME Overall, where GPT-4o mini scores 61.2 against 50.2.
| Benchmark | GPT-4o mini | LlaVA-Interleave-Qwen-7B |
|---|---|---|
| BLINK | 51.9 | 53.1 |
| Video-MME Overall | 61.2 | 50.2 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.