Across 2 shared benchmarks, Claude 3.5 Sonnet scores higher on 2 and LlaVA-Interleave-Qwen-7B on 0. The widest gap is Video-MME Overall, where Claude 3.5 Sonnet scores 55.9 against 50.2.
| Benchmark | Claude 3.5 Sonnet | LlaVA-Interleave-Qwen-7B |
|---|---|---|
| BLINK | 56.5 | 53.1 |
| Video-MME Overall | 55.9 | 50.2 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.