HumanEval-V (visual/vision-language code benchmark) — Distinct multimodal extension of HumanEval for VLMs.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Step3 VL 10B | StepFun | 66 | 2 | 2026-06-05 |
| 2 | MiMo VL RL 2508 (7B) | Xiaomi | 32 | 2 | 2026-06-05 |
| 3 | GLM-4.6V-Flash (9B) | Z.ai | 29.3 | 2 | 2026-06-05 |
| 4 | Qwen3 VL Thinking (8B) | Alibaba | 26.9 | 2 | 2026-06-05 |
| 5 | InternVL-3.5 (8B) | OpenGVLab | 24.3 | 2 | 2026-06-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.