Per-card average column (Granite embedding + Tencent audio cards) — an aggregate, not a standalone benchmark.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | LFM2.5-1.2B-JP-202606 | Liquid AI | 62.5 | 1 | 2026-08-09 |
| 2 | Qwen3 1.7B (instruct) | Alibaba | 51.2 | 1 | 2026-08-09 |
| 3 | Granite 4.0 1B | IBM | 34.1 | 1 | 2026-08-09 |
| 4 | Qwen2-VL-72B-Instruct | Alibaba | 30.9 | 1 | 2026-08-24 |
| 5 | GPT-4o | OpenAI | 27.8 | 1 | 2026-08-24 |
| 6 | Claude 3 Opus | Anthropic | 25.7 | 1 | 2026-08-24 |
| 7 | Gemma 3 1B IT | 24.6 | 1 | 2026-08-09 | |
| 8 | Gemini Ultra | 23.2 | 1 | 2026-08-24 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.