PopQA — Open-domain factual QA benchmark testing long-tail entity knowledge.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Llama 3.1 8B Instruct | Meta | 28.8 | 1 | 2026-06-11 |
| 2 | Granite 3.1 8B Instruct | IBM | 28.7 | 1 | 2026-06-11 |
| 3 | Granite 3.2 8B Instruct | IBM | 28 | 1 | 2026-06-11 |
| 4 | Granite 3.3 8B Instruct | IBM | 26.2 | 2 | 2026-08-23 |
| 5 | Granite 3.2 2B Instruct | IBM | 20.6 | 1 | 2026-06-11 |
| 6 | Granite 3.1 2B Instruct | IBM | 20.6 | 1 | 2026-06-11 |
| 7 | Granite 3.3 2B Instruct | IBM | 18.4 | 1 | 2026-06-11 |
| 8 | DeepSeek-R1-Distill-Llama-8B | DeepSeek | 13.3 | 1 | 2026-06-11 |
| 9 | DeepSeek-R1-Distill-Qwen-7B | DeepSeek | 9.9 | 1 | 2026-06-11 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.