CorpusQA (1M context) — Long-context (1M token) corpus QA benchmark from DeepSeek-V4 cards
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 71.7 | 4 | 2026-07-04 |
| 2 | DeepSeek-V4-Pro | DeepSeek | 62 | 9 | 2026-06-27 |
| 3 | DeepSeek-V4-Flash | DeepSeek | 60.5 | 4 | 2026-06-27 |
| 4 | Gemini 3.1 Pro | 53.8 | 4 | 2026-07-04 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.