MCP Atlas (agentic MCP tool-use benchmark), Public Set — DeepSeek-V4/Gemini Public Set subset; not merged w/ base MCP Atlas (ranges do not overlap).
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 73.8 | 4 | 2026-07-04 |
| 2 | DeepSeek-V4-Pro | DeepSeek | 73.6 | 4 | 2026-07-04 |
| 3 | GLM-5.1 | Z.ai | 71.8 | 2 | 2026-06-06 |
| 4 | Gemini 3.1 Pro | 69.2 | 4 | 2026-07-04 | |
| 5 | GPT-5.4 | OpenAI | 67.2 | 4 | 2026-07-04 |
| 6 | Kimi K2.6 | Moonshot | 66.6 | 3 | 2026-06-27 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.