AceBench — Published agent/tool-calling benchmark (2025)
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2 Instruct | Moonshot | 76.5 | 4 | 2026-08-23 |
| 2 | Claude Sonnet 4 | Anthropic | 76.2 | 2 | 2026-06-15 |
| 3 | Claude 4 Opus | Anthropic | 75.6 | 2 | 2026-06-15 |
| 4 | DeepSeek-V3 | DeepSeek | 72.7 | 2 | 2026-06-15 |
| 5 | Qwen3 235B A22B | Alibaba | 70.5 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.