BFCL v2 (tool-use category) — Berkeley Function-Calling Leaderboard v2, tool-use subset; Meta Llama card
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Llama 3.1 Instruct 405B | Meta | 81.1 | 1 | 2026-05-01 |
| 2 | Llama 3.1 70B Instruct | Meta | 77.5 | 1 | 2026-05-01 |
| 3 | Llama 3.3 70B Instruct | Meta | 77.3 | 2 | 2026-08-23 |
| 4 | Llama 3.1 Nemotron Ultra 253B V1 | NVIDIA | 74.1 | 1 | 2026-08-23 |
| 5 | Qwen3.5 2B | Alibaba | 68.9 | 1 | 2026-08-15 |
| 6 | Llama 3.2 3B Instruct | Meta | 67 | 1 | 2026-08-23 |
| 7 | Llama 3.1 8B Instruct | Meta | 65.4 | 1 | 2026-05-01 |
| 8 | Qwen3.5 0.8B | Alibaba | 64 | 1 | 2026-08-15 |
| 9 | Llama 3.1 Nemotron Nano 8B V1 | NVIDIA | 63.6 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.