Across 6 models scored on FLTEval pass@1, Claude Opus 4.6 leads at 39.6, ahead of Leanstral 1.5 at 28.9. The median tracked score is 23.7, and the field spans 16.6 to 39.6.
Data as of August 25, 2026| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 39.6 | 1 | 2026-08-25 |
| 2 | Leanstral 1.5 | Mistral | 28.9 | 1 | 2026-07-03 |
| 3 | Qwen3.5 397B A17B | Alibaba | 25.4 | 1 | 2026-07-08 |
| 4 | Claude Sonnet 4.6 | Anthropic | 23.7 | 1 | 2026-08-25 |
| 5 | Claude Haiku 4.5 | Anthropic | 23 | 1 | 2026-08-25 |
| 6 | GLM5-744B-A40B | Z.ai | 16.6 | 1 | 2026-07-10 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.