Across 6 shared benchmarks, Claude Opus 4.8 scores higher on 6 and laguna-s-2.1 on 0. The widest gap is DeepSWE, where Claude Opus 4.8 scores 59 against 40.4. laguna-s-2.1 is the cheaper of the two on tracked API pricing ($0.09 against $5.00 per million input tokens).
| Benchmark | Claude Opus 4.8 | laguna-s-2.1 |
|---|---|---|
| DeepSWE | 59 | 40.4 |
| DeepSWE 1.1 | 59 | 40.4 |
| SWE-bench Multilingual | 84.4 | 78.5 |
| SWE-bench Pro | 69.2 | 59.4 |
| Terminal-Bench 2.1 | 85 | 70.2 |
| toolathlon | 59.9 | 49.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, poolside), otherwise the lowest tracked offer.