Across 6 shared benchmarks, Gemini 3.1 Pro scores higher on 1 and laguna-s-2.1 on 5. The widest gap is DeepSWE, where laguna-s-2.1 scores 40.4 against 10. laguna-s-2.1 is the cheaper of the two on tracked API pricing ($0.09 against $2.00 per million input tokens).
| Benchmark | Gemini 3.1 Pro | laguna-s-2.1 |
|---|---|---|
| DeepSWE | 10 | 40.4 |
| DeepSWE 1.1 | 12 | 40.4 |
| SWE-bench Multilingual | 76.9 | 78.5 |
| SWE-bench Pro | 54.2 | 59.4 |
| Terminal-Bench 2.1 | 74 | 70.2 |
| toolathlon | 48.8 | 49.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, poolside), otherwise the lowest tracked offer.