Across 14 shared benchmarks, Claude 3.5 Haiku scores higher on 0 and Quasar 438B (max, based on GLM-5.2) on 14. The widest gap is Terminal-Bench 2.1, where Quasar 438B (max, based on GLM-5.2) scores 69.3 against 10.1.
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.