Across 22 shared benchmarks, Claude 3.5 Sonnet scores higher on 3 and Qwen3.5 122B A10B on 19. The widest gap is OCRBench, where Claude 3.5 Sonnet scores 790 against 92.1. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $3.00 per million input tokens).
| Benchmark | Claude 3.5 Sonnet | Qwen3.5 122B A10B |
|---|---|---|
| AA Intelligence | 10 | 32.8 |
| AI2D | 94.7 | 93.3 |
| arena_vision | 1146 | 1246 |
| Artificial Analysis Coding Index | 30.2 | 45.7 |
| C-Eval | 76.7 | 91.9 |
| GPQA Diamond | 67.2 | 86.6 |
| GSM8K | 96.9 | 94.5 |
| HLE | 3.9 | 47.5 |
| ifeval | 90.1 | 93.4 |
| LiveCodeBench v6 | 37.2 | 78.9 |
| longbench_v2 | 41 | 60.2 |
| MathVista | 67.7 | 87.4 |
| mmlu_redux | 88.9 | 94 |
| MMLU-Pro | 78 | 86.7 |
| MMMU | 72 | 83.9 |
| MMMU-Pro | 54.7 | 76.9 |
| MMStar | 62.2 | 82.9 |
| OCRBench | 790 | 92.1 |
| RealWorldQA | 60.1 | 85.1 |
| scicode | 36.6 | 42 |
| supergpqa | 48.2 | 67.1 |
| SWE-bench Verified | 50.8 | 72 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.