VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Llama 3.1 Instruct 405B or Qwen3.5 122B A10B?
Across 23 shared benchmarks, Llama 3.1 Instruct 405B scores higher on 3 and Qwen3.5 122B A10B on 20. The widest gap is τ²-Bench Telecom (AA run), where Qwen3.5 122B A10B scores 93.6 against 19.

Llama 3.1 Instruct 405B vs Qwen3.5 122B A10B

Across 23 shared benchmarks, Llama 3.1 Instruct 405B scores higher on 3 and Qwen3.5 122B A10B on 20. The widest gap is τ²-Bench Telecom (AA run), where Qwen3.5 122B A10B scores 93.6 against 19.

MetavsAlibaba23 shared benchmarks320 head-to-head
BenchmarkLlama 3.1 Instruct 405BQwen3.5 122B A10B
AA Agentic Index6.321.3
AA Intelligence8.532.8
AA-LCR24.770.3
AA-Omniscience-17.1-41.5
Artificial Analysis Coding Index14.545.7
C-Eval72.591.9
critpt00.9
gdpval024.3
GPQA Diamond51.586.6
GSM8K96.894.5
HLE4.247.5
IFBench3976.1
ifeval88.693.4
longbench_v236.160.2
mmlu_redux86.294
MMLU-Pro73.486.7
mmmlu73.886.7
OmniScience Accuracy23.224.4
OmniScience Non-Hallucination4912.9
scicode29.942
SWE-bench Verified24.572
Terminal-Bench Hard6.831.1
τ²-Bench Telecom (AA run)1993.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.