VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, GPT-5.1 or Qwen3.5 122B A10B?
Across 33 shared benchmarks, GPT-5.1 scores higher on 26 and Qwen3.5 122B A10B on 7. The widest gap is critpt, where GPT-5.1 scores 4.9 against 0.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $1.25 per million input tokens).

GPT-5.1 vs Qwen3.5 122B A10B

Across 33 shared benchmarks, GPT-5.1 scores higher on 26 and Qwen3.5 122B A10B on 7. The widest gap is critpt, where GPT-5.1 scores 4.9 against 0.9. Qwen3.5 122B A10B is the cheaper of the two on tracked API pricing ($0.40 against $1.25 per million input tokens).

OpenAIvsAlibaba33 shared benchmarks267 head-to-head
BenchmarkGPT-5.1Qwen3.5 122B A10B
AA Agentic Index32.221.3
AA Intelligence37.532.8
AA-LCR76.770.3
AA-Omniscience5.4-41.5
arena_vision12501246
Artificial Analysis Coding Index49.445.7
browsecomp50.863.8
coding_arena_elo14211358
critpt4.90.9
gdpval2524.3
GPQA Diamond88.186.6
HLE28.547.5
IFBench72.976.1
LiveCodeBench v68778.9
MMLU-Pro8786.7
mmmlu9186.7
MMMU85.483.9
MMMU-Pro75.576.9
multichallenge63.461.5
OmniScience Accuracy37.724.4
OmniScience Non-Hallucination48.112.9
scicode43.342
SWE-bench Verified76.372
TauBench V3 - Banking15.915.3
Terminal-Bench 2.047.649.4
Terminal-Bench 2.152.447.6
Terminal-Bench Hard45.531.1
vectara_answer_rate10099.8
vectara_avg_summary_length254.486.4
vectara_factual_consistency89.188.8
vectara_hallucination_rate12.111.2
τ²-Bench Telecom (AA run)81.993.6
τ³-Bench1413.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.