VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Gemma 4 31B or Qwen3.5 122B A10B?
Across 51 shared benchmarks, Gemma 4 31B scores higher on 16 and Qwen3.5 122B A10B on 34, with 1 level. The widest gap is RefSpatial-Bench, where Gemma 4 31B scores 4.7 against 0.7. Gemma 4 31B is the cheaper of the two on tracked API pricing ($0.15 against $0.40 per million input tokens).

Gemma 4 31B vs Qwen3.5 122B A10B

Across 51 shared benchmarks, Gemma 4 31B scores higher on 16 and Qwen3.5 122B A10B on 34, with 1 level. The widest gap is RefSpatial-Bench, where Gemma 4 31B scores 4.7 against 0.7. Gemma 4 31B is the cheaper of the two on tracked API pricing ($0.15 against $0.40 per million input tokens).

GooglevsAlibaba51 shared benchmarks1634 head-to-head
BenchmarkGemma 4 31BQwen3.5 122B A10B
AA Agentic Index14.421.3
AA Intelligence29.732.8
AA-LCR68.370.3
AA-Omniscience-47.9-41.5
AI2D8993.3
arena_vision12751246
Artificial Analysis Coding Index43.445.7
C-Eval82.691.9
CC-OCR75.781.8
coding_arena_elo13641358
critpt1.40.9
DeepPlanning2424.1
DynaMath79.585.9
ERQA57.562
gdpval15.524.3
GPQA Diamond85.786.6
HallusionBench67.467.6
HLE26.547.5
IFBench75.676.1
LiveCodeBench v68078.9
mathvision85.686.2
MathVista79.387.4
mmlu_redux93.794
MMLU-Pro85.286.7
mmmlu88.486.7
MMMU80.483.9
MMMU-Pro76.976.9
MMStar77.382.9
OCRBench86.192.1
OmniDocBench 1.580.189.8
OmniScience Accuracy2024.4
OmniScience Non-Hallucination18.112.9
RealWorldQA72.385.1
RefSpatial-Bench4.70.7
scicode43.442
SimpleVQA52.90.6
supergpqa65.767.1
SWE-bench Verified5272
t2-bench86.479.5
TauBench V3 - Banking14.815.3
Terminal-Bench 2.042.949.4
Terminal-Bench 2.143.447.6
Terminal-Bench Hard36.431.1
vectara_answer_rate10099.8
vectara_avg_summary_length75.886.4
vectara_factual_consistency92.688.8
vectara_hallucination_rate7.411.2
VITA-Bench4333.6
WideSearch35.260.5
τ²-Bench Telecom (AA run)65.593.6
τ³-Bench67.513.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.