VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which model leads t2-bench?
Across 19 models scored on t2-bench, Gemini 3.1 Pro leads at 99.3, ahead of Gemini 3 Flash Preview at 90.2. The median tracked score is 80.3, and the field spans 11.6 to 99.3.

t2-bench

Across 19 models scored on t2-bench, Gemini 3.1 Pro leads at 99.3, ahead of Gemini 3 Flash Preview at 90.2. The median tracked score is 80.3, and the field spans 11.6 to 99.3.

19 models tracked
Data as of August 25, 2026
#ModelVendorBest scoreRunsLast seen
1Gemini 3.1 ProGoogle99.322026-08-25
2Gemini 3 Flash PreviewGoogle90.212026-08-25
3GLM-5Z.ai89.712026-08-25
4Qwen3.5 397B A17BAlibaba86.712026-08-25
5Gemma 4 31BGoogle86.412026-08-25
6Gemma 4 26B A4BGoogle85.512026-08-25
7Gemini 3 ProGoogle85.412026-08-25
8Qwen3.5 35B A3BAlibaba81.212026-08-25
9DeepSeek-V3.2-SpecialeDeepSeek80.312026-08-25
10DeepSeek-V3.2DeepSeek80.312026-08-25
11Qwen3.5 4BAlibaba79.912026-08-25
12Qwen3.5 122B A10BAlibaba79.512026-08-25
13Qwen3.5 9BAlibaba79.112026-08-25
14Qwen3.5 27BAlibaba7912026-08-25
15Qwen3 Max (Reasoning)Alibaba74.812026-08-25
16Gemma 4 E4BGoogle57.512026-08-25
17Qwen3.5 2BAlibaba48.812026-08-25
18Gemma 4 E2BGoogle29.412026-08-25
19Qwen3.5 0.8BAlibaba11.612026-08-25

Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.